appresta.iq

Evidence review

How good is agentic AI at analytics, really?

A proof-point review of what autonomous AI actually delivers on business and contract analytics. It separates independent measurement from vendor marketing, and tests one premise: AI can only ever be as good as the data beneath it.

Independent studies weighted over vendor claims · figures current to mid-2026

01 / The headline

Accuracy is real — but it's bimodal, and it's data-bound

There are two accuracy numbers for agentic AI, and they're miles apart. Which one you get depends entirely on the task. Give it a narrow, standardized job over clean inputs and it matches or beats trained professionals. Point it at open-ended reasoning over messy enterprise data and accuracy falls off a cliff — often by 60 to 70 points versus the demo.

What moves that number is mostly the data: how clean, structured, and governed it is, plus the workflow around it. Model quality matters far less than the pitch decks suggest. So the framing that AI can only be as good as the data it's built on turns out to be the whole ballgame — and the evidence backs it hard.

The evidence, in four numbers

The number you're shown, and the one you'll live with

Every vendor demo runs on the vendor's own clean data. These are the independent measurements — the gap between demo and deployment is set by your data.

91%17%1

Text-to-SQL accuracy: same class of models on a clean academic benchmark vs. real enterprise data

~9%2

Of annual revenue lost to poor contract data and management — up to 15% on complex portfolios

17–33%3

Hallucination rate of legal-research AI marketed as “hallucination-free,” in the first independent evaluation

~95%4

Of enterprise generative-AI pilots showed no measurable P&L impact

02 / The cliff

The gap between the benchmark and the boardroom

The clearest example is text-to-SQL — the “ask your data a question in plain English” capability under most AI analytics pitches. The same class of models answered about 91% of questions correctly on a clean academic schema, and 17–21% against a real enterprise warehouse with thousands of columns, several SQL dialects, and the messy documentation any mature company actually has1.

The failure modes tell the story: schema-linking errors (the model can't reliably map “revenue” to the right column among thousands), dialect confusion, context overflow. Those are all data-environment problems, not model defects. A dashboard that's wrong 10% of the time, with no flag on which 10%, is worse than no dashboard — someone will make a call on it.

03 / Contract & legal AI

What the independent evidence says about contract work

The best case — narrow, standardized, clean.The landmark result put an AI against 20 experienced corporate lawyers spotting issues in NDAs: 94% accuracy against the lawyers' 85% average, in 26 seconds versus 92 minutes5. That result holds up — but NDAs are about the most standardized contract there is, the task was closed-form issue-spotting against a fixed rubric, the inputs were clean, and the vendor funded the study. It is the strongest honest case for contract AI, built on exactly the conditions that rarely survive a real portfolio.

The reality check — open-ended, novel, messy.The first pre-registered, independent evaluation of commercial legal AI tested the retrieval-augmented tools marketed as “hallucination-free”3:

ToolAccurateHallucinationNote
Lexis+ AI65%~17%Best performer tested; still needs verification
Westlaw AI-Assisted Research42%~33%Longer answers correlated with more errors
Ask Practical Law AIIncomplete or ungrounded on 60%+ of queries
GPT-4 (general model)~43%Baseline; the RAG tools beat it, none solved it

Grounding AI in a trusted document set cut hallucinations versus a raw model — it did not get rid of them. And when the vendors got measured, they disputed the methodology rather than publishing their own independent benchmarks. Vendor clause-extraction figures of 94%+ may hold for the clause types and clean inputs they tested on, but they're self-reported and measured on the vendor's terms, not your portfolio — treat them as claims, not facts.

What a wrong field costs. Poor contract management already erodes about 9% of annual revenue, up to 15% on large, complex portfolios2, and 71% of companies can't reliably locate at least 10% of their own contracts6. Hand that portfolio to an AI extractor and the errors cluster where they hurt most — renewal dates, amendments, party names, carved-out clauses. A blank renewal field gets flagged and checked; a wrong one looks complete, so the team trusts it and acts on it. Wrong data does more damage than missing data.

04 / The core question

“As good as the data” is measurable now

The link between input quality and AI output has moved from folk wisdom to something you can put a number on. In a controlled example, dropping placeholder zeros into a single column cut a model's predictive fit from R² 0.96 to 0.76 — no error thrown, the output just quietly got worse in step with the input9. That is the dangerous kind of failure: the kind that never raises its hand.

No prompt is clever enough to fix data that's incomplete, inconsistent, or undocumented. Agentic AI has no independent source of truth — only a probabilistic model of language plus whatever data you wire in. The ceiling on any factual task is set by that connected data. Improve it and the ceiling rises; leave it messy and no model, however sharp, gets past the ceiling the data sets.

05 / Deployment reality

What survives contact with an actual organization

Lab accuracy is necessary but not enough. About 95% of enterprise generative-AI pilots have shown no measurable P&L impact4; MIT blames a “learning gap” — poor integration, tools that don't adapt — more than raw model quality. Gartner predicts over 40% of agentic-AI projects will be canceled by the end of 20278. These are layers of the same wall: bad data caps accuracy, poor integration caps value even when accuracy is fine, and with no measurement you can't tell which problem you have.

The productivity mirage. In a randomized trial, experienced developers expected AI to make them ~24% faster and felt ~20% faster — but were measured 19% slower, most of it eaten by prompting, reviewing, and fixing the output7. Don't trust self-reported time savings: the verification work AI hands back to the human is real, usually invisible in a demo, and often big enough to wipe out the headline gain — doubly so in legal and contract work, where every output must be checked against a primary source.

06 / Synthesis

The position, and how it maps to contract analytics

Agentic AI is genuinely accurate on contract analytics inside one envelope: standardized instruments, common clause types, clean inputs, a defined rubric. Step outside it — novel terms, mixed legacy portfolios, open-ended judgment, dirty repositories — and accuracy degrades unpredictably, often without a sound. Since data integrity is the binding constraint and model horsepower matters much less, the order of operations is simple: diagnose data readiness, fix the gaps, then automate. Run analytics on an ungoverned repository and what comes out is confident, well-formatted error at scale.

Plotted against those two axes — task type and data condition — contract analytics sorts into a map you can defend to a GC:

judgmentWhat you ask AI to doextraction
ungovernedData conditionclean, structured
ReliableModerate — verifyUnreliable
Illustrative positioning — a conceptual map, not measured coordinates.
TaskData conditionRealistic reliability
1Issue-spotting on standardized agreements (NDAs, common forms)Clean, standardizedHigh — at or above human
2Extraction of common structured fields (dates, parties, renewal terms)Machine-readable, consistentHigh, with QA sampling
3Clause extraction across a heterogeneous legacy portfolioMixed formats, OCR'd, inconsistentModerate — human in the loop
4Obligation or risk reasoning requiring legal judgmentEven on clean dataLow without verification
5Open-ended portfolio analytics ("aggregate exposure to X?")Ungoverned repositoryUnreliable — data-bound

07 / Where this could be wrong

The honest counter-arguments

The models are moving fast. Most studies here tested 2024 and early-2025 models; agentic scaffolding targets exactly these failure modes. Expect the numbers to improve — the open question is whether they improve enoughto clear the binary reliability bar high-stakes contract work demands. Improvement is near-certain; enough improvement isn't.

It only needs to beat the human baseline.For high-volume, low-stakes review, an AI that's faster and no worse than an average human is a real win. But a well-run legal function isn't benchmarking against the average solo practitioner, and the errors are asymmetric — one hallucinated obligation in a board-level analysis costs more than a hundred correct extractions save.

MIT blames integration, not data.Fair — it means “fix your data” is necessary but not sufficient. Treat data quality as the accuracy ceiling and integration as the value ceiling: two separate constraints, both real.

08 / What to do with this

Insist on on-your-data benchmarking

Any accuracy figure that wasn't measured on a real sample of your own contracts is a demo number dressed up as a deployment number. Before you spend on AI, diagnose whether the estate is even in shape for AI to be accurate — a readiness check is both the honest first move and the step the failure data says decides ROI. Then bake a controlled pilot into every evaluation, and keep a human in the loop on anything that matters.

The pilot protocol

What to demand before you believe any headline number

  • 100+ of your own contracts, sampled across types and complexity — not the vendor's clean demo set.
  • Precision and recall reported per clause type, so you see where accuracy holds and where it breaks.
  • Money-and-deadline fields called out separately: renewal dates, amendments, party names, liability caps.
  • A verification budget: measure the human time to check outputs, and count it against the headline time saved.
  • A readiness baseline first — is the estate clean, complete, and structured enough for AI to be accurate at all?

Sources

  1. 1. In independent testing, the same class of AI models answered 91% of natural-language data questions correctly on a clean academic database — and 17–21% on real enterprise data. Lei et al., Spider 2.0 (ICLR)The gap is a data-environment problem — schema-linking, dialect confusion, context overflow — not a model defect.
  2. 2. Poor contract data and management erodes roughly 9% of annual revenue — and up to 15% on large, complex portfolios. World Commerce & Contracting benchmark
  3. 3. In the first independent, pre-registered evaluation, legal-research AI marketed as “hallucination-free” still hallucinated on roughly one in six to one in three queries. Magesh et al., Stanford RegLab / HAI, Journal of Empirical Legal Studies (2025)Grounding in a trusted document set reduced hallucinations versus a raw model — it did not eliminate them.
  4. 4. About 95% of enterprise generative-AI pilots have shown no measurable P&L impact. MIT NANDA, The GenAI Divide: State of AI in Business 2025MIT attributes this mostly to an integration and organizational “learning gap,” not raw model quality — this is the value ceiling, distinct from the data-quality accuracy ceiling. It is not evidence that AI does not work.
  5. 5. In a 2018 study, an AI matched experienced lawyers on spotting issues in NDAs — 94% accuracy vs. their 85% average, in 26 seconds vs. 92 minutes — on the most standardized contract there is, against a fixed rubric, on clean inputs. LawGeex (2018), vendor-sponsoredStudy funded by the vendor. This is AI's home-turf envelope — a standardized instrument, a closed-form task, clean inputs — not a general contract-review result. Never cite the bare 94%.
  6. 6. 71% of companies admit they can't reliably locate at least 10% of their own contracts. World Commerce & Contracting
  7. 7. In a 2025 randomized trial, experienced developers expected AI to make them ~24% faster and felt ~20% faster — but were measured 19% slower, most of it lost to reviewing and fixing AI output. Becker et al. (METR), arXiv 2507.09089METR flagged self-selection limits and an unusually hard setting (experts on code they knew cold); the durable finding is the perception gap, not the exact figure.
  8. 8. Gartner predicts over 40% of agentic-AI projects will be canceled by the end of 2027, citing cost, unclear value, and weak risk controls. Gartner press release, June 2025
  9. 9. In a controlled example, dropping placeholder zeros into a single column cut a model's predictive fit from R² 0.96 to 0.76 — no error raised, the output just quietly got worse. Illustrative controlled experiment (Vodworks)Illustrative worked example, not a benchmark. Consistent with the Spider 2.0 schema-linking failure modes and the BEAVER warehouse benchmark.

Independent, peer-reviewed, and pre-registered sources are weighted above vendor-sponsored ones; vendor figures are labeled as claims. Figures current to mid-2026.