Any AI accuracy figure that wasn't measured on a real sample of your contracts is a demo number dressed up as a deployment number. Before you spend on AI, demand the measurement the vendors won't run.
| Task | Data conditions | Realistic reliability |
|---|---|---|
| Issue-spotting on standardized agreements (NDAs) | Clean, standardized | High — at or above human |
| Extraction of common structured fields | Machine-readable, consistent | High, with QA sampling |
| Clause extraction across a legacy portfolio | Mixed formats, OCR'd | Moderate — human in the loop |
| Obligation / risk reasoning | Any | Low without verification |
| Open-ended portfolio analytics | Ungoverned repository | Unreliable — data-bound |
Drawn from Appresta.IQ's evidence review of agentic AI accuracy in business and contract analytics. Independent, peer-reviewed, and pre-registered sources are weighted above vendor-sponsored ones. Figures current to mid-2026. Start with a free readiness diagnostic at apprestaiq.com/legal-ai-readiness.