Appresta.IQ · Buyer's checklist

The on-your-data pilot protocol

Any AI accuracy figure that wasn't measured on a real sample of your contracts is a demo number dressed up as a deployment number. Before you spend on AI, demand the measurement the vendors won't run.


What to demand before you believe any headline number

  1. Test on your own portfolio.100+ of your own contracts, sampled across types and complexity — not the vendor's clean demo set.
  2. Precision and recall, per clause type.Reported separately for each clause type, so you see where accuracy holds and where it breaks.
  3. Isolate the money-and-deadline fields.Renewal dates, amendments, party names, liability caps — the fields that carry money and deadlines are where AI errs most, because that's where the language turns non-standard.
  4. Budget for verification.Measure the human time to check outputs against the source, and count it against the headline time saved. The checking work is real and usually invisible in a demo.
  5. Establish a readiness baseline first.Is the estate clean, complete, and structured enough for AI to be accurate at all? Accuracy is an output of data readiness — diagnose it before the model ever ships.

Set expectations by task tier

TaskData conditionsRealistic reliability
Issue-spotting on standardized agreements (NDAs)Clean, standardizedHigh — at or above human
Extraction of common structured fieldsMachine-readable, consistentHigh, with QA sampling
Clause extraction across a legacy portfolioMixed formats, OCR'dModerate — human in the loop
Obligation / risk reasoningAnyLow without verification
Open-ended portfolio analyticsUngoverned repositoryUnreliable — data-bound

Drawn from Appresta.IQ's evidence review of agentic AI accuracy in business and contract analytics. Independent, peer-reviewed, and pre-registered sources are weighted above vendor-sponsored ones. Figures current to mid-2026. Start with a free readiness diagnostic at apprestaiq.com/legal-ai-readiness.