Talk to an Expert

Legal AI That Shows Its Sources. Every Time.

Hallucination is the reason legal AI adoption stalled for years, and it's still the first question every serious legal tech founder asks a development partner. Here's exactly how we control for it — not a marketing promise, the actual methodology.

Annotated demo · Self-built
ContractIQ — Grounded Answer Demo
CONTRACT_v2.pdf
§3.1 Confidentiality — Each party shall maintain the confidentiality of all disclosed information...
§4.2 Termination — Either party may terminate this Agreement upon thirty (30) days' prior written notice to the other party. §7.1 Indemnification — Each party shall indemnify and hold harmless the other from claims arising out of...
What's the termination notice period in this agreement?
The termination notice period is 30 days, per the clause highlighted at left.
Source: §4.2, p.3 98% grounded

Why This Is the Question That Matters Most

Hallucination isn't a hypothetical risk in legal AI — it has already produced real, public consequences. In 2023, lawyers were sanctioned by a federal court after submitting a brief containing case citations fabricated by an AI tool, a widely reported incident that made 'unauthorized practice of law' and 'AI hallucination' live search concerns rather than abstract ones. Law firms rejected early legal AI tools specifically because hallucination rates were too high for professional use — and it's only as grounding techniques matured that production-grade legal AI became realistic.

That history is why we don't lead with a generic 'we use AI' pitch. We lead with the methodology that keeps outputs grounded, verifiable, and safe to put in front of a client or a court.

Our Five-Layer Accuracy Methodology

01
Layer 1 — Grounded Retrieval (RAG), Not Open-Web Generation
Every model response is generated from retrieval over your verified source documents — contracts, case law, firm precedent, statutes — never from the model's general training data alone. If the answer isn't in the retrieved source set, the system is designed to say so rather than generate a plausible-sounding guess.
02
Layer 2 — Citation-Level Verification
Every claim in an output is traced back to the exact paragraph, clause, or page it came from, with a confidence score attached. This is the same principle behind the demo on our main legal tech page: a contract clause extractor that shows the source paragraph next to every answer, not just the answer.
03
Layer 3 — Evaluation Benchmarks
Before launch, we build a golden dataset specific to your use case — known-correct question/answer pairs — and benchmark citation precision and faithfulness against it. We also run adversarial testing designed specifically to surface hallucination edge cases (ambiguous queries, out-of-scope questions, contradictory source documents) rather than only testing the easy cases.
04
Layer 4 — Human-in-the-Loop Review Gates
The review gate is configurable by risk level. Lower-stakes internal research queries might not need a human check on every response; a client-facing contract summary or litigation analytics output typically routes through attorney sign-off before it's delivered. We design the gate placement with you, not as an afterthought bolted onto a finished product.
05
Layer 5 — Confidentiality & Security Architecture
Privileged and confidential documents are handled under NDA, with encryption at rest and in transit, role-based access controls, and SOC 2-aligned data handling practices. We do not use client documents to train shared or third-party models, and we support VPC-isolated or on-premise deployment for clients with strict data-residency requirements.

Demo Video

The same annotated product screenshot featured on our legal tech MVP page: a self-built RAG-based contract clause extractor that cites its exact source paragraph for every answer, with a confidence score attached. A full walkthrough video is planned — see the production spec at the end of this deck — and will replace this screenshot once it's filmed.

What We Mean by 'Benchmark' — Honestly

We're not going to put an unearned universal hallucination-rate percentage on this page — any legal AI vendor who quotes one flat number without publishing methodology should be treated skeptically, and we'd rather earn that scrutiny than avoid it. What we do instead: build a benchmark specific to your document types and use case, measure against it before launch, and share the methodology and results directly with you as part of the engagement. Accuracy is use-case-specific; a benchmark that isn't should raise questions, not confidence.

Where This Fits In Your Evaluation

If you're comparing build-vs-buy-vs-outsource for a legal AI product, this methodology is the thing to press every vendor on, including us: ask what's grounded vs. generated, how citations are verified, what the evaluation benchmark actually measures, and where the human review gate sits. If a vendor can't answer those four questions concretely, that's the signal — not the sales deck.

Book a Technical Discovery Call — bring your hardest accuracy question.

Book a Technical Discovery Call