All insights
AI & Technology

The AI "Hallucination" Problem — and Why Contract Review Is a Different Case Than Legal Research

Headlines about AI inventing fake legal citations are real and growing — but they describe a different failure mode than reviewing a document you already have in hand. The distinction matters more than most coverage explains.

June 13, 20267 min readRedline Construction Solutions

Key takeaways

  • A public database has tracked over 1,300 documented cases of AI-generated fake legal citations in court filings as of April 2026, up from under 100 a year earlier.
  • Every documented sanctions case involves AI legal research or drafting — inventing case law that doesn't exist — not AI contract review of a document already in hand.
  • Legal research asks AI to recall facts from memory; contract review asks AI to analyze text you've already given it — a fundamentally different, more checkable task.
  • The safeguard that actually prevents the contract-review failure mode is verification: confirming a quoted clause is really in the document before presenting it as a finding.
  • Not all legal AI is built the same way — the difference between 'hallucination-prone' and 'grounded' is a design choice, not an unavoidable property of AI.
  • Understanding this distinction lets you evaluate a specific tool on the right question: does it verify against your document, or does it just sound confident?

The headlines are real — and worth taking seriously

The stories are not exaggerated: lawyers have been sanctioned for submitting briefs with AI-invented case citations that don't exist. A public tracking project, the AI Hallucination Cases Database, had documented more than 1,300 such cases globally by April 2026 — up from under 100 in mid-2025. Federal courts have issued sanctions ranging from a few thousand dollars to tens of thousands, and at least one attorney has been suspended over it.

That's a genuine, serious, and rapidly growing problem. Anyone using AI in a legal context should take it seriously — which is exactly why it's worth understanding precisely what kind of task produces this failure, and what kind doesn't.

The 13x growth in documented cases in under a year is itself worth sitting with — it suggests this isn't a problem that's being solved as the underlying models improve, at least not yet, in the specific context of unverified legal research use.

It's a research problem, not a review problem

Look closely at the documented cases and a pattern emerges: every one involves AI legal research or brief drafting — asking a model to recall or generate case law, then citing what it produced without checking that the cases are real. That's a fundamentally different task from contract review. Legal research asks AI to pull facts from its training memory, where it has no way to verify against an authoritative source in real time. It's answering from recall, and recall can be confidently wrong.

Contract review is a different shape of problem entirely: the AI isn't being asked to remember a court case from months of training data — it's being asked to analyze a document that's sitting right in front of it, provided in the same request. That's a task built for verification, not recall.

Put another way: legal research asks "what does the law say," a question the model can only answer from what it happened to absorb during training. Contract review asks "what does this specific document say," a question the model can answer by actually reading the text you gave it — a categorically easier and more checkable task.

Why that distinction actually matters

Because the contract is provided as input rather than pulled from memory, a well-built system can do something legal research tools structurally cannot: check its own answer against the source. Before presenting a finding — "this clause says X" — the system can confirm that the exact quoted language actually appears in the document you uploaded. If it doesn't, that's a red flag the system can catch and surface, rather than a confident error that slips through.

This is exactly the failure mode a contract-review tool should be engineered to close off. It's not a guarantee that every finding is legally correct — a real attorney's judgment call is still a judgment call — but it does eliminate the specific, well-documented "invented the source material" failure that's driving the hallucination headlines.

This distinction is also why comparing a contract-review tool's reliability to a legal-research tool's documented hallucination rate is genuinely apples-to-oranges — they're different tasks with different failure modes, and treating them as equivalent risk categories obscures more than it clarifies.

Not all AI tools are built this way

This is the part that gets lost in general "AI hallucinates" coverage: whether a tool verifies its findings against source text is a design decision, not an unavoidable property of AI itself. A general chatbot with no verification step and a purpose-built contract tool with a document-anchoring and second-pass verification layer are running the same underlying model technology but producing very different reliability profiles.

That's the practical question worth asking about any contract-review tool you're evaluating: does it show you where in your document a finding came from, and has it checked that the quote is real? If the answer is no, or if you can't tell, treat its output the same way you'd treat an unverified web search result — a starting point, not a finding.

Vendors who've built this verification layer tend to be happy to explain it in specific, technical terms, since it's a genuine differentiator. Vague reassurance ("our AI is very accurate") without a specific explanation of how it checks its own work is itself a useful signal.

What this means for how you use AI on contracts

The headline risk with AI right now is real, but it's concentrated in legal research and drafting, not in reviewing a document you already have. That doesn't mean contract-review AI is risk-free — it means the risk is different and more manageable, provided the tool is built to verify rather than just generate.

If you're evaluating any AI tool for contract work, ask specifically how it handles this. See how RCS anchors and verifies every finding against your actual contract before you decide how much to trust any tool's output.

The broader lesson generalizes well beyond contract review, too: whenever you're evaluating an AI tool for any task, ask specifically whether that task involves the model recalling facts from memory, or analyzing material you've handed it directly — because that single distinction predicts a great deal about how much you should actually trust its output.

As the documented hallucination-case count keeps climbing in the legal-research context, expect the industry conversation to keep conflating the two failure modes — which is exactly why it's worth being the person in the room who can explain, specifically, why contract review isn't the same risk category.

This article is general information about construction contracting and law, not legal advice. Construction law varies significantly by jurisdiction and project. Consult qualified counsel about your specific contract and circumstances.

Put this into practice on your own contracts.

Redline Construction Solutions applies your firm's non-negotiables and jurisdiction-aware standards to mark up a contract automatically — and returns it ready for your team to review.

See how it works
Redline Construction Solutions

Redline construction contracts in minutes — not weeks. Reviewed against current law and your standards, in your own private cloud.