

Everybody watches the pile an AI first pass flags. Almost nobody looks at the pile it clears. That second pile is where a missed deviation actually costs you, because nobody's eyes ever land on it again.
Here's the setup at most teams running an AI first pass on incoming contracts: the tool reads the document, checks it against the playbook, and sends back a short list of things that don't match. A reviewer works through that list, makes calls, and moves on. The contracts that came back with nothing flagged just get filed. That's the whole point of running a first pass, to stop reading things that already look fine.
The trouble is that "nothing flagged" and "actually fine" are two different claims, and only one of them has been checked.
A first pass can clear a contract for reasons that have nothing to do with the contract being clean. The clause exists but uses phrasing the tool wasn't tuned to recognize as a match to your playbook position. A defined term shifts meaning three pages later than the clause it should be read against. An exhibit reference points to a document that was swapped after signature. None of that shows up as a flag, because a flag only fires when something looks like a deviation. A miss looks exactly like a pass.
General-purpose AI models make this kind of error more often than most people assume when a task looks like reading text and applying rules to it. A widely cited Stanford study found that large language models produced hallucinated or inaccurate answers between 69% and 88% of the time when asked general legal questions, and got it wrong roughly a third of the time even on a narrower set of legal queries built specifically to test accuracy. Contract clause matching against a playbook is a much narrower job than open legal research, and a purpose-built review tool checking defined positions is not the same as a general chat model answering an open question. But the general lesson holds. Text that reads as confidently correct on a screen is not the same thing as text that's actually been checked.
Pull a random sample of the contracts your first pass cleared. Not the ones it flagged, the ones it didn't touch. Have a reviewer who wasn't involved in building the playbook actually read them against it. Ten or fifteen contracts a month is enough to start.
Count how many of those "clean" contracts had something in them that should have been caught. That number is your false-clear rate, and it's the only honest measure of whether the first pass is doing what you think it's doing. A playbook that catches every deviation on the contracts it flags but misses a chunk of what it clears has a blind spot nobody's measured yet, no matter how clean the flagged pile looks.
About 60% of legal teams still don't operate from a written playbook at all, according to a 2024 survey of in-house counsel and legal ops professionals. If that's your team, sampling the cleared pile won't tell you much, because there's no defined standard for "clean" to be checked against yet. How to build a contract playbook an AI can enforce is the piece to read first in that case, because a sampling program on a vague playbook just proves the playbook is vague.
The contracts worth sampling first aren't a random cross-section. Start with contract types where the counterparty usually sends their own paper rather than yours, since that's where phrasing drifts furthest from what a tool was tuned to recognize. Pull extra from whichever reviewer's queue clears the most contracts with zero flags, because a suspiciously clean queue is either a great sign or a sign nobody's checking. And don't skip anything reviewed in the first month after a playbook update, before you know whether the new positions actually get applied the way they were written.
None of this replaces the reviewer's judgment on what's worth escalating. It just tells you, with actual numbers instead of a hunch, whether the tool clearing a contract is a decision you can stand behind or a gap nobody's found yet. If you're running an AI first pass on incoming contracts, that sampling loop is worth building before you trust the volume it's clearing.
Try goHeather free and run a batch of contracts through a first pass, then pull a few of the clean ones and read them yourself.
This is legal information, not legal advice; consult a lawyer for legal advice.
Jeff Dutton is a lawyer who advises on technology, corporate, privacy, commercial, employment and real estate law.
Jeff founded his own small law firm, Dutton Law, in 2016 (and merged it with a larger firm in 2019). Before that, Jeff was a prosecutor and a commercial law lawyer at a national boutique law firm.
Jeffrey is a frequent lecturer on legal matters and has been published in newspapers and trade journals. In addition, Jeff was the editor and co-author of a leading employment law text for lawyers for many years.
Education:
Western University, BA (2009)
University of Ottawa, Faculty of Law, JD (2012)

Get the latest contract tips, updates, and exclusive content straight to your inbox. Subscribe now and never miss out on what's new in contract law or at goHeather!
Our AI sifts through each clause, identifying potential risks. This enables us to provide quick yet comprehensive contract reviews, equipping you with the legal information you need to make informed decisions.