

Once your AI review starts clearing contracts with no findings, somebody on your team has to decide whether to trust that pile or reread it. Teams tend to land on one of two options, and both have a problem: reread everything, which erases the time you just saved, or trust the clear stamp completely, which means nobody notices when a rule quietly stops catching what it was written to catch. There's a third option, and it doesn't take more work than either of the other two: sample the cleared pile, but sample by rule instead of by contract.
A contract that comes back with zero findings didn't necessarily fail to have problems. It just didn't trip anything in the playbook that ran against it. Those are different things, and the gap between them gets wider as volume goes up, because nobody is reading behind the AI line by line anymore. That's the whole point of running review at volume. It also means the queue is producing a claim, "nothing here," that nobody checks unless you build a way to check it. We've written before about what's worth checking in the pile a first pass clears, and spot-checking a handful of random contracts is a weak test for the specific failure worth worrying about.
The instinct is to pull a random 5 or 10 percent of cleared contracts each week and have someone read them cold. That catches obvious problems, but it's a poor way to find a rule that has quietly stopped firing, because a small sample can run for months without ever landing on the narrow slice of contracts that rule is supposed to cover.
Auditors ran into this exact math before AI review existed. Internal control testing used to lean on small fixed samples, and the numbers behind that approach are worse than most people assume: testing two examples of a monthly control leaves an 83% chance of missing a single control failure occurring during the year, and testing thirty examples of a control that runs more than once a day can still leave close to a 99% chance of missing one. A handful of randomly picked cleared contracts runs the same risk. Which contracts you read matters more than how many.
Pull the finding log for every rule in your playbook over the last quarter, or the last month at high volume, and look at hit rates instead of contract counts. A rule that fires on 40% of your MSAs and a rule that hasn't fired once in six months deserve two different responses. The first is doing its job. The second is either genuinely rare, which is fine, or it's a rule the AI has stopped recognizing in current drafting, which a random spot check will almost never catch, since you'd need to happen to pull one of the few contracts where it should have applied.
Once you have that hit-rate list, find contracts you're fairly confident should have tripped a silent rule, an MSA with an uncapped indemnity clause when your uncapped-indemnity rule hasn't fired all quarter, for instance, and read those specifically against that rule. That's a targeted sample built around a hypothesis, not a random draw, and it tests the one thing a random sample is worst at finding.
Whoever wrote the rule shouldn't be the one auditing it, for the same reason a second set of eyes helps anywhere else: they already believe the language is airtight. Rotate this through the team, or hand it to whoever currently owns the playbook, and log the results in the same place you log playbook edits. A rule that turns out to be quietly broken should show up in the same record as the fix, not in a private note nobody else sees.
This isn't a one-time cleanup exercise. NIST's framework for managing AI risk treats measurement as something that keeps running after a system is live, not a check you do once before launch, and calls for testing to continue as conditions change. A contract playbook works the same way. Rules get edited, counterparties change how they draft, and a rule that caught everything in March can go quiet by July without anyone deciding to break it.
If you want a sanity check on this before you build a whole audit calendar around it, run your contract review workflow against a stack of contracts you've already reviewed by hand and see which known issues get caught and which slip through. That tells you roughly where your rule-level blind spots are before you spend a quarter building a formal process around guesses.
Rule-level auditing won't catch everything a careful human reader would. It's built to catch one specific thing: a rule that used to work and quietly doesn't anymore, which is exactly what a general spot check is least likely to find.
Try goHeather free if you want to see what your own playbook catches, and misses, on a contract you've already reviewed by hand.
This is legal information, not legal advice; consult a lawyer for legal advice.
Jeff Dutton is a lawyer who advises on technology, corporate, privacy, commercial, employment and real estate law.
Jeff founded his own small law firm, Dutton Law, in 2016 (and merged it with a larger firm in 2019). Before that, Jeff was a prosecutor and a commercial law lawyer at a national boutique law firm.
Jeffrey is a frequent lecturer on legal matters and has been published in newspapers and trade journals. In addition, Jeff was the editor and co-author of a leading employment law text for lawyers for many years.
Education:
Western University, BA (2009)
University of Ottawa, Faculty of Law, JD (2012)

Get the latest contract tips, updates, and exclusive content straight to your inbox. Subscribe now and never miss out on what's new in contract law or at goHeather!
Our AI sifts through each clause, identifying potential risks. This enables us to provide quick yet comprehensive contract reviews, equipping you with the legal information you need to make informed decisions.