4
min. read

Your playbook decides what an AI contract review flags

Jeff Dutton
By
Jeff Dutton
Lawyer
Last update:
September 4, 2026

Review any Contract With AI Before you Sign it

The rules you attach to an AI review decide what it flags. Change those rules and the same contract comes back with a different list of problems, and neither list is wrong.

That matters if you run a review queue, because the finding count is the first number anyone looks at. A reviewer opens a file, sees three flags, and reads that as easy paper. Someone else opens the same file under a different rule set, sees seventeen, and escalates it. The contract is identical in both cases. What differs is the rule set each of them had attached to the run.

What happened when one MSA ran against five playbooks

One master services agreement in our own account got reviewed five separate times, each run against a different playbook, using the same contract review workflow. The finding counts came back 17, 16, 8, 8 and 8.

Nothing about the document changed between runs. What changed was the standard it was being measured against. A playbook that takes a hard position on liability caps and data indemnity turns up things that a playbook built around payment terms and termination never goes looking for.

Run a contract with no playbook attached and you get a general read: unusual language, one-sided terms, drafting problems. That's useful. It is also completely unopinionated about what your company will actually sign.

Vague instructions give you wobbly output

There is a machine learning reason underneath this.

A 2026 research paper on prompt sensitivity compared two ways of instructing a language model on classification work: vague prompts that barely describe the job, and prompts that spell out specific instructions. The vague ones swung around far more from run to run. Spelling the instructions out cut the swing.

Your playbook is the specific instruction. "Review this and flag the risks" is the vague one. Ask the broad question and you get back an answer shaped by whatever the model treats as a reasonable concern in general. Ask it against forty numbered rules that each carry a primary position, a fallback and a walk-away line, and the question is narrow enough to get a steady answer.

A finding count is not a quality score

If you track review output across a team, a finding count only means something next to the playbook that produced it. Playbooks vary hugely in size; the ones we see run from a dozen rules up past eighty. Twelve findings off an eighty-rule playbook and twelve findings off a twelve-rule playbook are not the same event, and averaging them together tells you nothing.

Comparing two reviewers runs into the same thing. If one had the shared team playbook attached and the other was running a private set they built themselves, those numbers were never comparable. A playbook can sit with one person or be shared across a team, which is fine right up until nobody can say which one ran on which file. The same thing happens to contracts sitting in the queue when a playbook rule changes underneath them, where the version that ran depends on what day the file arrived.

The part that isn't your playbook

Even holding the playbook fixed, you won't get identical output every time. Thinking Machines Lab ran one prompt a thousand times at temperature zero, which is the setting meant to make output repeatable, and got eighty different completions. Their explanation has nothing to do with the model being creative. It comes down to how inference servers batch requests together under changing load. They also showed it can be engineered away, which isn't how a typical setup runs today.

For contract review that means the wording of a finding can shift between runs, and a borderline item can show up one time and not the next. None of that makes the review unreliable. It does mean finding count is a soft number, and reading a move from nine to eight as a signal about anything is reading noise.

What is actually under review here

When a review comes back with a list you disagree with, the instinct is to argue with the software. The rule that produced the finding is the more useful thing to look at, because somebody on your team wrote that rule, and it is going to run on every contract after this one.

If you keep dismissing the same finding, the rule behind it is wrong, and it will keep coming up until someone edits it. A problem the review never raises is usually one nobody wrote a rule for. Either way the argument is about your own standards, and the software is showing you where they currently sit.

Try goHeather free if you want to see what your playbook flags on a contract you have already reviewed by hand.

This is legal information, not legal advice; consult a lawyer for legal advice.

About the author

Jeff Dutton is a lawyer who advises on technology, corporate, privacy, commercial, employment and real estate law.

Jeff founded his own small law firm, Dutton Law, in 2016 (and merged it with a larger firm in 2019). Before that, Jeff was a prosecutor and a commercial law lawyer at a national boutique law firm.

Jeffrey is a frequent lecturer on legal matters and has been published in newspapers and trade journals. In addition, Jeff was the editor and co-author of a leading employment law text for lawyers for many years.

Education:

Western University, BA (2009)
University of Ottawa, Faculty of Law, JD (2012)

Jeff Dutton
By
Jeff Dutton
Lawyer

Stay Updated on All Things Contract Law with goHeather

Get the latest contract tips, updates, and exclusive content straight to your inbox. Subscribe now and never miss out on what's new in contract law or at goHeather!

Thank you! You will receive an email to confirm your subscription.
Oops! Something went wrong while submitting the form. Try again later.

Review any Contract With AI Before you Sign it

Our AI sifts through each clause, identifying potential risks. This enables us to provide quick yet comprehensive contract reviews, equipping you with the legal information you need to make informed decisions.

Related articles