Human review step in an AI workflow: how to design it
How to design human review of AI-drafted work: what a person must check, what they should see, and how to tell whether the review catches anything.
King & Company
In short
Design the review as part of what the workflow produces. Decide by consequence which fields a person checks every time, have the workflow show the source quote and page beside each of them, let it report what it could not find, and move arithmetic and completeness checks into code. Then log what reviewers change, because that record tells you whether the review is working.
The human review step in an AI workflow works when it is designed as part of what the workflow produces: a short list of items a person must check, each one shown beside the source text that supports it. A reviewer who is handed a finished draft and asked to approve it will usually approve it, so the design work is deciding what they check and what they see while they check it.
This article is for the partner, team lead, or senior practitioner whose name goes on the work, and who has to decide what a person looks at before an AI-drafted abstract, memo, schedule, or proposal goes to a client.
Why "keep a human in the loop" is not a design
Guidance on this topic often comes down to two instructions: keep a human in the loop, and send low-confidence output to a person. Both are reasonable, and neither tells a reviewer what to do on a Tuesday afternoon with eleven abstracts in the queue.
Confidence thresholds are a weak foundation for document work. An evaluation of confidence elicitation in language models, published at ICLR 2024, found that models "tend to be overconfident" when they state their confidence in words. A workflow that routes only the items the model says it is unsure about will pass along every error the model was sure of.
Automation bias: what happens when the draft looks finished
A clean, well-formatted draft reads as correct. The reviewer skims it, finds nothing odd, and signs. This tendency is called automation bias, and the EU AI Act names it in its text. Article 14 says the people assigned to oversee a high-risk AI system must be enabled, as appropriate and proportionate, "to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias)," and to "disregard, override or reverse the output."
That provision applies to systems the Act classifies as high-risk. We cite it as an authoritative statement of the problem, and it should not be read as a rule that binds a U.S. brokerage or accounting firm. Whether any regulation applies to your firm's use of AI is a question for your own counsel.
Awareness helps, and the more dependable fix is to change what the reviewer is given, which is what the rest of this article covers.
Decide what a person must check, by consequence
Start with the document your team sends out and ask of each field or claim what happens if it is wrong. Sort the answers into three groups.
| Group | Test | Review rule | Example in a lease abstract |
|---|---|---|---|
| Must check | An error reaches a client or moves money | A person checks it against the source every time | Option notice date, base rent, escalations |
| Skim | An error is embarrassing but gets caught downstream | A person reads it once | Permitted use, description of the premises |
| Sample | An error has little consequence | Spot check a few per batch | Formatting, party addresses |
The first group should be short. In a lease abstract it is the dates, the dollars, and the options. In a tax workpaper it is the figures that flow to the return. In a diligence memo it is every number and every statement about what a contract permits. The same sorting applies to any of these, and our article on abstracting a commercial lease with AI works through one case field by field.
Make the workflow show its evidence
A reviewer who has to reopen a long lease to confirm a date is doing the original work a second time, and under deadline they will stop doing it. The workflow should put the supporting quote and its page number beside each value, so that review is a comparison of two things on one screen.
This is what Anthropic's guidance on grounding is for. Its documentation on reducing hallucinations recommends asking Claude to extract word-for-word quotes before it performs a task on a long document. It also suggests having Claude find a supporting quote for each claim and retract any claim it cannot support. For teams building on the API, the citations feature returns the exact passages that support each claim, and the documentation says those citations "are guaranteed to contain valid pointers to the provided documents." The same page notes a limit worth knowing: PDFs that are scans of documents and contain no extractable text are not citable. In our view, a scanned document should be converted to text before it enters the workflow and should get a heavier review.
The same Anthropic page is direct about what these techniques do not do. It says they "don't eliminate them entirely" and tells builders to always validate critical information. A quote beside a value makes checking fast. A person still has to do the checking.
Let the workflow say what it could not find
A blank field with a note is more useful to a reviewer than a plausible guess. Anthropic's documentation says that explicitly giving Claude permission to admit uncertainty "can drastically reduce false information."
Build two outputs into the workflow alongside the draft:
- A not-found list: every required field the source documents did not answer, stated as "not found," with nothing filled in.
- A conflicts list: every place two sources disagree, such as a base lease and a later amendment, with both quotes and both page numbers.
The reviewer reads both lists in full. They are short, and they point at the places where judgment is needed.
Put the mechanical checks in code
Some checks need no judgment, and a person should not spend attention on them. Have the workflow run them and report only the failures:
- Totals that must tie, such as a rent schedule that sums to the stated annual rent.
- Dates in a sensible order, such as commencement before expiration.
- Required fields present and in the right format.
- Figures that appear in two places and must match.
These are ordinary rules written in code, and they return the same answer every time. They belong in the build from the first version, which is one reason we treat review as part of how an AI workflow gets built from the start.
Who should review, and how long it should take
The reviewer should be the person who owns the work and would have produced it by hand. They know what wrong looks like: a notice period that is unusual for that landlord, a deduction that does not fit that client. A reviewer without that knowledge can confirm that a value matches its quote and cannot tell that the quote came from the wrong section.
Recent research on oversight describes the reviewer's side of this. A 2025 preprint on human oversight design notes that domain experts' roles are shifting from performing tasks to overseeing AI output. It is a small study, four co-design workshops in which experts from psychology and computer science oversaw an AI-based grading system, so read it as a description and not as a measurement. The authors report four things those reviewers needed: to understand their tasks and responsibilities, to gain insight into the AI's decision-making, to contribute meaningfully to the process, and to collaborate with peers and the AI.
Length matters as much as the choice of person. In our view, a checklist of twenty-five items that are nearly always right trains people to stop reading. Keep the mandatory list to what the consequence test put in the first group, and let the code and the not-found list carry the rest. A junior colleague can do a first pass, but the sign-off stays with the owner.
Track what reviewers change
Record every correction a reviewer makes: the field, the value the workflow produced, the value the reviewer entered, and a short reason. That log does two jobs.
It measures the review. If reviewers have changed nothing for weeks, either the workflow is very good on those fields or nobody is reading them, and you can find out which by having a second person recheck a sample against the source. If the same field is corrected often, the review is working and the workflow needs attention.
It also improves the workflow. A repeated correction usually traces to an unclear definition in the template or the instructions, and fixing that definition removes the error for every later run.
Writing the review step down
The NIST AI Risk Management Framework, which NIST describes as intended for voluntary use, asks in its Core that "processes for human oversight are defined, assessed, and documented." The framework does not prescribe a format. One way a small firm can act on that line is a single page per workflow:
- What the workflow produces and who receives it.
- The must-check fields and who checks them.
- The checks that run in code.
- Where corrections are logged and who reads the log.
- Who can stop the workflow or override its output.
That page also gives a new team member the review standard on their first day. It covers review only, and whether any framework or rule applies to your firm is a question for your own counsel or compliance lead. What data may enter the workflow and what an AI agent is allowed to send on its own are separate decisions, covered in our articles on secure AI workflows for confidential client data and AI agent permissions and governance.
We design this step into every workflow we build with a client's team, and if you want to work through it for one of yours, you can get in touch.
Common questions
What does human in the loop mean in an AI workflow?
It means a person checks or approves the workflow's output before anyone relies on it. The phrase describes an intention and leaves the design open: which items the person checks, what evidence they see, and what happens to their corrections all still have to be decided.
How much of an AI workflow's output should a person review?
Review every item where an error would reach a client or move money, every time, and skim or sample the rest. For a lease abstract that usually means dates, dollar amounts, and options. Keep the mandatory list short enough that the reviewer reads each item properly.
What is automation bias?
The EU AI Act, in Article 14, describes it as the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system. In practice it is the reviewer who approves a finished-looking draft without checking it against the source.
Can I rely on the AI's confidence score to decide what to review?
It should not be the only rule. A peer-reviewed evaluation published at ICLR 2024 found that language models tend to be overconfident when they state their confidence in words. Route by consequence first, and treat a low confidence statement as one extra reason to look.
Is human review of AI output legally required?
It depends on the jurisdiction, the work, and your professional obligations, so confirm with your own counsel or compliance lead. Article 14 of the EU AI Act requires human oversight for high-risk AI systems, and the NIST AI Risk Management Framework, which is voluntary, asks that oversight processes be defined, assessed, and documented.