How to build an AI workflow your team will use
A practitioner's order of work: pick one deliverable, collect past examples, build small, test against finished work, and design the review step first.
King & Company
In short
Pick one deliverable the team already produces every week and one person who owns it. Gather ten to twenty past examples with the version the firm sent out, write down what the owner corrects in a draft, build the smallest version inside the AI workspace you already pay for, and test it against those past examples before anyone relies on it. Design the review step first and have the owner run the workflow in the working session.
To build an AI workflow your team will use every day, start from one deliverable the firm already produces every week, collect the finished versions you sent out, and test the workflow against them before anyone relies on it. Which tool to build it in is the last decision, and for most first workflows the AI workspace the firm already pays for is enough.
The order below needs a person who owns the work, a folder of past examples, and about an hour a week of that person's attention.
What counts as an AI workflow, and what does not
Anthropic's engineering team defines workflows as "systems where LLMs and tools are orchestrated through predefined code paths" and agents as systems where the model directs its own process and tool use. The same article says workflows offer predictability and consistency for well-defined tasks.
For a firm that runs on documents, the practical definition is narrower. A workflow is a fixed way of producing one named deliverable. It has the same source documents, the same written instructions, the same output format, and the same reviewer every time. A lease abstract produced from an executed lease and its amendments is a workflow. So is a proposal first draft assembled from the firm's precedent library, or a weekly market report built from the same data pulls.
One person pasting a document into a chat and asking for a summary is useful, and it falls short of a workflow because the result depends on who asked and how. Our comparison of workflows, agents, skills, and custom software covers where each one fits.
Pick the first workflow: one deliverable, one owner, real documents
Choose a piece of work that meets four tests:
- The team produces it at least weekly, so there are plenty of past examples.
- It has a clear finished product, such as an abstract, a workbook, a memo, or a report.
- It is produced from documents the firm already holds.
- One person owns it and reviews it today.
The fourth test matters most. The owner is the person who knows what a good draft looks like, and nobody else can tell you when the workflow is finished.
Avoid starting with work that is rare, work whose standard changes with every client, and work where nobody reviews the output now.
Collect past examples before you build anything
Gather ten to twenty real past examples. Each one needs the source documents and the final version the firm sent out. Include the awkward ones on purpose, such as the lease with six amendments, the scanned document, and the client who wanted a different format.
Then sit with the owner and ask what they correct when a junior person hands them a draft. Write every correction down. Those corrections are the specification, and they are more precise than any description of the task written from scratch.
This step comes before the build because it defines what success means. Anthropic's documentation for developers says success criteria should be specific, measurable, achievable, and relevant, and that evaluations should mirror your real-world task distribution and include edge cases. That guidance is written for developers, and the principle carries over directly to a firm: your past deliverables are the task distribution, and the hard ones are the edge cases.
Build the smallest version inside the workspace you already have
Most first workflows need neither an agent nor an automation platform. Anthropic's engineering team writes that when building applications with language models, "we recommend finding the simplest solution possible, and only increasing complexity when needed. This might mean not building agentic systems at all."
The smallest version is usually three things kept in one shared place: written instructions that capture the owner's corrections, the firm's own template for the output, and two or three strong past examples. In Claude, a project is a self-contained workspace with its own chat history and knowledge base, with project instructions that shape the responses, and Team and Enterprise plans can share a project with colleagues. If the firm uses a different vendor, look in its documentation for the equivalent feature.
Build against the firm's own template. A workflow that produces the firm's abstract, in the firm's column order, with the firm's naming, needs no translation step before someone can use it.
Test it against work the firm has already finished
Run the workflow on the past examples and put each draft next to the version that went out. Have the owner mark every difference and sort each one into three groups: the draft was wrong, the draft was acceptable but different, or the draft caught something the original missed.
Decide the pass standard before you run the test, and write it per field or per section. Anthropic's documentation notes that most use cases need evaluation along several success criteria. For an abstract, that could mean dates and dollar amounts must match the source exactly, while a narrative summary only has to be accurate and complete in the owner's judgment.
Fix the instructions, rerun the same set, and repeat until the owner would accept the draft as a starting point. Keep the test set. It is what lets you change the workflow later without guessing whether you broke it.
Design the review step before the first live run
A passing test does not remove the reviewer. Decide on day one who checks each output, what they check, and how they see where each figure came from. A draft that cites the page and clause for each extracted term can be checked against the source directly, and a draft that does not leaves the reviewer to find each term again.
Put the review where the firm already reviews. If a senior person signs off on the deliverable today, they sign off on the AI-assisted version too. We cover the design choices in how to design the human review step.
Hand it to the person who owns the work
Adoption is decided while the workflow is being built, so the owner should run the workflow themselves, on a live piece of work, during the working session where it is built. A training held weeks afterwards asks people to adopt something they had no hand in shaping.
A workflow built from the owner's corrections, on the firm's template, and tested on the firm's past work is already familiar when it arrives. The owner knows why each instruction is there because they supplied it.
When a workflow should become a skill, an integration, or software
Keep the workflow simple until a specific need forces a change.
| What you notice | What to add |
|---|---|
| Several people need the same workflow, and the instructions are being copied around | A skill the whole team can call |
| The output has to be retyped into Excel, a CRM, or another system | An integration that writes the data where it belongs |
| The work should start on its own when a document arrives, or needs an approval handoff | Workflow automation |
| The job needs its own interface, its own data store, or many users with different permissions | Custom software |
Anthropic describes skills as folders of instructions, scripts, and resources that Claude loads dynamically, and says that on Team and Enterprise plans an organization's owners can provision them for all users. Our guide to what Claude skills are explains how a tested workflow gets packaged into one.
Before client documents go into any of this, read the vendor's data terms for the plan you are on. Anthropic, for example, says that by default it will not use inputs or outputs from its commercial products to train its models, with exceptions when a customer reports feedback or chooses to allow it. That article covers the commercial products, and the consumer plans have separate terms. Confirm what your client agreements and any regulator require with your own counsel or compliance lead.
How many firms have got this far?
The answer depends on who is asked. The Census Bureau's Business Trends and Outlook Survey asks firms whether they used AI in the past two weeks, and overall usage hovered between 17% and 20% from December 2025 to May 2026. The same release puts use at 37% for firms with at least 250 employees and 33.9% in finance and insurance.
Individuals report more. A Federal Reserve note from April 2026 puts work-related generative AI adoption at about 41 percent of the workforce as of November 2025, against about 18 percent of firms at the end of 2025, and it traces part of the gap to who is sampled, how answers are weighted, and how the question is framed. A Census household survey conducted in March 2026 found that about 56% of U.S. workers have used AI on the job for at least one of 11 tasks. Among those who used it in the prior week, 31% said it saved one to two hours, 25% said it saved less than an hour, and 10% said it saved no time.
Our reading is that many people are using chat on their own and getting modest returns, and far fewer firms report AI as part of how the business produces its work. If that describes your firm, you have plenty of company. The firms that turn individual use into a repeatable system first are the ones that pull ahead.
What eight to twelve weeks of this adds up to
The method repeats from one deliverable to the next. In our engagements we bring a working draft of one workflow to a one-hour session each week, the owner corrects it, and most of the effort happens on our side between sessions. Over eight to twelve weeks that produces a named set of working systems, a team that runs them, and documentation the firm owns.
A firm can do this with its own people if someone has the time to build between sessions. If nobody does, a forward deployed engineer is one way to buy that capacity, and you can get in touch if you would like to talk through which deliverable to start with.
Common questions
What is the difference between an AI workflow and just using ChatGPT or Claude?
A chat is one person asking for help with whatever is in front of them, and the result depends on how that person asked. A workflow is a fixed way of producing one named deliverable: the same source documents, the same written instructions, the same output format, and the same review step each time, so that anyone on the team gets a comparable draft.
Do I need an automation platform like Zapier, Make, or n8n to build an AI workflow?
Usually not for the first one. Anthropic's engineering guidance recommends finding the simplest solution possible and adding complexity only when needed, and a first workflow that turns documents into a draft for one reviewer can normally live inside the AI workspace the firm already has. An automation platform earns its place later, when the workflow has to start on its own or pass data between systems.
Which workflow should a firm build first?
Choose a deliverable the team produces at least weekly, that has a clear finished product, that is made from documents the firm already holds, and that one person owns and reviews today. Frequency gives you past examples to test against, and a named owner gives you someone who can say what a good draft looks like.
How do I know if an AI workflow is accurate enough to rely on?
Run it on past work the firm has already finished and compare each draft with the version that went out, field by field or section by section. Write the pass standard down before you run the test, have the owner of the work do the comparison, and keep a person reviewing every live output regardless of the score.
How long does it take to build an AI workflow?
It depends on how many source documents are involved and how much judgment the deliverable carries. In our engagements we bring a working draft of one workflow to a one-hour session each week, the owner corrects it, and over eight to twelve weeks that adds up to a named set of working systems.