AI noise versus what matters: a test for your firm
One question sorts AI news for a document-heavy firm, five kinds of change deserve attention, and a small test set of your own files settles the rest.
King & Company
In short
An AI announcement matters to a firm when it changes a workflow the firm already runs or has named as next, and almost nothing else does. Five kinds of change pass that test: what the model can read, what systems it can reach, whether instructions can be packaged for the whole team, the data handling terms, and where a person has to review. The way to settle any of them is to rerun a small set of your own documents with known right answers.
Sorting AI noise from what matters for your business comes down to one question: does this change something about a workflow we already run, or one we have named as next? If the answer is no, you can ignore the announcement, and for a firm that runs on leases, workpapers, policies or proposals the answer is no most weeks.
The rest of this article is the working version of that test. It covers the five kinds of change that do pass it, the things you can stop reading about, and a way to check any claim against your own documents in an afternoon.
Why AI news feels impossible to keep up with
A firm owner's week already has clients, deals and deadlines in it. On top of that sits a steady run of model launches, vendor emails and posts from people who say a new release changed everything. Nobody at a twenty-person brokerage or accounting firm is paid to sort that stream.
The numbers behind the stream do move fast. Stanford's 2026 AI Index reports that performance on one coding benchmark, SWE-bench Verified, rose from 60% to near 100% in a single year, and that AI agents went from 12% to roughly 66% task success on OSWorld, a test of real computer tasks. The same report puts organizational adoption at 88%.
Adoption and usefulness are different things. Gallup reported in 2025 that even among employees who use AI, only 16% strongly agree that the AI tools their organization provides are useful for their work, and that only 22% of employees say their organization has communicated a clear plan or strategy for integrating it. Buying a tool and connecting it to a piece of work are separate steps, and those figures suggest the second one is where organizations stall.
The one question that sorts signal from noise
Ask whether the announcement changes a workflow you already run or one you have named as next. To use the question you need a short written list of those workflows, with the document that goes in, the output that comes out and the person who reviews it. For a brokerage team that list might be lease abstraction, the first pass of a broker opinion of value and the weekly market report. For an accounting firm it might be workpaper preparation, engagement letters and proposal drafting from a precedent library.
A firm without that list has nothing to test news against, which is why every launch feels equally urgent. If you have not built the list yet, how to build an AI workflow walks through choosing and scoping the first one.
Five kinds of change that deserve a firm's attention
Each of these can change a step in a document-heavy workflow. Each also comes with a question you can answer by testing.
| Kind of change | Why it matters to a document-heavy firm | What to check |
|---|---|---|
| What the model can read | Your work arrives as long scanned leases, spreadsheets with many tabs, policy forms and marked-up PDFs | Does it now read the file types and lengths you work in, without you splitting or retyping them? |
| What systems it can reach | The data lives in your CRM, your document store, your practice management system and Excel | Can it read from and write to those systems directly, or does someone still copy and paste? |
| Whether instructions can be packaged | A workflow is only dependable when the analyst and the partner get the same result from the same input | Can the steps, templates and house rules be saved once and used by the whole team? |
| Data handling terms | You hold confidential client and deal information | What do the terms say about retention, training on your data and who can access it, on the plan you would be buying? |
| Where a person needs to review | The review step is what makes output safe to send to a client | Does the change move, shrink or add a point where someone has to check the work? |
Packaged instructions are the least understood of the five, and what Claude skills are explains one form of them. Data handling terms are a subject of their own, covered in what happens to data you send to an AI model. Treat any change in those terms as a reason to read the vendor's current documents and to confirm with your own counsel or compliance lead before client data goes in.
What you can safely ignore
Most of the stream fails the test because it describes performance on somebody else's work.
- Benchmark scores and leaderboards. They are scored on the benchmark's own tasks. The AI Index figures above show how quickly those scores move, and none of them was measured on your leases or workpapers.
- Demo videos. A demo shows a document the vendor chose, and it tells you nothing about the documents you would choose.
- Model version numbers and launch-day rankings. A new version matters only if it changes one of the five things in the table.
- Features for work you do not do. Image generation, video and coding tools are real advances that have nothing to do with an insurance agency's renewal work.
- Posts that report how much faster someone feels. The section on perceived speed below explains why.
- Predictions about where AI will be in a few years. You can act on what ships when it ships.
Ignoring these costs you very little, because anything that matters will show up again as a concrete change to a file type, a connection, a packaging feature, the terms or the review step.
How to test a new model or tool on your own documents
The dependable way to judge a change is to rerun real work whose right answer you already know, and a small firm can do that with a folder and a spreadsheet.
- Pick one workflow from your list. Start with the one the team runs most often.
- Collect five to ten real documents for it. Include the awkward ones: the scanned lease with three amendments, the trial balance with an odd chart of accounts, the policy with handwritten endorsements.
- Record the right answer for each. Use the abstract, schedule or comparison that a senior person already reviewed and signed off on. This is your answer key.
- Run the set through your current setup and score it. Count the fields that were right, wrong and missing, and note how long the review took. That is your baseline.
- When something relevant ships, rerun the same set. Keep the instructions and the documents identical so the only thing that changed is the model or tool.
- Compare against the baseline and decide. A change that fixes the errors you kept seeing is worth adopting. A change that scores the same is not worth the disruption of switching.
Keep the folder where the team can reach it and keep client confidentiality in mind when you choose the documents. Run the set only in tools whose terms you have already accepted for that kind of data.
The same answer key does a second job. It tells you where the model still gets things wrong, which is where the human review step belongs.
Why feeling faster is not proof
How fast a tool feels and how fast it is can differ, even for experts. In a randomized controlled trial run by METR in early 2025, 16 experienced open-source developers worked 246 real issues, and when they were allowed to use AI tools they took 19% longer. They had expected a 24% speedup beforehand, and afterward they still believed AI had sped them up by 20%.
That study needs to be read with its limits. METR says plainly that it is not evidence about fields other than software development or about what newer tools will do, and it has since marked those results as out of date. Its February 2026 follow-up reports that the newer data gives an unreliable signal of the current effect, partly because 30% to 50% of developers said they were choosing not to submit some tasks that they did not want to do without AI. The part worth carrying into your own firm is the gap the first study recorded between what skilled people believed and what the clock showed.
Survey figures on time saved have the same weakness, because they are self-reported. In a St. Louis Fed analysis of a November 2024 survey, 33.0% of workers who used generative AI in the previous week reported saving an hour or less and 20.5% reported saving four hours or more. Average savings among users came to 5.4% of work hours, about 2.2 hours in a 40-hour week, and to 1.4% of total hours once nonusers were counted.
So when a colleague or a vendor says a tool feels twice as fast, treat it as a reason to run the test set and time the review.
A monthly review that takes an hour
Put one hour on the calendar each month and give it to one named person who does the work the workflows cover. The agenda is short.
- Ten minutes: read the release notes for the tools the firm already pays for. Skip the commentary about them.
- Ten minutes: mark anything that touches one of the five categories and one of your listed workflows.
- Thirty minutes: if something was marked, rerun the test set and score it. If nothing was marked, skip this.
- Ten minutes: write three lines for the team covering what changed, what was tested and what, if anything, the team should do differently.
Most months the note will say that nothing needs to change, and that is a useful thing for a team to hear. It also gives everyone else permission to stop following the news themselves.
Questions to ask a vendor before a demo
These questions turn a pitch into something you can check.
- Which of our workflows does this change, and at which step?
- Can we run the demo on our own documents, using a set we choose?
- What file types and lengths does it read, and what happens to a scanned or very long document?
- Which of our systems does it connect to today, and does it read, write or both?
- Can our own instructions and templates be saved so every person on the team gets the same result?
- What are the data retention and training terms on the plan you are quoting, and where are they written down?
- Where in the output should a person review, and how does the tool show its source for each figure?
- If we stop paying, what do we keep?
A vendor who answers these directly and agrees to run your documents is worth an hour. Signs that a claim is hype include a headline benchmark with no mention of your kind of document, a percentage improvement with no description of how it was measured, and reluctance to test on files you supply.
We build workflows for brokerage, accounting, insurance and advisory teams against their own documents, which is why we recommend testing this way. If you would like help building your test set, get in touch.
Common questions
How do I know if a new AI tool is worth trying?
Name the workflow it would change and the step inside that workflow. If you can, run your own test documents through it and score the output against answers you already know are right. If you cannot name the workflow, the tool can wait.
Do AI benchmark scores matter for my business?
They tell you the field is moving, and they tell you very little about your own work. A benchmark is scored on its authors' tasks, and your firm runs on its own leases, workpapers or policies. A score on ten of your own documents is the number that should drive a decision.
How often should we revisit which AI tools we use?
A monthly review of about an hour is enough for most firms, with a rerun of your test documents only when something on your short list of relevant changes has shipped. Switching tools is a separate and rarer decision, and it should follow a test result on your own files.
Is it a mistake to wait before adopting new AI features?
Waiting on a feature that touches no workflow you run costs you nothing. Waiting on a change that removes a step your team does by hand every week does cost you, which is why the review is tied to named workflows and a fixed calendar slot.
Who on the team should be responsible for following AI news?
One named person who does the work the workflows cover, with an hour a month set aside for it and the test documents in hand. The qualification that matters is being able to tell a right abstract or schedule from a wrong one, so a technical hire is not required.