How to measure AI adoption on a team beyond login counts
A monthly scorecard a partner can keep by hand: five numbers per workflow, what vendor dashboards really report, and an honest way to state time saved.
King & Company
In short
Count workflows, and leave the login report as background. For each piece of work the team has rebuilt with AI, record once a month who owns it, whether it ran on live work, what share of eligible jobs went through it, how long the first pass and the review took against the old way, and what the reviewer corrected. Vendor dashboards report activity and estimates of time, so treat them as a check on access and get your time figures from a small sample of jobs you timed yourself.
To measure AI adoption on a team, count the pieces of work that are now done a different way, and treat the number of people who logged in as background. A firm of forty can do this in a spreadsheet once a month with five numbers per workflow, and no analytics team is needed.
The sections below set out those five numbers, what the vendor dashboards say about their own figures, and a way to state time saved that holds up when a partner or an investor asks how you know.
Why do logins and seat counts tell you so little?
A login shows that a person opened the tool. It does not show whether the lease abstract, the proposal, or the month-end package that went out the door was produced any differently from last year.
The national surveys show the same gap between use and change. Gallup reported in June 2025 that 44% of employees say their organization has begun integrating AI, while 22% say it has communicated a clear plan or strategy. Pew, in a survey of 5,273 employed U.S. adults, found that among workers who have used AI chatbots for work, 40% say the tools are extremely or very helpful for doing things more quickly, and 29% say the same about improving the quality of their work. People using a tool, a firm changing how it works, and the work getting better are three separate things, and a count of seats in use speaks only to the first.
A platform's dashboard defines adoption as what the product can count. That is useful for confirming that people have access, and it answers a different question from the one a managing partner is asked at the partner meeting.
Measure the workflow instead of the person
A workflow, as we use the word, is a fixed way of producing one named deliverable: the same source documents, the same instructions, the same output format, and the same review step each time. If your team has not yet built one, start with a single deliverable and come back to measurement afterward, because there is nothing to count yet.
The workflow is the right unit for three reasons. It ties to work the firm already tracks, such as leases abstracted, proposals sent, or returns reviewed. It has an owner who can answer for it. And it avoids scoring individuals on how often they type into a chat window, which measures activity and tells you nothing about the result.
The five numbers worth tracking each month
Keep one row per workflow. The owner of the workflow fills it in, and it should take a few minutes.
| Number | What to record | What it tells you |
|---|---|---|
| Owner | The one named person responsible for the workflow | A blank here means nobody will notice when it stops working |
| Ran on live work | Yes or no for this month, on a real client or deal file | Separates a working system from a demonstration |
| Coverage | Jobs that went through the workflow, over all jobs of that kind this month | Whether the team has switched, or only one person has |
| Time | Minutes for the first pass and for the review, against the old way on the same kind of file | The only time figure you can defend |
| Corrections | What the reviewer had to fix, in a few words, and how many items | Whether the draft is trustworthy and getting better |
Coverage is the number to read first. A workflow can exist, be demonstrated well, and have an enthusiastic owner, and still carry, say, three of a month's twenty jobs. The scorecard makes that visible, and the next question is whether the team trusts the output, which is something you can fix.
What do vendor dashboards report, and how should you read them?
The dashboards are worth opening. Read them as a record of activity and as a starting point for questions.
Microsoft Copilot Dashboard
Microsoft's documentation says the dashboard covers readiness, adoption, impact, and sentiment. An active Copilot user is an employee who performed at least one Copilot activity in the previous 28 days, and a returning user is one who took at least one Copilot action in both the current and the preceding period. One action in four weeks is a low bar, so read the active user figure as a count of people who have tried it.
The same page defines Copilot assisted hours as an estimate that is "computed based on your employees' actions in Copilot and multipliers derived from Microsoft's research on Copilot users." The method is to group actions into categories and multiply each count by an assistance factor. The documentation lists 6 minutes per search or summary action and 6 minutes per creation action, and says these are "broad approximations based on the best available research rather than precise calculations." Copilot assisted value multiplies those hours by an hourly rate that defaults to $72 and that you can change.
So assisted hours is a count of actions expressed in hours. As documented, the method counts a draft action the same way whether the email was sent, rewritten, or discarded. Microsoft calls the figure a general estimate, and that caveat should travel with the number whenever you report it.
Two details matter for a small firm. The documentation's feature table lists benchmarks, week and month trendlines, and the survey-based sentiment metrics as not included for tenants with between 1 and 49 Copilot licenses. It also says metrics are not shown for groups smaller than the minimum group size, to protect individual privacy.
Claude usage analytics
Anthropic's help center says Owners and Primary Owners on Team plans, and Owners, Primary Owners, and Admins on Enterprise plans, can view usage analytics. The listed metrics include weekly active members, active members and assigned seats, adoption level, product stickiness, skills with cost per use and number of uses, connectors with the number of users and counts of read and write actions, and estimated time saved. For Claude Code, the page describes a separate Value tab and says every formula on it is shown inline and that you can adjust the inputs to match your organization's assumptions.
The skills count is the most useful of these for the scorecard, because a skill usually corresponds to a workflow. If, for example, a call prep skill shows a handful of uses in a month when the team made far more calls than that, you have a coverage question to ask. The page lists estimated time saved without explaining how it is calculated, so label it as the vendor's estimate, as you would assisted hours, and keep it apart from the time you measured yourself.
How do you measure time saved honestly?
Do not ask people how much time they saved. In a randomized trial of 16 experienced open-source developers across 246 tasks, METR found that developers took 19% longer when allowed to use AI tools, and afterward still believed AI had sped them up by 20%. METR states that the study does not show AI fails to speed up people in other fields, and its February 2026 update reports that later data points toward a speedup while calling that data an unreliable signal. The part that carries over to any firm is the gap between what people felt and what the clock recorded.
Survey figures on time saved are self-reports as well. The St. Louis Fed's estimate from a November 2024 survey, that workers who used generative AI in the previous week saved an average of 5.4% of their work hours, comes from asking respondents how many additional hours they would have needed without the tool. That is a fair way to run a national survey, and it is weak support for a claim about your own firm.
Timing a sample is not much work:
- Pick one workflow and five to ten jobs of the same kind, for example office leases of similar length.
- For each job run through the workflow, write down the minutes spent on the first pass and the minutes spent on review, as two separate figures.
- For the comparison, use time records you already keep for the same kind of job, or have the same person do two or three jobs the old way and time them.
- Report the result with its sample size, as in "eight leases, timed in September," and resist turning it into an annual figure for the whole firm.
Keeping first pass and review separate matters. A workflow that produces a draft quickly and then needs a long review has moved the work to a more senior person, and a single combined number hides that.
Tracking review corrections as a quality signal
The correction log is the part no dashboard can produce. Each time a reviewer fixes something in a draft, the owner notes what it was: a wrong rent escalation date, a missing renewal option, a paragraph in the wrong tone for that client.
Over a few months the log shows whether the same errors recur, which means the instructions need work, or whether errors are scattered and falling, which means the workflow is settling. It also tells you whether the human review step is doing its job. A month with zero corrections logged across many jobs is worth a question, since it can mean the draft is excellent or that nobody is checking.
Signs that a workflow is not really adopted
- The owner has changed roles and nobody was named in their place.
- It ran this month only when someone asked for a demonstration.
- Coverage has stayed at one person's jobs for three months.
- People run the workflow and then redo the work by hand before sending it.
- The correction log is empty, or lists the same fix every month.
Redoing the work by hand is the clearest sign of a workflow that exists but that nobody trusts, and it still counts as activity in a usage report.
A one page scorecard for a partner meeting
One page is enough. List each workflow as a row with the five numbers, add last month's coverage beside this month's, and put the dashboard figures (active members against seats, and any estimated hours, labelled as the vendor's estimate) in a single line at the bottom. Close with two sentences: which workflow moved most and why, and which one is stuck and what will change next month.
On surveillance, the scorecard should contain no individual's prompt count. In our view a team that feels watched will tell you less, and the corrections and complaints are the information you most need. If your firm does look at per-user data, settle who can see it and for what purpose in your acceptable use policy, and confirm the approach with your HR lead and counsel.
What should you do when the numbers are flat?
Read the row before you buy anything. Flat coverage with many corrections points to draft quality, so the owner and the builder should rework the instructions against the logged fixes. Flat coverage with few corrections points to habit or awareness, and the remedy is usually to train on the team's own files until the people who do that job have each run it on live work. A workflow with no owner needs an owner before it needs anything else.
If several workflows are flat for different reasons, the scorecard has still done its job, because it tells you which one to work on first and what kind of work it needs.
Common questions
What is a good AI adoption rate for a team?
We have not found a published benchmark specific to a firm of ten to a few hundred people doing professional work, so we do not recommend a target percentage. A more useful standard is internal: for each workflow the team has built, the share of eligible jobs that went through it this month compared with last month, and whether the reviewer is correcting less over time.
How do I know if my employees are actually using AI?
Ask about the work instead of the person. Pick a named deliverable, count how many were produced this month, and count how many of those started from the workflow the team built. A login report tells you someone opened the tool, and the job count tells you whether the firm's work is being done a different way.
How do you measure time saved by AI?
Time a small sample of real jobs. Record the first pass and the review separately for five to ten jobs done through the workflow, and compare them with the same kind of job done the old way, using time records you already keep or a few jobs timed for the purpose. Report the sample size alongside the figure, and do not rely on people's own impression of how much faster they are.
Are Copilot assisted hours an accurate measure?
Microsoft's documentation describes Copilot assisted hours as a general estimate, computed by multiplying counts of Copilot actions by assistance factors drawn from its research, and calls the results broad approximations. Read it as an indicator of how much activity there is. Nobody at your firm observed those hours being saved.
Should I track individual employees' AI usage?
We recommend measuring workflows and outcomes and leaving individual prompt counts alone. Counting prompts rewards activity over useful work, and in our view a team that feels watched is less willing to report what is not working. If your firm does review individual usage data, confirm the approach with your HR lead and counsel first.