Usage is counted, output is not
Licences issued, weekly active users, prompts run. Every one of those numbers can rise while the work leaving the team is exactly what it was.
You are measuring the tool.How we measure AI-assisted work
We measure AI-assisted work by what it produces: accepted output, the review it created, the rework it caused, the time it saved once both of those are counted, the risk it carried and the business value behind it. All of it against how the work ran before AI touched it.
AI adoption by anecdote
Each one is survivable alone. Together they are why an organization can be two years into AI adoption and unable to say what it returned.
Licences issued, weekly active users, prompts run. Every one of those numbers can rise while the work leaving the team is exactly what it was.
You are measuring the tool.Drafting time falls and checking time rises somewhere downstream, usually onto a more senior person who was not part of the pilot.
The saving is real for one desk and paid for at another.Fluent, correctly formatted, on brand, and thin. It passes a glance and fails at the next stage, where someone has to rebuild it.
Trust in AI-assisted work drops faster than it was earned.Performance is strong across a set of tasks and then falls off sharply on one that looks identical from the outside.
The failure arrives confident and nobody catches it.Nobody recorded how long the work took, what it cost or how often it was wrong, so the only comparison available is memory.
No baseline means no improvement, only a claim.Throughput is reported to the executive team. Policy exceptions, corrections and near misses are handled quietly by the people who find them.
Exposure scales with adoption and nobody is counting it.What the research shows
Three controlled studies, three different outcomes. What separates them is the work the tool was pointed at and who was using it, which is what measurement is for.
Consultants given an AI assistant completed more work, faster, and at higher rated quality on tasks inside the capability the model has. On a task built to sit just outside it, the group using AI performed worse than the group without it.
Experienced open-source developers working real issues in their own repositories took about 19 percent longer to finish them with early-2025 AI tools, and reported afterwards that they had been quicker.
A large support deployment produced an average productivity gain of roughly 14 percent. Almost all of it went to newer and lower-performing agents. The most experienced saw little effect.
What gets measured
Speed on its own is the measure most likely to mislead, so it sits fourth and only counts once review and rework are inside it.
These are the six that answer what a leadership team asks: how will we know AI is working, how do we stop spending on tools that return nothing, how do we control the risk without slowing the business down, how do we prove the return after implementation, and how do we get teams using this inside real work. If the question you are holding is the last one, that is an AI-enabled workforce.
Before anything is built
This is the part that gets skipped. Every item below is cheap to decide before the build and impossible to reconstruct after it. A workflow that cannot supply them is telling you something useful about itself.
By AI surface
An assistant, an automation and an agent create different value and fail in different ways. Measuring them the same way hides both. Each card names what we track and the way that surface usually goes wrong.
General-purpose help with writing, research and thinking through a problem. The failure mode is confident error: fluent output carrying a wrong fact, a wrong number or a source that does not exist.
Suggestions inside the editor, the document, the spreadsheet, the CRM or the inbox. The failure mode is over-acceptance: the suggestion is plausible, the context is subtly wrong, and it is accepted anyway.
Rule-based sequences that run without anyone starting them. The failure mode is the exception: the ninety percent runs cleanly and the remaining ten arrive as duplicates, misroutes and silent stalls.
Work that takes actions in systems rather than producing a draft for a person. The failure mode is accountability: an action was taken, it was wrong, and no one had agreed in advance who reverses it.
Front-line responses at volume. The failure mode is the confident wrong answer to a customer, which is a commitment your organization may be held to whether or not a person made it.
Retrieval over internal documents and records. The failure mode is quieter than the others: the answer is well written, it is grounded in a document that was superseded a year ago, and nothing about it looks wrong.
The method
The three phases before the pilot are what make it mean something. The two after it keep the result from decaying. A pilot that starts at phase four has neither, which is why so many of them end in a debate about whether the numbers count.
Where is AI already in use and where is the work friction worst? Interviews, a tool review, workflow mapping and a data sensitivity pass produce the audit and the current-state baseline.
Which work is worth redesigning? Candidates are scored on structure, whether the output can be judged, context availability, risk, reviewer capacity and whether a baseline can be measured at all.
How should the AI-assisted workflow run? Inputs, outputs, done standard, review path, escalation, and the telemetry that has to be captured while the work is happening.
Does the redesigned workflow beat the baseline? A controlled run against the recorded before, with review and rework counted inside the result rather than reported separately.
Can it scale safely? Policy, approvals, logging, incident handling, security and vendor controls, sized to the risk level the work was given in design.
How does it get better? A standing review of what failed, why, and what changed because of it, with an improvement backlog that carries owners and dates.
In the work
There is nothing to buy on this page. An engagement that cannot show what it changed is unfinished work, so this is carried in all five rather than sold beside them.
The three findings above are from Dell'Acqua and colleagues on knowledge work inside and outside the capability frontier, the METR randomized controlled trial of experienced open-source developers, and the Brynjolfsson, Li and Raymond study of a large customer support deployment. We cite them because two of the three are unflattering, and a measurement page that only quoted the wins would be proving the point it is arguing against.
Tell us what it produces, who checks it and how you would know if it improved. If those three answers do not exist yet, that is where the work starts.