How we measure AI-assisted work

Prove the work got better

We measure AI-assisted work by what it produces: accepted output, the review it created, the rework it caused, the time it saved once both of those are counted, the risk it carried and the business value behind it. All of it against how the work ran before AI touched it.

  • Six measures
  • Six AI surfaces
  • A six-phase method
  • Built into every engagement

Six ways an AI program reports a win it cannot show

Each one is survivable alone. Together they are why an organization can be two years into AI adoption and unable to say what it returned.

Usage is counted, output is not

Licences issued, weekly active users, prompts run. Every one of those numbers can rise while the work leaving the team is exactly what it was.

You are measuring the tool.

The review moved downstream

Drafting time falls and checking time rises somewhere downstream, usually onto a more senior person who was not part of the pilot.

The saving is real for one desk and paid for at another.

Polished output that says very little

Fluent, correctly formatted, on brand, and thin. It passes a glance and fails at the next stage, where someone has to rebuild it.

Trust in AI-assisted work drops faster than it was earned.

It holds until the task shifts slightly

Performance is strong across a set of tasks and then falls off sharply on one that looks identical from the outside.

The failure arrives confident and nobody catches it.

There is no before

Nobody recorded how long the work took, what it cost or how often it was wrong, so the only comparison available is memory.

No baseline means no improvement, only a claim.

Speed is reported, risk is absorbed

Throughput is reported to the executive team. Policy exceptions, corrections and near misses are handled quietly by the people who find them.

Exposure scales with adoption and nobody is counting it.

The same tools produced all three of these results

Three controlled studies, three different outcomes. What separates them is the work the tool was pointed at and who was using it, which is what measurement is for.

  1. Field experiment, consultants

    Help inside the task, harm just outside it

    Consultants given an AI assistant completed more work, faster, and at higher rated quality on tasks inside the capability the model has. On a task built to sit just outside it, the group using AI performed worse than the group without it.

  2. Randomized trial, developers

    Faster and slower at the same time

    Experienced open-source developers working real issues in their own repositories took about 19 percent longer to finish them with early-2025 AI tools, and reported afterwards that they had been quicker.

  3. Field study, customer support

    The average hid who got the gain

    A large support deployment produced an average productivity gain of roughly 14 percent. Almost all of it went to newer and lower-performing agents. The most experienced saw little effect.

Each measure is a question your team can answer out loud

Speed on its own is the measure most likely to mislead, so it sits fourth and only counts once review and rework are inside it.

Accepted output
Did the work meet the agreed standard and move to the next stage without a material fix?
Review burden
How many minutes of human checking, correction, approval and escalation did the AI-assisted version create, and whose minutes were they?
Rework
How often did an output need material revision before anyone could use it, and for which reason: accuracy, completeness, policy, tone, customer impact or downstream fit?
Cycle time
Did the workflow get faster once review and rework are counted inside it rather than beside it?
Risk
Did the work stay inside policy, data, customer, legal and security boundaries, and what evidence shows that it did?
Business value
Did the change produce capacity, quality, revenue, cost or risk impact that someone outside the project would recognize as real?

These are the six that answer what a leadership team asks: how will we know AI is working, how do we stop spending on tools that return nothing, how do we control the risk without slowing the business down, how do we prove the return after implementation, and how do we get teams using this inside real work. If the question you are holding is the last one, that is an AI-enabled workforce.

Eight things we define first, because none can be recovered later

This is the part that gets skipped. Every item below is cheap to decide before the build and impossible to reconstruct after it. A workflow that cannot supply them is telling you something useful about itself.

The unit of work
The bounded piece of work being measured. Work that cannot be given an owner, an input and an output is not ready to be automated, and finding that out is itself a result.
The owner and the reviewer
Who is accountable for the output, and who has the time, the expertise and the authority to reject it. Those are usually two different people and occasionally nobody.
The done standard
What an acceptable output looks like, written down before the first one is produced. Without it, every item is judged against whatever the person receiving it happened to expect.
The baseline
How the work was done before: how long it took, what it cost, how often it was wrong and how much of it came back.
The risk level
Set from data sensitivity, customer impact, legal exposure and how reversible a mistake is. It decides how much review the work has to carry, so it is set before the build rather than after an incident.
The telemetry
Which numbers are captured while the work runs: cycle time, human time, review time, errors, rework, approvals and escalations. Instrumented at the start, because it cannot be recovered later.
The evidence retained
What is kept for each completed item: the source documents, the output, the decision the reviewer recorded and the final artifact. This is what an auditor, a client or an insurer asks for.
The review rhythm
When the numbers get read, by whom, and what happens to the workflow when they are bad. A measure nobody is scheduled to look at is not a control.

One scorecard for all AI use is the mistake

An assistant, an automation and an agent create different value and fail in different ways. Measuring them the same way hides both. Each card names what we track and the way that surface usually goes wrong.

  1. Assistant

    Drafting, summarizing and analysis support

    General-purpose help with writing, research and thinking through a problem. The failure mode is confident error: fluent output carrying a wrong fact, a wrong number or a source that does not exist.

    • Accepted output rate
    • Review burden
    • Rework rate
    • Whether sources were verified
  2. In-tool copilot

    Assistance inside the system the work already runs in

    Suggestions inside the editor, the document, the spreadsheet, the CRM or the inbox. The failure mode is over-acceptance: the suggestion is plausible, the context is subtly wrong, and it is accepted anyway.

    • First-pass acceptance
    • How much editing the output needed
    • Error rate
    • Debt created downstream
  3. Workflow automation

    Repeatable work with a defined trigger and output

    Rule-based sequences that run without anyone starting them. The failure mode is the exception: the ninety percent runs cleanly and the remaining ten arrive as duplicates, misroutes and silent stalls.

    • Cycle time
    • Handoffs removed
    • Exception rate
    • Automation failure rate
  4. Agent

    Multi-step work that plans, retrieves and acts

    Work that takes actions in systems rather than producing a draft for a person. The failure mode is accountability: an action was taken, it was wrong, and no one had agreed in advance who reverses it.

    • Goal completion
    • Tool error rate
    • How often a person had to step in
    • Actions reversed
  5. Support bot

    Customer or employee answers on bounded topics

    Front-line responses at volume. The failure mode is the confident wrong answer to a customer, which is a commitment your organization may be held to whether or not a person made it.

    • Resolution rate
    • Escalation rate
    • Customer satisfaction
    • Wrong-answer incidents
  6. Knowledge assistant

    Answers drawn from your own approved sources

    Retrieval over internal documents and records. The failure mode is quieter than the others: the answer is well written, it is grounded in a document that was superseded a year ago, and nothing about it looks wrong.

    • Grounding rate
    • Citation accuracy
    • Correct refusals
    • Permission boundaries held

The pilot is the fourth of six phases

The three phases before the pilot are what make it mean something. The two after it keep the result from decaying. A pilot that starts at phase four has neither, which is why so many of them end in a debate about whether the numbers count.

  1. Diagnose

    Where is AI already in use and where is the work friction worst? Interviews, a tool review, workflow mapping and a data sensitivity pass produce the audit and the current-state baseline.

  2. Select

    Which work is worth redesigning? Candidates are scored on structure, whether the output can be judged, context availability, risk, reviewer capacity and whether a baseline can be measured at all.

  3. Design

    How should the AI-assisted workflow run? Inputs, outputs, done standard, review path, escalation, and the telemetry that has to be captured while the work is happening.

  4. Pilot

    Does the redesigned workflow beat the baseline? A controlled run against the recorded before, with review and rework counted inside the result rather than reported separately.

  5. Govern

    Can it scale safely? Policy, approvals, logging, incident handling, security and vendor controls, sized to the risk level the work was given in design.

  6. Improve

    How does it get better? A standing review of what failed, why, and what changed because of it, with an improvement backlog that carries owners and dates.

Measurement sits inside the price of every engagement

There is nothing to buy on this page. An engagement that cannot show what it changed is unfinished work, so this is carried in all five rather than sold beside them.

Baselines, AI fit, and what a useful output would be

FusionMap

Risk measures, control checks and audit evidence

FusionGuard

Before-and-after proof that the build improved the work

FusionBuild

Output judgment, review standards and rework reduction

AI Systems Mastery

Live dashboards, ROI review and an improvement backlog

AI Operating System

The three findings above are from Dell'Acqua and colleagues on knowledge work inside and outside the capability frontier, the METR randomized controlled trial of experienced open-source developers, and the Brynjolfsson, Li and Raymond study of a large customer support deployment. We cite them because two of the three are unflattering, and a measurement page that only quoted the wins would be proving the point it is arguing against.

Bring us a workflow you already run

Tell us what it produces, who checks it and how you would know if it improved. If those three answers do not exist yet, that is where the work starts.