Analysis · AI & Productivity
Whose Time Did the AI Actually Save?
A faster first draft can become someone else’s longer working day. Measure the handover.
A team may generate proposals faster while finance spends longer checking the assumptions. Developers may submit more changes while reviewers accumulate a larger queue. Managers may produce more analysis while decision-makers spend longer sorting through it.
These scenarios illustrate a measurement problem. AI can create value for the person using it while changing the workload of everyone who receives the result.
For builders and buyers, the question is straightforward:
Whose time did the AI actually save—and what happened to the rest of the work?
Follow the work beyond the user
At D82Labs, we propose a practical starting point: measure an AI-assisted task through to acceptance, including the people downstream.
Define “accepted” before the trial. A report might require checked figures and supported recommendations. A code change might require passing tests and review against specified requirements. A proposal might need enough detail for its recipient to make a decision.
The person accountable for the task should establish these requirements with the people who will use the result.
Then account for the human effort:
Total human effort = preparation + production + review + correction + handover + later rework.
Treat these as distinct categories so the same work is not counted twice. Sum effort across contributors, and use the same categories for the existing process and the AI-assisted process.
Measure elapsed time separately. Three people working for twenty minutes consume sixty person-minutes, even if they finish together. A task requiring ten minutes of attention can also wait two days in a review queue.
Effort measures labour consumed. Elapsed time measures how long the recipient waits. An improvement in one does not necessarily improve the other.
The bottleneck can move
Consider a simplified, illustrative example.
A team can prepare four items per hour. Its reviewer can process six. Preparation limits the workflow to four items per hour.
An AI tool raises preparation capacity to twelve items per hour. Review capacity remains six.
Preparation capacity has tripled. Maximum workflow throughput has increased from four to six items per hour—a 50% increase, assuming sufficient demand, unchanged quality, unchanged review effort and no other bottleneck.
If twelve items continue arriving every hour while six are reviewed, the queue grows by six items per hour.
The tool has created useful capacity. Realising more of that capacity requires changes elsewhere: improving review capacity, making outputs easier to check or controlling how much work enters the queue.
A demonstration that ends at generation cannot show which of these conditions applies.
Faster acceptance can conceal missed errors
Review time alone is insufficient.
A reviewer who overlooks a defect may approve an output faster than one who finds it. Measuring only approval speed could reward weaker checking.
A useful comparison therefore needs a quality assessment beyond the original user’s confidence that the task is complete.
Use defined criteria. Where appropriate, have another qualified reviewer examine a sample without knowing how it was produced. Record disagreements, required corrections and defects discovered after handover.
Set a follow-up period suited to the task. Report its length and acknowledge that defects appearing later fall outside the measurement.
Include the existing process’s normal checking effort. Counting every minute of AI review as an additional cost would distort the comparison just as surely as excluding review altogether.
Run a trial that can change your mind
Choose one recurring task and define the result that would justify adopting the tool.
Use comparable task instances for the existing and AI-assisted processes. Vary which method goes first to reduce order effects, and avoid simply repeating the same task when familiarity would make the second attempt easier.
Record:
- Human effort by stage, role and expertise.
- Elapsed time to acceptance, including queue delays.
- The proportion accepted without correction.
- Required corrections and unresolved defects.
- Problems discovered during the follow-up period.
- Whether the recipient could use the result for its intended purpose.
Separate initial setup and learning from later performance, while retaining both in the adoption decision. Repeat across enough task instances to examine variation. A small trial can guide the next decision without establishing a general result.
If performance disappoints, locate the source of friction before replacing the model. Unclear requests, missing inputs, overloaded reviewers and poor handovers may persist across tools.
Change one major factor at a time when investigating. Otherwise, an improvement may be difficult to attribute.
Build for the person who must check the output
This approach creates concrete product choices.
Place relevant evidence beside consequential claims. Make calculations inspectable. Show what changed. Identify missing inputs. Automate repeatable checks with clear pass conditions.
Ask for clarification when an incorrect assumption is likely to create more work than the interruption. Remove unnecessary output when it increases selection effort without improving the decision.
These features also need evaluation. Additional explanations, citations and dashboards can become another workload. Their value depends on whether they help people assess and use the result.
Product demonstrations can make this visible by showing the review and correction stage, including what the first response still requires.
Give additional value its own category
Some AI-assisted work takes longer and is still worthwhile because it produces a better result. Other work becomes feasible only because AI lowers the cost of attempting it.
Evaluate those benefits directly. Record the added value and full cost rather than describing every benefit as time saved.
Distinguish hours by role and expertise: saving an hour at one stage may have different operational consequences from adding an hour at another. When evaluating cost, include software, integration, maintenance and supervision alongside labour.
The adoption question is whether the benefit is meaningful, repeatable and worth its cost at the required quality.
Evidence of consistently lower total effort at equivalent or better quality should count in the tool’s favour. The evaluation must remain open to genuine gains as well as hidden costs.
Keep the clock running after generation
The first user’s experience is only part of the evidence.
Follow the output to the reviewer, the recipient and the point of use. Look for work removed, work transferred and work added. Track whether more usable results reach the end of the process.
For serious builders, that is a demanding and useful design target: make the complete job easier.
If your product saves its user an hour, can you show what happened to everyone who received the output?