Measuring the ROI of an AI Agent Without Fooling Yourself
Most agent ROI figures do not survive a careful reading, and the reason is almost always the same: the baseline was reconstructed after the fact, and the errors were left out of the arithmetic.
Getting an honest number is not hard, but it has to start before launch. Write down what today costs, in the units you will use afterwards. Everything else follows from that one discipline.
Write the baseline down first#
- Volume: how many of these tasks happen per week?
- Handling time: how long does a person take, measured on a sample rather than remembered?
- Fully loaded cost per hour for the people doing it.
- Current quality: error or rework rate today, because the agent will be compared against it.
- Waiting time: how long the requester currently waits, if that matters to them.
A baseline reconstructed after launch will always favour the project, and everyone reviewing it knows that.
Metrics that hold up, and the ones that flatter#
| Flattering metric | Why it misleads | Use instead |
|---|---|---|
| Deflection rate | Counts unanswered as resolved | Completed tasks needing no human |
| Messages handled | Volume is not value | Tasks completed end to end |
| Satisfaction on agent chats | Survivor bias; the frustrated leave | Satisfaction across all contacts |
| Time saved per response | Ignores review time | Net minutes after human review |
| Cost per token | Not a business number | Cost per completed task |
The formula, including the part people omit#
Annual benefit equals tasks completed without a human, times minutes saved per task, times loaded cost per minute — minus the cost of errors the agent introduced, minus the review time it created. That subtraction is the honest part. An agent that completes 70% of tasks but requires a human to check every one has saved review time, not handling time, and the difference is usually a factor of three.
Estimate error cost explicitly, even roughly: correction time, plus goodwill, plus any refund or credit issued because of a mistake.
Benefits that are real but not on the invoice#
Some genuine value never appears in the cost model. Faster responses at three in the morning. Consistency — the same answer to the same question regardless of who is on shift. A written trace of why something was decided, which is worth a great deal in a regulated environment. Staff spending their time on the interesting half of the work. Report these separately and honestly rather than converting them into invented currency; a finance reviewer trusts a stated qualitative benefit far more than a suspiciously precise number.
When the honest answer is no#
Sometimes the arithmetic says stop, and saying so is the most valuable thing an ROI exercise does. Low volume tasks rarely pay back a build. Tasks where every output must be checked anyway save review time only. And tasks where errors are expensive can produce a negative return even at high accuracy — 3% wrong on ten thousand high-value decisions is three hundred problems. Publish that result too. A team that has killed one agent on evidence is much more credible the next time it proposes one.
Frequently asked questions
What is a realistic completion rate for a first agent?
Sixty to eighty percent of a well-scoped, high-volume task, with the remainder escalated. Anyone promising ninety-five percent before seeing your data is describing a benchmark, not your Tuesday afternoon.
How long until payback?
For a well-chosen internal task at reasonable volume, commonly six to twelve months including maintenance. If your model shows payback in six weeks, check whether error cost and review time are in the arithmetic.
How do I count value when the agent only drafts?
Measure the time from blank page to approved output, before and after. Drafting agents frequently deliver most of the saving with a fraction of the risk, and they are much easier to get approved in the first place.
ai agent roimeasuring ai valueautomation paybacksupport deflection metricai business case