AI ROI: How to Measure the Return on an AI Agent (Worked Example)

AI ROI in 2026: the formula we use to measure return on an AI agent, a worked support-agent example with real numbers, and the four costs most cases miss.

September 2, 2026
Huzaifa Shad, Founder & Managing Director

Huzaifa founded DevEntia and leads its strategic direction, project execution and client delivery. His background spans software engineering, project management, software testing, design analysis and QA engineering. LinkedIn

AI ROI: How to Measure the Return on an AI Agent (Worked Example)

McKinsey's State of AI survey, published on 25 August 2026, found that 80% of companies report productivity gains from AI, yet 63% see no measurable impact on earnings, and only 6% qualify as high performers (The Register). That gap between "feels faster" and "shows up in the P&L" is the whole AI ROI problem in one sentence. Individual output rose. Business results did not, because nobody measured the right thing.

We build AI agents for clients, and the first question we ask in every scoping call is "what number will move, and how will you know?" This post gives you the formula we use to answer it, a worked example with real 2026 cost benchmarks, and the four costs that most AI business cases leave out. It does not cover the ROI of chat assistants for individual staff; that is a productivity story, and the McKinsey data shows productivity stories rarely reach the income statement.

What is AI ROI, in one definition

AI ROI is the net financial return from an AI system divided by its total cost over the same period, where "return" means a measurable change in one of three things: cost removed, revenue added, or risk avoided. Expressed as a formula: AI ROI = (measured gain minus total cost) divided by total cost. A result of 100% means the system paid for itself once over. Anything you cannot attach to one of those three buckets is not return; it is a benefit, and benefits do not survive a CFO review.

Two numbers put the difficulty in context. IBM's 2026 CEO study found that only about 25% of AI initiatives delivered their expected ROI, and just 29% of CEOs said they could measure AI ROI with confidence (IBM). MIT's NANDA project put the failure rate of generative AI pilots at 95% for measurable P&L impact (Fortune). The projects in the successful minority have one thing in common: a baseline measured before the build started.

Why most AI ROI calculations are wrong before they start

The typical AI business case counts hours saved, multiplies by an hourly rate, and calls the result savings. That number is fiction unless headcount or overtime actually falls, or the freed hours produce billable output. In our experience the four errors below account for most of the gap between projected and realized AI ROI.

  • Counting time saved that nobody reclaims. A drafting assistant that saves 40 minutes a day per person produces zero financial return if the team simply absorbs the slack. Return exists only when the hours convert to reduced cost or additional revenue.
  • Ignoring inference cost. Gartner reported in August 2026 that inference now consumes 55 cents of every enterprise cloud AI dollar, ahead of training for the first time (Tech Times). Agentic workloads use 5x to 30x the tokens of a simple chat call per task, and CIO.com documented a single runaway three-hour agent loop that cost roughly $3,700 (CIO).
  • Assuming 100% adoption on day one. Sinch's 2026 survey of 2,527 companies found that 74% had rolled back a live AI communications agent at least once (Sinch). Model the ramp, and model the rollback.
  • Leaving out the humans who keep it running. Every production agent needs someone reviewing escalations, updating prompts and tools when the business changes, and watching the eval dashboard. That is a fraction of an FTE, but it is not zero.

The AI ROI formula we use with clients

Our version has five inputs. Three are gains, two are costs, and every one of them must be measured, not estimated, before the project is approved.

InputWhat it isHow to measure it
Cost removedWork the agent does that a person no longer doesBaseline unit cost (per ticket, per invoice, per document) x units handled by the agent x share resolved without a human touch
Revenue addedSales or retention the agent causesConversion, response time, or churn measured against a holdout group that the agent does not touch
Risk avoidedErrors, fines, or outages preventedHistorical incident rate x average cost per incident x measured reduction
Build costOne-time engineering, integration, evaluation and change managementVendor quote plus internal time; integration is usually the largest line
Run costInference tokens, hosting, monitoring, human oversight, model updatesPer-task token cost from a two-week pilot x expected volume, plus a fixed monthly line

Then: Annual AI ROI = (cost removed + revenue added + risk avoided minus annual run cost minus build cost) divided by (build cost + annual run cost). Report it alongside payback period in months, because payback is what most boards actually decide on.

Worked example: a customer support agent

This example uses public 2026 benchmarks and round numbers so you can substitute your own. A mid-sized SaaS company handles 4,000 support tickets a month. Fully loaded human cost per resolved ticket is $8.00, in line with the $6 to $12 range in published 2026 benchmarks, while AI resolution costs run $0.99 to $2.00 per ticket (Fin). We will use $1.40 per AI-resolved ticket, including tokens.

LineAssumptionMonthly figure
Tickets per monthMeasured baseline4,000
Resolved by agent without a human55% after a 3-month ramp2,200
Saving per deflected ticket$8.00 minus $1.40$6.60
Gross cost removed2,200 x $6.60$14,520
Fixed run costHosting, monitoring, 0.15 FTE oversight$1,200
Net monthly gainGross minus fixed run cost$13,320
Build cost (one-time)Agent, integrations, eval suite$45,000

Payback is $45,000 divided by $13,320, or about 3.4 months after the ramp. Year-one return is ($13,320 x 12 minus $45,000) divided by $45,000, roughly 255%. Those are attractive numbers, which is exactly why you should stress them before believing them.

Stress test the two assumptions that break

Deflection rate and cost per ticket are where AI ROI projections die. If deflection lands at 35% instead of 55%, gross cost removed falls to $9,240 and payback stretches to 5.6 months. If a prompt-injection incident or a bad release forces a two-month rollback, the year loses two months of gain and adds remediation cost. The project still clears the bar in both scenarios, which is what "robust ROI" means. A project that only works at the optimistic number should not be built.

What we deliberately did not count

Faster response times, happier customers, and the support team's improved morale are real. They are also unmeasurable in the first year, so they stay out of the ROI figure and go into a separate "expected secondary effects" section. If churn measurably drops in the agent cohort against a holdout, it moves into revenue added the following quarter with a number attached.

How do you measure AI ROI step by step?

  1. Pick one process and one metric. "Support" is not a process. "Tier-1 billing tickets, measured as cost per resolution" is.
  2. Measure the baseline for at least four weeks. Volume, unit cost, error rate, cycle time. Without this you cannot prove anything later.
  3. Run a bounded pilot with a holdout. Route a fixed share of volume to the agent, keep the rest as a control, and log per-task token spend from day one.
  4. Price the run cost from pilot data, not vendor decks. Multiply measured tokens per task by list price, add hosting, monitoring, and the oversight fraction of an FTE.
  5. Calculate ROI and payback at three deflection rates. Pessimistic, measured, and optimistic. Approve on the pessimistic case.
  6. Re-measure every quarter. Model prices, volumes, and the process itself change. An ROI number older than a quarter is a guess.

Which AI projects have the clearest ROI in 2026?

Return is easiest to prove where there is a unit cost you already track and a volume large enough to matter. The pattern we see across client work is consistent, and it matches the published benchmarks.

Use caseReturn typeTypical measurabilityWatch out for
Tier-1 support deflectionCost removedHigh: cost per ticket existsRollback risk, escalation quality
Document extraction (invoices, claims, KYC)Cost removed + risk avoidedHigh: cost per document, error rateException handling volume
Lead qualification and follow-upRevenue addedMedium: needs a holdoutAttribution disputes with sales
Internal knowledge assistantProductivityLow: hours rarely convertBecomes a "benefit," not a return
Code generation for engineeringProductivityLow to mediumReview and rework time offsets gains

If your candidate project sits in the bottom two rows, it can still be worth doing, but do not sell it internally as ROI. Sell it as capability, with a cost cap. We go deeper on which agent shapes earn their cost in our guide to AI agents for business that actually work, and on why so many pilots stall in why AI pilots fail.

Build cost benchmarks to plug into the formula

The build side of the equation is the part you can get quoted, so get it quoted properly. Our published tiers, which match what we see across the market, run from roughly $15k to $50k for a single well-scoped agent with two or three integrations, $40k to $120k for a multi-step system with retrieval and an evaluation harness, and $10k to $40k for a document-processing pipeline. The full breakdown, including what drives a quote up, is in our AI agent development cost guide. If you are choosing between a chatbot, an agent, or plain automation for the job, the decision rules are in chatbot vs agent vs automation.

One rule of thumb for the run side: if the pilot shows per-task token cost above 20% of the unit cost you are trying to remove, the architecture is wrong, usually because a workflow was built as a multi-agent swarm. We covered that failure mode and its 15x token multiplier in multi-agent AI systems for business.

Frequently asked questions

What is a good ROI for an AI project?

For a cost-removal agent with a measurable unit cost, we expect payback inside 12 months and a first-year ROI above 100% on the pessimistic deflection case. Published enterprise studies report lower averages, such as Forrester's 120% ROI with a 15-month payback for one vendor's agentic solutions, because they blend well-scoped projects with poorly scoped ones. Judge your project against its own pessimistic case, not against an industry average.

How do you calculate the ROI of an AI agent?

Measure the baseline unit cost of the process, multiply by the volume the agent resolves without a human, subtract the agent's per-task cost and fixed run cost, and divide the annual net gain minus build cost by total cost. Always compute payback in months as well, and run the calculation at a pessimistic, measured, and optimistic resolution rate.

Why do 63% of companies see no earnings impact from AI?

Because most deployments target individual productivity rather than a process with a unit cost. Time saved per person does not become money unless it is reclaimed as lower cost or higher output. McKinsey's 2026 data shows the 6% of high performers redesign workflows and measure outcomes at the process level, which is the approach described in this post.

What costs do AI business cases usually miss?

Inference tokens at production volume, human oversight time, integration and data-cleanup work, evaluation and monitoring tooling, and the cost of a rollback. Inference alone is now the majority of enterprise cloud AI spend according to Gartner, and it scales with usage rather than sitting still like a licence fee.

Should AI ROI include soft benefits like customer satisfaction?

Not in the headline number. Track them separately and promote them into the formula only once you can attach a measured financial effect, such as a churn reduction in the agent cohort versus a holdout. Mixing soft benefits into ROI is the fastest way to lose a CFO's trust.

Key takeaways

  • AI ROI equals measured gain minus total cost, divided by total cost. If the gain is not cost removed, revenue added, or risk avoided, it is not return.
  • 80% of companies feel productivity gains and 63% see no earnings impact. Measure at the process level with a baseline and a holdout, not at the individual level.
  • Inference is now 55% of enterprise cloud AI spend. Price run cost from pilot token data, never from a vendor deck.
  • Worked example: 4,000 tickets a month, 55% deflection, $45k build, gives a 3.4-month payback. At 35% deflection it is still 5.6 months. Approve on the pessimistic case.
  • 74% of companies have rolled back a live AI agent. Model the ramp and the rollback, and re-measure quarterly.
  • Support deflection and document extraction are the easiest returns to prove. Knowledge assistants and coding tools are capability buys, not ROI plays.

Get the number before you get the build

If you have an AI project in front of you and no baseline yet, send us the process and the volume. We will return a one-page ROI model with pessimistic, measured, and optimistic cases, using the benchmarks in this post and your real numbers. If the pessimistic case does not clear payback inside a year, we will say so and suggest a cheaper shape, which is how our AI development practice keeps its clients out of the 63%.

Sources

Share this post

By subscribing you agree to our Privacy Policy.

Continue Reading

Blog & News

Learn, Grow, and Stay Ahead

Stay updated on tech, product development, and marketing insights.