The Agent ROI Numbers Don't Add Up — and That's Fixable
Introduction
On August 8, 2026, during the sourcing for this piece, AI-generated search summaries twice asserted as fact that Forrester Research had analyzed 287 enterprise AI agent deployments across 14 industries and found an average ROI of 540%, with a median payback of 7.3 months. Neither attached a title, date, or document. The one page that carried it, on the SEO site ajentik.com, returned a 404. The most-quoted agent ROI statistic in circulation exists without a source.
What matters is where it lands. Forrester predicts that enterprises will defer 25% of planned AI spend into 2027, and KPMG's Q2 2026 pulse survey of 204 US leaders at billion-dollar companies records $202 million average planned AI spend for the coming year. Numbers like 540% are arriving on the slides that justify those approvals.
Tracing the year's headline figures to their instruments supports two conclusions. The numbers do not add up as commonly quoted: the loudest figure is untraceable, and the genuine ones measure different populations with different questions under a quantified self-report bias. And the numbers problem is fixable: operators can know what their agents return. Whether the return is there is a separate question, and nothing below promises it.
The 540% Study Forrester Never Published
A claim that specific should be trivial to verify. It is not. Searches restricted to forrester.com and tei.forrester.com return zero hits for a 540% figure, as of August 8, 2026; the reputable press has never covered it. The only "287" on any Forrester property is the 287% ROI in a Total Economic Impact study of Portnox Cloud, a network access control product, a plausible seed for a mangled "287 deployments."
Nobody has publicly debunked the figure either, which is partly why it propagates. Its parameters read as pastiche: Forrester's TEI program produces single-vendor composite models, never a cross-industry meta-analysis of 287 deployments. The figure is unverifiable at best, most likely fabricated, and appears here only under that label.
What Forrester actually publishes: 120%, modeled, and commissioned
"The Total Economic Impact of Microsoft's Agentic AI Solutions", published in January 2026 and commissioned by Microsoft, reports 120% ROI, payback under 15 months, and $24.2 million in net present value. Each number is modeled for a composite organization with 10,000 employees and $2.5 billion in revenue, built from eight interviews across six organizations plus 420 survey respondents; the study states results "may not be representative of all experiences." Under Forrester's own policy, TEI studies are commissioned consulting, excluded from the firm's syndicated research.

Sibling studies fill the class: compilations report 106–314% for Copilot Studio and 116% for Microsoft 365 Copilot, GitLab announced roughly 400% for its Duo platform, and a ~396% Agentforce figure circulates through secondary compilations. Quoting "Forrester says agents return N percent" without the words modeled, composite organization, and commissioned misdescribes the evidence class. The second tell is the publisher's own voice.
The house predicting a reckoning was not promising 540%
Forrester's October 28, 2025 predictions release warns that "the gap between inflated vendor promises and the value delivered to enterprises is widening, forcing a market correction," and finds that fewer than one-third of decision-makers can tie AI's value to financial growth. By June 3, 2026, its analysts described a small minority in meaningful agentic production and wrote that "ROI uncertainty traps enterprise ambition in pilot mode." The 540% attribution inverts the stance of the firm it names. The under-one-third figure returns below from an unrelated instrument.
"80% ROI proven" and the phantom Gartner pulse: the laundering layer
One fabricated statistic could be an accident; the pattern is not. agenticaiinstitute.org sells a report headlined "Enterprise AI Agent Deployment 2026: 80% ROI Proven," repackaging Anthropic's finding that 80% of organizations report measurable returns. Two distortions at once: self-report became "proven," and a share of respondents became an implied return magnitude. digitalapplied.com cites a "Gartner Agentic AI Pulse 2026" for the claim that 41% of rollouts reach positive ROI within a year; no such Gartner instrument can be located.
The layer feeds itself. The AI summaries asserting the 540% study cited only this ring: a statistic invented on an SEO page becomes a machine-generated answer, then a slide, then a budget line. Tracing every number to a named, dated instrument is the antidote, and the genuine surveys demand it too.
Real Surveys, Same Arithmetic Problem: Deployment Everywhere, Value Somewhere
Strip out the fabrications and an obvious rejoinder remains: the real numbers still show agents working. The strongest counter-evidence sits inside single instruments: one population, one questionnaire, one quarter, and arithmetic that refused to close.
Inside one survey: 97% deployed, 29% seeing significant ROI
Writer's second annual enterprise AI survey, run with research firm Workplace Intelligence and released April 7, 2026, sampled 2,400 C-suite executives and employees, fielded December 17, 2025 through January 25, 2026. The filter: employees already used generative AI at work, and executives' companies permitted it. Within that AI-forward population, 97% of executives say they deployed AI agents in the past year, and 29% report significant ROI. 54% say AI adoption is "tearing their company apart," up from 42% in the 2025 wave.
Writer is an enterprise-AI vendor whose commercial narrative the chaos findings serve; its figures get the same audit further down. The 68-point internal gap survives the caveat because it needs no cross-survey comparison: the executives reporting near-universal deployment report significant returns 29% of the time. That 29% also independently re-confirms Forrester's under-one-third attribution finding.
Nine in ten CEOs see benefits; 14% can define the P&L impact
A BCG title circulating in agent-ROI roundups, a "cost center to value engine" framing, matches no BCG publication and traces to a contact-center marketing blog. BCG's actual July 22, 2026 publication, "CEOs Are Starting to See Value from AI. Now Comes Execution.", outdoes the phantom as evidence. Among 152 CEOs of $500 million-plus companies, nine in ten see cost or revenue benefits in targeted areas, more than half name linking AI initiatives to the P&L as a key barrier, and 14% have clearly defined the P&L impact of all their AI initiatives. High performers are seven times more likely to have redesigned workflows.
From a source with every incentive to report value materializing, the sharpest finding is a measurement symptom: benefits nine in ten CEOs perceive, only 14% can locate on a financial statement. The 14% carries the fix at the end.
Anthropic's 80% and PwC's 56% are answers to different questions
The genuine ancestor of the laundered "80% proven" headline is Anthropic's 2026 State of AI Agents report, published December 9, 2025: of the 500+ technical leaders surveyed with research firm Material, 80% report that agent investments already deliver measurable economic returns. The public pages disclose no sampling frame, margin of error, or field dates, an unknown, not an accusation, but the respondents are by construction organizations already building agents. Weeks later, PwC's 29th Annual CEO Survey reported that 56% of CEOs saw neither higher revenue nor lower costs from AI in the prior twelve months; 12% saw both.
These figures, routinely presented as one mid-2026 wave, actually span September 2025 through July 2026, and each can accurately record what its respondents said while the set stays incompatible as measurements of one quantity. What lets them coexist is the best-quantified mechanism in the evidence.
The Optimism Tax: Self-Report Runs ~40 Points Hot
In July 2025, METR published a randomized controlled trial: 16 experienced open-source developers, 246 real tasks, with and without AI tools. With AI, they took 19% longer. Asked afterward, they estimated AI had made them 20% faster. METR labels the result historical, tooling having moved on, and warns against generalizing from 16 developers. The durable finding is not the slowdown but the inversion: perception and measurement pointed in opposite directions.

METR's May 11, 2026 survey of 349 technical workers, a disclosed convenience sample, found median self-reported speed gains of 3x against value gains of 1.4–2x, and restated its earlier finding: people overestimated AI's effect on their time by roughly 40 percentage points. Executives supply the confession. In a July 15, 2026 Wakefield Research survey for KUNGFU.AI, 74% of 300 US C-level executives admitted projecting unwarranted confidence in their AI strategy; 75% of Writer's executives called their own AI strategy more for show than guidance.
The same respondents forecast a 2.5x value gain by March 2027; the instrument class measurement keeps correcting is forecasting its own future. Survey-based ROI headlines carry a known positive bias of unknown per-study size; that bias turns the year's contradictions into one picture.
Ten Percent to Ninety-Seven: The Same Year's Numbers, Reconciled
Line up the year's deployment claims. Writer: 97% of executives at AI-permitting firms deployed agents in the past year. Google Cloud, September 4, 2025: 52% of senior leaders' organizations actively use or have deployed agents. KPMG, June 2026: 53% of US leaders at billion-dollar firms deploy agents. McKinsey's State of AI, late 2025, as reported: 23% scaling an agentic system. Gartner's 2026 CIO survey, as reported: 17% of organizations have deployed. DigitalOcean's Currents survey of 1,100+ practitioners: 10% scaling agents in production. Stanford's AI Index 2026: agent deployment in single digits across nearly every business function. Salesforce's telemetry counts activity rather than value.

The value side spreads the same way: Anthropic's 80% of builders reporting measurable returns, Google Cloud's 74% of adopters reporting first-year ROI, McKinsey's 39% attributing any EBIT impact with about 6% high performers, Writer's 29%, PwC's 56% reporting nothing. Pin each figure to who was asked, what counted as deployment, and how value was defined, and the chaos orders itself. Builders outrank large-firm leaders, who outrank all CEOs, who outrank measured functions; "any deployment in the past year" outranks production, then scaling, then per-function use. Restored to their instruments, most of these numbers can be true at once; stripped of them, as commonly quoted, they cannot.
Four different quantities, all named "ROI"
The reconciliation runs on one definitional key: at least four non-equivalent quantities travel under "ROI." Modeled ROI is the TEI construct met earlier, a risk-adjusted projection for a composite organization in a vendor-commissioned study. Survey-reported ROI is an executive checkbox with no audit and no counterfactual, the Writer, Google Cloud, Anthropic, and KPMG class. Measured impact is the class of randomized trials and telemetry. Predicted outcome, the analyst forecast, measures nothing yet. Setting the fabricated 540% against Writer's 29% set a fiction beside a checkbox and called the distance a controversy. The same key now has to turn against the numbers this argument might prefer to keep.
The Gloomy Numbers Fail the Same Audit
MIT Project NANDA's July 2025 preprint supplied the year's most-quoted gloomy statistic: 95% of GenAI pilots show no measurable P&L return, against $30–40 billion invested. The provenance check bites: the report is preliminary and not peer-reviewed, its P&L window runs six months or less, its base mixes 52 interviews, some 153 leader surveys, and 300 public deployments, and the underlying data was never released. Futuriom called the picture "irresponsible and unfounded." The 95% enters the record as contested, both positions attached, not as fact.
Gartner's figure that more than 40% of agentic AI projects will be canceled by end-2027 is a prediction, not a measurement, and its January 2025 evidence was a poll of 3,412 webinar attendees, a convenience sample. Writer's "54% tearing apart" is vendor self-report with theatrical question framing, completing the audit deferred earlier. BCG's client-outcome claims, 3x productivity and 80% cycle-time reductions, are unauditable consultancy numbers in the same class. The audit is symmetric or it is cherry-picking with extra steps. What symmetry does not do is settle the question underneath.
Measurement Artifact or Real Shortfall: The Question the Data Can't Settle
Two readings of the deployment-to-value gap survive the audit. The measurement reading holds that value cannot be seen because almost nobody instruments it: BCG's 14%, KPMG's 26% cost visibility, Forrester's under-one-third attribution, and METR's quantified optimism bias all describe enterprises that could be gaining or losing without knowing which. The value reading holds that much of the deployed estate has nothing to measure: Gartner attributes its predicted cancellations to escalating costs, unclear business value, and inadequate risk controls, while PwC's 56% of CEOs report no financial effect at all.
The dispute stays unresolved for a reason that is the thesis in another form: both positions rest on self-report at scale; no economy-wide measured dataset exists. The strongest supportable reading is narrower: honest measurement determines which condition an operator is facing, portfolio by portfolio. That is what "fixable" is entitled to mean, and all it is entitled to mean.
The Honest-Metrics Stack Only 14% Have Built
The practices that make agent numbers decision-grade converge across sources that agree on almost nothing else; the BCG and KPMG figures establish they are rare, not impossible. Three layers do the work.
Five provenance questions before any number reaches a slide
The checklist the 540% case validates is cheap to run. Who paid for the study? What population was sampled, through what filter? Is the figure self-reported or measured? Is it a realized result, a modeled composite, or a projection? What time window, and under what definitions of "agent," "deployed," and "ROI"? The fabricated Forrester figure fails the first question outright; NANDA's 95% stumbles at the window; the Microsoft TEI's 120% passes and gets priced correctly: modeled, for a composite organization, commissioned by the vendor.
Unit economics first: ROI has a denominator
Most organizations lack the cost base a return calculation needs. KPMG finds 26% of large US companies with full real-time visibility into AI operating costs and 36% with token or usage controls. DigitalOcean's data shows the stakes: 49% name inference cost as the top blocker to scaling AI, and 44% of teams spend 76–100% of their AI budget on inference. An organization that cannot see its inference bill cannot compute agent ROI, whatever its dashboards claim. Cost-to-serve, inference included, is the denominator, and it comes first.
Sentiment, activity, value: rank the evidence classes
Vendor-survey percentages are sentiment, never grounds for a scaling decision. Telemetry is activity: Salesforce's 1.9-day mean time-to-production and fleets growing from 5 agents to 13 count motion, not value. Only baselined outcome deltas are value, for the reason METR demonstrated: recollection cannot serve as the numerator. An honest metric looks like eSentire's, from Anthropic's report: security investigations cut from five hours to seven minutes, a cycle-time delta against a baseline. Cost-to-serve, cycle time, and error or rework rate form the stack, and the redesign multiples, BCG's 7x and McKinsey's high performers, tie outcomes to workflow change rather than model choice. Anthropic itself advises measuring business outcomes over token costs, sound advice from an interested party.
One bound remains. Larridin's five-link chain, running from spend through adoption, proficiency, and productivity to business outcome, states honestly how hard attribution stays even with good instruments; most organizations measure only the two ends. The stack makes the gap visible and decision-grade. It does not conjure the return.
Conclusion
The 2027 budget cycle is under argument now, against Forrester's predicted 25% spend deferral and KPMG's $202 million average planned investment. The numbers arriving with those requests include a statistic nobody published, several laundered past their instruments, and genuine figures answering different questions under a roughly 40-point optimism bias. The fix is as large as the evidence allows: five provenance questions, a real cost denominator, and a ranked evidence hierarchy let an operator know what their agents return. The ROI itself is not assured; honest measurement is what reveals whether a portfolio holds a measurement artifact or a real shortfall.
One absence completes the picture. No vendor, analyst house, or consultancy publishes a distribution of realized agent paybacks; modeled paybacks, expected paybacks, and one fabricated 7.3-month median circulate instead. The field's most decision-relevant dataset does not publicly exist, and the operator who builds honest metrics acquires it privately. Whether even those metrics need a further quality term, net of rework and review overhead, is the question METR's speed-versus-value gap leaves open.