Loading home page
Research Note: Why Enterprise Accounting Cannot See What AI Investment Produces
Curated by

Dr. Stéphane Niango
Research LeadershipExpert in DCOs & Strategic Transformation
DigitalQatalyst
Dr Stéphane Niango is a globally recognised digital transformation architect, strategy consultant and organisational design expert specialising in the evolution of Digital Cognitive Organizations (DCOs).
AI investment ROI looks invisible at the enterprise level not because the return is absent, but because the accounting infrastructure enterprises use was built to measure physical capital, not compounding, non-rival digital value.
Enterprise AI budgets keep climbing at the same time that finance leaders keep struggling to defend the return. MIT's Project NANDA reviewed more than 300 public AI initiatives and conducted 52 structured interviews in 2025, and found that roughly 95 percent of generative AI pilots showed no measurable effect on profit and loss, despite an estimated $30-40 billion in enterprise investment (Challapally et al., 2025). McKinsey's contemporaneous global survey found a similar shape at larger scale: more than 80 percent of respondents reported no tangible enterprise-level EBIT impact from generative AI, despite near-universal adoption (Singla et al., 2025).
This is not simply a story about immature technology. The research question this note investigates is why enterprise ROI measurement so consistently fails to register AI's return, and what that persistent gap implies for how executives should govern AI investment. The evidence points to a structural answer: the accounting and reporting infrastructure built to measure physical capital cannot see the way digital, compounding, non-rival value from AI accumulates, so the absence of a number is not proof of an absent return.
Economy 4.0, in the 6xD framework's D1 (Digital Economy) lens, describes an economic regime in which value increasingly derives from digital assets, data, and software-mediated capability rather than physical capital alone. Its defining features for this note are that digital value can compound with use, can be reused across functions at near-zero marginal cost, and remains partly intangible in a way traditional accounting was not built to record.
"ROI measurement," as used here, refers to an enterprise's ability to attribute a portion of financial or operating performance to a specific AI investment through its standard finance, planning, and reporting systems, rather than through bespoke research exercises. This note is bounded to enterprise-level, financially reportable measurement. It does not address task-level productivity studies or macroeconomic estimates of AI's aggregate effect on GDP, which follow different measurement logics entirely.
The analysis assumes that a persistent, cross-survey pattern of unmeasured value, observed independently by different research organizations using different methods, is more likely to reflect a measurement limitation than a coincidence of underperformance repeated at scale. This note does not assume AI investment is delivering positive returns everywhere. It asks why the enterprise's own instruments cannot settle the question either way.
The evidence converges on a consistent pattern across four independent research streams, despite different sponsors and methods.
Adoption is outpacing any registered financial return. MIT's Project NANDA combined a systematic review of more than 300 publicly disclosed AI initiatives with 52 structured interviews and 153 senior-leader survey responses, and found that only about 5 percent of generative AI pilots reached measurable profit-and-loss impact (Challapally et al., 2025). McKinsey's global executive survey found the same shape at larger scale: 78 percent of organizations reported using AI in at least one function, but more than 80 percent reported no tangible enterprise-level EBIT effect, and only about 17 percent attributed 5 percent or more of EBIT to generative AI (Singla et al., 2025). Two independently constructed studies, using different sampling and different questions, describe the same divide between adoption and registered value.
Finance functions cannot agree on a single measurement approach, and are increasingly abandoning the attempt to force one. Gartner's 2026 CFO research argues that the search for one uniform ROI formula is itself the wrong ambition: AI investments span routine productivity tools, targeted process changes, and transformational bets that do not share a cost curve or a value profile, and CFOs applying one metric across all three routinely misjudge the portfolio (Gartner, 2026). This follows an earlier Gartner finding that 30 percent of generative AI projects were expected to be abandoned after proof of concept by the end of 2025, commonly because a value case could not be demonstrated in the terms the organization was demanding (Gartner, 2024).
Where organizations do report progress, it correlates with redesigning the measurement approach, not only the technology. Deloitte's 2025 survey of more than 1,850 European and Middle Eastern executives found that only about one in five organizations qualify as AI "ROI leaders," and that the strongest performers are redefining what counts as return, beyond direct cost savings, to include resilience, innovation capacity, and process reinvention (Deloitte, 2025). McKinsey's analysis of the practices distinguishing higher performers found that redesigning workflows around AI, rather than deploying it into an unchanged process, was the strongest correlate of EBIT impact among the organizational attributes it tested (Singla et al., 2025).
Counterevidence deserves acknowledgment. Not every gap is a measurement artifact. Gartner's abandonment finding shows that some pilots are cancelled for ordinary reasons, cost, poor fit, or weak execution, that have nothing to do with accounting categories. The evidence does not show that all unmeasured AI investment is secretly productive. It shows that the enterprise's standard instruments are demonstrably unable to distinguish between the two explanations.
Three sources point at the same underlying mechanism from different disciplines, and read together, they explain why the gap persists rather than closing as AI matures.
Accounting theory has a name for part of this problem. Corrado, Hulten, and Sichel's foundational work on intangible capital found that national accounts have historically excluded hundreds of billions of dollars of intangible investment, in software, organizational process, and knowledge capital, because it does not fit categories built for physical plant and equipment (Corrado et al., 2009). Enterprise accounting inherits the same categories. Under IAS 38, the international standard governing intangible assets, costs incurred in the research phase of an internally generated intangible must be expensed as incurred, and only a narrow band of development costs meeting strict criteria can be recognized as an asset (IFRS Foundation, n.d.). Most enterprise AI investment, prompt engineering, data preparation, workflow redesign, model fine-tuning, falls on the expensed side of that line. It reduces reported profit in the period it is incurred and creates no balance-sheet asset against which a later return can be measured. The absence is not a reporting failure. It is the standard working as designed, for a different kind of economy.
Brynjolfsson, Rock, and Syverson's Productivity J-Curve supplies the mechanism's time dimension. Their model of general-purpose technologies shows that early adoption periods systematically understate output and productivity, because firms are investing heavily in unmeasured complementary intangibles, retraining, process redesign, new organizational routines, that do not show up as capital in the accounts. Only later, once those investments mature into visible output, does measured productivity catch up, and can even overshoot (Brynjolfsson et al., 2021). Applied to the present pattern, this suggests at least part of today's invisible AI return is not permanently invisible. It is temporarily unmeasured, and the accounting system is not built to distinguish a temporary lag from a permanent absence.
That distinction is the analytical contribution this note adds to the evidence. The research reviewed here typically treats "AI investment isn't showing a return" as one condition. It is more useful, and more actionable, to treat it as two different conditions that look identical on a P&L: measurement lag, where real value exists but has not yet compounded into recognized output, and measurement blindness, where the value form itself, non-rival reuse of a model across business units, compounding accuracy gains, avoided downside risk, will never register cleanly in accounts built around discrete, depreciable, single-use assets. Confusing the two produces exactly the pattern Gartner documents: some pilots reporting no return are cancelled for underperforming when they are, in fact, still in the lag phase; others are kept running on faith when their value is structurally blind and no future accounting cycle will resolve the question.
Read together, the three research streams describing organizations that do show results, MIT NANDA's roughly 5 percent extracting real value, Deloitte's one-in-five ROI leaders, and McKinsey's workflow-redesign cohort, are plausibly describing an overlapping population: firms that have built parallel, purpose-built measurement (reuse rate, capability trajectory, cycle time) alongside standard accounting, rather than waiting for GAAP or IFRS categories to catch up. The redesign is not incidental to their results. On this reading, it is close to a precondition for them.
Diagnose before defunding. When a program shows no measurable return, leaders should first classify which condition they are looking at: measurement lag, where the underlying capability is real but not yet captured by standard reporting cycles, or measurement blindness, where the value form will never be visible to the existing chart of accounts. Cancelling a lagging investment on the same evidence used to correctly cancel a genuinely unproductive one destroys value; funding a structurally blind investment indefinitely, with no parallel measure, does the opposite.
Build a parallel measurement layer before scaling, not after. Following Gartner's portfolio logic, treat productivity tools, targeted process changes, and transformational bets as different value profiles requiring different metrics, not one ROI formula (Gartner, 2026). Reuse rate across business units, model accuracy trajectory, cycle-time reduction, and avoided exception cost sit outside standard financial reporting but can be tracked consistently and audited internally.
Treat workflow redesign as a measurement precondition, not a follow-on step. The organizations showing registered results redesigned the process the AI sits inside before, or alongside, deployment, not after (Singla et al., 2025). Boards should ask whether a proposed AI investment includes an explicit change to the workflow and the metric that will register its effect, not only a description of the model or use case.
Expect the gap to close unevenly, and hold the accounting system accountable for saying so. Finance functions should be explicit, in board reporting, about which AI investments are being tracked on a lag basis pending future recognition, and which are being tracked on parallel non-financial metrics because they are structurally unlikely to appear in standard accounts. Silence on this distinction is itself a governance risk.
AI investment ROI looks absent at the enterprise level not because AI investment fails to produce value, but because the accounting infrastructure enterprises use to measure it was built for a different kind of capital. Evidence from MIT, McKinsey, Deloitte, and Gartner converges on the same divide between adoption and registered return, and academic work on intangible capital and general-purpose technology mismeasurement gives that divide a credible structural explanation, distinguishing a temporary measurement lag from a more durable measurement blindness.
The main limitation is that no dataset yet directly tests this note's central distinction. No study has separated "lagging" AI investments from "structurally blind" ones and tracked what each looks like several years later. The next research question is empirical: which AI investments currently showing no return are still invisible five years on, and which resolve into recognized value once the intangible investment behind them matures?
Financial services firms adopt MACH to become composable. Evidence from banking and insurance shows the application layer decomposes while the data layer stays monolithic, so the architecture looks composable but does not behave as one.

A review of enterprise AI deployment research finds that most pilots that work in testing never reach production, and the strongest predictor of failure is not model quality but whether the data and platform architecture around it can support it at scale.

Financial services firms adopt MACH to become composable. Evidence from banking and insurance shows the application layer decomposes while the data layer stays monolithic, so the architecture looks composable but does not behave as one.

Capability-sequencing research suggests DCO competency areas are not parallel investment tracks. Governance capacity behaves as a binding constraint that automation, experience, and workforce competencies depend on to compound.