Loading home page
This Research Note examines why technically successful enterprise AI pilots so often fail to reach production by reviewing enterprise survey research, applied-research findings, and analyst data on AI deployment outcomes. It finds that the dominant failure point sits at the data and platform layer rather than the model layer, with important implications for how leaders sequence AI and platform investment.
Curated by

Dr. Stéphane Niango
Research LeadershipExpert in DCOs & Strategic Transformation
DigitalQatalyst
Dr Stéphane Niango is a globally recognised digital transformation architect, strategy consultant and organisational design expert specialising in the evolution of Digital Cognitive Organizations (DCOs).
The pilot-to-production gap is rarely a model problem. It is a data and platform problem: pilots run on curated, bounded data assembled specifically to make the demonstration work, and that very design choice is what makes the pipeline unable to generalize to production.
Why do enterprise artificial intelligence pilots that perform well under test conditions so often fail to become durable production capabilities, and where in the transition does the failure concentrate? The question has become urgent because AI experimentation inside large organizations has grown far faster than AI actually running in production. Executives increasingly confront a version of the same puzzle after a demonstration: the model clearly worked, so why is the capability not live a year later, operating on real transactions across the business?
Recent enterprise-scale research offers a direct answer. Independent studies of AI project outcomes converge on the same pattern: pilots are judged competent far more often than they are judged deployable. The evidence points to failure concentrating not at the model layer, where pilots are typically evaluated and often succeed, but at the data and platform layer, where the enterprise's ability to grant a system reliable, governed, and repeatable access to live information becomes the binding constraint. Model quality is necessary but does not explain the gap; data and platform readiness does.
A pilot, for the purposes of this note, is a bounded technical exercise: a model is applied to a defined task using a curated dataset, a limited user group, and close technical support, with the goal of establishing whether the underlying approach can work. Production, by contrast, means the same capability operating continuously against live enterprise systems, serving a broad user population, under normal governance, security, and support conditions, without the manual scaffolding that made the pilot possible.
The "pilot-to-production gap" refers to the difference between a use case that has been technically validated and one that has become an operating capability. This note treats the gap as primarily an architectural condition rather than a talent, budget, or model-selection problem: the central question is whether the enterprise's data and integration architecture can support the same task once the pilot's temporary scaffolding is removed. That scope deliberately excludes organizational change management and workforce adoption, which are real but separate constraints addressed elsewhere in this research program.
Most AI projects fail, and the failure modes cluster around deployment infrastructure and data, not model capability. RAND Corporation's synthesis of interviews with 65 experienced data scientists and engineers found that more than 80 percent of AI projects fail, roughly twice the failure rate of comparable non-AI information-technology projects. The research identified five recurring root causes: misunderstanding between business and technical stakeholders about the problem being solved, poor data quality, technology-driven rather than problem-driven project selection, underinvestment in the infrastructure needed to deploy and operate a model, and attempts to apply AI beyond the current state of the art. Three of the five causes are architectural or infrastructural rather than about the model itself (Ryseff et al., 2024).
Very few pilots convert to measurable value, and the shortfall is attributed to an inability to adapt to enterprise context. MIT NANDA's 2025 study of the state of AI in business, drawing on a review of more than 300 disclosed AI initiatives, 52 structured organizational interviews, and 153 senior-leader survey responses, found that roughly 95 percent of generative AI pilots delivered no measurable return, while a narrow band of adopters captured most of the value. The researchers attributed the divide to a "learning gap": most deployed systems do not retain context, integrate with existing workflows, or improve through use in the way production systems require (Challapally et al., 2025).
Data architecture is a leading, not lagging, indicator of AI outcomes. Gartner's February 2025 survey of 248 data-management leaders found that 63 percent of organizations either lacked, or were unsure whether they had, the data-management practices AI requires. Gartner has projected that a majority of AI initiatives lacking AI-ready data will be abandoned before reaching durable production status (Gartner, 2025).
Maturity gaps are widening, not narrowing, and they concentrate in a small group of well-prepared organizations. Boston Consulting Group's 2025 assessment of 1,250 senior executives across more than 25 sectors, evaluated against 41 capability dimensions, found that only about 5 percent of organizations qualified as "future-built" for AI. Those organizations reported materially higher revenue growth, shareholder return, and margin improvement than the remaining majority, which continued to report minimal measurable value from AI investment (Boston Consulting Group, 2025).
Adoption is outrunning workflow and platform readiness at the enterprise level. McKinsey's 2025 global survey found that 78 percent of organizations reported using AI in at least one business function, yet more than 80 percent reported no material enterprise-level profit impact, and only 21 percent had redesigned any workflow around the capability (Singla et al., 2025). Separately, Deloitte's 2024 survey of more than 2,700 director-to-C-suite respondents found that 55 percent had avoided specific generative AI use cases because of data-related concerns, alongside regulatory compliance and governance gaps cited as leading deployment barriers (Deloitte, 2024).
The individual studies describe different symptoms; read together, they describe a single structural mechanism worth naming explicitly, since none of the cited sources states it directly. A pilot is judged successful by a narrow criterion: does the model produce a good answer on the data it was given? To meet that criterion efficiently, pilot teams build the smallest, cleanest, most hand-curated data pathway that will make the demonstration work. That pathway is, by design, bespoke: assembled for one use case, manually verified, and often maintained by the people running the pilot rather than by a platform team. This is what can be called the pilot paradox: the very design choices that make a pilot convincing (a narrow, hand-built, tightly scoped data pathway) are the opposite of what production requires (a broad, reusable, governed data pathway that other consumers can also depend on). A cleaner, more impressive pilot does not predict an easier production transition; it can predict a wider gap, because a highly optimized bespoke pipeline has less in common with a reusable platform service than a rougher one built closer to real operating conditions.
This reframes what RAND's "underinvestment in deployment infrastructure," MIT NANDA's "learning gap," and Gartner's "AI-ready data" findings have in common: each names the same missing layer from a different vantage point (an infrastructure lens, an adaptation lens, and a data-management lens), and each is describing the conversion of a boutique pilot integration into a utility-grade platform capability.
The 6xD interpretation places this squarely in D3, Digital Business Platforms: production readiness is a property of the platform an AI system runs on, not a property of the model. Consistent APIs, documented data ownership, and governed access are what let a capability be repeated, monitored, and reused, rather than rebuilt for every new use case. This is a boundary condition worth stating precisely: platform readiness is a prerequisite for workflow redesign, not a substitute for it or a consequence of it. An enterprise can redesign a workflow around an AI capability that still cannot reliably reach the data it needs; the redesigned workflow will fail for a different, less visible reason. Platform and workflow are sequenced conditions, not interchangeable investments, which is also why BCG's small population of "future-built" firms pulls away from the rest: they appear to have addressed the platform layer early enough that later workflow and adoption investment could actually take effect.
Assess data and platform readiness before approving pilot investment, not after. A pilot's success on curated data says nothing about whether the same data pathway can be made consistent, documented, and accessible at production scale. Leaders should require an explicit statement of what would need to change in the data and integration architecture before a pilot could operate as a production service, priced and scoped alongside the pilot itself.
Track the ratio of pilots to production deployments, not the count of pilots running. A large pilot portfolio can coexist with almost no production conversion if every pilot reproduces the same unresolved data-access problem. The conversion ratio is a more informative signal of programme health than pilot volume or user enthusiasm.
Fund reusable data and platform services, not one-off pilot integrations. Where multiple use cases depend on the same underlying data domain, treat that domain as a shared platform product with an accountable owner, rather than allowing each pilot team to rebuild its own bespoke pathway into the same data.
Sequence platform investment ahead of, or alongside, workflow redesign. Workflow redesign initiatives depend on a capability that can actually reach the data it needs; committing to redesign before the underlying data pathway is production-grade risks solving the wrong constraint first.
The evidence indicates that enterprise AI pilots fail to scale primarily because the data and platform conditions that make a pilot convincing are structurally different from the conditions production requires. Model quality is rarely the limiting factor; the ability of the enterprise to grant a system consistent, governed, repeatable access to real data is. The main limitation of this evidence base is that it rests on cross-sectional surveys, vendor- and consultancy-collected samples, and a research design not built to isolate data architecture from other contributing factors such as leadership alignment or workforce readiness; the pattern is consistent across independent sources but does not establish a single causal pathway. The next question this note leaves open is a measurement one: what specific, observable data-architecture indicators predict production conversion early enough to change an investment decision before the pilot begins, rather than after it has already succeeded on its own narrow terms.
The cited research relies on self-reported executive and practitioner surveys, consultancy engagement data, and a research design not built to isolate data-architecture maturity from other contributing variables such as leadership alignment, funding model, or regulatory context. Sample populations vary by industry, region, and organizational size, and figures such as RAND's 80 percent and MIT NANDA's 95 percent are not directly comparable, since they measure different outcomes (project failure versus absence of measurable return) using different methods. This note identifies a recurring, cross-source pattern rather than a precise, universal failure rate.
Enterprise AI spending keeps rising while ROI measurement stays stuck. The evidence suggests the failure is structural: accounting built for physical capital cannot register compounding, non-rival digital value, and closing that gap is a governance task.

Financial services firms adopt MACH to become composable. Evidence from banking and insurance shows the application layer decomposes while the data layer stays monolithic, so the architecture looks composable but does not behave as one.

Enterprise AI spending keeps rising while ROI measurement stays stuck. The evidence suggests the failure is structural: accounting built for physical capital cannot register compounding, non-rival digital value, and closing that gap is a governance task.

Capability-sequencing research suggests DCO competency areas are not parallel investment tracks. Governance capacity behaves as a binding constraint that automation, experience, and workforce competencies depend on to compound.