Executive Summary
Klarna deployed AI in its customer operations in February 2024 and publicly claimed it was performing the equivalent work of 700 human agents. By May 2025, the company had begun rehiring human agents: citing customer experience quality as the reason. The deflection trap is the pattern Klarna illustrated: optimize for ticket avoidance (AI handles the query, no human involved) and the metric improves while the actual outcome: did the customer get a resolution that served their needs?: deteriorates. Gartner research finds that 64% of customers who interact with AI-first customer service report being frustrated by lack of resolution on complex issues. Brynjolfsson's peer-reviewed QJE study on AI in call centre operations (+14% average productivity, +34% for lower-skilled workers, approximately 0% for expert workers) shows the real picture: AI amplifies novice workers' performance dramatically but does nothing for experts.
Klarna optimized for deflection, then had to rehire humans
Klarna deployed AI in its customer operations in February 2024 and publicly claimed it was performing the equivalent work of 700 human agents. By May 2025, the company had begun rehiring human agents: citing customer experience quality as the reason. The deflection trap is the pattern Klarna illustrated: optimize for ticket avoidance (AI handles the query, no human involved) and the metric improves while the actual outcome: did the customer get a resolution that served their needs?: deteriorates.
Gartner research finds that 64% of customers who interact with AI-first customer service report being frustrated by lack of resolution on complex issues. Brynjolfsson's peer-reviewed QJE study on AI in call centre operations (+14% average productivity, +34% for lower-skilled workers, approximately 0% for expert workers) shows the real picture: AI amplifies novice workers' performance dramatically but does nothing for experts.
AI amplifies novices but does nothing for experts
D5: Digital Worker and Workspace: addresses exactly this challenge. The digital worker's performance is not determined by whether AI is present, but by how AI is integrated into their work design. Brynjolfsson's finding is a D5 result: AI improved novice workers because it gave them access to expertise they didn't have; it produced no uplift for experts because their value is in judgment and relationship, not information retrieval.
The D5 lens requires organizations to ask: at what tasks and for which worker profiles does AI amplify human performance, and at what tasks does it substitute for human interaction in ways that destroy the value that made the interaction worth having? The deflection trap arises when this question is not asked: when AI deployment is measured by call volume reduction rather than by human-AI work design effectiveness.
Replace deflection rate with resolution rate and recovery cost
Operations leaders in financial services customer functions should replace deflection rate as a success metric with resolution rate and downstream recovery cost (the cost of complaints, escalations, and relationship damage that follows an unresolved AI interaction). The Klarna case quantifies the lesson: public claims of AI efficiency followed by a rehiring cycle and a customer experience correction is a costly way to learn that deflection is not resolution.
The practical design question is: for each customer interaction category, is the right design (a) AI handles it fully, (b) AI handles it with a human available on request, or (c) a human handles it with AI providing support? The Brynjolfsson skill-differentiation finding suggests that categories (b) and (c) are far more prevalent in financial services customer operations than most organizations' current AI deployment assumes.
What will reveal whether the field is learning
Two signals will indicate whether financial services operations is absorbing the deflection-trap lesson or repeating it.
- Klarna operations model public reporting (2025-2026): Klarna's investor communications and operational data will confirm whether the rehiring cycle was a temporary correction or a permanent rebalancing: and at what human-AI ratio the customer operations function has stabilised.
- Gartner customer operations AI maturity benchmarks: Gartner's 2026 customer service AI maturity research will indicate whether the deflection-trap pattern is sector-wide or concentrated in early-adopter organizations.
- Brynjolfsson follow-up research on AI-worker productivity patterns: the QJE paper's skill-differentiation finding is now widely cited; follow-up research examining whether organizations have redesigned their AI deployment in response to the finding will show whether the field is learning from the evidence.
Sector Context: Customer Operations Is a Work-Design Problem
Financial-services customer operations combine high-volume routine interactions with moments of significant financial and emotional consequence. Balance checks, status queries, and simple servicing requests are structurally different from fraud disputes, vulnerable-customer cases, hardship conversations, complaints, or complex product decisions. Treating all of them as one automation opportunity obscures the actual design problem.
The first wave of conversational AI emphasized containment and deflection because those metrics translated directly into contact-center cost. The next wave must optimize the complete service outcome: whether the issue was resolved, whether the customer understood the answer, whether risk was managed, and whether downstream rework was created.
Four Forces Exposing the Deflection Trap
AI makes routine handling dramatically cheaper. That creates strong incentives to maximize automation even when interaction categories differ materially in complexity and consequence.
Complex cases contain more context than scripts capture. Financial distress, fraud, bereavement, complaints, and exceptions often require judgment, empathy, negotiation, and accountability rather than information retrieval.
Customer effort moves downstream when resolution fails. A deflected interaction can reappear as a repeat contact, complaint, escalation, churn risk, or regulatory issue. The apparent saving at the first contact may therefore be false economy.
AI changes worker performance unevenly. The productivity evidence cited in this brief suggests that AI can transfer elements of expert practice to less-experienced workers. That makes augmentation design at least as important as substitution.
The Structural Shift: From Channel Automation to Human-AI Service Architecture
D5 reframes the problem around tasks, skills, decision rights, and workflow. The design unit is not “the contact center” or “the chatbot.” It is the interaction category. For each category, the organization should determine whether AI can resolve independently, should operate with human escalation, or should support a human who remains the primary decision-maker.
D2 supports the learning loop by connecting customer outcomes, complaints, escalations, and resolution data back into service design. D3 provides the shared customer context and workflow services required for seamless handoff. D4 ensures automation is scaled according to evidence rather than channel-level cost targets.
Opportunities and Risks
A calibrated model can reduce routine workload, improve response speed, raise novice performance, and reserve experienced staff for interactions where judgment creates the most value. It can also improve consistency by giving workers better access to knowledge during the interaction.
The risks include inaccessible escalation, incorrect financial guidance, loss of empathy, automation bias, and metrics that hide customer harm. Poorly designed AI can also deskill the workforce if humans are left only with the hardest cases but receive less exposure to the routine interactions through which expertise was previously built.
Five Executive Priorities
Segment interactions by complexity and consequence. Do not apply one automation target across all customer journeys.
Replace deflection with end-to-end resolution economics. Measure first-contact resolution, repeat contact, escalation, complaint, recovery cost, and customer outcome together.
Design escalation as a core capability. Customers and AI systems should be able to move to an accountable human without losing context.
Use AI to build worker capability. Instrument whether AI improves less-experienced employees and design coaching, knowledge, and supervision around that effect.
Govern high-consequence interactions separately. Apply stronger testing, monitoring, human oversight, and auditability where financial vulnerability, complaints, fraud, or regulated advice are involved.
Executive Decision Test
The practical test for leaders is whether the next investment strengthens an enduring sector capability or merely improves one local initiative. Before approval, executives should be able to identify the operating-model dependency being changed, the reusable capability being created, the owner of the cross-functional decision, the outcome metric that will demonstrate value, and the governance mechanism that will remain after the implementation team leaves. If those answers are missing, the organization is still funding activity rather than redesign.
The sequencing principle is equally important. Leaders do not need to replace the entire estate before value can emerge. They do need each increment to move toward a coherent target architecture. That means using current initiatives to establish shared data, interfaces, decision rights, measurement, and reusable controls that subsequent initiatives can consume. The result should be cumulative: every deployment should make the next deployment easier, faster, safer, or cheaper.
This is also the distinction between adoption and capability. Adoption measures whether a technology or service is being used. Capability measures whether the organization can repeatedly produce the intended outcome under changing conditions. Industry Brief decisions should therefore be evaluated against capability compounding, not launch completion.
Closing Perspective
The central issue is structural rather than technological. The organizations that create durable advantage will be those that turn the capability described in this brief into part of the operating model, with clear ownership, reusable architecture, measurable outcomes, and governance that persists beyond an individual project. The leadership question is therefore not whether to adopt another tool or launch another initiative. It is whether the sector's operating architecture is being redesigned so that each investment strengthens the next one.



