Why do AI pilots succeed while production fails? See the 6 deployment gaps leaders must close before scaling conversational AI in customer service.
For leaders running or sponsoring an AI customer service deployment, and anyone about to greenlight the next phase after a promising pilot.
Most conversational AI pilots succeed. Most production deployments don't deliver what the pilot promised.
If that sounds like a contradiction, it isn't.
You've probably seen it happen, it's the most predictable pattern in enterprise AI today, and it showed up again and again in a conversation we hosted with senior leaders at a Directors Club breakfast roundtable in London.
The gap between pilot and production isn't a technology problem, it's a structural one. Pilots are built around conditions that don't exist at scale: curated data, limited scope, dedicated teams, and success metrics often set only after the early results are already in. Production strips all of that away.
Why we’re facing this struggle
AI in customer service isn't a bet you're weighing anymore. It's already here, and the numbers back that up, according to Microsoft's 2025 Work Trend Index:

The gap between leaders and laggards is widening fast. The question isn't whether you deploy AI in customer service anymore. It's how you make it stick.
Why pilots win by design
Pilots are built to succeed. None of the four conditions that make them look good will survive contact with production:
Limited scope: you test the best-case scenarios, high volume, well-documented flows. The messy, ambiguous, emotionally charged interactions rarely make it into the test.
Curated data: no edge cases, no legacy inconsistencies, none of the knowledge-base contradictions that exist in your live system.
A controlled environment: a dedicated team with full attention, while production means shared resources and a dozen competing priorities pulling at the same people.
Favorable measurement: you set success metrics after seeing early results, so the pilot ends up measured against a bar it helped set.
A pilot exists to help you learn, and it's the one window where you still have dedicated people, senior attention, and budget to course-correct. Every problem you find during the pilot is a problem you won't discover in production, when the team has dispersed and fixing things costs an order of magnitude more. The discomfort of finding problems early is the entire point. A pilot that finds nothing wrong has almost certainly been looking in the wrong places.
The pattern every leader recognizes
The pilot looks strong. Automation rates improve, service levels hold, stakeholders are excited, and the board approves the next phase.
Then production arrives. Escalation rates rise, complexity increases, operations struggle to keep pace, costs creep back up, and quietly, the project gets shelved.
This isn't bad luck, it's what happens by default when a pilot is designed to validate a decision you've already made, instead of testing whether that decision holds up once real conditions hit.
Six things that break between pilot and production
We've seen the same six gaps show up repeatedly, across insurance, e-commerce, FMCG, energy, and financial services deployments. It doesn't seem to matter what the sector is.
1. Pilot design. Pilots run on structured hours and cooperative customers, skipping peak load and emotional intensity. One insurance pilot looked strong until production brought volume and emotional range the pilot had never tested against.
The fix: test peak hours and real emotional range before go-live, not after.
2. Knowledge base quality. Pilots run on curated content, while production surfaces every outdated article and contradiction. One e-commerce retailer had returns-policy articles that contradicted each other and the agent's own logic, a knowledge base that had simply never been audited against live process.
The fix: audit the full knowledge base against live agent logic before go-live and treat it as ongoing hygiene.
3. Systems integration and data readiness. Pilots often run against mocked or disconnected systems, so nothing slows the conversation down. One FMCG company avoided this by integrating with its Salesforce sandbox during the pilot, catching a query bottleneck early enough to fix it cheaply.
The fix: integrate with real systems during the pilot, not after; bottlenecks found early are a success.
4. Use-case fitness and assumed portability. What works in one market rarely transfers automatically to another. A B2B ordering use case built around one European market's buying structure had no basis in the UK, where buyer relationships work differently.
The fix: validate fit locally before assuming portability across markets or business units.
5. Compliance, regulation, and infrastructure. Pilots typically run below the compliance radar. One banking deployment proved the concept fine at pilot scale, then hit a wall the moment production-scale data handling and security questions activated the bank's full infosec machinery, a conversation both sides had quietly deferred.
The fix: involve compliance and infosec early, even when it slows the pilot down.
6. Operational ownership and maintenance. Pilots run with senior, engaged people on both sides, but production needs a structured handover. Without a named owner for optimisation, knowledge, and escalation, degradation sets in fast once the project team moves on.
The fix: define RACI before the pilot closes, with named owners for optimisation, knowledge, and escalation.
What production-grade deployment looks like: leading European bancassurance provider
This client's inbound call routing is a case where the hard decisions, compliance requirements under DORA, integration complexity, and infrastructure constraints, were confronted during the project, not after.
The traditional IVR forced customers into generic menu options, driving up transfers, cost, and frustration, while manual call distribution inflated handling times across the board. The solution replaced the IVR entirely with real-time intent classification, routing every call to the right team without menus or transfers, integrated with Talkdesk CTI and Salesforce CRM for automated call distribution and case creation, and hosted on Microsoft Azure private cloud for security and governance by design.
The result: Over 90% intent classification accuracy in call routing, reduced call transfers, faster issue resolution, and improved customer satisfaction and agent productivity, because the gaps that usually surface in production had already been closed during the project.
What the organisations that scale AI do differently
They don't try to automate everything at once to justify the investment. That's exactly why most others fail. Instead, they:
Choose the right entry use case: highest volume, clearest scope, fastest measurable impact. Prove it before you scale it.
Start small but design for expansion, documenting where the roadmap goes after the first launch.
Measure impact from day one. If you can't measure it within 30 days, you picked the wrong use case.
Assign operational ownership, with named owners for optimisation, knowledge, and escalation. Not the vendor, and not IT alone.
Prepare the ground on data, systems, and change management before kick-off, instead of finding out what's missing once you're live.
They treat the pilot as a learning tool, not a box to tick. The goal isn't a tidy scorecard. It's a complete list of everything that needs to be true before production, tested and fixed while the team and the budget are still in the room.
Questions worth taking back to your own organisation
On pilot design: How do we design pilots that genuinely predict production performance, not just prove a concept?
On operational ownership: Who owns AI-driven customer service once it's live, and what does a sustainable operating model actually look like?
On use-case prioritisation: How do we balance automation targets with customer experience, and how do we make sure the right use cases, and the right expertise, are in place from the start?
Want to go deeper?
This post draws on our Directors Club Breakfast Roundtable briefing, "Why Conversational AI Wins Pilots and Loses in Production" we did in London. Full document here.



