Language poses particular challenges for AI in customer service: whilst brief delays in chat go largely unnoticed, on the telephone, seconds can make or break the natural flow of a conversation. In this guest article, Gordian Braun from ElevenLabs explains why successful voice agents need far more than just good speech recognition. He explains the role played by response speed, system integrations, escalation logic and agent capabilities – and why the future of telephone-based customer service lies not in rigid conversation scripts, but in intelligent, proactive systems.
We’ve all heard the message: “Your call is important to us.” Then some music, followed by a menu, then a request to enter your customer number – which, four minutes later, a human representative reads back to you. The telephone is the channel through which service organisations handle the bulk of their volume, and yet it is also the one where the least has changed in the last twenty years.
Progress is evident in writing. Assistants answer standard enquiries, consolidate information from multiple systems, and hand over to human agents when matters become complex. On the phone, however, the same projects fail one after another. Not because of speech recognition. But because of something far more mundane: time.
A chat allows for pauses. The user sees three dots, waits, reads at their own pace, scrolls back. A phone call has none of these buffers. Anyone who hesitates for longer than the blink of an eye during a conversation doesn’t come across as thoughtful, but as having a problem. And people on the phone don’t behave like forms: they interrupt, carry on talking, mumble, start a sentence mid-flow, and only get round to the actual point at the end.
This shifts the crucial question. It is no longer ‘Does the system understand the intention?’, because many systems can do that nowadays. It is: How quickly does it respond, what happens when someone interrupts, and how does it keep a conversation going whilst a contract system is being queried in the background?
Rule-based voice dialogue systems have elegantly circumvented this problem by passing the burden onto the caller: “Say ‘contract’, ‘bill’ or ‘problem’.” As long as the case fits into the decision tree, this works. But as soon as someone has two issues at once, adds a detail or phrases a question differently from what was expected, the path ends where it always ends: on hold.
Agent-based systems turn this logic on its head. Instead of scripting every dialogue path in advance, you define target states, available tools, data access and guidelines. The agent finds the path between them itself. In a conversational context, this means, in concrete terms: it asks a follow-up question without losing the thread, creates a ticket whilst speaking, retrieves a delivery status and incorporates the result into the same sentence.
Importantly, ‘agent-based’ does not mean uncontrolled. In production environments, such systems operate with defined authorisations, verified knowledge sources and clear handover points to human operators. Providers and operators are subject to the transparency obligations set out in Article 50 of the EU AI Regulation: callers must be able to recognise who or what they are speaking to.
In practice, four factors determine whether a voice agent is successful.
Integration with the existing phone system. An agent that works only within an app misses out on exactly the customer group that calls because they don't want to use an app.
The Logic of Escalation. A good agent recognizes early on when they’ve reached a dead end and hands the case off with the full context. The worst moment in any automation process is when the system says, “Please describe your issue again.”
Measurable quality. Without defined evaluation criteria, monitoring during live operations, and regular analysis of actual conversations, any statement about automation rates remains nothing more than a claim with a decimal place.
The channel. Customers rarely stick to a single channel. They call, then send a message, and later attach a photo to a ticket. The obvious solution is to connect systems. The better solution is to avoid separating them in the first place. An agent who is familiar with the case doesn’t need to know which channel the most recent information came from. They listen, read, and write—for them, the phone, email, messenger, and ticket are all different forms of the same process. For the customer, this simply means: They don’t have to repeat themselves, no matter where they continue.
With the Magenta AI Call Assistant, Deutsche Telekom has integrated a voice assistant directly into the mobile network. It answers calls when the customer is unable to do so, works in around 50 languages and does not require an app on the device. What is interesting is not so much the function as the design: here, voice intelligence is not an application on the end device, but a layer within the infrastructure. You don’t install it. It’s just there.
The breakthrough does not lie in systems providing better responses. It lies in their ability to take action before anyone even picks up the phone. A callback regarding a known fault, rescheduling an appointment before a complaint is made, or providing information that renders the call unnecessary. The best service call is the one that never takes place.
Anyone who takes speech as an interface seriously therefore no longer writes conversation scripts. Instead, they build capabilities: tools, permissions, knowledge and decision-making logic. The conversation then flows naturally from these.
ElevenLabs is an AI research and product company that is transforming how businesses and individuals communicate with their customers, employees, and the world. We develop the industry’s leading AI language and audio models and deliver them through a platform that enables interactions across the entire enterprise—from AI agents for sales, support, and operations to creative tools for marketing and media.
No Comments