It’s tempting to think we’ve already solved realistic voice.
Listen to a modern text-to-speech model and that conclusion is easy to reach. The voices sound natural, expressive, and increasingly difficult to distinguish from human speech.
Then put those same models on real customer calls over noisy phone lines, in poor network conditions, and they’ll likely get interrupted mid-conversation. Suddenly, the gap between an impressive demo and a dependable voice agent becomes painfully obvious.
That gap was at the center of my conversation with Ankur Edkie, CEO of Murf AI. Ankur has spent more than five years building voice AI across the industry’s evolution—from the period before ChatGPT to today’s conversational AI boom. One lesson has stayed consistent throughout that journey: strong models still need strong production systems around them.
From Voice Realism to Production Reality
When Murf launched, text-to-speech was largely confined to IVRs and robotic assistants. Ankur’s ambition was broader. He wanted synthetic voices to be convincing enough for advertising, media, and real conversations—good enough that experienced creative professionals would struggle to distinguish them from human recordings.
That ambition quickly revealed a bigger challenge. Every new product, whether for content creation, dubbing, or conversational AI, followed the same pattern: internal demos looked impressive until they reached experienced users and real-world deployments. Professional directors, voice artists, and enterprise customers consistently identified nuances that benchmarks couldn't capture.
Production introduced an entirely new level of complexity, from noisy phone lines and poor microphones to unstable networks and unpredictable environments, where even a small transcription error could cascade through the entire voice stack. Ankur's point was that production exposes a range of conditions no controlled demo can fully reproduce. A dependable system has to keep working through that variability.
Why Cascaded Systems Still Matter
There's growing excitement around end-to-end speech-to-speech models that promise lower latency, richer context, and more natural conversations. While Ankur believes that's ultimately where the industry is headed, he argues that today's enterprise deployments still benefit from cascaded architectures, where RTC, ASR, LLM, and TTS each solve a distinct part of the problem.
There’s a practical reason for that separation. Speech recognition interprets messy human input, language models reason over intent, and text-to-speech focuses on delivering natural expression. Combining all of that into a single model dramatically increases complexity, compute requirements, and latency. For now, specialized systems working together remain the most practical way to deliver reliable, production-grade voice AI while the industry continues moving toward more unified architectures.
Latency Is a Budget and Consistency Sets the Rhythm
That architecture also shapes how Ankur thinks about latency. He treats latency as a shared budget across the entire pipeline: faster speech recognition and synthesis leave more room for the language model to reason and generate a higher-quality response. Consistency matters just as much. A system that responds in 120 milliseconds, then 400 milliseconds, then 180 milliseconds often feels less natural than one that reliably responds in 200 milliseconds. Humans notice rhythm as much as responsiveness, which is why Murf measures both response time and how consistently it is maintained across conversations.
The Enterprise Benchmark Is Trust
What stood out to me just as much was Ankur’s emphasis on trust. Enterprise teams judge a voice system by the experience it creates for customers. A bank may need a voice that makes people comfortable discussing sensitive financial information; a healthcare provider may prioritize a professional, dependable tone. That puts reliability, consistency, and voice selection at the center of the product decision. Hyper-realism can help, but it is only one part of an interaction that has to feel appropriate to the situation and dependable from one call to the next. For Ankur, trust is ultimately the standard that matters.
Build Around the Full Conversation
My biggest takeaway from Ankur was straightforward: evaluate the complete customer experience. Every component influences the next, so optimizing APIs in isolation can produce a technically impressive stack that still fails in practice. Production readiness comes from testing the whole system under the conditions in which people will actually use it.
That means paying attention to reliability across noisy inputs, consistent turn timing, the fit between the voice and the use case, and whether the agent achieves the customer’s intended outcome. Taken together, those details determine whether a voice experience feels dependable in everyday use.
The full conversation dives deeper into enterprise voice AI, latency, cascaded architectures, turn-taking, Falcon's efficient TTS architecture, and where Ankur believes conversational AI is headed next.
Watch the full episode here: https://www.youtube.com/watch?v=alTbaNDxQ5o
Ready to build?
- Explore: Agora’s Conversational AI Engine
- Learn: The Anatomy of Voice AI Agents
- Join: Our Developer Community on Discord


