Back to Blog

How Turn-to-Turn Latency Shapes Voice AI

Low latency sounds like the obvious way to improve voice AI. Make transcription faster, and the conversation should feel faster too. AssemblyAI reduced emission latency to roughly 150 milliseconds, close to the threshold of human perception, yet some customers still felt it was slow.

That result led me to ask Luka Chkhetiani, Staff Researcher and Voice Agents Team Lead at AssemblyAI, why a model with such low emission latency can still feel slow in practice. His answer came back to the overall experience. Users never encounter an ASR model, endpoint detector, or LLM on its own. They notice the pause between finishing a thought and hearing the agent respond. Improving one component can move a benchmark without fixing the conversation.

That same end-to-end view shapes how AssemblyAI builds its products. Model quality and developer experience must go hand in hand. Even a highly accurate speech model is not very useful if teams struggle to integrate, configure, or deploy it.

Why Turn-to-Turn Latency Matters in Voice AI

Emission latency measures how quickly recognized words appear. Turn-to-turn latency covers everything that happens next: deciding that the speaker has finished, handling interruptions, passing context to the model, generating a response, and starting playback. The distinction sounds technical, but the experience is familiar. A fast transcription model cannot compensate for slow endpoint detection or an agent that hesitates at the wrong moment.

For Luka, that changes the engineering question. Teams need to examine the full path from the end of the user's turn to the start of the agent's reply and identify where the conversation loses its rhythm.

What Reliable Speech AI Looks Like at Scale

Clean audio can make speech models appear remarkably accurate. Real conversations are messier, with background noise, overlapping speakers, inconsistent microphones, unstable networks, accents, and interruptions. Benchmarks often underrepresent these conditions.

When I asked what reliability means at scale, Luka focused on predictability. Developers need to know how a model will behave across thousands of calls, not only whether it can perform well once. Predictable behavior gives product teams something they can design around. Strong aggregate benchmarks help, but unpredictable failures still make a model difficult to use in production.

Speech AI Benchmarks Should Show Where Models Fail

Word Error Rate, latency, and leaderboard rankings remain useful, but comparisons can be misleading when datasets, test conditions, and evaluation methods vary. A single score can obscure the failure patterns a production team most needs to understand.

I appreciated Luka's emphasis on transparent evaluation. Developers need to know where a model is weak and where it performs well. Each product balances accuracy, latency, cost, model size, infrastructure, and ease of integration differently. The right trade-off depends on the customer problem.

How AssemblyAI Gives Speech Models More Context

One part of AssemblyAI's recent work stood out to me: giving speech models more of the conversation. Agent context provides access to both sides of an exchange. Longer adaptive context windows help the model interpret what has already been said, while speaker revision improves attribution after the conversation.

This is also where voice agents still have room to grow. People naturally adjust their responses when someone sounds frustrated, confused, or hesitant. Most voice agents continue with the same tone and response style. Recognizing those signals and adapting in real time, while remaining predictable, may matter more than another incremental gain on a narrow metric.

I came away with a simple takeaway: production speech AI has to be evaluated as a complete system. It needs to respond at the right moment, handle messy audio, use context intelligently, and make its limits clear to developers.

In the full episode, Luka and I also discuss real-time reasoning, multimodal speech systems, developer experience, and the future of speech AI.

Watch the full episode here: https://www.youtube.com/watch?v=4gayz_8CBvo

Ready to build?

RTE Telehealth 2023
Join us for RTE Telehealth - a virtual webinar where we’ll explore how AI and AR/VR technologies are shaping the future of healthcare delivery.

Learn more about Agora's video and voice solutions

Ready to chat through your real-time video and voice needs? We're here to help! Current Twilio customers get up to 2 months FREE.

Complete the form, and one of our experts will be in touch.

Try Agora for Free

Sign up and start building! You don’t pay until you scale.
Try for Free