A voice can change the feel of an entire conversation. The same words can sound reassuring, excited, uncertain, or sincere depending on how they are delivered. For developers building voice experiences, getting that delivery right has often meant working around the limits of text to speech: choosing from a small set of voices, adjusting scripts to coax out the right tone, and hoping the voice stays consistent through a longer exchange.
The launch of Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS marks a step forward. This is a new generation of text to speech that developers can direct with greater precision. Gemini 3.8 Flash TTS also brings an extended library of more than 20,000 voices, opening up far more choice across languages and regions.
At Agora, we’re excited about what this means for the people building real-time voice experiences. A more expressive voice gives an AI agent more ways to meet the moment, whether it’s helping a customer, guiding someone through an app, or bringing a story to life.
Control Speech Emotion, Pace, and Accent
Traditional text to speech starts with a script and produces audio. The words may be correct, but the delivery does not always match the intent. A welcome can sound flat. An apology can sound cheerful. A change in mood halfway through a sentence can be difficult to convey.
Gemini 3.8 Flash TTS gives developers more control over how speech is delivered. They can guide qualities such as emotion, pace, character, and accent, and shift the style as the moment changes. A character might begin a line quietly and end it with excitement. A voice agent might respond with warmth when a customer is frustrated, then sound more upbeat when the issue is resolved.
That control matters because conversation is dynamic. People adjust their tone as they listen and respond. Giving developers a way to shape those changes brings generated speech closer to the experience they are trying to create.
More Than 20,000 Voices Across Languages and Regions
Voice selection is another part of that creative process. Gemini 3.8 Flash TTS introduces an extended library of more than 20,000 voices, including options localized to specific languages and regions. The original named voices remain available as well.
For developers, that means more opportunity to find a voice that fits the experience and the people using it. A learning app, a game character, and a customer service agent may each call for a different sound. Regional voice options can also help a product feel more familiar to audiences in different places.
The size of the library is exciting, but the real value is choice. Teams can think beyond a handful of default voices and consider what their product should sound like for each audience and use case.
Voice Consistency and Multispeaker Speech
A convincing voice experience has to work beyond a single sentence. In a longer conversation, small changes in vocal quality or accent can become distracting. Exchanges between speakers also need to feel like a conversation, with natural pacing and responses.
Gemini 3.8 Flash TTS is designed to improve consistency through extended speech and to support more natural multi-speaker dialogue. It brings greater realism to elements such as breathing, pacing, and turn-taking. Those details can make a meaningful difference in an interactive story, a guided lesson, or a voice agent that spends several minutes helping someone complete a task.
For real-time applications, the goal is an experience people can stay engaged with. The voice should support the interaction, from the first greeting through the rest of the conversation.
What Developers Can Build with Gemini TTS
We see a broad range of possibilities for more expressive generated speech: voice agents that respond with appropriate tone, learning experiences that make dialogue engaging, characters that remain recognizable across a story, and applications that can speak to people in voices suited to their language and region.
Developers will decide what to build with Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. What makes this launch compelling is the additional room they now have to design the voice itself. They can think about delivery, identity, and the flow of a conversation alongside the words on the page.
At Agora, we believe the next wave of voice AI will be defined by experiences that feel natural to participate in. More expressive speech and a much wider range of voices give builders new tools to make those experiences possible.


