Back to Blog

Building AI Avatars Worth Talking To

‍"Why am I talking to this character? Why do I care?"

Jia Shen posed that question while describing AKA Virtual's work on conversational characters. It stayed with us because it underlies every impressive avatar demo. An avatar earns sustained attention when its appearance, voice, behavior, and role give the interaction purpose.

In this AI Avatar Masterclass, we bring together three approaches to that problem. Jeff Lu, founder and CEO of Akool, is building photorealistic people and real-time video experiences. Jia Shen, CEO of AKA Virtual, favors animated characters that set clearer expectations. Richard Bowdler of Trulience focuses on the visual front end while allowing developers to connect their preferred language and voice technologies behind it.

What surprised us was how many decisions sit behind a convincing digital human. Visual style matters, but so do conversational timing, lip sync, character-specific voices, model architecture, and the job the character is expected to perform.

Realism Is a Product Decision

Akool puts the person at the center of visual storytelling. Jeff described a technical stack built to generate photorealistic characters, support live interaction, and move more inference onto laptops and mobile devices. Running more inference closer to the user reduces dependence on costly cloud compute and makes face-to-face interaction more responsive.

Jia's experience with VTubers led AKA Virtual toward a different visual language. A cartoon character does not invite the same scrutiny as a digital clone. Its stylization tells the audience that this is an AI character with a defined set of abilities. The design can then emphasize the traits that matter for the experience instead of trying to reproduce an entire person.

Give Every Character a Clear Job

AKA Virtual separates avatar experiences into two broad categories: functional assistants and entertainment characters. A concierge in a shopping mall or train station has a direct task: answer questions, provide directions, or guide a purchase. An entertainment character needs a premise, a personality, and an interaction the audience already understands.

One project shows how a clear premise gives an entertainment character direction. AKA Virtual created a fortune-teller experience that makes the distinction concrete. Visitors know what questions to answer, the character can stay within a specific role, and the session ends with a fortune they can take away. What we liked about this example is that Jia's team measures whether the character attracts people, holds their attention, and gives them a worthwhile experience. Those signals say more than visual novelty alone.

Voice and Timing Shape Believability

Trulience takes a modular approach, integrating its avatars with external language models, speech recognition, and text-to-speech systems. Richard calls this approach “bring your own brain.” Developers can choose the models that suit their language, privacy, and deployment requirements while Trulience handles the visual character.

That modularity does not remove the real-time engineering work. Richard said responses arriving in less than roughly 500 milliseconds can feel unnaturally quick. Lip sync presents another challenge because mouth shapes change across sounds, speakers, and languages. Emotion has to align with the response as well as the face. A surprised expression delivered at the wrong moment can be as distracting as a slow reply.

Voice design becomes even more specific in entertainment. Jia explained that a generally strong Japanese voice model may still fail an anime character whose pitch, vocabulary, and comic timing define the role. The voice has to sound like that character, not simply like fluent speech.

Build an Avatar Stack That Can Change

Each company expects the underlying models to keep evolving. Akool develops its core avatar technology in-house while using selected open-source foundation models elsewhere in the video stack. Jeff's team continually balances output quality, generation speed, compute cost, usability, and the control professional customers need.

AKA Virtual also keeps its infrastructure flexible. Its Shisa model can respond quickly while a larger model handles a more demanding request, helping the character maintain conversational flow. Trulience takes a similar modular view by allowing customers to combine its visual layer with different language and voice providers. This separation lets teams change models without rebuilding the character experience around every new release.

Digital Humans Already Have Practical Jobs

Akool customers use avatars for marketing, film production, AI agents, and internal communications. Its translation workflow combines voice cloning, language translation, timing, and whole-face reanimation. Jeff said the system supports more than 150 languages, giving global teams a way to distribute the same message without recording a separate version for every audience.

Richard described a Trulience deployment with a state government in India that placed screens around a village. People who could not rely on written interfaces could speak in Hindi with an avatar and access healthcare information. Trulience renders avatars in the browser, using the end user's device rather than allocating a dedicated server to every interaction. That approach can make large deployments more practical.

We came away with a clear starting point for anyone building an AI avatar: decide what the character is for. That choice should shape how it looks, how it speaks, how quickly it responds, and which models sit behind it. Visual novelty may get someone’s attention, but a clear purpose gives them a reason to stay.

The full episode explores photorealistic and animated characters, real-time translation, edge inference, specialized voices, healthcare access, and the product decisions behind useful digital humans.

Watch the AI Avatar Masterclass here: https://www.youtube.com/watch?v=ZNVFGBLvNG0  

Ready to build?

RTE Telehealth 2023
Join us for RTE Telehealth - a virtual webinar where we’ll explore how AI and AR/VR technologies are shaping the future of healthcare delivery.

Learn more about Agora's video and voice solutions

Ready to chat through your real-time video and voice needs? We're here to help! Current Twilio customers get up to 2 months FREE.

Complete the form, and one of our experts will be in touch.

Try Agora for Free

Sign up and start building! You don’t pay until you scale.
Try for Free