Back to Blog
AISeptember 24, 20266 min read

Google's Talking Avatars and AI Voiceover in 100+ Languages

Google's Talking Avatars and AI Voiceover in 100+ Languages
Image: Google (Gemini Live Avatar)

If you want to release a brand film in five languages today, you’re still looking at a long checklist: voice actors, studio hours, a separate edit for every language and another round of approvals. Three announcements this week hand a big part of that checklist over to software. Voice can now be directed, a face can now hold a live conversation, and a livestream can now dub itself.

Every shortcut comes with a bill, though. This time it arrives in the form of consent, watermarking, likeness rights and copyright. Here’s what actually changed, where the real opportunity lies for exporting brands, and where you need to tread carefully.

What changed?

Voice got a director’s chair. On September 23, Google announced Gemini 3.8 Flash TTS and Flash-Lite TTS. According to the announcement, the models cover more than 100 languages with over 2,000 ready-made voices. You can design a new voice by describing it in plain language, direct the performance line by line and stage two-speaker scenes. The headline feature is voice cloning: a 30-second sample is enough to replicate a voice, but only with a verbal consent recording from the voice’s owner. Cloning is not available in Illinois, Texas, the European Economic Area, the UK, Switzerland or India. Every clip carries a SynthID watermark.

The face started talking in real time. A day later came Gemini 3.8 Live with Live Avatar, which gives AI agents a lip-synced, near real-time talking avatar. Google cites support for 97 languages. Brand-specific avatars can be generated from a high-quality reference image, but for now only for companies on an enterprise allowlist. The feature lives in Gemini Enterprise, pricing hasn’t been disclosed, and the output is again marked with SynthID.

The stream started dubbing itself. At Made on YouTube 2026, YouTube announced real-time auto dubbing for livestreams and said likeness detection is coming to mobile. Likeness detection lets you enroll, get match alerts when your face shows up in someone else’s video, and take action; YouTube also plans to add speaking-voice detection. The event also brought video A/B testing across up to three different cuts.

What it means for your brand

In one line: voice and face are no longer assets you produce once. They’re assets you manage. Just like your logo and colour palette, your brand’s voice and its digital spokesperson are becoming part of a system.

For exporting brands, that’s a concrete opportunity. A trade-fair film, a product explainer or an investor deck no longer has to live in a single language. German, Spanish and Arabic versions of the same edit, with the same tone and the same character, can be ready in a fraction of the time. A live avatar puts a “brand face” on the table: one that greets visitors on your website, answers questions at a trade-fair stand or walks customers through product training in their own language.

Right next to the opportunity sit four serious risks:

  • Consent. Cloning a voice takes 30 seconds technically. Legally, it takes explicit, written permission with a clear scope. Google’s verbal-consent requirement is a floor, not a contract. Whose voice, in which languages, on which channels, for how long? Put it in writing first.
  • Watermarks and transparency. SynthID makes it possible to detect that content was AI-generated. That’s a good thing, but don’t build a campaign on the assumption that nobody will notice. Telling your audience up front is always cheaper than being found out later.
  • Likeness rights. Building an avatar from an employee’s, a founder’s or an actor’s face means renting that face to your brand. What happens when that person leaves? YouTube rolling out likeness detection more widely is a sign that platforms are starting to police this space too.
  • Music licensing. You’ve sorted voice and face; don’t forget the soundtrack. As Engadget reports, Universal and Sony have also sued over Suno’s v6 model, which Suno released with licences from Warner, BMG and Believe. The labels allege v6 was trained on outputs and preference signals from earlier models that the labels say were built on their recordings without permission; the suit covers at least 60,202 recordings, and Suno rejects the claims. The lesson is simple: even a “licensed model” label isn’t, on its own, enough cover for the music in your commercial.

Then there’s geography. With voice cloning switched off in the EEA and the UK, a brand targeting Europe may find the feature unavailable exactly where it needs it most. Don’t hang your plan on a single feature of a single tool.

Practical steps

  1. Define your voice identity first. Is your brand warm, authoritative, youthful? Write it into your brand guidelines as clearly as your typography. When voices can be designed from a description, a good description is your most valuable input.
  2. Sign the consent agreement before production. For anyone whose voice or face you’ll use, get written permission covering languages, channels, duration and what happens if they leave.
  3. Separate stock voices from cloned ones. For most jobs, choosing from 2,000-plus ready-made voices is faster and far less risky than cloning a real person.
  4. Have a native speaker review every language. AI can nail pronunciation, but only someone who lives the language can tell you how a tagline lands in that market.
  5. License music as its own line item. Libraries with documented licensing or an original score, not generative music of unclear origin.
  6. Disclose your AI use. One line in the description or credits prevents a loss of trust.
  7. Start small with live avatars. One scenario (say, product FAQs), one language and a measurable goal. Then scale.

How we approach it at Klyne

For Gala Fresh (MBA Tarım), we produced AI animated films for two international trade fairs: Fruit Attraction in Madrid and Fruit Logistica Asia in Hong Kong. For films made for an audience abroad, the natural next step is telling the same story with the same character in more than one language. This week’s announcements make that step far more accessible.

Our approach is simple: tools change, the system stays. When we plan an AI animated film, we think through character, voice, music and language versions together from day one. We treat consent and licensing as part of production, and adapt each language version for its market in the edit and in motion graphics. Where a story needs a real face, we shoot it: video production and AI production sit in the same team, under one roof.

Let’s work out together where an AI voice is the right call and where a human voice actor is. If you’re considering a multilingual film or a digital spokesperson for your brand, get in touch.

Frequently asked questions

Can I clone my own voice or an employee’s voice with AI?

According to Google, Gemini 3.8 TTS can replicate a voice from a 30-second sample, but it requires a verbal consent recording from the voice’s owner, and the feature is unavailable in the EEA, the UK, Switzerland, India, Illinois and Texas. Beyond that technical consent, we recommend a written agreement with a clearly defined scope.

Can AI-generated voice and video be detected?

Google marks both TTS and Live Avatar output with a SynthID watermark, which helps identify content as AI-generated. That’s one more reason to disclose AI use openly rather than hide it.

Is it safe to use AI-generated music in a commercial?

For now, caution is warranted. Universal and Sony’s second lawsuit against Suno shows that even a model backed by licensing deals isn’t clear of legal dispute. For commercial work, we recommend music with documented licensing or an original score.

Sources

  1. Gemini 3.8 text-to-speech says hello · Google Blog
  2. Introducing Gemini 3.8 Live with Live Avatar · Google Blog
  3. Made On YouTube 2026: All Announcements & New Features · YouTube Blog
  4. Sony Music and UMG say Suno's new models still violate their copyright · Engadget