Google is turning its latest AI advances into something that sounds almost mundane: people talking across languages without having to slow down, pause or stare at their screens. With Gemini 3.5 Live Translate, a new audio model now rolling out across Google Translate, Google Meet and developer tools, the company says it can deliver near real‑time voice‑to‑voice translation in more than 70 languages while staying just a few seconds behind the speaker, and preserving tone, pacing and pitch.

How Gemini 3.5 Live Translate works
Gemini 3.5 Live Translate is an audio‑first model that processes speech as it is streamed, automatically detecting more than 70 languages, and generating translated speech on the fly. Google says the system is designed to balance two competing demands: waiting long enough to gather context for accurate translations and speaking quickly enough to keep up with natural conversation.
Unlike traditional “turn‑by‑turn” translation apps, which wait for a speaker to finish an utterance before generating output, Gemini 3.5 Live Translate continuously listens, interprets, and speaks, aiming to stay only a few seconds behind the source audio. This streaming approach reduces the long, awkward pauses that have made live machine translation feel stilted in the past, and it can handle overlapping speech and casual back‑and‑forth dialogue better than earlier systems.
The model is built to be noise‑robust, with Google highlighting its performance in loud, unpredictable environments like street tours, classrooms, or call centers, where background sound and cross‑talk can easily trip up older tools.
Where you can use it: Translate, Meet and APIs
Google is not keeping Gemini 3.5 Live Translate in the lab. The company says the model is rolling out in three main directions:
Google Translate (Android and iOS)
Gemini 3.5 Live Translate now powers the Live translate feature in the Translate app globally. Users can tap “Live translate” and plug in any pair of headphones to hear a more seamless translation that mirrors the speaker’s tone across 70+ languages. On Android, a new “listening mode” lets users hold the phone to their ear like a call and hear the translated audio through the earpiece, no headphones required.
Google Meet (enterprise preview)
Speech translation in Google Meet previously supported just five languages and only in and out of English. With Gemini 3.5 Live Translate, Meet can now handle 70+ languages and more than 2,000 language pairs in a single meeting, allowing, for example, Spanish–Japanese or Arabic–French translations without routing everything through English. A new button in the Meet controls lets users start speech translation instantly. This upgrade is in private preview for select Google Workspace business customers, with a broader rollout promised later this year.
Gemini Live API and Google AI Studio (developers)
For developers, Gemini 3.5 Live Translate is available in public preview via the Gemini Live API and Google AI Studio, enabling integration into third‑party apps and services that need low‑latency speech‑to‑speech or speech‑to‑text translation. Pricing data compiled by Cloudprice shows Google and Vertex AI charging $3.50 per million audio‑input tokens and $21 per million audio‑output tokens for the model.
In all cases, Google says it uses its SynthID watermarking technology to embed subtle signals in generated audio, aiming to make AI‑translated voice output detectable without changing how it sounds.
What’s new compared with older live translation
Gemini 3.5 Live Translate represents a clear step up from Google’s previous live‑translation offerings, both in coverage and in feel.
- Language and pairing expansion: Meet jumps from a small set of languages and only English‑centric pairs to 70+ languages and over 2,000 possible combinations, significantly broadening who can join a multilingual call without human interpreters.
- Continuous streaming vs. turn‑taking: Older systems tended to chunk speech into segments, translate, then play back, creating a staccato rhythm of speak‑pause‑listen. Gemini 3.5 Live Translate instead streams translations continuously, which CNET notes can make a bilingual conversation feel “closer to natural speech rhythms” with only minimal delay.
- Prosody preservation: Rather than using a flat, generic synthetic voice, the new model aims to preserve aspects of the original speaker’s delivery, including rhythm, intonation, and emotional inflection, in the translated audio. That can make it easier to follow a speaker’s intent and engagement, not just their literal words.
- Noise handling: The model is explicitly tuned for noisy, real‑world settings, with Google saying it can manage background chatter, overlapping voices and informal speech patterns better than prior versions.
In practical terms, that means applications like customer‑service calls, guided tours, classrooms, ride‑sharing trips, and live broadcasts could use the technology without forcing participants to wait for long segments of translated speech.
Examples: From tour groups to global meetings
Google’s own examples sketch out how Gemini 3.5 Live Translate might show up in everyday life.
- A traveler on a Spanish‑language museum tour in Mexico City can hold their phone to their ear and hear a near real‑time English translation streamed through the earpiece, staying in sync with the guide’s cadence instead of hearing delayed summaries.
- A multinational team on Google Meet can carry on a discussion where each participant speaks their own language while hearing translations in another — say, German, Japanese and Portuguese in the same call — thanks to the new language‑pair flexibility.
- A call‑center platform built on the Gemini Live API could route callers and agents who don’t share a language into a single conversation, with Live Translate running in the background to bridge the gap.
In each scenario, the aim is not just to translate words but to keep the rhythm of interaction intact, so speakers don’t have to change how they talk to accommodate the technology.
Limits, safeguards and what comes next
Despite the upbeat launch, Google acknowledges that Gemini 3.5 Live Translate has limits. Like other AI translation systems, it can still struggle with idioms, technical jargon, low‑resource languages, and context where small errors carry high stakes. That’s one reason the initial Meet rollout is business‑only and in private preview, giving Google room to refine performance before exposing the feature to millions of consumer calls.
On the safety side, the use of SynthID watermarks is meant to ensure that AI‑generated audio remains detectable as such, a growing concern in an era of deepfake voices and impersonation scams. Gemini 3.5 Live Translate also runs inside Google’s broader policy framework for Gemini models, which includes filters on hate speech and other harmful content, though how those behave in live translation scenarios will attract scrutiny.
For now, the launch of Gemini 3.5 Live Translate signals Google’s confidence that its latest AI can handle the messy, overlapping, imperfect way people actually talk, and that there is a large market for tools that make cross‑language communication feel less like a demo and more like a normal conversation.
