Gemini 3.8 Flash TTS
Gemini 3.8 Flash TTS converts scripts into directed performances — voice design, two-speaker dialogue and 130 languages in one API call.
Support
Pro AI Tools
Explore elite tools including the Suno AI Music Generator.
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1
FLUX 3 Video Generator

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

How Gemini 3.8 Flash TTS Redefines Voice Production
Launched September 23, 2026, Google's Gemini TTS line pairs a creative flagship with a high-throughput engine built for bulk audio.
- A Single Launch, Two WorkloadsGemini 3.8 Flash TTS acts as the creative flagship, while Flash-Lite TTS serves as the lower-cost engine for high-volume speech.
- Direct the Delivery Instead of Choosing a PresetPer-turn style notes, structured speech metadata and inline vocal events steer tone, pacing, emotion and accent.
- Design New Voices or Replicate with ConsentDescribe any voice in plain language, or reproduce a real speaker from a reference clip plus a matching consent recording.
How to Prompt Gemini 3.8 Flash TTS the Right Way
Four habits that keep your transcript clean and let the performance metadata do the work.
Inside Gemini 3.8 Flash TTS: Core Capabilities
From hands-on performance direction to 130-language coverage, here is what the flagship Gemini TTS model really delivers.
Direction-Level Expressive Control
Per-turn style plus inline laughs, sighs, coughs, breaths and pauses — closer to coaching a voice actor than picking a preset.
Voice Design in Plain Language
Prompts can define age range, personality, accent, vocal texture and role, drawing on 2,000+ production voices from the Voices endpoint.
Voice Replication Behind a Consent Gate
A clean reference recording plus a matching consent clip from the same adult speaker, secured by SynthID watermarking and C2PA credentials.
Built for Two-Speaker Dialogue
Scripts carry podcasts, educational exchanges, product demos and game scenes with no manual line stitching required.
Stable Long-Form Narration, No Voice Drift
Google documents consistent voice identity, timbre, volume and room tone across multi-minute narration and extended dialogue.
130 Languages with Regional Accents
Flash TTS spans 130 languages against Flash-Lite's 101, covering regional accents, minority dialects and IPA pronunciation overrides.
Gemini 3.8 Flash TTS: Frequently Asked Questions
Quick answers on Gemini 3.8 Flash TTS pricing, model selection, benchmark scores and built-in safeguards.
What does Gemini 3.8 Flash TTS charge per minute of audio?
At launch rates of $0.50 per million input tokens and $9 per million output tokens, roughly 1.35 cents per audio minute.
Which model should I pick — Flash TTS or Flash-Lite TTS?
Choose Flash TTS when acting nuance and long-form audio matter; choose Flash-Lite TTS for bulk output and low latency.
How does it score against rival voice models?
Google reports 71.4 on Hume's Voice Design Benchmark, while Voice Arena ranks it second with 1,260 Elo.
What changed from Gemini 3.1 Flash TTS Preview?
Flash-Lite TTS supersedes the 3.1 preview and lowers audio output pricing from $20 to $6 per million tokens.
Can I clone a voice, and what safeguards apply?
Replication requires a reference clip alongside a matching consent recording from the same adult speaker.
Why does the model read my stage directions out loud?
Your input is treated as a verbatim transcript, so move any sustained directions into the speech metadata.
Run Gemini 3.8 Flash TTS on Your Own Scripts
Try both tiers in the Gemini API or Google AI Studio — moving between Gemini 3.8 Flash TTS and Flash-Lite TTS takes nothing more than swapping a model identifier. Weigh batch against priority inference before you lock in a production budget.
