Gemini 3.8 Flash TTS

Gemini 3.8 Flash TTS converts scripts into directed performances — voice design, two-speaker dialogue and 130 languages in one API call.

Gemini 3.8 Flash TTS
Craft expressive speech with Gemini 3.8 Flash TTS, or scale affordably on Flash-Lite TTS
AI Video Prompt Generator

Support

Pro AI Tools

Explore elite tools including the Suno AI Music Generator.

placeholder hero

How Gemini 3.8 Flash TTS Redefines Voice Production

Launched September 23, 2026, Google's Gemini TTS line pairs a creative flagship with a high-throughput engine built for bulk audio.

  • A Single Launch, Two Workloads
    Gemini 3.8 Flash TTS acts as the creative flagship, while Flash-Lite TTS serves as the lower-cost engine for high-volume speech.
  • Direct the Delivery Instead of Choosing a Preset
    Per-turn style notes, structured speech metadata and inline vocal events steer tone, pacing, emotion and accent.
  • Design New Voices or Replicate with Consent
    Describe any voice in plain language, or reproduce a real speaker from a reference clip plus a matching consent recording.

How to Prompt Gemini 3.8 Flash TTS the Right Way

Four habits that keep your transcript clean and let the performance metadata do the work.

Inside Gemini 3.8 Flash TTS: Core Capabilities

From hands-on performance direction to 130-language coverage, here is what the flagship Gemini TTS model really delivers.

Direction-Level Expressive Control

Per-turn style plus inline laughs, sighs, coughs, breaths and pauses — closer to coaching a voice actor than picking a preset.

Voice Design in Plain Language

Prompts can define age range, personality, accent, vocal texture and role, drawing on 2,000+ production voices from the Voices endpoint.

Voice Replication Behind a Consent Gate

A clean reference recording plus a matching consent clip from the same adult speaker, secured by SynthID watermarking and C2PA credentials.

Built for Two-Speaker Dialogue

Scripts carry podcasts, educational exchanges, product demos and game scenes with no manual line stitching required.

Stable Long-Form Narration, No Voice Drift

Google documents consistent voice identity, timbre, volume and room tone across multi-minute narration and extended dialogue.

130 Languages with Regional Accents

Flash TTS spans 130 languages against Flash-Lite's 101, covering regional accents, minority dialects and IPA pronunciation overrides.

FAQ

Gemini 3.8 Flash TTS: Frequently Asked Questions

Quick answers on Gemini 3.8 Flash TTS pricing, model selection, benchmark scores and built-in safeguards.

1

What does Gemini 3.8 Flash TTS charge per minute of audio?

At launch rates of $0.50 per million input tokens and $9 per million output tokens, roughly 1.35 cents per audio minute.

2

Which model should I pick — Flash TTS or Flash-Lite TTS?

Choose Flash TTS when acting nuance and long-form audio matter; choose Flash-Lite TTS for bulk output and low latency.

3

How does it score against rival voice models?

Google reports 71.4 on Hume's Voice Design Benchmark, while Voice Arena ranks it second with 1,260 Elo.

4

What changed from Gemini 3.1 Flash TTS Preview?

Flash-Lite TTS supersedes the 3.1 preview and lowers audio output pricing from $20 to $6 per million tokens.

5

Can I clone a voice, and what safeguards apply?

Replication requires a reference clip alongside a matching consent recording from the same adult speaker.

6

Why does the model read my stage directions out loud?

Your input is treated as a verbatim transcript, so move any sustained directions into the speech metadata.

Run Gemini 3.8 Flash TTS on Your Own Scripts

Try both tiers in the Gemini API or Google AI Studio — moving between Gemini 3.8 Flash TTS and Flash-Lite TTS takes nothing more than swapping a model identifier. Weigh batch against priority inference before you lock in a production budget.