Google introduces Gemini 3.8 Flash TTS voice models
Google has released Flash TTS and Flash-Lite TTS for creative speech generation and high-volume audio workflows, with 2,000+ preset voices, text-directed control, and safety measures.
Google has introduced two dedicated speech-generation models: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. According to AI News, the larger model targets interactive entertainment, game development, and long-form narration, while Flash-Lite is aimed at automated dubbing, conversational agents, and high-throughput translation.
The models provide access to more than 2,000 preset voice profiles across over 100 languages, including regional varieties such as Quebec French, Scots English, and Mexican Spanish. Google has also announced a forthcoming voice-remixing feature intended to adjust timbre, pitch, pace, and accent through text instructions.
In benchmark results cited from Hume AI, Flash TTS scored 71.4 overall and 60.8 for accent modelling. The report says the two new systems ranked first and second in Hume's Overall Quality Index. Those figures describe the cited evaluations and should not be treated as proof of superiority for every use case.
Production features include two-speaker dialogue, stable voice character over multi-hour files, and script tags for non-verbal sounds such as laughs, sighs, and gasps. For voice cloning, Google requires a 30-second reference recording plus spoken consent from the voice owner. Generated audio includes imperceptible SynthID watermarks and C2PA provenance metadata.
Developers can access both models through Google AI Studio and the Gemini API. Flash TTS is also available in Gemini Notebook, while Flash-Lite TTS is being integrated into Google Vids; administrative API access for Gemini Enterprise customers is planned for a later rollout.

Source: AI News