[short pause], [whisper], [yawn], [softly], or Say excitedly: "..." to control expression.
Gemini TTS Studio converts text to speech using Google's Gemini API, with single-speaker and multi-speaker conversation modes and a choice of voice profiles. It requires your own Gemini API key.
Single-speaker mode reads text in one voice. Multi-speaker mode assigns different voices to named speakers, which is what makes dialogue, interviews and narrated conversations sound like more than one person reading a script.
Because the synthesis happens through Google's API rather than your browser's built-in voices, quality is considerably higher and consistent across devices. The cost is that it needs network access and an API key.
v1.0