IT & AI News 2 min read

Google Launches Gemini 3.5 Transcribe for Enhanced Speech-to-Text

Victoria Sterling

Key Takeaways

  • Gemini 3.5 Transcribe enhances voice input by removing filler words.
  • It boasts a 70% increase in speed and a 5.5% error rate.
  • Available on Pixel 11 and soon to expand to other devices.
  • Developers can access the model through Gemini API and AI Studio.

Introduction of Gemini 3.5 Transcribe

Google has unveiled Gemini 3.5 Transcribe, a new AI model aimed at improving voice input by eliminating unnecessary filler words like “ums” and “uhs.” This model, which enhances the transcription process, is currently integrated into the Gboard’s “Rambler” feature on Pixel 11 devices and is set to roll out across the broader Google ecosystem.

Performance Enhancements

According to Google, the new model significantly outperforms its predecessor, Chirp 3, in both speed and accuracy. Gemini 3.5 Transcribe is reported to be 70% faster in converting voice to text, with a live-speech error rate reduced to 5.5%. While this is an improvement over Chirp 3’s 7.32% error rate, the ability to minimize typing errors during voice input remains a key advantage.

Advanced Features

This model not only captures spoken words more effectively but also interprets the intended meaning behind them. It can seamlessly edit out verbal pauses and make real-time corrections, utilizing a custom vocabulary for specialized terms. Gemini 3.5 Transcribe supports 85 languages and can handle audio from up to three speakers in pre-recorded formats.

Availability and Future Plans

Currently, Gemini 3.5 Transcribe is available in select applications, including Rambler on Gboard, but is limited to Pixel 11 users for now. Google plans to extend this feature to more devices later this year, including macOS users of the Gemini app.

Developer Access

Developers will also benefit from Gemini 3.5 Transcribe, as it is now integrated into Antigravity and AI Studio, allowing for enhanced voice-to-text capabilities in app development. The model will also be accessible via the Gemini API.

Upcoming Features

For those not using compatible devices or applications, Gemini 3.5 Transcribe will soon be available in the Chrome browser, enabling users to input text into any web field by voice. This feature will facilitate various tasks, from composing emails to interacting with AI chatbots.