Google has officially announced the launch of Gemini 3.5 Transcribe, its latest speech-to-text model designed to deliver precise and intelligent real-time transcription. Unlike traditional speech recognition models, this model can effectively handle background noise, specialized terminology, and variations in speech fluency, converting raw audio directly into accurate and formatted text.
Core Features of Gemini 3.5 Transcribe
According to Google, Gemini 3.5 Transcribe can capture the user's natural language style, better understand intent, and recognize custom vocabulary, thereby enabling the ability to execute tasks through speech. The model offers several advanced features, including:
“Smart transcription: Seamlessly handles self-corrections, removes filler words, and automatically formats text.”
Google
In addition, Gemini 3.5 Transcribe supports multi-language recognition, automatically detecting and transcribing over 85 languages, and can accurately identify up to three speakers in audio, providing timestamps for pre-recorded audio.
Potential Applications for Developers and Enterprises
The model can be accessed through Google AI Studio's Gemini API and the Gemini Enterprise Agent Platform, allowing developers to seamlessly integrate it into their workflows. Google also noted that the model offers two different API services for real-time streaming and pre-recorded audio processing, namely:
“Real-time streaming: Provides continuous two-way streaming with latency below one second, suitable for interactive voice applications.”
Google
“Pre-recorded audio processing: Transcribes recorded audio, meetings, call logs, etc., and provides speaker attribution and verbatim timestamps.”
Google
This enables developers to easily build voice agents, real-time captioning tools, or post-call analysis pipelines, enhancing the user experience.
Performance Improvements of Gemini 3.5 Transcribe
Gemini 3.5 Transcribe shows significant performance improvements over its predecessor, Chirp 3. According to human analysis, its average Word Error Rate (WER) is 4.0% in streaming mode and 2.6% in non-streaming mode. The model performs excellently in noisy real-world environments, accurately capturing alphanumeric entities such as postal codes and order IDs.
“In the FLEURS benchmark test, the model outperformed Chirp 3 across multiple languages and regions, with a WER of 5.50% in streaming mode and 5.04% in non-streaming mode.”
Google
These improvements make Gemini 3.5 Transcribe a powerful tool for developers and enterprises in the field of speech transcription.
Source: Google Official Announcement

