Google Launches Gemini 3.5 Transcribe Speech-to-Text Model, Debuting in Gboard Rambler and Chrome

·by Henderson
Google Launches Gemini 3.5 Transcribe Speech-to-Text Model, Debuting in Gboard Rambler and Chrome

Google has officially launched Gemini 3.5 Transcribe, touted as its "most accurate speech-to-text model to date," and has already integrated it into several of its products. Unlike traditional speech recognition systems that are often disrupted by background noise, industry jargon, and filler words, the new model can directly convert raw audio into precise, organized, and properly formatted text. According to the company, this model "aims to preserve the natural tone of speech, thereby more accurately understanding user intent and identifying custom vocabulary."

Gemini 3.5 Transcribe can handle real-time self-corrections, such as "See you on Tuesday—no, Wednesday," and automatically removes filler words like "um" and "uh," while also supporting automatic formatting and natural speech editing features. In terms of accuracy, according to tests by Artificial Analysis, the average Word Error Rate (WER) in streaming mode is 4.0%, dropping to 2.6% in non-streaming mode. The model performs stably in noisy real-world environments, accurately recognizing strings of letters and numbers such as postal codes and order numbers.

Gemini 3.5 Transcribe Supports Custom Vocabulary and Multilingual Recognition

In terms of custom vocabulary, the model can recognize specific industry terms and unique spellings, and automatically adjust transcription results based on the vocabulary list provided by the user. In terms of language support, the system can automatically detect and transcribe over 85 languages, covering different regional accents and dialects. The multi-speaker recognition feature can timestamp and attribute statements to up to three speakers in pre-recorded audio, while scenarios with more than three speakers are still experimental.

In terms of performance, Google claims that Gemini 3.5 Transcribe has made "significant progress" in all capabilities, word error rates, and latency compared to the Chirp 3 transcription model launched in 2025. According to data from Artificial Analysis, the time required to complete the final transcription is 70% shorter than that of Chirp 3. In the FLEURS benchmark test, the streaming mode word error rate for multiple mainstream languages and regions is 5.50%, and the non-streaming mode is 5.04%, with multilingual performance also surpassing that of the previous model.

Voice Agent and Cross-Platform Integration

Another core goal of the new model is to enable users to "perform tasks using voice." Through the function call feature, Gemini 3.5 Transcribe can "hand over complex tasks such as image generation and file analysis to other Gemini models," with related experiences visible in the Speak to Window feature of the macOS version of the Gemini application. In addition to the macOS version of the Gemini application and the Android platform's Gboard Rambler, Gemini 3.5 Transcribe has also been built into the recording microphone of the Google Antigravity prompt box, which can, with user authorization, combine screen content and conversation history,

ensuring accurate transcription in scenarios such as file names, intelligent agent thoughts, and the current document.

The model is expected to be available on the Chrome browser later, allowing users to use voice input in any web field, whether drafting replies, drafting posts, or issuing voice commands to Gemini in Chrome, making operations more natural and smooth. Gemini 3.5 Transcribe is currently available to two types of users: developers can try the public preview version through the Gemini API (located in Google AI Studio and Google Antigravity); enterprise users can obtain the public preview version on the Gemini Enterprise Agent Platform, and it will soon be available on Gemini

Enterprise for Customer Experience.

H
About the author
Henderson