2026年8月26日
|
我们最新的语音转文字模型,专为精准、智能的实时转写而设计。
Diego Melendo Casado
Gemini Audio 工程高级总监
Luke Leonhard
Gemini Audio 幕僚长,代表 Gemini Audio 团队

收听文章
[[duration]] 分钟
此内容由 Google AI 生成。生成式 AI 处于实验阶段。
今天,我们推出 Gemini 3.5 Transcribe,这是我们迄今为止最精准的语音转文字模型,专为智能语音交互而设计。与在背景噪音、复杂术语和流畅度清理方面表现不佳的传统语音识别模型不同,Gemini 3.5 Transcribe 可将原始音频直接转换为准确、精炼、格式化的文本。
在我们如 Gemini 应用及 Android 等产品中,我们已经看到消费者通过新的语音功能受益于该转写模型,例如 Android 上的 Rambler 以及 macOS 上的 Gemini 应用。现在,开发者可以通过 Google AI Studio 中的 Gemini API 和 Gemini Enterprise Agent Platform 使用 Gemini 3.5 Transcribe 构建类似功能。
我们构建的 3.5 Transcribe 可无缝集成到您的开发者工作流程中,无论您是在构建语音代理、实时字幕工具还是通话后分析管道。该模型可通过两个独立的 API 使用:
- 实时流式传输: 通过 Live API 使用
gemini-3.5-transcribe-live提供亚秒级延迟的连续双向流式传输,适用于交互式语音应用。 - 预录制音频处理: 通过 Interactions API 使用
gemini-3.5-transcribe转写录制的音频、会议、通话记录等,并支持说话人归属和词级时间戳。
获得更精准、更智能的转写
Gemini 3.5 Transcribe 旨在捕捉您自然的说话风格,以更好地理解您的意图并识别自定义词汇,从而让您通过语音执行任务。
- 智能转写: 无缝处理自我纠正(例如“我们周二见面——不,周三”),移除填充词(“嗯”和“啊”),并自动格式化您的文本。
- 函数调用: 该模型可通过函数调用将复杂任务(如图像生成和文件分析)委派给其他 Gemini 模型。目前已在 Gemini macOS 应用 中可用。
- 更精准的转写: 根据 Artificial Analysis 的测量,流式用例的平均词错误率(WER)为 4.0%,非流式用例为 2.6%。在嘈杂的真实环境中表现出色,能准确捕捉邮政编码和订单 ID 等字母数字实体。
- 自定义词汇: 通过无缝调整转写结果以匹配您提供的自定义词汇,识别专业术语和独特拼写。
- 全球语言支持: 自动检测并转写超过 85 种语言,无缝处理地区口音和多样方言。
- 多说话人识别: 在预录制音频中准确归属语音,并为最多三位说话人提供时间戳(支持 3 位以上说话人处于实验阶段)。
Gemini 3.5 Transcribe 的性能相比我们之前的转写模型 Chirp 3 是一次重大进步,提供了新功能、更优的词错误率和显著改善的延迟。根据 Artificial Analysis 的测量,例如,最终转写时间改善了 70%。在一组主要语言和地区的 FLEURS 基准测试中,该模型提供了精准的多语言性能,优于 Chirp 3,在流式模式下实现了 5.50% 的 WER,在非流式用例中实现了 5.04% 的 WER。


体验智能转写与高级听写
除了 Google AI Studio 中的 Gemini API 和 Gemini Enterprise Agent Platform 之外,3.5 Transcribe 超越了标准语音转文字的功能,使跨 Google 的工作更加自然直观。通过将上下文感知理解直接带入 Gboard、Antigravity、Gemini 应用和 Chrome 等日常使用场景,它能够轻松捕捉细微差别、意图和内联编辑。
- 在 Android 上的 Gboard 中,通过新的 Rambler 功能,3.5 Transcribe 将口述想法转换为格式良好的文本,并过滤掉填充词。您还可以使用语音进行编辑、纠正拼写错误以及更改写作风格。
- 在 Google Antigravity 中,3.5 Transcribe 在获得您的许可后,结合屏幕上下文和聊天历史,确保在文件名、代理想法和活动文档中实现精准的转写准确性。
- 在 Google AI Studio 中,您可以在 Build 模式下访问 3.5 Transcribe,随时通过语音快速编写应用。
- 在 macOS 上的 Gemini 应用 中,3.5 Transcribe 不仅将您自由自然的语音转写为干净的格式化文本,还支持语音命令,可与屏幕上下文无缝配合以驱动复杂工作流程。通过在后台调用其他 Gemini 模型来处理繁重任务,该模型让您仅凭语音即可轻松总结本地文件、跨应用重用文本或在光标处生成图像。
- 即将在 Chrome 中推出,您将能够在任何网页字段中通过语音输入——让您更自然、更轻松地通过语音口述回复、撰写帖子或提示 Gemini。
阅读早期评测
通过利用 Gemini Live API,诸如 Agora、Fishjam、LangChain、LiveKit、Pipecat、Vercel 和 Vision Agents 等开发者平台,使开发者能够轻松构建和部署高性能的语音驱动界面。这些平台在幕后管理复杂的实时媒体流基础设施,让开发者能够完全专注于打造用户体验。
Vivo、Intellitek Health 和 Lingopal 等公司也对 3.5 Transcribe 给予了积极反馈,称赞其出色的延迟表现、准确性和广泛的语言支持。






立即开始使用 3.5 Transcribe
- 面向开发者:通过 Google AI Studio 中的 Gemini API 和 Google Antigravity 提供公开预览。
- 面向企业:通过 Gemini Enterprise Agent Platform 提供公开预览,并即将在 Gemini Enterprise for Customer Experience 中推出。
- 面向所有用户:在 macOS 版 Gemini 应用中提供英文版本,在 Android 版 Rambler 中面向部分国家和地区提供,并即将在 Chrome 中推出。
在收件箱中获取 Google 最新资讯
订阅我们的新闻通讯,获取产品更新、活动信息、特别优惠等更多内容。
您的信息将按照 Google 隐私政策 使用。您可以随时选择退出。
Aug 26, 2026
|
Our latest speech-to-text model designed for precise and intelligent real-time transcription.
Diego Melendo Casado
Senior Director, Engineering, Gemini Audio
Luke Leonhard
Chief of Staff, Gemini Audio, on behalf of Gemini Audio Team

Listen to article
[[duration]] minutes
This content is generated by Google AI. Generative AI is experimental
Today, we’re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions. Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.
Across our products like the Gemini app and on Android, we’ve seen consumers already benefiting from this transcription model with new voice capabilities like Rambler on Android and in the Gemini app on macOS. Now, developers can build similar capabilities with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.
We've built 3.5 Transcribe to plug seamlessly into your developer workflows, whether you’re building voice agents, real-time captioning tools, or post-call analytics pipelines. The model is available across two separate APIs:
- Real-time streaming: Delivers continuous, bidirectional streaming with sub-second latency for interactive voice apps via the Live API using
gemini-3.5-transcribe-live****. - Pre-recorded audio processing: Transcribes recorded audio, meetings, call logs, and more with speaker attribution and word-level timestamps via the Interactions API using
gemini-3.5-transcribe****.
Get more precise and intelligent transcription
Gemini 3.5 Transcribe is designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary, so you can execute tasks with your voice.
- Smart transcription: Seamlessly handles self-corrections (like "let’s meet Tuesday—no, Wednesday"), removes filler words (“ums” and ‘“ahs"), auto-formats your text.
- Function calling: The model can delegate complex tasks (such as image generation and file analysis) to other Gemini models via function calls. Currently available in the Gemini macOS app.
- More precise transcription: As measured by Artificial Analysis, achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases. It shows strong performance across noisy, real-world environments, accurately capturing alphanumeric entities like postal codes and order IDs.
- Custom vocabulary: Recognizes specialized jargon and unique spellings by seamlessly adapting transcriptions to your provided custom vocabulary.
- Global language support: Automatically detects and transcribes over 85 languages, seamlessly handling regional accents and diverse dialects.
- Multi-speaker identification: Accurately attributes speech in pre-recorded audio with timestamps for up to three speakers (support for 3+ speakers is experimental).
Gemini 3.5 Transcribe’s performance represents a major advancement from our previous transcription model, Chirp 3, offering new capabilities, improved word error rates, and significantly better latency. As measured by Artificial Analysis, time to final transcription, for example, improves by 70%. On the FLEURS benchmark across a set of top languages and locales, the model delivers precise multilingual performance, improving over Chirp 3, and achieving a 5.50% WER in streaming mode and 5.04% WER in non-streaming use-cases.


Experience smart transcription and advanced dictation
In addition to the Gemini API in the Google AI Studio and Gemini Enterprise Agent Platform, 3.5 Transcribe goes further than standard speech-to-text to make working across Google feel more natural and intuitive. By bringing context-aware understanding directly into everyday surfaces like Gboard, Antigravity, the Gemini app, and Chrome, it captures nuances, intent, and inline edits with ease.
- On Gboard on Android, through the new Rambler feature, 3.5 Transcribe transforms spoken thoughts into well-formatted text, filtering out filler words. You can also use your voice to make edits, correct misspellings, and change the writing style.
- On Google Antigravity, 3.5 Transcribe pairs screen context and chat history, with your permission, to ensure pinpoint transcription accuracy across file names, agent thoughts, and active documents.
- In Google AI Studio, you can access 3.5 Transcribe in Build mode to vibe code apps with your voice on the fly.
- In the Gemini app on macOS, 3.5 Transcribe not only transcribes your free natural speech into clean formatted text, but also enables voice commands that can pair seamlessly with screen context to power complex workflows. By calling on other Gemini models in the background to handle the heavy lifting, the model makes it effortless to summarize local files, repurpose text across apps, or generate images right at your cursor—using just your voice.
- Coming soon to Chrome, you’ll be able to talk to type in any web field — making it effortless to dictate replies, draft posts, or prompt Gemini in Chrome more naturally and easily with your voice.
Read the early reviews
By leveraging the Gemini Live API, developer platforms such as Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents enable developers to build and deploy high-performance voice-driven interfaces with ease. These platforms manage complex real-time media streaming infrastructure behind the scenes, allowing developers to focus entirely on crafting the user experience.
Companies like Vivo, Intellitek Health, and Lingopal have also shared positive feedback on 3.5 Transcribe, highlighting its impressive latency, accuracy, and expansive language support.






Start using 3.5 Transcribe today
- For developers: In public preview in the Gemini API via Google AI Studio and Google Antigravity.
- For enterprises: In public preview via Gemini Enterprise Agent Platform and coming soon to Gemini Enterprise for Customer Experience.
- For everyone: In Gemini app on macOS in English, Rambler on Android in select countries and languages, and coming soon to Chrome.
Get the latest news from Google in your inbox
Sign up for our newsletters with product updates, event information, special offers, and more.
Your information will be used in accordance with Google's privacy policy. You may opt out at any time.
本文内容采集自官方网站,排版和翻译可能与原页面存在差异。
阅读官方全文