Gemini 3.5 Transcribe 正式发布(GA):推出了两款基于 Gemini 音频理解的专用语音转文本模型:Gemini 3.5 Transcribe(gemini-3.5-transcribe):高精度、低延迟的非流式语音转文本,支持基于话语的语种检测,覆盖 85 种以上语言,提供说话人分离、词级时间戳,以及自定义词汇偏置(最多 1,000 个术语)。Gemini 3.5 Transcribe Live(gemini-3.5-transcribe-live):通过 Live API 在 WebSockets 上实现低延迟、双向流式语音转文本,支持临时和最终转录事件、智能转录模式,以及多种语音活动检测(VAD)策略。要开始使用,请参阅音频转录指南、实时转录指南和 Gemini 3.5 Transcribe 模型页面。
Gemini 3.5 Transcribe generally available (GA): Released two dedicated speech-to-text models based on Gemini's audio understanding: Gemini 3.5 Transcribe (gemini-3.5-transcribe): High-accuracy, low-latency non-streaming speech-to-text with utterance-based language detection across 85+ languages, speaker diarization, word-level timestamps, and custom vocabulary biasing (up to 1,000 terms). Gemini 3.5 Transcribe Live (gemini-3.5-transcribe-live): Low-latency, bidirectional streaming speech-to-text over WebSockets using the Live API, supporting interim and finalized transcription events, Smart transcription mode, and multiple Voice Activity Detection (VAD) strategies. To get started, see the Audio transcription guide, the Live transcription guide, and the Gemini 3.5 Transcribe model page.
本文内容采集自官方网站,排版和翻译可能与原页面存在差异。
阅读官方全文