Measuring benchmark optimization in speech recognition
(翻译)衡量语音识别中的基准优化
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
(翻译)衡量语音识别中的基准优化
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
(翻译)使用 Amazon Connect 构建餐厅电话 AI 接待员
Learn how to build a voice ordering system for restaurants that answers a phone call and takes an order end to end, with no app, no website, and no sign-in. It uses Amazon Connect for telephony, Amazon Connect Agentic Voice for real-time speech, an Amazon Connect AI agent for reasoning, and Amazon B
(翻译)构建低延迟多语言语音智能体:借助 NVIDIA Magpie TTS 实现开放权重与完整部署控制
A Blog post by NVIDIA on Hugging Face

(翻译)Gemini 3.5 Live Translate:流畅自然的语音翻译
Gemini 3.5 Live Translate brings near real-time, natural speech translation to Google AI Studio, Google Translate and Google Meet.

(翻译)使用 Granite Speech 5.0 Turbo CTC 实现极其快速且准确的转录
A Blog post by IBM Granite on Hugging Face


9月2日,腾讯 WorkBuddy 生态发布会上宣布,与专业音视频硬件品牌HOLLYLAND猛玛达成深度生态合作,推出猛玛LARK A2无线麦克风WorkBuddy联名款。作为首款原生联动WorkBuddy智能体的AI语音入口硬件,该产品在专业收音基础上,通过机身按键联动WorkBuddy语音输入、内容删除与指令发送,为电脑端AI语音交互提供新的操作方式。

(翻译)DuplexSpeechBench-IFEval:评估全双工语音智能体中的隐式指令遵循
Abstract page for arXiv paper 2609.03423: DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents
最近,OpenAI 发布了一篇关于 GPT-Live 的工程报告。该报告详细描述了他们如何设计这个系统,将对延迟敏感的媒体处理与更广泛的应用工作分离,同时保持语音交互的连续性。

(翻译)RESCUE-BENCH:面向关系感知的多方情感支持对话系统
Abstract page for arXiv paper 2609.09657: RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems
(翻译)LLM在草稿-验证-修订流水线中能否解决指示语歧义?
Abstract page for arXiv paper 2609.12162: Can LLMs in Draft-Verify-Revise Pipelines Resolve Deictic Ambiguity?
(翻译)GLARE:面向社会动态预测的对抗奖励估计生成学习
Abstract page for arXiv paper 2609.12165: GLARE: Generative Learning via Adversarial Reward Estimation For Social Dynamics Forecasting
#欢迎关注爱范儿官方微信公众号:爱范儿(微信号:ifanr),更多精彩内容第一时间为您奉上。

9月15日,阶跃发布新一代语音大模型 StepAudio 3 系列,一次性推出 StepAudio 3 Realtime、StepAudio 3 ASR、StepAudio 3 TTS、StepAudio 3 Gen 和 StepAudio 3 Music 五款模型,覆盖实时语音交互、语音识别、真人级语音生成和音乐创作等场景。其中,多款模型在 权威 第三方 AI 模型评测机构 Artificial Analysis 榜单中位列全球第一。

(翻译)让 AI 覆盖每一种语言和每一个人
We’re moving beyond traditional text translation to build models that understand the world’s rich, living languages exactly as they are expressed.
