Voice AI Research

Week 2026-W38
16announcements

Alibaba / Qwen@alibaba_cloud 3 in corpus · first 2026-07-21

2026-09-20 · 2026-W38 · model · model · Qwen3.8-LiveTranslate (qwen3.8-livetranslate-flash-realtime)

Alibaba ships Qwen3.8-LiveTranslate realtime interpretation

Alibaba Cloud and the Qwen team released Qwen3.8-LiveTranslate, a WebSocket realtime interpretation model that understands 60 languages and speaks 29, with average lagging cut from 2.8 s to 2.3 s versus the prior generation.

Puts a dated 2.3 s LAAL claim and speaker-attributed cloned speech on a 60/29-language realtime API — a direct peer for Gemini Live Translate-style interpretation workloads.

2026-09-20 · 2026-W38 · model · model · Qwen-Audio 3.1 Realtime Plus

Alibaba ships Qwen-Audio 3.1 Realtime Plus for duplex voice agents

Alibaba Cloud Model Studio documents qwen-audio-3.1-realtime-plus as the recommended end-to-end speech-to-speech model for voice assistants and customer-service conversations, available in Singapore and China (Beijing).

Makes 3.1 the documented default for Bailian realtime voice agents while leaving Flash as the cost-sensitive option; migration is a model ID and voice swap on the existing 3.0 protocol.

xAI@xai 2 in corpus · first 2026-07-29

2026-09-18 · 2026-W38 · model · model · Grok Voice Transcribe 2.0

xAI releases Grok Voice Transcribe 2.0 speech-to-text

xAI shipped Grok Voice Transcribe 2.0, a batch and streaming STT model it says is twice as accurate as Transcribe 1.0 at the same $0.10/hr batch and $0.20/hr streaming price.

Puts a hyperscaler STT refresh into the same price band as Transcribe 1.0, with a named Loom deployment and a timed 1.0 sunset — peers will pressure-test the vendor benches.

Speechmatics@Speechmatics 1 in corpus · first 2026-09-17

2026-09-17 · 2026-W38 · model · model · Agent STT (Linden 1)

Speechmatics launches Agent STT powered by Linden 1

Speechmatics released Agent STT, a speech-to-text API for production voice agents, powered by its new Linden 1 model with sub-350 ms finalisation, 55+ languages, live diarisation, and conversational events over `/v2/agent`.

A dedicated agent STT surface from a long-running ASR vendor, with Pipecat/LiveKit plugins on day one and a public latency–accuracy claim that peers will pressure-test.

Hume AI@hume_ai 2 in corpus · first 2026-09-10

2026-09-17 · 2026-W38 · model · other · Voice Controllability Leaderboard

Hume publishes Voice Controllability Leaderboard for TTS direction-following

Hume published a Voice Controllability Leaderboard scoring 17 TTS models on voice design, instruction-following, inline tags, and role fit with blind human raters across 13 languages.

Separates identity cloning from controllability. Same model can win motivational casting and lose meditation; Fish's voice-design swings are the clearest example in the board.

A result from the Voice Controllability leaderboard we published today: Fish's voice design model won nearly every head to head for motivational speaker voices and lost nearly every one for meditation voices. Same model, same prompt format, and the job you're casting for flipped

https://x.com/hume_ai/status/2100584836864024785

Deepgram@DeepgramAI 3 in corpus · first 2026-08-12

2026-09-17 · 2026-W38 · model · model · Nova-3 Pharma

Deepgram launches Nova-3 Pharma speech-to-text

Deepgram released Nova-3 Pharma, a speech-to-text model trained for drug names and pharmaceutical vocabulary, available for batch and streaming on hosted and self-hosted deployments via model=nova-3-pharma.

Splits pharma entity accuracy from general medical STT — a measurable KRR claim competitors in clinical voice will be asked to match.

Introducing Nova-3 Pharma, the first speech-to-text model purpose-built for the pharmaceutical industry.

https://x.com/DeepgramAI/status/2100614929573093694
2026-09-15 · 2026-W38 · model · api · India endpoint (api.in.deepgram.com)

Deepgram makes India voice endpoint generally available

Deepgram opened a generally available India regional endpoint at api.in.deepgram.com in AWS ap-south-2 (Hyderabad), with in-country storage and inference for STT, TTS, Voice Agent, and text intelligence at Global/EU/Australia pricing.

Adds a managed onshore option after EU and Australia — relevant for Indian BFSI voice workloads that previously faced self-host or cross-border trade-offs.

The Deepgram India endpoint is now generally available.

https://x.com/DeepgramAI/status/2099894045023641666

Cartesia@cartesia 2 in corpus · first 2026-08-27

2026-09-17 · 2026-W38 · model · launch · Multilingual Voices

Cartesia ships Multilingual Voices for Sonic TTS

Cartesia launched Multilingual Voices so one cloned or library voice can speak multiple languages and accents natively while keeping the same voice ID, demonstrated with SF bakery Baklavastory.

Moves brand-voice localisation from one-voice-per-locale cloning to accent adds on a single ID — useful for agents that already run Sonic-3.6.

Assort Health@assort_health 2 in corpus · first 2026-08-12

2026-09-17 · 2026-W38 · platform · launch · NextGen Enterprise EHR direct integration (expanded) · vendor NextGen Healthcare

Assort expands direct NextGen Enterprise EHR actions for specialty practices

Assort Health announced expanded Platinum API Tier access so its AI agents can read appointment slots and create, cancel, or reschedule appointments directly in NextGen Enterprise EHR, plus referrals, notes, and payment workflows.

Moves vertical voice agents past intake into write-back scheduling and referrals on a major ambulatory EHR — the bottleneck after the phone is answered.

Tavus@tavus 2 in corpus · first 2026-09-10

2026-09-16 · 2026-W38 · platform · launch · Memories (for PALs)

Tavus ships Memories for long-term PAL relationships

Memories gives each PAL–person pair a Profile and Timeline that update after calls, so returning conversations continue with prior goals, events, and preferences without re-explaining.

After Phoenix-4.5, CVI gets relationship continuity: post-call consolidation plus inspectable/editable stores, with Pinned Memories for handoff context.

SoundHound AI@SoundHound 2 in corpus · first 2026-07-23

2026-09-16 · 2026-W38 · platform · launch · Human Assisted Resolution (HAR)

SoundHound ships Human Assisted Resolution inside OASYS

SoundHound launched Human Assisted Resolution (HAR), an opt-in OASYS feature that lets an AI agent ask a colleague a specific question in real time and finish the conversation without a full handoff.

Separates brief human judgment from escalation — a concrete HITL pattern for branded voice/chat agents that otherwise double-pay for transfer.

Today we're launching Human Assisted Resolution (HAR).

https://x.com/SoundHound/status/2100253157674582423

ElevenLabs@ElevenLabs 4 in corpus · first 2026-08-06

2026-09-16 · 2026-W38 · platform · launch · Reception (Reception.ai by ElevenAgents)

ElevenLabs launches Reception, an AI receptionist for small businesses

Reception by ElevenAgents answers inbound calls, books appointments from a website scan, and goes live with a phone number in minutes; plans start at $29/month with 70+ languages.

Packages ElevenAgents as an SMB receptionist SKU (calendar + after-hours) against dedicated receptionist vendors, not only API buyers; HIPAA stays on the broader platform.

AssemblyAI@AssemblyAI 2 in corpus · first 2026-08-19

2026-09-16 · 2026-W38 · model · api · Dictation API

AssemblyAI launches Dictation API for finished-text STT

AssemblyAI released the Dictation API, a sync endpoint that returns a verbatim transcript and an LLM-cleaned finished text from one short clip on Universal-3.5 Pro, priced at $0.62 per audio hour all-in.

Bundles transcript-plus-rewrite into one billable call for push-to-talk UIs — a different request shape from streaming agent STT already in the corpus.

Dear developers, Forms are broken.

https://x.com/AssemblyAI/status/2100710329282101434

Blurt: push-to-talk dictation on the AssemblyAI Dictation API.

https://x.com/AssemblyAI/status/2100971811584516172

Hello Patient@HelloPatient 1 in corpus · first 2026-09-15

2026-09-15 · 2026-W38 · platform · other · Converse Health acquisition (back-office agents)

Hello Patient acquires Converse Health for referral and back-office AI agents

Hello Patient bought Converse Health so its patient-conversation agents can also do referral and fax intake, chart-driven follow-up, authorisation paperwork, and records work inside outpatient EHRs.

Moves a voice-first patient-access vendor into EHR document workflows, so outbound referral calls sit on the same agent stack as inbound phones rather than a bolted-on fax tool.

Google@Google 3 in corpus · first 2026-08-26

2026-09-15 · 2026-W38 · model · model · Gemini 3.8 Live and 3.8 Live Extended Thinking

Google ships Gemini 3.8 Live and Extended Thinking on the Live API

On 15 Sep Google released Gemini 3.8 Live and 3.8 Live Extended Thinking for native speech-to-speech on the Gemini Live API and AI Studio, with async tool calls, visual grounding, and mid-conversation switching across 97 languages.

A hyperscaler S2S drop with async tools + visual grounding, same day LiveKit/Pipecat wire it — the competitive reference for GPT-Live-1-class agent stacks, with AA / τ-Voice numbers that need independent checking.

DeepL@DeepLcom 1 in corpus · first 2026-09-15

2026-09-15 · 2026-W38 · model · model · DeepL Voice (voice preservation)

DeepL Voice adds real-time voice preservation across languages

DeepL shipped new Voice models that preserve each speaker's voice, tone, rhythm, and expression during live multilingual translation, initially across 14 languages, plus a desktop app for Zoom, Teams, and Meet.

Moves meeting translation from shared synthetic voices to per-speaker identity — a different product axis from agent STT/TTS already tracked here.

Introducing voice preservation in DeepL Voice.

https://x.com/DeepLcom/status/2099831527404154912