Mistral Releases Voxtral TTS, an Open-Source Voice Model Targeting ElevenLabs

Mistral's Voxtral TTS clones a voice from three seconds of audio, runs on a smartphone, and costs $0.016 per 1,000 characters — a direct challenge to ElevenLabs' $11 billion proprietary empire.

Mistral Releases Voxtral TTS, an Open-Source Voice Model Targeting ElevenLabs

Mistral released Voxtral TTS on Thursday — a 4-billion-parameter, open-weight text-to-speech model that runs on a smartphone, clones a voice from three seconds of audio, and costs a fraction of what ElevenLabs charges. It supports nine languages, delivers audio in under 100 milliseconds, and is available for free on Hugging Face. It is a direct attack on the business model that made ElevenLabs worth $11 billion.

To understand why this matters, you have to understand how the current voice AI market works. ElevenLabs, Deepgram, OpenAI — every major player runs a proprietary, API-first business. You rent access to their voice generation. You send your audio data to their servers. You pay per character, per request, per month. The cost scales with usage, the vendor controls the infrastructure, and you have no ownership over the core technology. For enterprises, this means voice AI is a line item on a cloud bill, not a capability they control.

Voxtral TTS flips that. Open weights means you download the model, run it on your own hardware, and never send a single audio frame to a third party. For a company processing millions of support calls — or a startup building voice agents for healthcare, finance, or legal — that changes the math entirely. On-premise inference means no vendor risk, no data exposure, and no per-call fees compounding at scale. Pierre Stock, Mistral's VP of science operations, described a world where audio becomes the primary interface for AI agents: "We see audio as a big bet and as a critical and maybe the only future interface with all the AI models."

The timing is deliberate. Voice AI as a market crossed $22 billion globally in 2026. ElevenLabs closed a $500 million Series D in February and locked in an enterprise partnership with IBM the day before Mistral dropped Voxtral. The incumbent saw something coming. The open-source bet is Mistral's answer to a market where the dominant players have moats built on proprietary data and cloud lock-in. Their argument is that frontier quality doesn't have to mean closed systems.

Voxtral TTS is available now via API at $0.016 per 1,000 characters, or as open weights on Hugging Face under a CC BY-NC 4.0 license. The model works with vLLM and can run on any GPU with 16GB of memory or more. For developers who have already moved their LLM inference on-premise, Voxtral closes the last gap in the stack. A fully local voice-to-voice AI pipeline — no cloud dependencies, no external providers — is now buildable.

Comments

Get tomorrow's roundup. Free.

One email each morning. Sneakers, sports, culture, tech.

Link copied