ElevenLabs is the $11B voice-AI flagship - the company that turned a 2022 text-to-speech demo into the de-facto English-speaking TTS standard by 2026. Scribe v2 sits at #1 on the Artificial Analysis speech-to-text benchmark with a 2.3% word error rate; Eleven v3 clones a voice from 5 seconds of audio and speaks 32 languages; and in May 2026 they slashed API prices up to 55%. If you've heard an AI voice this year that made you do a double-take, there's a decent chance it came from here. Here's the unvarnished deep dive.
What Is ElevenLabs?
Vendor: ElevenLabs, founded by Mati Staniszewski and Piotr Dabkowski, headquartered in London and New York. The headline number everyone cites is the February 4, 2026 Series D - $500M at an $11B valuation, more than 3x the $3.3B from their January 2025 Series C. Sequoia led; a16z quadrupled its stake; ICONIQ tripled down. ARR crossed $500M in the first four months of 2026, up from $330M+ at the end of 2025.
That's venture math, not product. The product positioning is simpler: ElevenLabs wants to be the audio layer of AI - not just TTS, but the whole stack: synthesis, transcription, cloning, dubbing, music, sound effects, and conversational agents. Enterprise names like Deutsche Telekom, Square, Revolut, and the Ukrainian Government are already running it in production for customer support and citizen services. When a telco and a national government both pick the same voice vendor, that's a tell.
Core Capabilities
What it actually does, as of the July 2026 snapshot - no marketing fog:
- Eleven v3 TTS - the flagship synthesis model. 32 languages, emotion tags, and 5-second voice cloning. This is the "de-facto English-speaking TTS standard" line you'll see in every 2026 comparison, and it's earned: the prosody on dramatic English reads is still a step above everyone else.
- Multilingual v2 - the most stable, realistic multilingual model. Pick this when you need a voice to hold character across languages.
- Turbo v2 / Flash v2.5 - low- and ultra-low-latency variants. Flash v2.5 hits roughly 75ms time-to-first-word, which is what makes real-time agent conversations feel human instead of laggy.
- Scribe v2 (ASR) - launched January 9, 2026. 90+ languages, 98% stated accuracy, speaker diarization, and character-level timestamps. New in v2: keyterm prompting (feed it domain terms and it transcribes them correctly) and built-in entity detection. Scribe v2 Realtime is the ultra-low-latency sibling for agents.
- Voice cloning - minutes of audio (down to ~5 seconds for an instant clone on v3) reproduces any voice, and it carries that voice across languages naturally. This is the feature that broke the dubbing industry's economics.
- AI dubbing - translate to 30+ languages while preserving the original speaker's timbre and delivery.
- Music API - text-to-music, studio quality, vocal or instrumental, trained on licensed data with commercial rights.
- Sound effects, voice isolation, ElevenAgents - the supporting cast. ElevenAgents is the enterprise conversational-AI platform the Series D money is explicitly funding.
The honest gaps: Eleven v3's emotional range is best in English and weakens on lower-resource languages; the cheaper Flash/Turbo tiers audibly trade realism for speed; and music, while good, isn't Suno-tier for full song structure yet.
Pricing
ElevenLabs' pricing confuses everyone because it mixes subscription tiers, character quotas, and a separate API. Collected from the platform in June 2026, after the May 7 price cuts:
| Plan | Price | Chars/mo | Key unlock | |------|-------|----------|-----------| | Free | $0 | 10K (~10 min) | Non-commercial, attribution required | | Starter | $5/mo | 30K (~30 min) | Commercial license, instant cloning | | Creator | $22/mo | 100K (~100 min) | Pro voice cloning | | Pro | $99/mo | 500K (~500 min) | Priority render, fine-tuning | | Scale | $330/mo | 2M (~2,000 min) | Team seats, API | | Business | $1,320/mo | 11M | Enterprise SLAs |
The May 7, 2026 update is the part most people missed: they introduced pay-as-you-go across the API and cut prices hard. TTS on the Flash model dropped 55% to $0.05 per 1,000 tokens; Scribe v2 STT dropped 45% to $0.22 per 1,000 tokens; ElevenAgents dropped 20% to $0.08 per minute. For music, Pro runs $9.99/mo for 500 songs with commercial rights, plus a 7 free songs/day tier.
The catch: characters bill on request, not success - a 422 error still burns quota. And the free tier is a sandbox, not something you'd ship. The honest sweet spot: Creator at $22/mo for a working solo creator; Pro at $99/mo once you're past one audiobook chapter a month.
Leaderboard Performance
Cross-checked against Artificial Analysis and independent evals, July 2026:
- Artificial Analysis AA-WER v2.0 (speech-to-text, released March 2026): Scribe v2 #1 at 2.3% WER. Google's models are the only ones in the same zip code. Open-source sits mid-pack at ~4.2%; Qwen3ASR Flash at 5.9%, Amazon Nova2Omni at 6.0%, Rev AI at 6.1% bring up the rear.
- AA-AgentTalk (voice-assistant commands): Scribe v2 #1 at 1.6% WER, with Gemini3Pro close behind at 1.7%. This is the benchmark that matters for agent builders.
- TTS has no single Elo-style arena, but every independent roundup puts ElevenLabs at #1. The DEV community 2026 test scored it 9.3/10, top of seven tools; the Digital Applied comparison called it the quality leader against Voxtral and OpenAI TTS despite Voxtral winning on cost.
The headline: ElevenLabs doesn't win on one benchmark, it wins across the whole audio surface - synthesis quality and transcription accuracy at the same time. Almost nobody else is competitive in both.
How It Compares
Same modality (voice AI), July 2026:
| Model | Strength | Latency | Cloning | Best at | |-------|----------|---------|---------|---------| | ElevenLabs (v3) | Quality + breadth | 75ms (Flash) | 5-sec instant | English drama, multilingual, ASR | | Cartesia Sonic | Speed | 75ms TTFW | Limited | LiveKit real-time agents | | OpenAI TTS | Simplicity, price | real-time stream | No | Fastest code-to-speech | | PlayHT | Stability, batch | moderate | Yes | Long-form narration, volume | | Hume EVI 2 / Sesame | Emotion | - | - | Affective conversational voice | | Voxtral (open-source) | Cost, self-host | 70ms TTFA | Limited | On-prem, $0.016/1K chars |
The pattern: ElevenLabs is the quality-and-breadth leader, OpenAI TTS is the cheap-and-simple one, Cartesia is the latency king, and the open-source crowd (Voxtral, Fish Audio, CosyVoice) wins on cost and self-hosting. Voxtral is 73% cheaper than ElevenLabs Flash and self-hosts on a single 16GB GPU - but it covers 9 languages to ElevenLabs' 70+, and quality tests still favor ElevenLabs.
Here's what no comparison table says out loud: ElevenLabs has the highest uptime in the industry - higher than OpenAI, per Regal's agent-infrastructure testing. For production voice agents, that matters more than a half-point of WER.
Who Should Use ElevenLabs
Straight answer:
- Use it for English-language dramatic content - ads, trailers, narrative. The prosody is still a tier above.
- Use it when you need one vendor for both TTS and ASR - Scribe v2 plus Eleven v3 in one stack.
- Use it for voice cloning across languages - dubbing pipelines, localization.
- Use it for enterprise voice agents where uptime is non-negotiable.
- Don't use it if raw cost-per-character is the whole equation - Voxtral or OpenAI TTS undercut it badly.
- Don't use it if you need sub-50ms latency or self-hosting - Cartesia or Voxtral.
- Don't use it for full-song music generation - that's still Suno's lane.
The mental model: ElevenLabs is the studio. OpenAI TTS is the quick-and-dirty. Cartesia is the real-time engine. Pick by what your pipeline actually bottlenecks on.
x-rush: 顶级模型 + 智能路由
Here's the part that matters if you're trying to actually use these models instead of bookmarking leaderboard screenshots. x-rush 接入 ElevenLabs 等顶级大模型 - alongside Cartesia Sonic, OpenAI TTS, PlayHT, and the rest of the frontier voice stack - and 智能路由到最符合任务的模型. An English dramatic read goes to Eleven v3; a real-time agent conversation goes to Cartesia's 75ms latency; a cost-sensitive batch narration goes to whichever model gives you the most minutes per dollar that week; a multilingual dubbing job goes to ElevenLabs' clone-preserving dubbing. The router reads language, latency budget, fidelity needs, and cost, then picks the model that's actually best for that job - not a fixed default.
And the model pool isn't frozen. 接入随世界潮流随时更新. When Scribe v2 took the #1 WER spot in March 2026, it entered the ASR pool. When Eleven v3 shipped 5-second cloning, the router learned to weight it for quick-clone briefs. When the May price cuts dropped TTS to $0.05/1K tokens, the cost curve for high-volume narration rerouted through Flash. The voice frontier moves every month in 2026; the router moves with it.
That's the pitch: you stop betting on a single voice model, and start getting the best model for every task, updated as the world updates.
How to Use It on x-rush
The audio workbench is where ElevenLabs lives in practice. Open the Audio workbench, drop in your script - a podcast intro, a multilingual dub, an agent voice - and the router decides whether ElevenLabs v3 is the right call or whether Cartesia's latency or PlayHT's batch stability serves the job better. You don't pick the model. You describe the audio; the router picks.
That's the point of the abstraction: ElevenLabs' quality and 70-language breadth when you need them, and you don't pay for v3-tier fidelity when a cheaper model gets the job done.
The Bottom Line
ElevenLabs is, by every independent measure I can find, the most complete voice-AI platform shipping in July 2026 - Scribe v2 at #1 WER on the Artificial Analysis benchmark, Eleven v3 with 32-language 5-second cloning, an $11B valuation backing it, and API prices that just dropped up to 55%. It's also not the cheapest (Voxtral, OpenAI TTS), not the lowest-latency (Cartesia), and not the best at full music (Suno). None of that is a contradiction; it's a platform with a clear shape - breadth and quality first, cost and edge-latency second.
If your work is English-dramatic, multilingual, cloning-heavy, or agent-critical, ElevenLabs is the voice you want in the room. If your work is pure cost-sensitive batch or self-hosted, there are cheaper answers. And the smart move in 2026 isn't to pick one - it's to route. That's what x-rush does: 接入顶级大模型,智能路由到最符合任务的,随世界潮流随时更新.
Pick the voice for the job. Or let a router do it for you.
Try the audio workbench - it runs on x-rush's smart-routed model pool, ElevenLabs included. The frontier, and every voice model that challenges it next, is already in the pool.