Verdict
Eleven v3 is the expressive benchmark of its generation. It still holds the #1 spot on Vapi’s Humanness Index at 97/100 and ranks best overall on nonverbal-vocalization expressiveness in independent evaluation. But the same boards that praise its voice also expose its limits: 758ms latency, a slide from ~#2 to ~#15 on Artificial Analysis through 2026, and non-English output that native speakers describe as unusable. It is the right tool for pre-rendered, narrative-grade voice where emotional range is the point — and the wrong tool for real-time, low-resource languages, or fine-grained control.
What it is
Eleven v3 is ElevenLabs’ flagship text-to-speech model, released mid-2025. Of the four models in this review series, it is the best-covered — its Hacker News launch thread alone hit 293 points. That thread set the shape of the coverage that followed: near-professional US-English, and a much harder time everywhere else.
Where it sits across the boards
texttospeech.com aggregates three public blind-test leaderboards into a single 0–100 score via a published Bayesian Pairwise Rank Aggregation method. Eleven v3’s score is the output of that method, not a unanimous verdict — the three boards disagree about it in ways worth spelling out.
- Vapi Humanness Index (as of August 2026) puts Eleven v3 at 97/100, tied for the top spot. The same index records 758ms latency, so this ranking measures how human it sounds, not how fast it responds.
- Artificial Analysis Speech Arena (August 2026) lists Eleven v3 at Elo ~1,180, rank ~15, down from roughly #2 earlier in 2026. The “Eleven v3 Conversational” variant sits higher at ~1,215. It is one of the new models in our August new-models roundup.
- Voice Arena (voicearena.com) lists Eleven v3 at Elo 1,000, ~#10 of 16 models.
That is the split in one line: tied for #1 on human-likeness, ~#10 on Voice Arena, and mid-teens on Artificial Analysis after a year of decay. The aggregate reflects that tension rather than papering over it.
What wider independent evidence shows
Away from the leaderboards, practitioner reports and academic evaluations are broadly consistent: expressive, sometimes too expressive, and weaker than the older v2 on control.
- Duke’s DDMC, a first-hand practitioner test, called v3 “too expressive” for narration, noted fewer control variables than v2, and reported clicking artifacts.
- Oğuzhan Koçaklı’s hands-on review found v3 less forgiving than v2: the Creative setting “goes off the rails,” while Robust is “reliable but boring” and ignores tags.
- Reedsy, an audiobook-industry platform, flagged intermittent glitches — uneven volume, skipped syllables — that force credit-burning regenerations, plus pronunciation failures. They price v3 at ~$99 per 80,000 words against $2,000–5,000 for professional narration, but warn on 10–20 hours of editing and credit expiry.
- arXiv 2605.27383, a Thai (low-resource) evaluation, measured NMOS 4.21 — good naturalness — against a WER of 40.6–42.3%, which is poor intelligibility.
- arXiv 2604.16211 (NVV-SuperBench) found ElevenLabs best overall across languages on nonverbal-vocalization expressiveness and naturalness.
The Hacker News launch thread supplies the raw version of the same pattern. US-English was called near-professional. Native speakers were harsher on non-English: Russian was “glass in your ears,” Romanian “like 15 years ago,” and the accent was reported to “snap” into the target language mid-clip. There were stability complaints — loudness jumps, gibberish — and pricing gripes at ~$0.08/min against OpenAI’s ~$0.015/min.
Strengths
- Expressiveness. Best overall across languages on nonverbal-vocalization expressiveness and naturalness per NVV-SuperBench.
- Human-likeness. #1 on the Vapi Humanness Index at 97/100.
- US-English quality. Near-professional by community judgment.
Weaknesses
- Latency. 758ms on Vapi’s index, and Coval states plainly that v3 is not a real-time model. Coval also notes P50 latency looks identical across vendors while P95/P99 diverge by 200ms+ — the tail is where v3 shows up badly.
- Elo decay. From ~#2 to ~#15 on Artificial Analysis through 2026.
- Non-English. Native speakers on the launch thread called Russian, Romanian, and others out as unusable; Thai evaluation shows natural but unintelligible output.
- Less control than v2. Fewer control variables (DDMC), ignored tags (Koçaklı), “too expressive” for narration.
- Stability. Loudness jumps, gibberish, skipped syllables, uneven volume.
- Cost and caps. ~5¢/min post-May Business pricing (Coval), a 5,000-character cap, and credit expiry (Reedsy).
Who it’s for — and who it isn’t
For: pre-rendered, narrative-grade voice where emotional range matters — audiobooks with a real editing budget, trailers, character work, advertising reads. Anyone who needs the most human-sounding voice and can absorb the latency and a cleanup pass.
Not for: real-time voice agents, IVR, or anything interactive; non-English-heavy or low-resource production; fine-grained control, where v2 remains the better tool; budget-sensitive long-form where credit expiry and regen burn matter.
Bottom line
Eleven v3 is still the expressive benchmark and still the most human-sounding model on Vapi’s index. But it is a 2025 flagship measured against 2026 real-time models, and the boards show the gap: #1 on humanness, mid-teens on Artificial Analysis, latency the faster competitors have largely closed. If your output is pre-rendered and emotional range is the point, it remains the pick. If you need real-time response or fine control, v2 or a faster competitor is the better tool.