The TTS Meta-Leaderboard
The best text-to-speech APIs, ranked
Simba 3.2 by SpeechifyAI is currently the best text-to-speech API, scoring 96/100 across all 3 public TTS benchmarks at $10.00 per 1M characters. We rank 110 models from 46 providers by aggregating artificialanalysis.com, voicearena.com and humannessindex.vapi.ai into one score — a fixed formula, no editorial adjustments, no vendor weighting. The cheapest scored model is Kokoro 82M v1.0 at $0.70 per 1M characters.
Quick answers
Model rankings
Per-board cells show the board's own normalized score; "—" = not listed. The Score is a 0–100 Bayesian pairwise rank aggregate: each board contributes a fractional rank (#1 = 1.0, scaled by board size), breadth pulls single-board models toward the field average, and the Borda aggregate rewards models that are strong across every board — Schulze-tie-broken, with a High/Med/Low confidence label on each model. Carets show score-rank movement over the trailing week. Snapshot: Aug 30, 2026.
Text-to-speech APIs: common questions
- What is the best text-to-speech API right now?
- Simba 3.2 by SpeechifyAI is the top-ranked text-to-speech model on the texttospeech.com meta-leaderboard, scoring 96/100 with high confidence as of Aug 30, 2026. The score aggregates 3 independent public benchmarks — artificialanalysis.com, voicearena.com and humannessindex.vapi.ai — so it reflects agreement between boards rather than one board's view. It is priced at $10.00 per 1M characters. The ranking is recomputed weekly and the full standings are on the leaderboard.
- What is the cheapest text-to-speech API?
- Kokoro 82M v1.0 by Kokoro is the cheapest scored text-to-speech model at $0.70 per 1M characters, ranked #59 of 110 with a score of 31/100. The cheapest model inside the top ten is Simba 3.2 by SpeechifyAI at $10.00 per 1M characters, so that is the cheapest option we would call competitive on quality. Prices are $ per 1M characters; per-minute pricing is converted to the same unit so every row compares like with like.
- What is the best open-weights text-to-speech model?
- Fish Audio S2 Pro by Fish Audio is the highest-ranked open-weights model, #29 overall with 54/100 as of Aug 30, 2026. Open-weights models can be self-hosted instead of called through a vendor API, so their running cost depends on your own hardware rather than a per-character rate. The leaderboard can be filtered to open weights only.
- How does texttospeech.com score text-to-speech models?
- Every score is a Bayesian pairwise rank aggregation (BPRA) over 3 public boards. Each board contributes a model's rank as a fraction — #1 is 1.0, scaled by how many models the board ranks — a Bayesian coverage term pulls models that appear on few boards toward the field average, and an ecosystem-normalized Borda count turns the result into a 0–100 score: #1 on every board is exactly 100, and missing a board caps the ceiling. Ties break by Schulze path strength. We do not run our own listening tests; we aggregate boards that do, and publish the formula and every source snapshot.
- How often is the text-to-speech leaderboard updated?
- The source boards are polled weekly and the ranking recomputes whenever any of them changes. The current snapshot was captured Aug 30, 2026. Every snapshot is archived and downloadable as JSON, so any past ranking can be reproduced from the same data and the same published formula.
- Is the ranking sponsored or influenced by vendors?
- No. There is no vendor sponsorship, referral fee, or paid placement, and no code path that can elevate or demote a named vendor: every model goes through the identical formula, the formula is open source, and each published ranking links the exact source snapshots it was computed from. The source boards are also independent of the vendors they rank — that is one of the criteria a board must meet to be aggregated at all.