Speko's Arena

Leaderboard · English

Which AI voices sound most human

Ranked by blind A/B votes with a Bradley-Terry model, measured against a real human baseline.

Read the methodology & get the full 2026 report
#VoiceHuman%
·
Human reference
100
1
MiniMax Speechspeech-2.6-hd
364ms · $0.012/min
95
292
3
ElevenLabseleven_v3
526ms · $0.100/min
89
4
Hume Octaveoctave-2
613ms · $0.150/min
88
5
OpenAI Speechgpt-4o-mini-tts
695ms · $0.015/min
86
Cb
Chatterbox0.5B open weights
· free
77
7
Alibaba Qwen-TTSqwen3-tts-instruct-flash
1908ms · $0.012/min
76
8
Rimearcanav3
2213ms · $0.020/min
73
9
Inworld TTSinworld-tts-2
129ms · $0.020/min
71
10
Cartesiasonic-3.5
126ms · $0.025/min
70
11
450ms · $0.028/min
64

Humanness rescales the blind win-rate vs a real human: 50% (indistinguishable) = 100, so 95 ≈ picked 47% of the time. Latency is warm first-audio at the Speko gateway (n=30); price is the provider's published rate. Each row is one fixed male+female voice pair. Vote counts are live rounds; the fit also carries a disclosed Vapi-index prior. Dimmed rows: too few live rounds yet. Open data: JSON · CSV (CC BY 4.0).

The tradeoff

Naturalness vs. speed

The most human voice isn't always the one you can ship in a live call. The closer to the dashed human line and the further right, the better the deal.

FAST & HUMAN556065707580859095100100ms250ms500ms1s2.5sResponse latency · faster →↑ more humanHuman · 100MiniMaxxAI GrokElevenLabsHumeOpenAIAlibabaRimeInworldCartesiaGradium

Every voice here runs on Speko

Speko is the voice-AI gateway — one API for ElevenLabs, OpenAI, Cartesia and every provider above, auto-picked by benchmark for latency, cost and naturalness.

Build with Speko

Help decide the ranking

Every blind vote sharpens these numbers.

Play a round