Speko's Arena

Methodology & report

How the ranking works

A plain-language account of how anonymous taps become a humanness ranking — and the honest limits of what it means. The full write-up and the complete 2026 results are below.

Frequently asked questions

Which AI voice sounds most human?
It's ranked live on the leaderboard and updates as listeners vote. Current top contenders are MiniMax, xAI Grok, ElevenLabs and Hume, each scored against a real human baseline rather than against each other in the abstract; top contenders shift as votes land.
How is a text-to-speech voice's humanness measured?
Listeners take blind A/B tests — two voices read the same line and they pick which sounds more human, with names hidden until after the vote. A real human recording is mixed in as a hidden voice, and a voice's humanness (0–100) is how often listeners prefer it over that human, fit with a Bradley-Terry model. When listeners pick a voice over the human about half the time it's statistically indistinguishable — a score of 100, which is the ceiling.
Can listeners tell AI voices apart from a real human?
Often, but not always — that's the whole test. 3 of your first 10 rounds hide a real human reading the same line; about 1 in 4 after that. The arena tracks how often the most human AI is picked over the human, and the strongest voices now get mistaken for, or preferred over, a real person a meaningful share of the time. The current human reference is a female studio read, so human rounds appear in female pairings; a male reference is being added.
Why do some AI voices still sound robotic?
After voting, listeners can flag what felt off — robotic delivery, mispronunciation, wrong emotion, awkward pacing, or audio glitches. Those tags are being collected now and will appear as a per-voice weakness map on the leaderboard once each voice has enough tags.
Does a higher humanness score mean a voice is better?
Not necessarily. Humanness only measures whether a voice is mistaken for a real person in a blind test — it says nothing about intelligibility, latency, cost, or language coverage. The leaderboard pairs humanness with gateway-measured latency and published price so you can weigh all three for a use case.

Blind A/B preference

Each round shows the prompt text and two clips, A and B, rendered by two different systems from the same text. You play both and pick the one that sounds closer to a real person. If neither does, a de-emphasized “Neither sounds human” option lets you say exactly that — a directed claim, not a tie. Neither-votes never enter the ranking fit; they are recorded as a signal about which voices fail to clear the human bar at all. Vendor and model names stay hidden until the moment after your vote lands, and voting unlocks only once both clips have actually played.

Same gender, same script, same loudness

Both sides read the same prompt, and which side is shown left/right is randomized — so the comparison is voice-vs-voice, never script-vs-script or position-vs-position.

Voices are only ever paired within the same gender (male vs male, female vs female). A male and a female read aren't a fair “which is more human” test, so they never meet.

Every clip — AI and the human — is normalized to a common −16 LUFS, spectrally matched to one shared frequency profile (reference-EQ, Matchering), and re-encoded to one uniform format (44.1 kHz mono MP3); the human recordings are additionally dereverberated. Louder, brighter, or roomier audio is reliably judged by its mastering rather than its voice — equalizing the whole chain means a vote can only be about how the voice speaks, not how it was recorded.

What each vote records

On every vote your browser sends a small JSON payload to POST /api/vote: your pick (a / b / neither), reaction time in ms, how many ms of each clip actually played, which side was A, and which clip you played first. A neither vote is stored like any other, but it never enters the ranking fit.

Rows recorded before 2026-07-01 can also carry a tie verdict, from the since-removed Tie option; those rows are excluded from the ranking fit.

The server adds only an anonymous cookie id and a hashed IP (the raw address is never stored), then writes one row. No names, emails, or accounts — nothing that identifies you. The reaction-time and played-ms fields are an effort signal; the side/order fields let position and order bias be measured.

Humanness, measured against a real human

A real human recording — public-domain studio reads — is mixed into the arena as a hidden voice. A voice's humanness is how often listeners pick it over the human: picked about half the time, it's statistically indistinguishable — humanness 100. Nothing scores above 100.

The human reference is currently a female voice, so it appears in female rounds; a male reference is being added so male voices anchor to a human too.

From a seeded benchmark to live votes

The board doesn't start empty. Opening scores are a seed: Vapi's published Humanness Index results (humannessindex.vapi.ai, retrieved 2026-06) encoded as prior comparisons, roughly 940 votes per voice. Voices Vapi doesn't rank yet enter as estimates placed on the same scale, and live votes take over from the seed as they land. Latency is real — warm first-audio measured network-free at the Speko gateway (the same dataset as benchmarks.speko.dev, which measures latency, not humanness) — and price is each provider's published list rate.

As real votes accumulate they take over from the seed: a Bradley-Terry maximum-likelihood fit (Hunter's MM) on an Elo-style scale. The confidence intervals shown today are a conservative count-based heuristic — 360/sqrt(votes), clamped — that narrows with sample size; a percentile bootstrap is planned. Every ranked vote is a forced A/B choice — play both, pick the more human one. “Neither sounds human” votes never enter the fit — excluded exactly like the tie votes cast before the 2026-07-01 protocol change — so they can't compress the gap between two voices. A higher score is not automatically a win — two voices are distinguishable only when their intervals don't overlap.

Anti-gaming

A per-source rate limit — keyed on a hashed IP, never the raw address — blunts ballot-stuffing without collecting personal data.

About 1 in 15 rounds is a gold attention-check: the real human against an obviously synthetic voice, with a known answer. Those rounds never enter the ranking fit. The first 10 rounds also schedule exactly 3 human pairs for the intro game, and every round grades right or wrong: right means you spotted the human or correctly called an all-AI round; wrong means an AI tricked you — you picked a voice as the more human one when no human was there, missed the hidden human, or said “neither” when one was real. Grading is the game layer only — ranked votes are unaffected — and the shareable score stays humans-spotted, so no option can be farmed for score. Trick events are recorded per voice, which is how the fooled-rate stat is measured once enough human rounds land.

Honest caveats

Today's scores are seeded and will shift as votes land — the confidence intervals are the point. Preference is also relative and population-dependent: it measures which of two voices sounds more human to these listeners, not a ground-truth score, and says nothing about intelligibility, cost, or fit for a use case.

Protocol changelog

2026-07-02
Added a 'Neither sounds human' option after tester feedback: forcing a pick when both clips sound synthetic put dishonest votes into the ranking. Neither-votes never enter the pairwise fit - they are recorded as a signal about the human bar. A/B picks remain the only ranked votes.
2026-07-01
Removed the Tie option after tester feedback showed it read as “both are AI” rather than “equally human”. Forced choice from here; the 24 tie votes cast before the change are excluded from the ranking fit.

Get the 2026 humanness report

Enter your email to get the report when it ships, plus a heads-up when the ranking moves.

  • The full 2026 humanness report (PDF) when it ships
  • An email when the ranking shifts — new voices, protocol changes
  • Per-voice failure maps from listener tags, once collected