Leaderboard · Model
Cartesia
Ranked #10 of 11 models by blind A/B votes against a real human recording. Humanness 100 means listeners can no longer tell it from the person.
Standings as of August 31, 2026 · refreshed every 5 minutes
Hear it
“Thanks for calling — I can get that rescheduled for Thursday morning. Does that work for you?”
Background
Cartesia builds speech models on state-space architectures (the founding team behind Mamba-style SSMs); Sonic is known for ultra-low streaming latency. On this board it pairs near-instant first audio with a mid-field humanness score.
On the board
Live results — the numbers move as votes land. Latency is measured at the Speko gateway (warm first-audio, n=30); price is the provider's published rate. Method details on the methodology page; raw standings as open data (CC BY 4.0).