Sub-90ms text-to-speech API built for real-time voice agents.
Cartesia Sonic 3 is built around a single, unglamorous number that turns out to matter more than almost anything else in voice AI: time-to-first-audio. Sonic 3 hits roughly 90 milliseconds, and the Sonic Turbo variant drops that to around 40ms. In a voice agent, that delay is the gap between a person finishing their sentence and the machine starting to reply. Above ~200ms it reads as a stutter; below ~100ms it reads as a conversation. Sonic 3 lives in the comfortable side of that line, which is why it keeps showing up at the top of voice-agent latency rankings in 2026.
Speed is only useful if the voice is worth listening to, and here Sonic 3 also delivers: it is ranked #1 for naturalness in independent comparisons, with native support for 40-plus languages. You can clone a voice and localise it across 42 languages, and the API exposes fine control over pitch, speed, and emotion — enough to keep a brand voice consistent from English support calls to Spanish onboarding to Japanese IVR. That combination of low latency and high naturalness is the hard part; plenty of models nail one and miss the other.
The natural comparison is ElevenLabs, which PixlRun already covers and which remains the reference point for expressive, studio-grade narration. The distinction is use case. ElevenLabs is built for produced audio — audiobooks, video voiceover, polished long-form — where you can afford a few hundred milliseconds because nobody is waiting live. Cartesia is built for the phone call and the live agent, where every millisecond is perceptible. If your output is a rendered MP4, ElevenLabs’ expressiveness is the draw; if your output is a real-time conversation, Sonic 3’s latency is the draw. They are increasingly two different products, and Sonic 3 is the sharper tool for the conversational job.
Pricing is usage-based and starts genuinely free. The free plan is $0/month with no time limit, bundling 20,000 model credits and $1 of prepaid voice-agent usage — enough to prototype a real agent before paying a cent. Every plan includes unlimited workspace seats and unlimited voice slots, which is refreshing in a category that loves to meter collaborators. At scale, Sonic 3 runs around $35 per effective million characters, billed at roughly 15 credits per second of audio. That is premium pricing, and it is the main thing to model carefully before committing.
The honest caution is exactly that cost curve. At low volume the free tier and pay-as-you-go rates are easy to live with. At high call volume — a support line handling thousands of minutes a day — $35 per million characters compounds, and the credit-plus-agent-minute billing model takes real effort to forecast. You will want to instrument your character counts and run the math against your expected concurrency before you wire Sonic 3 into a production phone tree. The latency advantage is real, but you are paying for it.
There is also a scope point worth being clear about: Sonic 3 is an API, not an app. It is a building block for developers assembling a voice agent, not a turnkey product like a Synthesia or a HeyGen where you log in and click. If you do not have engineers, this is not the tool — you want a packaged voice-agent platform instead. If you do have engineers building conversational AI, Sonic 3 is one of the cleanest, fastest TTS layers you can drop in, with SDKs and a developer experience designed for exactly that.
The verdict: for real-time voice agents, Cartesia Sonic 3 is the model to beat in 2026. It pairs the lowest credible latency in the field with top-ranked naturalness and broad multilingual reach, and it lets you start for free. The two things keeping it from a higher score are the cost at scale and the forecasting friction of its credit model — neither a flaw in the technology, both a flaw in your spreadsheet if you skip the math. Build a real-time agent and Sonic 3 belongs on your shortlist; build batch narration and a cheaper model will serve you just as well.
How we would use it. Sonic 3 is the TTS layer we would reach for the moment a project crosses from “produced audio” into “live conversation.” Prototype on the free tier, instrument character counts early, and run the $35-per-million-characters math against expected concurrency before wiring it into a production phone tree — the technology is excellent, but the bill is where this decision is actually made. For anything pre-rendered, a cheaper model gives away nothing a listener will notice. For a real-time agent where the 90ms gap is the difference between a conversation and a stutter, Sonic 3 is the model to beat in 2026, and the free starting tier means there is no reason not to benchmark it against ElevenLabs and the rest on your own traffic before committing.
Every verified price, limit, and model change we have tracked for Cartesia Sonic 3.
One email when Cartesia Sonic 3 changes price or limits. No account, no spam.