Skip to main content
  1. Home
  2. API
  3. Text-to-speech API latency in 2026: time to first audio, measured independently
Published on •API

Text-to-speech API latency in 2026: time to first audio, measured independently

Speechify Team

Speechify Team


Speechify API delivers 300ms 
latency, human-quality voices, 
and 50+ languages

Apple2025 Apple Design Award
50M+ Users

Two benchmarks that Speechify does not run measured the time to first audio of SpeechifyAI's Simba 3.2 model in September 2026. Coval recorded a median of 106 ms over the 24 hours to 24 Sep 2026. Voice Arena recorded 123 ms on its US English board the same day, the lowest median of the 14 models it timed. This post explains what those numbers measure, why time to first audio is the latency figure worth comparing, and how to measure it from your own code.

The numbers

Benchmark

What it measures

Simba 3.2

When

Voice Arena, US English board

Median time to first audio on the real-time streaming path

123 ms, the lowest of 14 models timed

24 Sep 2026

Coval

Median time to first audio, counting any silence before the first audible sample. Open-source harness, runs every 30 minutes from US East

106 ms

24 hours to 24 Sep 2026

SpeechifyAI production traffic (our own measurement)

Time for our API to write the first audio byte on real customer requests in US East

56 ms p50, 102 ms p90

15 Sep 2026

The two independent readings sit above our own because they include the network between their runner and our servers, and any silence at the start of the audio. Each is that benchmark's reading on that date. Both boards keep running, so check them for today's figure.

Voice Arena's US English board, 24 Sep 2026

Model

Median time to first audio

SpeechifyAI Simba 3.2 Real-time

123 ms

Inworld Realtime TTS 2 Research Preview

168.5 ms

Gradium TTS Real-time

236 ms

Cartesia Sonic-3.5 Real-time

250 ms

Smallest AI Lightning 3.1 Pro Real-time

263.5 ms

Fish Audio S2 Pro Real-time

267 ms

Source: Voice Arena, read on 24 Sep 2026. The board timed 14 models; the six fastest are shown.

Why time to first audio is the number to compare

When a voice agent replies or an app reads text aloud, the listener waits from the moment the text is sent until they hear sound. That wait is time to first audio.

  • Time to first byte can look better than what a listener hears. A stream can open with silence, and the listener hears nothing until the first audible sample. Time to first audio counts that silence.
  • Total generation time is the wait for the last byte. It matters when you save a file, not when someone listens while the audio streams.
  • Conversation leaves little room. People usually answer each other within a few hundred milliseconds. A voice agent has to fit speech recognition, the language model and speech synthesis into that pause, so every millisecond the voice takes is one the rest of the pipeline cannot use.

How Simba 3.2 got faster without a smaller model

In Coval's data, Simba 3.2's daily median fell from 379 ms on 13 Sep 2026 to 116 ms on 18 Sep 2026, about 70% lower in five days.

The model did not change. It has the same weights, with no distillation, no quantization and no cut-down "turbo" variant. The time came off the serving path around the model:

  • A new serving build on a different GPU class, with low-level work in the inference stack so audio leaves sooner.
  • Plan, entitlement and rate-limit checks resolve from a cache instead of live lookups before the first byte.
  • The API gateway runs in the same region as the GPUs, on connections that stay warm between requests.
  • About 100 ms of leading silence is gone. Coval's own leading-silence measure for Simba 3.2 went from 102 ms on 14 Sep to 13 ms.

Quality held. On the Artificial Analysis text-to-speech leaderboard, Simba 3.2 had an Elo of 1,237 (±14) on 23 Sep 2026, in the top five of 92 provider voices, and every model rated above it costs more.

How to get the lowest latency in your app

  1. Stream. Call POST /v1/audio/stream for anything played while it generates. POST /v1/audio/speech responds only once the whole clip exists.
  2. Ask for Simba 3.2. Set model to simba-3.2 for English. When model is omitted, the API uses simba-3.0.
  3. Reuse one connection. Create one HTTP client, open the connection at startup and keep it for every request, so no request pays for the DNS lookup and the TCP and TLS handshakes.
  4. Send sentences as they arrive. With a language model in front, send each complete sentence as soon as it exists and play the streams in order.
  5. Pick the format your player needs. audio/pcm at 24 kHz needs no encoding step. Use MP3 or Ogg Opus when bandwidth matters more.
  6. Run close to the API. The API is served from US East, so a server in or near US East spends the least time on the network.

Building on LiveKit Agents? This guide builds a voice agent with the Speechify plugin, which streams each sentence as the language model writes it. The reasoning behind each step, with code, is in the latency documentation.

Measure it yourself

Time the first chunk of a streaming response from where your code runs. For a fair number:

  • Discard the first request or two. They include opening the connection.
  • Vary the text between requests, as your real traffic does.
  • Send requests one after another, not in parallel.
  • Take at least 20 requests and report the median (p50) and p90, not an average or the single best run.
  • Measure with PCM. With Ogg, the first bytes are headers rather than audio.

Each streaming response carries a Server-Timing header whose ttfb value is our share of the wait, so what remains is the network and your own stack. The latency documentation has scripts in Python and TypeScript that do this.

Frequently asked questions

SpeechifyAI's Simba 3.2 model wrote its first audio byte in 56 ms at the median and 102 ms at p90 on production traffic in US East on 15 Sep 2026. Independent benchmarks measured a median time to first audio of 106 ms (Coval, the 24 hours to 24 Sep 2026) and 123 ms (Voice Arena, 24 Sep 2026), which include their network and any silence at the start of the audio.

On Voice Arena's US English board on 24 Sep 2026, SpeechifyAI's Simba 3.2 had the lowest median time to first audio of the 14 models it timed: 123 ms, ahead of Inworld Realtime TTS 2 Research Preview at 168.5 ms. Benchmarks re-run continuously, so check the date on any ranking before you rely on it.

Yes. POST /v1/audio/stream sends audio over HTTP as it is generated, as PCM, MP3, Ogg Opus or AAC. PCM has the lowest latency because it needs no encoding step.

No. SpeechifyAI is Speechify's developer API, billed separately from the Speechify reading app, and a Reader subscription does not include API usage. API plans are monthly with no annual commitment: a free plan with 500,000 characters a month and no card, then Starter at $10 a month with additional usage at $10 per 1M characters, Pro at $99 with $8 and Scale at $499 with $6.

Try Simba 3.2 on SpeechifyAI: create a free API key and time your first request against the numbers above.

Access Speechify’s beloved voices via API fast, scalable, and developer-friendly

Get API Access
Sparse JavaScript keywords import, from, const, await on a white background

Share This Article

Speechify Team

Speechify Team


Speechify Team

Speechify

About Speechify

#1 Text to Speech Reader

Speechify is the world’s leading text to speech platform, trusted by over 50 million users and backed by more than 500,000 five-star reviews across its text to speech iOS, Android, Chrome Extension, web app, and Mac desktop apps. In 2025, Apple awarded Speechify the prestigious Apple Design Award at WWDC, calling it “a critical resource that helps people live their lives.” Speechify offers 1,000+ natural-sounding voices in 60+ languages and is used in nearly 200 countries. Celebrity voices include Snoop Dogg and Gwyneth Paltrow. For creators and businesses, Speechify Studio provides advanced tools, including AI Voice Generator, AI Dubbing, and its AI Voice Changer. Speechify also powers leading products with its high-quality, cost-effective text to speech API. Featured in The Wall Street Journal, CNBC, Forbes, TechCrunch, and other major news outlets, Speechify is the largest text to speech provider in the world. Visit speechify.com/news, speechify.com/blog, and speechify.com/press to learn more.