Monthly latency report

Voice stack latency report

Each month I measure turn latency across the major hosted voice stacks under one identical prompt set, then publish the results with the raw CSV. This page is the index.

How it is measured

This is the fixed method every report follows. The rows below describe how a run is set up, not any measured result.

Primary metric
Turn latency at p50 and p95.
Measured from
The moment the user stops speaking to the first audio byte back.
Also captured
Barge-in response time and call-setup time.
Stacks covered
The major hosted voice stacks.
Sample size
30 calls per stack (N = 30).
Prompt set
One identical prompt set, reused across every stack.
Raw data
A raw CSV published alongside each report.
Cadence
Published the same calendar day each month.

No measurements are published yet. Nothing here is estimated or modeled.

Reference budget, from published figures

Until the first measured report lands, here is what the stack costs on paper. Each stage below is a latency figure published by the vendor or an independent benchmark, cited. These are third-party numbers, not measurements of my line, and they are the starting point the monthly report will replace with measured ones. The full breakdown, and why the naive sum is a trap, is in the latency-budget note.

Human turn gap
~200ms is natural, comfortable under ~500ms (Prodinit).
Deepgram STT (Nova-3)
First token ~150ms in the US, sub-300ms streaming (Deepgram).
LLM first token
~100–180ms on fast models, 300–500ms common (Introl).
TTS first audio
~40–90ms vendor targets, ~190–310ms measured P50 (Gradium).
Telephony / network
20–40ms WebRTC; Twilio targets sub-600ms end to end (Twilio).

Published third-party figures, cited. Not measured on this line.

Reports

Awaiting first report

The first report publishes soon

Reports will appear here newest first, each linked to its raw CSV. The first run is being prepared. Come back on publish day to read it.

agent protocol