Monthly latency report
Voice stack latency report
Each month I measure turn latency across the major hosted voice stacks under one identical prompt set, then publish the results with the raw CSV. This page is the index.
How it is measured
This is the fixed method every report follows. The rows below describe how a run is set up, not any measured result.
- Primary metric
- Turn latency at p50 and p95.
- Measured from
- The moment the user stops speaking to the first audio byte back.
- Also captured
- Barge-in response time and call-setup time.
- Stacks covered
- The major hosted voice stacks.
- Sample size
- 30 calls per stack (N = 30).
- Prompt set
- One identical prompt set, reused across every stack.
- Raw data
- A raw CSV published alongside each report.
- Cadence
- Published the same calendar day each month.
No measurements are published yet. Nothing here is estimated or modeled.
Reference budget, from published figures
Until the first measured report lands, here is what the stack costs on paper. Each stage below is a latency figure published by the vendor or an independent benchmark, cited. These are third-party numbers, not measurements of my line, and they are the starting point the monthly report will replace with measured ones. The full breakdown, and why the naive sum is a trap, is in the latency-budget note.
- Human turn gap
- ~200ms is natural, comfortable under ~500ms (Prodinit).
- Deepgram STT (Nova-3)
- First token ~150ms in the US, sub-300ms streaming (Deepgram).
- LLM first token
- ~100–180ms on fast models, 300–500ms common (Introl).
- TTS first audio
- ~40–90ms vendor targets, ~190–310ms measured P50 (Gradium).
- Telephony / network
- 20–40ms WebRTC; Twilio targets sub-600ms end to end (Twilio).
Published third-party figures, cited. Not measured on this line.
Reports
The first report publishes soon
Reports will appear here newest first, each linked to its raw CSV. The first run is being prepared. Come back on publish day to read it.