← Back to all work
Voice agent · iGaming · Company withheld

From 3.4 seconds to 511 ms: p95 first-audio latency on a production voice agent.

In a voice experience, the pause between turns shapes the entire conversation. We optimized the real-time pipeline to make the agent respond faster while keeping the pipeline flexible and reliable.

iGaming · Voice AI · Company withheld

Faster responses, more natural conversations.

The work focused on reducing the silence between a caller finishing a sentence and the agent beginning its response. Each stage was measured independently, then optimized against the end-to-end customer experience.

Production benchmark
LiveKit CloudTransport & egress DeepgramSpeech to text GPT-4o-miniLanguage model CartesiaText to speech
Before and after

Production benchmark

p95 latency
Pipeline measurement Baseline Optimized Change
First audio 3,361 ms 511 ms −84.8%
Beyond speed

The work ran in two passes. The first tuned the existing pipeline stage by stage. The second restructured the response path so the model output streams straight into speech. The final figure includes the full agent logic, tool calls and guardrails. The review also fixed several configuration and retry issues in the call platform.

Have a voice flow that feels slow? Bring the recordings to a call.

Discuss your use case