Find out where the time goes first
Open any call and look at Average response in the header. Then open the Conversation tab — each assistant turn shows its own timings, split between transcribing what the caller said, the model composing a reply, and the voice starting to speak. Fix the biggest number. Changing anything else is guessing.The waiting-to-start gap
The assistant waits to be sure the caller has finished. That wait is the single biggest lever you have. In the assistant’s speech settings:
Go too low and the assistant starts talking over people who were pausing to
think. Go too high and every exchange feels sluggish. Short answers (“yes”,
“the fourteenth”) tolerate a shorter wait than open questions.
Make sure you’re on the Audio model turn detector. The text model waits for
the transcript before deciding, which adds a noticeable pause.
Model choice
The language model is usually the second biggest number. A smaller, faster model often makes a better phone assistant than a larger one — on a call, answering in a second beats answering perfectly in four. The cost estimate in the editor shows latency alongside price. Same for the voice: lower style on ElevenLabs voices renders faster. Heavy styling costs time on every single turn.Keep replies short
A long reply is slow to generate and slow to speak. “One or two sentences” in the prompt does more for perceived speed than most settings. See Writing prompts.Tools
A tool call is a real wait — your API, not us. Callers hear a soft keyboard sound while one runs, so the silence doesn’t read as a dropped call, but the fix is on your side:- Return less. Pick out the fields the assistant needs.
- Lower the timeout so a dead endpoint fails fast instead of holding the call.
- Move slow lookups to before the call, so the answer is already in the prompt when the assistant speaks.