Skip to main content
“Slow” on a phone call usually isn’t the model thinking. It’s the gap between the caller finishing and the assistant starting.

Find out where the time goes first

Open any call and look at Average response in the header. Then open the Conversation tab — each assistant turn shows its own timings, split between transcribing what the caller said, the model composing a reply, and the voice starting to speak. Fix the biggest number. Changing anything else is guessing.

The waiting-to-start gap

The assistant waits to be sure the caller has finished. That wait is the single biggest lever you have. In the assistant’s speech settings: Go too low and the assistant starts talking over people who were pausing to think. Go too high and every exchange feels sluggish. Short answers (“yes”, “the fourteenth”) tolerate a shorter wait than open questions. Make sure you’re on the Audio model turn detector. The text model waits for the transcript before deciding, which adds a noticeable pause.

Model choice

The language model is usually the second biggest number. A smaller, faster model often makes a better phone assistant than a larger one — on a call, answering in a second beats answering perfectly in four. The cost estimate in the editor shows latency alongside price. Same for the voice: lower style on ElevenLabs voices renders faster. Heavy styling costs time on every single turn.

Keep replies short

A long reply is slow to generate and slow to speak. “One or two sentences” in the prompt does more for perceived speed than most settings. See Writing prompts.

Tools

A tool call is a real wait — your API, not us. Callers hear a soft keyboard sound while one runs, so the silence doesn’t read as a dropped call, but the fix is on your side:
  • Return less. Pick out the fields the assistant needs.
  • Lower the timeout so a dead endpoint fails fast instead of holding the call.
  • Move slow lookups to before the call, so the answer is already in the prompt when the assistant speaks.

Knowledge

Documents set to always in the prompt are re-read on every turn of every call. One short document is fine; three long ones slow down every reply, whether or not they’re relevant. Use search unless you have a reason not to.

What won’t help

Shortening the system prompt. It’s processed once per turn and cached; trimming it rarely shows up in the timings. Blaming the transcription. It’s usually the smallest number of the three. Check before tuning it.