> ## Documentation Index
> Fetch the complete documentation index at: https://docs.qall.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Making an assistant feel fast

> Where the time goes on a call, and which settings actually move it.

"Slow" on a phone call usually isn't the model thinking. It's the gap between the
caller finishing and the assistant starting.

## Find out where the time goes first

Open any call and look at **Average response** in the header. Then open the
**Conversation** tab — each assistant turn shows its own timings, split between
transcribing what the caller said, the model composing a reply, and the voice
starting to speak.

Fix the biggest number. Changing anything else is guessing.

## The waiting-to-start gap

The assistant waits to be sure the caller has finished. That wait is the single
biggest lever you have.

In the assistant's speech settings:

| Setting | Try |
| - | - |
| **Min endpointing delay** | Lower it for a snappier feel |
| **Max endpointing delay** | Lower it if the assistant seems to hang after short answers |

Go too low and the assistant starts talking over people who were pausing to
think. Go too high and every exchange feels sluggish. Short answers ("yes",
"the fourteenth") tolerate a shorter wait than open questions.

Make sure you're on the **Audio model** turn detector. The text model waits for
the transcript before deciding, which adds a noticeable pause.

## Model choice

The language model is usually the second biggest number. A smaller, faster model
often makes a *better* phone assistant than a larger one — on a call, answering
in a second beats answering perfectly in four.

The [cost estimate](/assistants) in the editor shows latency alongside price.

Same for the voice: lower **style** on ElevenLabs voices renders faster. Heavy
styling costs time on every single turn.

## Keep replies short

A long reply is slow to generate and slow to speak. "One or two sentences" in the
prompt does more for perceived speed than most settings. See
[Writing prompts](/guides/writing-prompts).

## Tools

A tool call is a real wait — your API, not us. Callers hear a soft keyboard sound
while one runs, so the silence doesn't read as a dropped call, but the fix is on
your side:

* Return less. [Pick out the fields](/tools/http-tools) the assistant needs.
* Lower the timeout so a dead endpoint fails fast instead of holding the call.
* Move slow lookups to **before the call**, so the answer is already in the
  prompt when the assistant speaks.

## Knowledge

Documents set to **always in the prompt** are re-read on every turn of every
call. One short document is fine; three long ones slow down every reply, whether
or not they're relevant. Use **search** unless you have a reason not to.

## What won't help

**Shortening the system prompt.** It's processed once per turn and cached;
trimming it rarely shows up in the timings.

**Blaming the transcription.** It's usually the smallest number of the three.
Check before tuning it.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.