Wireface Desktop
Voice and hearing
How the face hears you and talks back: speech recognition on your PC, a realtime voice that follows everything the agent does, and the classic voice when there isn't one.
On this page
How a request flows
- You speak. Speech recognition on your PC turns it into text, and your words go straight to the agent as its prompt, word for word. (Or type, or drop text on the face.)
- The agent works, in a background session with all its tools.
- The voice follows. Every tool call, result and reply becomes a short note for the realtime voice model.
- The face tells you: replies, questions and progress, in its own words, lip-synced.
The realtime model never hears you, never does the work, and has no tools. It's the agent's voice, and it only speaks when the face asks it to. So every request is answered once, by the agent, and passed on by the voice.
Listening
Turn the microphone on with Listening: in the face's right-click menu, on the floating panel's toolbar, or on the Voice tab under Talking. It's off when a face is new. There's no wake word: while it's on, everything said nearby goes to the agent.
While the face speaks, the microphone rests, so it never hears itself. The exceptions are its name (below), and Interrupt by talking if you turn that on.
Typing still works whenever you like: a typed message goes to the agent, and the voice is told about it.
Say its name to cut in
Say the face's name ("Claude" by default) while it talks, and it stops within about half a second and listens. Then say what you want next: "Claude, ignore that and tell me the weather in Paris". What it was saying, and anything left of that answer, is dropped. If the agent is still working when you address it by name, what you say replaces that request instead of waiting behind it.
- A small keyword spotter on your PC listens for the name all the time. The face's own voice saying its name, or a word that sounds like it ("cloudy", "clause"), doesn't count.
- Just the name on its own isn't sent to the agent.
- Saying just "stop" stops it talking and stops the work. Esc and Stop do the same.
- The name is the face's Name (Agent tab, This face). If speech recognition mishears it, add other spellings under Also called in the Agent tab's Room group.
- Turn it off with Say its name to interrupt (Voice tab, Hearing).
Background noise and other voices
Two switches on the Voice tab, under Hearing, both on by default:
- Reduce background noise takes out steady noise (fans, hum, hiss, typing) before your words are recognised.
- Ignore distant voices drops speech much quieter than yours, like a TV or people across the room. It learns your level as you talk.
The realtime voice
By default the face talks with a realtime speech-to-speech model: Gemini Live or OpenAI Realtime. It needs a key for that service (see Add your keys). It reads silent notes about the agent's session and speaks when there's something for you:
- the agent's reply, with every fact, number and question kept, and code, paths and links described rather than read out;
- a permission question, which your "yes" or "no" answers;
- "on it", in a few words, if the agent hasn't answered a few seconds after you asked;
- an answer while the agent is busy: what you say then still goes to the agent (it waits behind the current work), and the voice answers at once, from its notes: how it's going, or that your request is next. When the agent gets to it, the voice passes on only what's new;
- a problem, a finished background task (if the agent doesn't report it itself), or progress while the agent works.
Measured with Gemini, the first sound comes about 0.5 to 0.8 seconds after the agent writes its reply or asks.
Its settings
On the Voice tab, under Realtime voice:
| Setting | What it does |
|---|---|
| Service | Gemini Live (the default) or OpenAI Realtime. A service without a key says "(no key)". |
| API key | Paste a key, saved for this face, or use the environment variable. Once a key is there it shows where it came from, with Replace or Use another, and Remove for a saved one. |
| Model | The speech-to-speech models your key can use, as the service lists them. Other... takes any model name, and Refresh the list asks again. The defaults are gemini-3.8-live and gpt-realtime-2.1. |
| Voice | The service's voice: Puck by default for Gemini, marin for OpenAI. |
| Updates | Quiet: replies, questions and problems, and no "on it". Normal (the default): also "on it" after 6 seconds, and progress after about 25 seconds of silent work. Chatty: "on it" after 3 seconds, and progress about every 10 seconds. You can always ask. |
| Speaks as | The agent ("I'm running the tests", the default), or its voice ("Claude Code is running the tests"). |
| Fall back to classic voice | On (the default): if the realtime voice can't connect, the voice on your PC reads the replies until it's back. Off: you get a message saying why, and replies show as text. |
| Interrupt by talking | Off by default. On: talk over it to cut it off. Use headphones, because through speakers it would hear itself. Off: the microphone rests while it speaks (Esc stops it). |
The top of the group says how it's doing: connected, connecting, or the problem and that it keeps trying.
When it can't connect
If the service can't be reached (no key, no credit, no network), the face says so once, in plain words. With Fall back to classic voice on, the classic voice reads the agent's replies until the service is back; with it off, the replies show as text meanwhile.
Sessions with the service end after a while: Gemini's about every ten minutes, OpenAI's after an hour. The face reconnects at once, resuming the conversation where the service allows it (Gemini) or telling the new session the latest notes (OpenAI).
Rooms always use the classic voice: the Realtime voice group then says it's not running.
The classic voice
The classic voice reads the agent's own replies aloud. It's used when the realtime voice isn't (while it can't
connect, with the fallback on), in rooms, and for a face set to it with
voice.mode: "classic" in its profile.json. Such a face shows
a Use the realtime voice button on the Voice tab to switch back.
The voice on your PC is Pocket, or Kokoro with the TTS_ENGINE=kokoro environment variable (from a
clone, install Kokoro with scripts\setup.ps1 -WithKokoro). Optionally, ElevenLabs speaks
instead. On the Voice tab, under Classic voice:
- Voice engine: Automatic uses ElevenLabs whenever a key is available and the voice on your PC otherwise; This PC never uses ElevenLabs; ElevenLabs always tries it first.
- Voice (or Voice on this PC): the local voice.
- ElevenLabs: its API key (or
ELEVENLABS_API_KEY/XI_API_KEY), the account's Voice, the Quality (faster models start speaking sooner; Flash is the default) and Test voice.
If ElevenLabs fails (a bad key, quota, no network), you're told once and that sentence is spoken by the voice on your PC. While ElevenLabs is used, the text of replies is sent to it; the microphone never is.
Spoken replies off
Turn off Spoken replies on the Voice tab (or untick Voice replies in the menu) and replies show as captions, without any voice.
The speech models
Speech recognition, the voice on your PC, the name spotter, noise reduction and the lip sync all run on models that are downloaded once: about 650 MB, the first time Wireface starts. The Voice tab shows them under Hearing, Speech models: All installed, or what's missing with an Install speech models button and the download's progress.
All faces share them. They live in deps\models in the install folder, unless the
MODELS_DIR environment variable points somewhere else. Uninstalling removes them.
What is sent where
| Goes to | What |
|---|---|
| Nowhere: it stays on your PC | Your voice. Recognition, the name spotter and noise reduction all run locally. |
| Your agent | Your words, exactly as recognised or typed, as its prompt. |
| The realtime voice service | The text of what you said, short notes about the agent's work, and its replies. |
| ElevenLabs, if used | The text of the replies it speaks. |