This system takes inputs streaming audio in one language and outputs streaming translated text and audio in another language. This document outlines it works, and what the main knobs, decisions, and tradeoffs are.
A cascaded pipeline of three models, with chunking and decision points
Two kinds of decision drive the pipeline, and they are made by different things:
In "fast" or "balanced" mode, we can start translating as the user is speaking. In "whole turn" mode, we wait for the STT to signal the end of a turn. Both systems are described below.
Analyze rolling transcriptions and chunk based on minimum word threshold or clause boundary.
Wait until the STT engine signals end of turn and then send the confirmed utterance through the pipeline
| component | measured | notes |
|---|---|---|
| STT partial updates | ~200 ms apart | Median 217 ms between partials during speech. How closely the transcript tracks the speaker. |
| STT end-of-turn confirmation | 600–1500 ms | The turn model deciding the turn is over — signalled from intonation and context, capped around 1.5 s after speech stops. Some models add a provisional early signal worth a few hundred ms. |
| LLM translation (per fragment) | 100–200 ms | Full completion for a clause-sized fragment (short outputs, so streaming would not help; Whole turn does stream its longer one). |
| Fish TTS Synthesis | ~150 ms | Socket is pre-warmed on session opening. Time is just flush -> audio |
| Speaking rate | 1.0–1.15× | At Fast the voice runs 1.15× to buy back the duration chunked synthesis adds. |
In the app this is the Pacing choice on the setup card: Fast (default), Balanced, Whole turn. Here's how they would differ in an example sentence.
| pacing | first translated word | synthesis seams |
|---|---|---|
| Fast (0.0) | 2396 ms after speech start | 5 |
| Balanced (0.5) | 4130 ms after speech start | 3 |
| Whole turn (1.0) | 1708 ms after speech stopped | 2 |
One position moves every timing knob at once. The knobs only make sense together, so they are not exposed separately:
| knob | Fast | Whole turn | what it governs |
|---|---|---|---|
| chunkWords | 0.6× base | 2.0× base | source gathered before a boundary-less cut, scaled by the target language's chunk base (below), so Japanese stays proportionally patient at every position |
| maxDelayMs | 200 ms | 900 ms | how long translated text may wait for clause punctuation before flushing anyway |
| minChunkWords | 2 | 5 | smallest fragment worth a round trip and a seam |
| speed | 1.15× | 1.0× | playback rate. Fast buys back the duration its extra seams add |
| contextPairs | 1 | 4 | prior source/translation pairs given to the translator for consistent terminology |
| chunking | live | sentence | at the top of the axis the pipeline stops speaking mid-turn entirely; that end-stop is the second race bar above |
The language logic lives in the chunk base. Every target locale carries one, set by its word order. SVO targets (Spanish, English, Chinese: base 8 words) can be spoken clause fragment by clause fragment. Verb-final targets (Japanese: base 24 words; German, Korean, Turkish, Hindi similar) need the whole clause before the verb can be placed, so their caps are safety valves rather than intended cuts.
Six scenarios: five at Fast, one at Whole turn. Step with the buttons or ← → arrow keys. White is locked at the voice; amber is still theoretically revisable.