Fish Audio

Try it live

How live translation works

This system takes inputs streaming audio in one language and outputs streaming translated text and audio in another language. This document outlines it works, and what the main knobs, decisions, and tradeoffs are.

The stack

A cascaded pipeline of three models, with chunking and decision points

MicWebsocket streaming audio input
STTLanguage(s) can be specified or auto-detected
ChunkerDeterministic rules decide when to chunk the transcript
LLMModel translates or rejects
ChunkerDeterministic rules decide when to chunk the translation
Fish TTSLow latency websocket streaming

Two kinds of decision drive the pipeline, and they are made by different things:

Latency breakdown

In "fast" or "balanced" mode, we can start translating as the user is speaking. In "whole turn" mode, we wait for the STT to signal the end of a turn. Both systems are described below.

Fast/balanced

Analyze rolling transcriptions and chunk based on minimum word threshold or clause boundary.

wait for min-word/clause rule~1–2 s of speech
STT partial~200 ms
LLM~150 ms
Fish TTS150 ms

Whole turn

Wait until the STT engine signals end of turn and then send the confirmed utterance through the pipeline

STT confirm~600 ms
LLM~150 ms
TTS~150 ms
componentmeasurednotes
STT partial updates~200 ms apart Median 217 ms between partials during speech. How closely the transcript tracks the speaker.
STT end-of-turn confirmation600–1500 ms The turn model deciding the turn is over — signalled from intonation and context, capped around 1.5 s after speech stops. Some models add a provisional early signal worth a few hundred ms.
LLM translation (per fragment)100–200 ms Full completion for a clause-sized fragment (short outputs, so streaming would not help; Whole turn does stream its longer one).
Fish TTS Synthesis~150 ms Socket is pre-warmed on session opening. Time is just flush -> audio
Speaking rate1.0–1.15× At Fast the voice runs 1.15× to buy back the duration chunked synthesis adds.

Fast vs Whole turn

Fast: speak while they speak
you speak · 4.6 s
translated voice · first word 2.4 s in
0 s6 s12 s
Whole turn: wait until they stop
you speak · 4.6 s
translated voice · starts 1.7 s after you stop
0 s6 s12 s

Pacing controls

In the app this is the Pacing choice on the setup card: Fast (default), Balanced, Whole turn. Here's how they would differ in an example sentence.

pacingfirst translated wordsynthesis seams
Fast (0.0)2396 ms after speech start5
Balanced (0.5)4130 ms after speech start3
Whole turn (1.0)1708 ms after speech stopped2

One position moves every timing knob at once. The knobs only make sense together, so they are not exposed separately:

knobFastWhole turnwhat it governs
chunkWords0.6× base2.0× base source gathered before a boundary-less cut, scaled by the target language's chunk base (below), so Japanese stays proportionally patient at every position
maxDelayMs200 ms900 ms how long translated text may wait for clause punctuation before flushing anyway
minChunkWords25 smallest fragment worth a round trip and a seam
speed1.15×1.0× playback rate. Fast buys back the duration its extra seams add
contextPairs14 prior source/translation pairs given to the translator for consistent terminology
chunkinglivesentence at the top of the axis the pipeline stops speaking mid-turn entirely; that end-stop is the second race bar above

The language logic lives in the chunk base. Every target locale carries one, set by its word order. SVO targets (Spanish, English, Chinese: base 8 words) can be spoken clause fragment by clause fragment. Verb-final targets (Japanese: base 24 words; German, Korean, Turkish, Hindi similar) need the whole clause before the verb can be placed, so their caps are safety valves rather than intended cuts.

Step through a real turn

Six scenarios: five at Fast, one at Whole turn. Step with the buttons or ← → arrow keys. White is locked at the voice; amber is still theoretically revisable.

Speaker
STT partial
Fragments
Translation