Live Captions

True streaming ASR on your microphone — conformer caches carried chunk to chunk, nothing re-decoded, nothing leaves your machine. EOU segments utterances from the model's own end-of-utterance tokens; Nemotron adds a ~1.6s right-context lookahead (that delay is the model looking ahead, not lag).

Idle.