Pick an audio file and the selected engines run on it — VAD, both ASRs, diarization
(TTS engines synthesize a fixed sentence, since they consume text). All engines run
by default; untick to skip. Results download as JSON. First run downloads model
weights (cached after). ?engines=asr-parakeet,vad-silero preselects.