Voice conversation test

Press Talk, speak, press it again. Your words go to Whisper, the text goes to the language model, and the reply is spoken in the voice you pick. Every delay is measured from the moment you press stop. Whisper, the speech models and the language model all run on the GPU box at home; nothing leaves the house.

Speech model
Changing the speech model reloads the GPU box (about 15 to 35 s).
Earlier first sound, mostly for VoxCPM2; the phrasing can suffer.
Loading models…
The browser asks once for the microphone. The first reply after a pause can be slow while the language model loads; the page asks it to load as soon as you press Talk.

Conversation

Nothing said yet.