Try it

Record something or upload a clip and the retrained model transcribes it here, in this page, without the audio leaving your machine. Four worked examples sit below if you would rather just see it.

Your own voice

The plan is for the model to run inside this page, so audio never leaves your machine. The browser build does not work yet, and saying so is cheaper than letting you find out: the quantised decoder I shipped is broken on its cache branch, so it produces the first token and then fails. The four worked examples below are real output from the same adapter, run properly, and they are the demo until the browser build is rebuilt.

Left in place rather than hidden, because the failure is diagnosed and the fix is known: the two decoder graphs were quantised separately and merged afterwards, which shrank the download from 739 MB to 186 MB and produced an invalid graph. Merging first and quantising after keeps it valid and keeps it large. That trade is the next piece of work.

Speak Awadhi or Hindi, for up to 15 seconds. This model learned from 8 kHz recordings of village speech, so a clean studio voice or a language it has never heard will make it loop on one syllable rather than transcribe. That failure is the thing the project is about, so the page names it when it happens instead of hiding it.

No Awadhi to record? Download any clip from the four examples below and upload it here. Those are the corpus recordings the model was trained against, so they are the fairest test of it.

Only the retrained model runs here. Carrying the stock model too would double the download to demonstrate a failure already shown above, with the recordings that produced it.

Four recordings

If you have no Awadhi to hand, these are from the held-out test set, picked across the range rather than from the wins: a total collapse, a solid gain, a small one, and one the adapter made worse.

Loading examples.