Method

How the adapter was trained, on what, and the four things the results do not tell you.

How it was built

CorpusSpeeD-IA Awadhi: 2,538 usable utterances, 3 hours 14 minutes, 18 speakers, split by the corpus authors into 2,029 for training and 509 held out.
Base modelopenai/whisper-small, 244 million parameters. Whisper has no Awadhi language id, so the hint is pinned to Hindi for every run, which caps what either model can reach.
AdapterLoRA on the attention projections only, rank 16, learning rate 1e-4, six epochs. 8.7 MB of trained weights.
HardwareOne RTX 4050 laptop GPU with 6 GB. 52 minutes, peaking at 3.7 GB. Batch 8 with gradient checkpointing beat both smaller batches and checkpointing turned off.
The browser buildThe model on the Try it page is quantised to int8 so it is 279 MB instead of 1.76 GB. That costs accuracy: on the 20 shortest held-out utterances the full-precision adapter scores 0.2553 character error and the browser build scores 0.2943, against 0.6170 for the stock model. It is a worse model than the one the Evidence page measures, and still a far better one than Whisper alone.
Scoringjiwer, corpus level rather than an average of per-utterance rates. Reference and hypothesis are normalised identically before scoring, and the harness is verified against known answers before it reports anything.

What these numbers do not say

The recordings are 8 kHz. Everything above 4 kHz was never captured, which is where much of the energy that separates fricatives lives. That ceiling applies to both models equally, so the comparison holds while the absolute figures stay low.
The same speakers appear in training and testing. The corpus authors split by utterance rather than by speaker. These results describe adaptation to known voices and say nothing about an unheard one.
Long answers improved least. Short prompted sentences gained roughly four times as much as the long spontaneous narratives, which had a fifth of the training data. The narratives are the part this page is built from.
Twenty of 509 utterances got worse. The stock model collapses into repeating one syllable until it runs out of room. Retraining mostly stops that. Mostly.