Loading the leaderboard…
Loading the leaderboard…
04 · Methodology
What every metric means, what each track allows, how much to trust a row, how to reproduce it — and what a dataset must clear to join the hub. Every number on this platform is a measurement; this page defines the measurement, identically for every dataset.
04.1 · Metrics
All metrics are computed by the platform harness from raw transcripts after normalization — WER-family values are percentages, lower is better unless marked otherwise.
Word error rate
WER = (S + D + I) / N × 100The headline metric: word-level edit distance against the human-verified reference, after normalizer v0.1. S, D, I are substitutions, deletions, insertions; N is reference words.
Character error rate
CER = (S + D + I) / N_chars × 100The same edit distance at character level — robust to tokenization choices and a fairer lens on Vietnamese diacritics.
Code-switch word error rate
WER over tokens ∈ code-switched spansWER restricted to code-switched English spans — the clinical failure mode this benchmark exists to expose. Vietnamese-specialized models win overall WER yet lose here.
Native word error rate
WER over Vietnamese-only tokensThe mirror image of CS-WER: WER over Vietnamese-only tokens with code-switched spans excluded. Multilingual models invert the pattern — strong here is rare when CS-WER is strong too.
Medical term recall
MTR = |terms recovered| / |terms in reference|Share of reference medical terms recovered in the hypothesis. The exact tier requires verbatim recovery; the fuzzy tier tolerates minor surface variation.
Real-time factor (inverted)
RTFx = audio seconds / wall-clock secondsThroughput during the platform run on the manifest’s declared hardware. Above 1.0 means faster than real time.
Headline metrics carry bootstrap 95% confidence intervals from utterance-level resampling, rendered as whisker micro-bars beside every value. The policy is strict: no SOTA badge without significance — a model is never promoted while its CI overlaps the incumbent’s.
Illustration — overlapping CIs, no badge awarded
Both reference and hypothesis pass through the same normalizer before scoring. The version is recorded in every run manifest.
WER is only comparable within a normalizer version. Cross-version numbers are never ranked against each other.
Frequency-bucketed MTR
Term recall split by training-frequency buckets — separates memorized vocabulary from genuine generalization.
Critical term error rate
Error rate over confusable drug-name pairs, where a substitution is a patient-safety event rather than a typo.
API latency & cost
p50/p90 latency and dollars per audio-hour for commercial APIs, captured during date-stamped snapshot runs.
04.2 · Tracks
A track defines exactly what a system may know about a dataset's audio and labels before it transcribes a single clip. Rows only compete within their track, on the same dataset.
| Track | Allowed signal | Status |
|---|---|---|
| AZero-shot | The model exactly as released. No dataset-specific signal of any kind — no prompts, no bias lists, no weight updates. | Live |
| BContextual biasing | Bias lists (medical-term hotwords or prompts) allowed at decode time. No weight updates; the bias list is recorded in the run manifest. | Live |
| CAdapted | Weights updated using the dataset's frozen Train split only; Valid for tuning. No other in-domain data of any kind. | Live |
| DOpen | Any external data or method, fully disclosed. The anything-goes track. | Roadmap |
| EEfficiency | Deployment-realistic constraints: under 1B parameters, CPU decode, under 8GB VRAM. | Config-ready |
| FConversational | WER + DER on multi-speaker dialogue. Requires conversational corpora (VietMed, clinic data) — the track closest to production. | Roadmap |
04.3 · Provenance
Every row declares where its numbers came from, and the chip follows the number everywhere it appears.
Run by the platform harness, on platform hardware, from a pinned model revision. The full run manifest is published beside the number. The only class eligible for a SOTA badge.
Submitted by the system’s authors with a manifest, reproduced on their hardware rather than ours. Honor-system until the private canary subset ships.
Run through the platform harness against the live commercial API on a given date. We hold the hypotheses and bootstrap CIs, but the service has no pinned revision and runs on the vendor’s hardware — so RTFx is omitted and the row is never SOTA-eligible. Re-run on a quarterly drift schedule.
Transcribed from a published paper. Normalizer and decode settings may differ from ours, so these rows wear a dashed outline everywhere they appear.
Row states — not trust classes
A platform run is executing right now. Metrics appear the moment the run completes and is verified.
Queued for a platform run. No numbers yet — the row exists so the roadmap is legible.
API systems change behind stable names, so their rows are date-stamped snapshots of the service on the day of the run. Quarterly drift re-runs are on the roadmap; superseded snapshots stay visible with their dates instead of being overwritten.
04.4 · Reproducibility
A result you cannot reproduce is an anecdote. Platform-verified runs publish the complete manifest below.
run_manifest.json — published with every platform-verified run
{
"run_id": "2026-06-09T21:14:02Z-a41f",
"model": { "id": "vinai/PhoWhisper-small", "revision": "84f3c1d" },
"harness": "medasr-bench v0.4",
"normalizer": "v0.1",
"dataset": { "id": "vimedcss", "split": "test" },
"hardware": "1x RTX 4090 24GB / CUDA 12.4",
"seeds": { "global": 1337 },
"decode": { "beam_size": 5, "temperature": 0.0, "bias_list": null }
}04.5 · Onboarding
MedASR Bench is a multi-dataset hub: any Vietnamese medical speech corpus clearing six criteria mounts into the same harness and becomes directly comparable. ViMedCSS cleared this bar first; VietMed and the private clinic set are next.
Redistribution rights documented up front — an open license, or a signed contract for held-out corpora like the private clinic set.
Where every clip came from, how transcripts were produced, and who verified them — published as part of the dataset profile.
Train, Valid, Test (and Hard where defined) frozen before the first run. Any revision creates a new dataset version — never a silent edit.
Systems that touched the labels — for example LLM-seeded transcripts — are declared and flagged permanently on that dataset's leaderboards.
References must score cleanly under the platform normalizer; orthographic quirks are documented and resolved before the first run.
WER, CER, CS-WER, N-WER, MTR, and RTFx must be computable from the references — or the gaps declared on the dataset profile.