Strong
Mean CER below 10% on this three-clip test. Still requires correction for names, punctuation, and real production audio.
A deliberately small smoke test designed to catch script failures, obvious model mismatch, and languages that need heavier correction before we publish search pages.
Three held-out FLEURS speech clips per language, explicit Whisper language code, normalized Unicode character edit distance. This is an engineering smoke benchmark, not a population-level accuracy claim.
Character error rate is the Unicode character edit distance divided by the reference length after normalization. Lower is better. It captures wording and script failures but does not independently score timestamp quality, typography, semantic severity, or readability.
5359861c739e955e79d9a303bcbc70fb988958b160ed5bc3dd14eea856493d334349b405782ddcaf0028d4b5df4088345fba2efe| language | mean CER | worst clip | decision band | evidence page |
|---|---|---|---|---|
| English · English | 1.0% | 2.9% | Strong smoke-test result | View sample and QA notes |
| Spanish · Español | 1.6% | 2.4% | Strong smoke-test result | View sample and QA notes |
| German · Deutsch | 7.4% | 16.3% | Strong smoke-test result | View sample and QA notes |
| Indonesian · Bahasa Indonesia | 8.7% | 12.5% | Strong smoke-test result | View sample and QA notes |
| Turkish · Türkçe | 9.4% | 13.1% | Strong smoke-test result | View sample and QA notes |
| Polish · Polski | 10.9% | 16.0% | Usable with close review | View sample and QA notes |
| Portuguese · Português | 11.3% | 17.9% | Usable with close review | View sample and QA notes |
| French · Français | 12.4% | 21.7% | Usable with close review | View sample and QA notes |
| Korean · 한국어 | 15.6% | 27.6% | Usable with close review | View sample and QA notes |
| Bulgarian · Български | 16.4% | 20.5% | Usable with close review | View sample and QA notes |
| Vietnamese · Tiếng Việt | 19.3% | 22.8% | Usable with close review | View sample and QA notes |
| Arabic · العربية | 21.1% | 38.4% | Usable with close review | View sample and QA notes |
| Japanese · 日本語 | 23.6% | 46.4% | Usable with close review | View sample and QA notes |
| Urdu · اردو | 25.7% | 37.0% | Limited; native review required | View sample and QA notes |
| Hindi · हिन्दी | 90.1% | 95.5% | Not approved | No indexable page |
Mean CER below 10% on this three-clip test. Still requires correction for names, punctuation, and real production audio.
Mean CER from 10% through 25%. Publish with the observed output and language-specific correction guidance visible.
Above 25% is limited; catastrophic script mismatch is not approved. Urdu is retained with a severe warning. Hindi gets no landing page.
The raw report preserves every source audio URL, reference transcript, observed output, duration, row index, and per-clip score. A larger benchmark should stratify dialect, speaker, noise, microphone, music, overlap, and domain vocabulary, then separately score word timing and correction time. This report does not claim superiority over another transcription product.