Skip to content

reproducible product evidence · 2026-08-02

Caption Plug local Whisper language benchmark

A deliberately small smoke test designed to catch script failures, obvious model mismatch, and languages that need heavier correction before we publish search pages.

method

Enough to reject failure—not enough to promise population accuracy

Three held-out FLEURS speech clips per language, explicit Whisper language code, normalized Unicode character edit distance. This is an engineering smoke benchmark, not a population-level accuracy claim.

Character error rate is the Unicode character edit distance divided by the reference length after normalization. Lower is better. It captures wording and script failures but does not independently score timestamp quality, typography, semantic severity, or readability.

Model
Whisper Base multilingual, revision 5359861c739e955e79d9a303bcbc70fb988958b1
SHA-256
60ed5bc3dd14eea856493d334349b405782ddcaf0028d4b5df4088345fba2efe
Runtime
v1.8.6-captionplug.2
Product
Caption Plug v1.1.3; Premiere Pro 2026
Speech dataset
Google FLEURS, CC-BY-4.0

all 15 tested languages

Results and publication decision

languagemean CERworst clipdecision bandevidence page
English · English1.0%2.9%Strong smoke-test resultView sample and QA notes
Spanish · Español1.6%2.4%Strong smoke-test resultView sample and QA notes
German · Deutsch7.4%16.3%Strong smoke-test resultView sample and QA notes
Indonesian · Bahasa Indonesia8.7%12.5%Strong smoke-test resultView sample and QA notes
Turkish · Türkçe9.4%13.1%Strong smoke-test resultView sample and QA notes
Polish · Polski10.9%16.0%Usable with close reviewView sample and QA notes
Portuguese · Português11.3%17.9%Usable with close reviewView sample and QA notes
French · Français12.4%21.7%Usable with close reviewView sample and QA notes
Korean · 한국어15.6%27.6%Usable with close reviewView sample and QA notes
Bulgarian · Български16.4%20.5%Usable with close reviewView sample and QA notes
Vietnamese · Tiếng Việt19.3%22.8%Usable with close reviewView sample and QA notes
Arabic · العربية21.1%38.4%Usable with close reviewView sample and QA notes
Japanese · 日本語23.6%46.4%Usable with close reviewView sample and QA notes
Urdu · اردو25.7%37.0%Limited; native review requiredView sample and QA notes
Hindi · हिन्दी90.1%95.5%Not approvedNo indexable page

threshold

Strong

Mean CER below 10% on this three-clip test. Still requires correction for names, punctuation, and real production audio.

threshold

Review

Mean CER from 10% through 25%. Publish with the observed output and language-specific correction guidance visible.

threshold

Limited / rejected

Above 25% is limited; catastrophic script mismatch is not approved. Urdu is retained with a severe warning. Hindi gets no landing page.

repeatability and limits

What to reproduce next

The raw report preserves every source audio URL, reference transcript, observed output, duration, row index, and per-clip score. A larger benchmark should stratify dialect, speaker, noise, microphone, music, overlap, and domain vocabulary, then separately score word timing and correction time. This report does not claim superiority over another transcription product.