monty
Try Monty free
← All articles
Guide2026-09-30· 7 min read

Captions in low-resource languages: the accuracy gap

Whisper covers 99 languages. On clean English it errs on 2.7% of words; on low-resource languages that figure passes 25%. What that means for a caption workflow.

Max GPTMax GPTBusiness development at Monty

Automatic captions are treated as a solved problem because they are solved in English. Move to Kazakh, Uzbek or any language with a thin training corpus and the same model produces output you cannot publish unread. The numbers are worth knowing before you promise a client multilingual captions.

2.7% vs 25%+
word error rate for Whisper large-v3 on clean English against low-resource languages. The workflow, not the model, has to absorb that difference

What the research reports

  • Whisper supports 99 languages, Kazakh and Uzbek among them.
  • Major Western European languages land close to English accuracy; low-resource ones can exceed 25% word error rate.
  • For Kazakh, the large model reports a character error rate around 4, and a fine-tuned version reaches CER 3.39 with WER 14.5.
  • Published work shows more than 10 percentage points of absolute WER reduction from targeted fine-tuning with unpaired speech and text.
  • Yandex SpeechKit covers 16 languages including Kazakh and Uzbek, and exports subtitles directly.

Where the errors cluster

Usually correct

  • Common verbs and function words
  • Short declarative sentences
  • Single-language speech
  • Studio-clean audio

Usually wrong

  • Names and place names
  • Code-switching mid-sentence
  • Numbers, dates, units
  • Domain jargon

That distribution is useful, because it means proofreading is targeted rather than total. A thirty-second clip holds seventy to ninety words; at 15% error that is a dozen corrections, most of them in the categories above.

A workflow that survives the gap

  1. Fix the recording first. A lavalier or close mic cuts more error than any model choice.
  2. Set the language explicitly instead of relying on auto-detection, which code-switching defeats.
  3. Keep a glossary of names and terms so the same corrections are not retyped every clip.
  4. Proofread the transcript, not the burned-in captions: fixing text is minutes, re-rendering is not.
  5. Check line breaks against the vertical safe zone before export.
Two caption tracks on screen at once do not fit a vertical frame if you respect the safe zone, and you should. Put the second language in the description or ship a separate cut.

Choosing a model

OptionStrengthTrade-off
Whisper large-v399 languages, open weights, fine-tunableneeds proofreading on thin-corpus languages
Fine-tuned Whispermaterially lower error on the target languagerequires data and a training pass
Yandex SpeechKit16 languages, subtitle export, hostedclosed and paid per usage

How Monty handles captions

Monty syncs captions to speech, keeps them inside the safe area, cleans and levels audio, builds a version per platform and posts to YouTube Shorts, Instagram Reels, TikTok and Telegram. On low-resource languages the proofreading step stays human, and we would rather say that than imply flawless Kazakh captions out of the box.

FAQ

Does Whisper support Kazakh and Uzbek?

Yes, both are among its 99 languages. Accuracy is the issue, not coverage: low-resource languages can exceed 25% word error rate against 2.7% for clean English.

How accurate are Kazakh captions in practice?

Reported character error rates sit near 4 for the large model and 3.39 for a fine-tuned one, with a word error rate of 14.5. Roughly one word in seven needs checking.

Whisper or Yandex SpeechKit?

SpeechKit is hosted, covers 16 languages including both, and exports subtitles. Whisper is open and can be fine-tuned on your own vocabulary. Test both on your own audio.

What reduces errors the most?

Clean audio, an explicit language setting, a glossary of names and terms, and a proofreading pass on the transcript before rendering.

Can I show two languages at once?

Not inside a vertical safe zone. Use the description or publish a separate cut per language.

Sources

  1. 1.How accurate is Whisper in 2026? WER data by language
  2. 2.Fine-tuning Whisper for Kazakh speech recognition
  3. 3.Improving Whisper for under-represented language Kazakh
  4. 4.Yandex SpeechKit supported languages and models
  5. 5.Whisper word error rate index (Artificial Analysis)
Share with a friendTelegramEmail

Hey Monty. Make me a reel.

Try Monty free

Keep reading