How to add Hindi captions to a video (and why most tools get it wrong)

A practical walkthrough for captioning Hindi and Hinglish video — what accuracy to actually expect, whether to use Devanagari or romanized text, and the font problem that breaks Indic captions in the exported file.

The short answer

Upload your video to a tool that routes Hindi audio to a Hindi speech model rather than a general multilingual one, generate a caption draft, then review it with the video playing — checking names, brands and English loanwords first, because that is where almost all the errors are. Choose Devanagari or romanized script based on who is watching, not on which is more correct. Then export a video with the captions burned into the frames rather than a subtitle file, because social platforms will not reliably display a sidecar.

Expect to do real review work. Clean Hindi audio comes back around 85-90% of words correct, against roughly 95% for clean English. On a one-minute clip that is somewhere around fifteen words needing attention. Code-mixed Hinglish is harder than either and should be treated as a draft, not output.

The rest of this is why each of those choices matters, and what goes wrong when you get them wrong.

Why generic captioning tools fail on Hindi

Nearly every captioning tool is built English-first. One multilingual speech model handles every upload, which is cheap to run and works acceptably across Western European languages because those are well represented in training data. Hindi is not represented on remotely the same scale, and India's other languages far less so.

What you get is captions that are broadly about the right subject and wrong in every detail that matters. Names mangled. Verb endings dropped, which in Hindi can invert who did what to whom. English loanwords — of which spoken Hindi is full — transliterated into Devanagari that nobody would ever write.

There is a nastier failure underneath. A model uncertain about Hindi tends to fall back on what it knows best. So you get English-shaped output from Hindi audio, or Devanagari that is phonetically plausible and orthographically wrong. Both look like confident, finished text. An obviously broken caption gets fixed; a confidently wrong one gets published.

The fix is not a better single model. It is routing: detect the language, then send the audio to a model chosen for that language. It costs more to operate, which is exactly why most tools do not do it.

What accuracy to actually expect

Be suspicious of any captioning tool quoting a single accuracy number. Accuracy is a property of the audio as much as the model, and it varies enormously by language.

Clean English — close mic, one speaker, quiet room, no music — runs around 95% of words correct. Clean Hindi under the same conditions is roughly 85-90%. Those figures come from benchmarks run on studio-quality audio, so your actual recording will do worse than both. Background music under speech, two people talking over each other, a phone mic held at arm's length and a room with hard parallel walls each cost you accuracy, and they compound.

Hinglish deserves its own paragraph because it is the register most Indian creators actually speak, and it is genuinely the hardest case. Code-mixing is not a third language a model can be pointed at. It is two languages alternating mid-sentence and often mid-phrase, with English words taking Hindi grammar and Hindi words written in Latin script. The model has to decide, word by word, which language it is hearing and which script to write it in. Every one of those decisions can go wrong.

If a tool advertises clean Hinglish captioning with no review step, that claim is not being measured on real code-mixed audio. Plan for a read-through and you will be fine; plan for finished output and you will publish mistakes.

Devanagari or romanized? Pick by audience

Hindi captions can be written in Devanagari (हिंदी) or romanized into Latin script (Hindi). Neither is more correct — they serve different viewers.

Devanagari reads naturally to anyone Hindi-literate and is the right default for an audience in India consuming Hindi content. Romanized Hindi reaches a large group that speaks Hindi fluently but reads Latin script faster, which includes a lot of younger urban and diaspora viewers. It also keeps captions visually consistent when half your sentence is already English — a Hinglish line in mixed scripts is harder to read than the same line in one.

One technical point worth knowing: good romanization is produced from recognised Devanagari rather than guessed straight from the audio. Transcribe to Devanagari first, then transliterate. Doing it in one step stacks two uncertain operations and the errors multiply.

Review in the right order

Reading every caption with equal attention is slow and catches less than working through error categories. The mistakes are not randomly distributed.

Start with proper nouns — people, companies, places, product names. These are what the model has least evidence for and what your audience notices instantly. Then the English words inside Hindi sentences, since that is where code-switching most often breaks. Then numbers, which are high-consequence and easy to skim past. Then anything field-specific.

Do this with the video playing rather than in a transcript box. Reading a caption while hearing the audio under it is how you catch the errors; reading text in a separate window means context-switching to the video for every suspicious line, and the switching is the real cost of captioning by hand. It also surfaces a problem a transcript cannot show you at all: a caption that is perfectly correct but on screen too briefly to read.

The font problem that breaks Indic captions

This is the part nobody warns you about, and it is where Indic video captioning actually fails.

Devanagari requires correct glyph shaping. Consonants combine into conjunct clusters, vowel marks attach above and below the line, and a character's drawn shape depends on its neighbours. A renderer without a font that supports this does not raise an error. It silently substitutes a default font and draws empty boxes, or disconnected fragments, or — worst of all — text that looks plausible with the conjuncts quietly wrong.

The expensive version of this failure is the inconsistent one: correct in the browser preview, broken in the exported video. The cause is that the browser and the rendering server resolve fonts independently. Your browser has a font that handles Devanagari; the machine doing the encode does not, substitutes something, and you only find out after waiting for a render.

So when you evaluate a captioning tool for Indic languages, export a real test clip with Devanagari in it and look at the output file. Not the preview — the downloaded video. If the glyphs survive the round trip, the tool ships its own fonts and names them explicitly to the renderer instead of trusting whatever the host happens to have installed. If they do not, no amount of transcription accuracy will save you.

Export a video, not a subtitle file

Finally: get a video out, with the words in the frames. A .srt file alongside your clip is the technically tidier answer and it does not survive contact with social platforms. A re-encoded upload, a download-and-reshare, an embedded player, a messaging-app preview — none of them carry your sidecar, and feeds autoplay muted, so the window that decides whether anyone watches is silent by default.

Burned-in captions, also called open captions, are drawn into the pixels. They cannot be turned off and cannot be stripped. For Hindi content this matters doubly, because a platform that does support subtitle files may still render Devanagari badly in its own player, and burning in takes that decision away from it.

Questions people ask

Can I add Hindi captions to a video automatically?

Yes, but choose a tool that routes Hindi audio to a model selected for Hindi rather than one general multilingual model. Expect around 85-90% word accuracy on clean Hindi audio and plan a review pass, concentrating on names, brands and English loanwords.

Should Hindi captions be in Devanagari or Latin script?

It depends on your audience, not on correctness. Devanagari suits Hindi-literate viewers. Romanized Hindi in Latin script suits viewers who speak Hindi but read Latin faster, and keeps mixed Hindi-English lines visually consistent.

Why do Hindi captions show boxes or broken letters in my exported video?

The renderer lacked a font supporting Devanagari glyph shaping and silently substituted a default. It fails without an error, which is why it often looks fine in a browser preview and broken in the exported file — the two resolve fonts independently. Always check a real exported clip, not just the preview.

How accurate are automatic Hinglish captions?

Less accurate than either Hindi or English alone, because the model must decide language and script word by word through a code-mixed sentence. Treat the output as a draft that saves you the typing rather than as publishable text.

Do I need a subtitle file or a captioned video?

For social platforms, a captioned video. Subtitle files are stripped by re-encoding, re-uploads and downloads, and feeds autoplay muted. Burned-in captions are part of the picture and are always visible.