Which AI Model Actually Follows a Hindi Script Brief? (2026)
An independent February 2026 benchmark scored GPT-5, Gemini 3, Gemma 3, Llama and Qwen 3 on verifiable instruction-following in Hindi. The rankings change completely depending on whether the brief was written in Hindi or translated into it.
Which AI Model Actually Follows a Hindi Script Brief? (2026)
By Ashok Sachdev, Founder of JustShoot · Published 28 September 2026 · Sources captured 28 September 2026
Short answer: On the only independent benchmark that publishes a per-model Hindi score, GPT-5 leads on both kinds of Hindi brief — 92.8% when the brief is an English instruction translated into Hindi, 77.9% when it was written in Hindi in the first place. The interesting part is what happens below it: on natively-written Hindi briefs, Gemini-3-Pro scores 58.9% and is beaten by Llama-3.2-1B-Instruct at 60.8% — a model roughly a thousandth of its size. The ranking you get depends on how the brief was written, not just on which model you picked.
That is a claim with a paper behind it, so here is the paper first, then what it means for a script brief.
This post is not a tool ranking. If you want the tools compared, we already did that in the best AI script writers for Hindi YouTube channels. This one is about the models underneath — named versions, published scores, and dates.
What was actually measured
The source is IndicIFEval: A Benchmark for Verifiable Instruction-Following Evaluation in 14 Indic Languages, arXiv:2602.22125v1, submitted 26 February 2026 by Thanmay Jayakumar, Mohammed Safi Ur Rahman Khan, Raj Dabre, Ratish Puduppully and Anoop Kunchukuttan.
"Verifiable" is the word that matters. This benchmark does not ask a panel whether the Hindi sounded good — that is unfalsifiable, and it is what most model comparisons actually do. It asks whether the model obeyed constraints a script can check: reply in Devanagari, keep it under four sentences, include this word at least eight times, end with this exact phrase, wrap the output in this format. The metric is prompt-level loose accuracy — the percentage of prompts where every constraint was satisfied. All models were run with greedy decoding through the Language Model Evaluation Harness.
The models are versioned, which is the second reason to use this source. Proprietary: openai/gpt-5-2025-08-07, openai/gpt-5-mini-2025-08-07, gemini-3-pro, gemini-3-flash. Open-weight: Gemma 3 (1B/4B/12B/27B), Llama 3.1 (8B/70B), Llama 3.2 (1B/3B), Llama 3.3 70B, Llama 4 17B-MoE, Qwen 3 (0.6B–32B) and Aya-Expanse (8B/32B).
And there are two suites, which is the third and most useful reason:
- INDICIFEVAL-TRANS — English IFEval prompts translated into each language, then verified by native speakers, cut down to a strictly parallel common subset of 321 prompts per language. This suite has an English column, so it can measure the gap against English.
- INDICIFEVAL-GROUND — prompts written natively in the Indic language rather than translated. For Hindi, annotators passed 457 of 527 candidate prompts. There is no English column, because there is no English original.
Those two suites are the difference between "can this model handle Hindi that started life as English" and "can this model handle Hindi that started life as Hindi." A creator typing a brief into a box is doing the second thing.
Hindi, translated briefs (INDICIFEVAL-TRANS)
Prompt-level loose accuracy on the Hindi (hi) column, with each model's own English (en) column beside it for reference. These are the printed table values, not averages I computed:
| Model (as named in the paper) | Hindi | English |
|---|---|---|
openai/gpt-5-2025-08-07 |
92.8 | 95.0 |
gemini/gemini-3-pro |
91.0 | 95.3 |
gemini/gemini-3-flash |
89.7 | 96.3 |
openai/gpt-5-mini-2025-08-07 |
88.8 | 91.9 |
google/gemma-3-27b-it |
85.0 | 88.2 |
google/gemma-3-12b-it |
81.9 | 86.0 |
meta-llama/Llama-3.3-70B-Instruct |
79.4 | 91.6 |
Qwen/Qwen3-14B |
78.8 | 87.8 |
Qwen/Qwen3-32B |
78.5 | 87.2 |
meta-llama/Llama-4-17B-Instruct |
77.3 | 89.4 |
google/gemma-3-4b-it |
75.1 | 81.9 |
CohereLabs/aya-expanse-32b |
70.4 | 77.3 |
meta-llama/Llama-3.2-1B-Instruct |
27.1 | 54.2 |
On this suite Hindi looks healthy. The paper says so in as many words: "Hindi remains the most robust language across all model families," and the Gemma family "consistently exhibits the smallest performance gap with English across all languages (e.g., ∆ < 0.15 for Hindi and Bengali)." Among open-weight models, Gemma-3-27B-IT wins overall, followed by Gemma-3-12B-IT and Llama-4-Scout-17B-16E-Instruct.
One warning that comes straight out of this table. Aya-Expanse-32B scores 70.4 on Hindi but its average across the 14 Indic languages is 38.8 — its Hindi is roughly thirty points above its own Indic average, because, as the paper puts it, the Aya family "performs moderately well for Hindi, Bengali, Urdu, and Tamil" while exceeding 0.40 disparity on most low-resource languages. If you average a model across Indian languages you get a number that describes none of them. That is why every figure on this page is a single language's own column.
Hindi, briefs written in Hindi (INDICIFEVAL-GROUND)
Same metric, same Hindi column, different suite — these prompts were authored in Hindi rather than translated:
| Model | Hindi |
|---|---|
openai/gpt-5-2025-08-07 |
77.9 |
openai/gpt-5-mini-2025-08-07 |
75.5 |
Qwen/Qwen3-14B |
73.1 |
meta-llama/Llama-3.1-8B-Instruct |
65.2 |
meta-llama/Llama-3.2-3B-Instruct |
64.8 |
Qwen/Qwen3-8B |
64.3 |
google/gemma-3-4b-it |
64.1 |
Qwen/Qwen3-4B |
63.7 |
google/gemma-3-12b-it |
63.5 |
Qwen/Qwen3-32B |
62.4 |
gemini/gemini-3-flash |
61.9 |
meta-llama/Llama-3.2-1B-Instruct |
60.8 |
google/gemma-3-27b-it |
60.0 |
gemini/gemini-3-pro |
58.9 |
CohereLabs/aya-expanse-32b |
54.9 |
GPT-5 stays first. Almost nothing else stays where it was. Gemini-3-Pro goes from second on translated Hindi to fourteenth of fifteen here. Gemma-3-27B, the best open-weight model on the translated suite, finishes below Gemma-3-4B. Qwen3-14B, mid-table before, is now third and ahead of every Gemini and Gemma entry.
The subtraction you are not allowed to do
It is tempting to read 91.0 → 58.9 as Gemini-3-Pro "losing 32 points on real Hindi." Do not. The two suites are different prompt sets with different constraint mixes and different difficulty — TRANS is a 321-prompt parallel subset, GROUND is a separately generated and separately verified set. The clearest proof that the gap is not a single directional effect is Llama-3.2-1B-Instruct, which goes the other way: 27.1 on translated Hindi, 60.8 on natively-written Hindi.
So the defensible claim is about ranking within each suite, not the delta between them. And within the natively-written suite, the ranking is genuinely surprising, which is the finding worth carrying: the model that tops a translated-Hindi leaderboard is not reliably the model that tops a written-in-Hindi one.
There is a per-language version of the same point. On the GROUND suite, Gemini-3-Pro's Hindi score of 58.9 is its second-lowest of all 14 languages, beaten only by Sanskrit at 50.9 — while its Telugu is 70.4 and its Odia 67.9. On the translated suite, Hindi was one of its strongest. "Good at Indian languages" is not a property a model has.
The one brief shape that fails hardest
The paper groups IFEval's 25 constraints into five categories — ENGLISH, KEYWORD, FORMAT, LENGTH and LANGUAGE — and reports the Indic-minus-English disparity (∆) for each. For Hindi, ∆ sits at 0.11–0.16 across most categories. In one category it jumps to 0.28: prompts written in Hindi that instruct the model to answer in English. The Llama and Aya families degrade past ∆ > 0.50 on that same category.
The paper's own manual inspection explains why: the model either keeps replying in Hindi and fails the English-output constraint, or switches to English and drops the other constraints on the way.
For a creator that is a concrete instruction. Do not write your brief in Hindi and ask for the English deliverables in the same request. Hindi script, then a separate request for the English title, description and tags. Two prompts, in the shape the models are measurably best at, instead of one prompt in the shape they are measurably worst at.
The KEYWORD category carries a second usable finding. Across all model families, models are good at avoiding a word when told to and significantly worse at including one a given number of times. "Never say 'guys'" is a constraint these models keep. "Say the channel name three times" is one they quietly drop — so check it yourself rather than trusting it.
Two smaller results worth knowing. Turning on Qwen3's thinking mode narrows the English-vs-Indic gap consistently across every model size tested. And scaling stops helping early: the paper finds performance "plateaus around the 12B to 14B parameter range with limited gains thereafter (< 5)." A bigger model is not the fix for Hindi instruction-following.
None of this tells you whether a model writes Hindi that sounds like you — instruction-following and voice are different problems. If the second one is what you are stuck on, our free tone fingerprint tool reads a transcript you have already published and reports your own Devanagari-to-Latin ratio, sentence rhythm and hook pattern, which is the thing you would otherwise be asking a model to guess.
What about the India-built models?
The obvious question: where is Sarvam, and where are the other Indian foundation models?
Not in this benchmark. Which leaves their own published numbers, so here is what they actually say. Sarvam AI's open-sourcing announcement (sarvam.ai, 6 March 2026) covers Sarvam-30B and Sarvam-105B, the latter a Mixture-of-Experts model with 10.3B active parameters, a 128K context via YaRN scaling, and an Apache-2.0 licence — with a custom tokenizer covering 22 scheduled Indian languages across 12 scripts, which genuinely matters for the cost of serving Devanagari.
Its reported scores are Math500 98.6, MMLU 90.6, AIME 25 88.3, GPQA Diamond 78.7, LiveCodeBench v6 71.7. Its Indic claim is that Sarvam-105B "wins on average 90% across all benchmarked dimensions" on an internal benchmark covering 22 languages in both native and romanised script. The Hugging Face card for sarvamai/sarvam-105b (accessed 28 September 2026) lists "IF Eval" at 84.8 inside a Knowledge & Coding table, not identified as an Indic variant.
So: a pairwise win-rate across 22 languages, no Hindi row, and no native-versus-romanised split published separately. That is not a criticism of the model — it is a statement about what you can and cannot compare. There is no axis on which Sarvam-105B and GPT-5 can be placed side by side on Hindi instruction-following today, and anyone who shows you one has built it themselves.
An August 2026 survey makes the same point more generally. Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models (arXiv:2608.11891, 12 August 2026) notes that "Sarvam AI reports the broadest coverage by a substantial margin" among Indian organisations, while observing that "Indian models achieve strong scores on established benchmarks such as MMLU and MATH-500. However, these are now widely regarded as saturated, and frontier developers no longer report them," and that Indian models "participate far less frequently in newer, agentic, and domain-specialized evaluations."
Sarvam is the most transparent Indian entrant and is still not comparable on this axis. That is the state of the evidence in September 2026, and it will date — which is the point of putting a date on it.
What this post does not do
We did not run these models ourselves. There is no self-run head-to-head here, no in-house scoring rubric, and no "we gave each model the same brief and judged the output" table — because a comparison scored by a company that sells a script tool is not evidence, whichever way it comes out. Every number above is from a named, dated, third-party source you can open and check.
One disclosure in one sentence: we sell a tool built on top of models like these, so treat our reading of the numbers with that in mind and check the tables yourself.
Two honest limits. First, verifiable instruction-following is not the same as writing quality — a model can obey every constraint and still produce flat Hindi, and the translationese problem is a separate one we cover in why AI Hindi scripts read as translated. Second, this is a snapshot of specific model versions in February 2026. Our own earlier hands-on comparison, the best AI for Hinglish script writing, tested GPT-4o, Claude 3.5 and Gemini 2.0 — a full model generation ago. A post like this is wrong within about a quarter and should say so.
Where JustShoot sits
JustShoot is not a model. It is a nine-agent pipeline that sits on top of one, carrying your channel profile and tone fingerprint through research, scripting, fact-check, storyboard, thumbnail, SEO, Shorts and distribution — which is the part the benchmark above cannot measure, because a benchmark has no memory of your channel.
Pricing, since people ask: Trial ₹0 (7 days, 2 scripts lifetime, no card) · Starter ₹499/mo (3 scripts) · Creator ₹999/mo (7 scripts, most popular) · Pro ₹1999/mo (15 scripts) · Studio custom. That is a fixed number of scripts a month with no rollover, GST included, monthly only. At Creator that works out to about ₹143 a script; at Pro, about ₹133. Every plan gets the full pipeline — the only difference is how many scripts. Details on the pricing page.
FAQ
Which AI model is best for writing Hindi YouTube scripts in 2026?
On the only independent benchmark publishing a per-model Hindi score, GPT-5 (gpt-5-2025-08-07) leads both suites — 92.8% on translated Hindi briefs and 77.9% on briefs written in Hindi (IndicIFEval, arXiv:2602.22125v1, 26 February 2026). But that measures constraint-following, not whether the Hindi sounds like a person, and it is a February 2026 snapshot of specific versions.
Is Gemini 3 bad at Hindi? Not straightforwardly. Gemini-3-Pro scores 91.0% on Hindi briefs translated from English — second only to GPT-5. On briefs written natively in Hindi it scores 58.9%, its second-lowest of the 14 languages tested after Sanskrit. Those are two different prompt sets, so the two numbers are not a before-and-after, but the ranking difference within each suite is real.
Are Indian models like Sarvam better at Hindi than GPT-5? There is no published comparison that answers this. Sarvam-105B is not in IndicIFEval, and Sarvam's own announcement (6 March 2026) reports a 90% pairwise win-rate across 22 Indian languages rather than a Hindi-specific score, with no separate native-versus-romanised breakdown. Anyone claiming a head-to-head has constructed it themselves.
Should I write my prompt in Hindi or in English? Write it in the language you want back. The worst-measured shape is a Hindi prompt asking for English output: for Hindi the disparity against English jumps to ∆ = 0.28 in that category against 0.11–0.16 elsewhere, and the Llama and Aya families pass ∆ > 0.50. Ask for the Hindi script in one request and the English title, description and tags in another.
Does a bigger model write better Hindi? Not past a point. IndicIFEval finds performance "plateaus around the 12B to 14B parameter range with limited gains thereafter (< 5)," and on natively-written Hindi prompts Gemma-3-4B (64.1%) outscores Gemma-3-27B (60.0%) and Llama-3.2-1B (60.8%) outscores Gemini-3-Pro (58.9%). Enabling a thinking mode narrowed the Hindi-versus-English gap more reliably than adding parameters did.
Sources. IndicIFEval: A Benchmark for Verifiable Instruction-Following Evaluation in 14 Indic Languages, arXiv:2602.22125v1, 26 February 2026 (Tables 1 and 2, Hindi column; prompt-level loose accuracy). Sarvam AI, Open-Sourcing Sarvam 30B and 105B, sarvam.ai, 6 March 2026. Hugging Face model card sarvamai/sarvam-105b, accessed 28 September 2026. Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework, arXiv:2608.11891, 12 August 2026. All figures are the values printed in those sources for a single language; no score on this page is an average across languages or across models.
How Accurate Are YouTube's Auto-Captions in Hindi? (2026)
Google publishes no Hindi accuracy figure for YouTube captions. The nearest dated benchmarks put Hindi ASR between 14% and 60% word error rate, depending entirely on the dataset.
How to Check If Your YouTube Script Sounds AI-Generated (2026)
The same 91 human-written essays were scored by AI detectors twice: 61.22% were flagged as AI in 2023, 23.1% in 2026. Native-speaker essays scored 5.19% and 0.0%.
Best AI Malayalam YouTube Script Generator (2026)
Which AI writes Malayalam YouTube scripts natively — spoken register, Manglish code-mix, dialect held steady? A buying guide for Kerala creators, with Kerala's real connectivity numbers and five tests to run before you pay.