4,900 UI translations from our June 2026 pipeline, AI-judged by a three-model panel across nine language families. Methodology, full results, and the languages where we were weakest.
Update, September 2026: these results come from our June 2026 pipeline. We've since changed models; treat them as June's results.
We aim for translations that read like native UI copy, with your interpolation variables intact, in every language you ship. An aim like that deserves numbers. This report documents how we measured it in June 2026 — methodology first, results second, and the parts that didn't go perfectly included, because a benchmark that only reports good news isn't a benchmark.
TL;DR (June 2026 pipeline, AI-judged): 4.96/5 overall quality across 49
languages spanning nine language families, scored by a panel of three
frontier-model judges. Variable preservation — do {placeholders} survive
translation? — scored 5.0/5 in all 49 languages. The weakest language
(Hindi) still scored 4.81; the spread across the entire catalog is 0.19.
Polyglot's production translation pipeline as it ran in June 2026: the same prompt structure, the same strict JSON output schema, the same batching. No benchmark-only configuration. (We don't disclose the underlying model — it's one input to a pipeline we tune as a whole.)
100 English UI strings across 12 categories, designed to mirror what production codebases actually contain:
Welcome back, {name}!, Showing {start} to {end} of {total} resultsThe full dataset ships in our open-source repository, along with every June translation, so you can re-judge the outputs behind every number below.
49 targets chosen to span nine language families — including the scripts and structures where machine translation traditionally degrades (RTL, agglutinative morphology, low-resource South Asian languages):
| Family | Languages |
|---|---|
| Germanic | Afrikaans, Danish, German, Icelandic, Norwegian Bokmål, Dutch, Swedish |
| Romance | Catalan, Spanish, French, Italian, Portuguese, Portuguese (Brazil), Romanian |
| Slavic | Bulgarian, Czech, Croatian, Polish, Russian, Slovak, Slovenian, Serbian, Ukrainian |
| Baltic, Finnic & Hellenic | Greek, Estonian, Finnish, Hungarian, Lithuanian, Latvian |
| East Asian | Japanese, Korean, Chinese (Simplified), Chinese (Traditional) |
| South Asian | Bengali, Hindi, Marathi, Tamil, Telugu, Urdu |
| Middle Eastern (RTL) | Arabic, Persian, Hebrew |
| Southeast Asian | Indonesian, Malay, Thai, Tagalog, Vietnamese |
| Turkic & African | Turkish, Swahili |
That's 4,900 translations per scoring pass.
Every translation was graded 1–5 on three dimensions:
{name}, {{count}}, %s preserved
exactly? (A mistranslated placeholder isn't a style problem; it's a
runtime crash.)Scoring was done by LLM judges from different model families than the June engine, to reduce the documented self-preference bias where models rate their relatives' output higher. We used a three-judge panel and scored the full 4,900-translation set with each judge in a separate pass:
The primary judge's per-string scores are what we publish. The other two re-scored the entire set from scratch as a cross-check — a result that holds up across three judges is worth more than any one judge's opinion.
Aggregate (June 2026 pipeline, AI-judged): 4.95 accuracy · 4.94 fluency · 5.00 variable preservation → 4.96/5 overall.
What a 4.96 average actually means, in plain percentages across all 4,900 translations: 94.5% received the top score (5/5) on every dimension. 99.1% needed at most minor stylistic polish (nothing below 4 on any dimension). And 100% preserved every interpolation placeholder exactly — the failure mode that turns into a runtime bug, not just an awkward phrase.
| Language | Accuracy | Fluency | Variables | Overall |
|---|---|---|---|---|
| Afrikaans | 4.97 | 4.97 | 5.00 | 4.98 |
| Danish | 4.81 | 4.78 | 5.00 | 4.86 |
| German | 4.94 | 4.95 | 5.00 | 4.96 |
| Icelandic | 4.97 | 4.97 | 5.00 | 4.98 |
| Norwegian Bokmål | 4.98 | 4.99 | 5.00 | 4.99 |
| Dutch | 4.98 | 4.94 | 5.00 | 4.97 |
| Swedish | 4.93 | 4.93 | 5.00 | 4.95 |
| Catalan | 5.00 | 5.00 | 5.00 | 5.00 |
| Spanish | 4.97 | 4.98 | 5.00 | 4.98 |
| French | 4.93 | 4.93 | 5.00 | 4.95 |
| Italian | 4.94 | 4.94 | 5.00 | 4.96 |
| Portuguese | 4.99 | 4.99 | 5.00 | 4.99 |
| Portuguese (Brazil) | 4.98 | 4.99 | 5.00 | 4.99 |
| Romanian | 4.96 | 4.97 | 5.00 | 4.98 |
| Bulgarian | 4.83 | 4.79 | 5.00 | 4.87 |
| Czech | 4.93 | 4.81 | 5.00 | 4.91 |
| Croatian | 4.97 | 4.95 | 5.00 | 4.97 |
| Polish | 4.95 | 4.90 | 5.00 | 4.95 |
| Russian | 4.94 | 4.91 | 5.00 | 4.95 |
| Slovak | 4.90 | 4.89 | 5.00 | 4.93 |
| Slovenian | 4.97 | 4.96 | 5.00 | 4.98 |
| Serbian | 5.00 | 4.99 | 5.00 | 5.00 |
| Ukrainian | 4.95 | 4.95 | 5.00 | 4.97 |
| Greek | 4.97 | 4.92 | 5.00 | 4.96 |
| Estonian | 4.97 | 4.97 | 5.00 | 4.98 |
| Finnish | 4.83 | 4.78 | 5.00 | 4.87 |
| Hungarian | 4.96 | 4.93 | 5.00 | 4.96 |
| Lithuanian | 5.00 | 5.00 | 5.00 | 5.00 |
| Latvian | 4.95 | 4.95 | 5.00 | 4.97 |
| Japanese | 4.98 | 4.98 | 5.00 | 4.99 |
| Korean | 4.98 | 4.98 | 5.00 | 4.99 |
| Chinese (Simplified) | 5.00 | 5.00 | 5.00 | 5.00 |
| Chinese (Traditional) | 5.00 | 5.00 | 5.00 | 5.00 |
| Bengali | 5.00 | 4.98 | 5.00 | 4.99 |
| Hindi | 4.81 | 4.63 | 5.00 | 4.81 |
| Marathi | 4.89 | 4.88 | 5.00 | 4.92 |
| Tamil | 4.98 | 4.97 | 5.00 | 4.98 |
| Telugu | 4.98 | 4.94 | 5.00 | 4.97 |
| Urdu | 4.93 | 4.90 | 5.00 | 4.94 |
| Arabic | 4.98 | 4.98 | 5.00 | 4.99 |
| Persian | 4.93 | 4.91 | 5.00 | 4.95 |
| Hebrew | 4.87 | 4.87 | 5.00 | 4.91 |
| Indonesian | 5.00 | 5.00 | 5.00 | 5.00 |
| Malay | 4.95 | 4.95 | 5.00 | 4.97 |
| Thai | 4.99 | 4.99 | 5.00 | 4.99 |
| Tagalog | 5.00 | 4.94 | 5.00 | 4.98 |
| Vietnamese | 5.00 | 5.00 | 5.00 | 5.00 |
| Turkish | 4.97 | 4.96 | 5.00 | 4.98 |
| Swahili | 4.99 | 4.99 | 5.00 | 4.99 |
The spread between the best languages (Catalan, Lithuanian, Serbian, Chinese, Indonesian, and Vietnamese at 5.00) and the weakest (Hindi at 4.81) is 0.19 — and even the weakest scored a perfect 5.00 on variable preservation.
The two cross-check judges — Kimi-K2.6 and GLM-5.2 — re-scored all 4,900 translations from scratch, blind to the primary's marks. They landed at 4.95 overall, within 0.01 of the published primary, and agreed language-by-language to a mean absolute difference of about 0.04. Across all 49 languages, every language scored 4.8 or higher with all three judges. That's agreement between AI judges, not a substitute for human review.
Four languages landed below 4.9, and they're in the table above, not hidden: Hindi (4.81), Danish (4.86), Bulgarian (4.87), and Finnish (4.87). The single lowest dimension anywhere is Hindi fluency at 4.63 — the output is accurate and keeps every variable, but reads a notch more stiffly than a native writer would phrase it.
This isn't random. The languages where the three judges diverged most are the same ones at the bottom of the table — lower-resource languages where fluency is harder to nail and harder to grade. These are the languages where native-speaker review matters most, because when a benchmark and a human reviewer disagree, the human gets the final word. A benchmark you can trust is one that shows you its outliers instead of rounding them away.
The string dataset, the scoring rubric, the verbatim judge prompts, and
every translation with its per-string scores are published at
polyglot-i18n/polyglot-bench.
We include a small rejudge.py script (standard library only) that re-scores
our exact outputs with any OpenAI-compatible judge you choose — try one we
didn't use:
export JUDGE_API_KEY=...
python3 rejudge.py results/de.json \
--model your-favorite-judge \
--endpoint https://your-provider/v1/chat/completions
It prints your judge's per-dimension means next to ours, with deltas. If you run it and get materially different numbers, we want to hear about it.
Re-judging scores the exact June outputs. The repo can also regenerate translations, but that runs through today's production model, so fresh outputs won't match June's.
A condensed version of these results, with per-language tables, lives at /benchmark.
Start in your terminal
Install the CLI, run a scan, and see exactly what you're missing. Free, no account required.