[Benchmark · June 2026]

4.96/5 translation quality. June 2026 pipeline. AI-judged.

We ran 4,900 translations of 100 real UI strings through our June 2026 pipeline in 49 languages spanning 9 language groups, then had three AI judges from other model families score every translation. The dataset, the judge prompts, and every primary-judge score are public.

[01]The headline
4.96/5

Overall quality

AI-judged, June 2026

100%

Placeholders preserved

every variable kept

94.5%

Top score (5/5)

top marks, every dimension

±0.04

Cross-judge agreement

mean gap, three AI judges

What that means:94.5% of translations got the top score on every dimension, 99.1% needed at most minor stylistic polish, and 100% preserved every code placeholder exactly.

Per-string rating distribution · 4,900 translations
5 · top score
94.5%
4
4.6%
3
0.8%
≤2
0.1%
[02]Every score

Every language, every dimension.

Scored 1–5 per string by DeepSeek-V4-Pro, then re-scored by two more AI judges. Every language lands at 4.81 or higher; the spread between the best and weakest is just 0.19.

Overall score · all 49 languages spread 0.19
4.64.74.84.95.0
Germanic
LanguageAccFluVarOverall
Afrikaans4.974.975.004.98
Danish4.814.785.004.86
German4.944.955.004.96
Icelandic4.974.975.004.98
Norwegian Bokmål4.984.995.004.99
Dutch4.984.945.004.97
Swedish4.934.935.004.95
Romance
LanguageAccFluVarOverall
Catalan5.005.005.005.00
Spanish4.974.985.004.98
French4.934.935.004.95
Italian4.944.945.004.96
Portuguese4.994.995.004.99
Portuguese (Brazil)4.984.995.004.99
Romanian4.964.975.004.98
Slavic
LanguageAccFluVarOverall
Bulgarian4.834.795.004.87
Czech4.934.815.004.91
Croatian4.974.955.004.97
Polish4.954.905.004.95
Russian4.944.915.004.95
Slovak4.904.895.004.93
Slovenian4.974.965.004.98
Serbian5.004.995.005.00
Ukrainian4.954.955.004.97
Baltic, Finnic & Hellenic
LanguageAccFluVarOverall
Greek4.974.925.004.96
Estonian4.974.975.004.98
Finnish4.834.785.004.87
Hungarian4.964.935.004.96
Lithuanian5.005.005.005.00
Latvian4.954.955.004.97
East Asian
LanguageAccFluVarOverall
Japanese4.984.985.004.99
Korean4.984.985.004.99
Chinese (Simplified)5.005.005.005.00
Chinese (Traditional)5.005.005.005.00
South Asian
LanguageAccFluVarOverall
Bengali5.004.985.004.99
Hindi4.814.635.004.81
Marathi4.894.885.004.92
Tamil4.984.975.004.98
Telugu4.984.945.004.97
Urdu4.934.905.004.94
Middle Eastern
LanguageAccFluVarOverall
Arabic4.984.985.004.99
Persian4.934.915.004.95
Hebrew4.874.875.004.91
Southeast Asian
LanguageAccFluVarOverall
Indonesian5.005.005.005.00
Malay4.954.955.004.97
Thai4.994.995.004.99
Tagalog5.004.945.004.98
Vietnamese5.005.005.005.00
Turkic & African
LanguageAccFluVarOverall
Turkish4.974.965.004.98
Swahili4.994.995.004.99
[03]Methodology

Built to be doubted.

Three AI judges, other model families

Every translation was scored by DeepSeek-V4-Pro, then re-scored from scratch by Kimi-K2.6 and GLM-5.2 — three AI judges from different model families than our June engine, so no model graded its own relatives. Across all 49 languages, the three agreed to within a 0.04 mean delta.

We publish the disagreement

We don't bury the spread. Four languages land under 4.9 — Hindi, Danish, Bulgarian, Finnish — right there in the per-language table above. And where the three judges diverge most (GLM-5.2 and the others split a few tenths on a handful of lower-resource languages), we publish that too. Benchmarks that show their outliers are the ones worth trusting.

Real strings, production pipeline

100 UI strings across 12 categories — buttons, error messages, plurals, and interpolated variables like {name} — translated by our production pipeline as it ran in June 2026. We’ve since changed models, so read these as June’s results.

[04]Reproduce it

Don't take our word for it.

The dataset, the verbatim judge prompts, and every translation with its per-string scores are open source. Re-judge our outputs with any LLM you like — including ones we didn't use — and compare against the published numbers.

$ export JUDGE_API_KEY=...
$ python3 rejudge.py results/de.json \
--model your-favorite-judge \
--endpoint https://your-provider/v1/chat/completions
dimension published your judge delta
────────────────────────────────────────────
accuracy 4.94 4.95 +0.01
fluency 4.95 4.96 +0.01
variables 5.00 5.00 +0.00▋

See it on your strings

Quality you can check yourself.

Don’t take a June number on faith. Your first 50 translations are free, no account required. Run a scan and judge today’s output yourself.

Get started
Translation Quality Benchmark — 4.96/5 across 49 languages | Polyglot