We ran 4,900 translations of 100 real UI strings through our June 2026 pipeline in 49 languages spanning 9 language groups, then had three AI judges from other model families score every translation. The dataset, the judge prompts, and every primary-judge score are public.
Overall quality
Placeholders preserved
Top score (5/5)
Cross-judge agreement
94.5% of translations got the top score on every dimension, 99.1% needed at most minor stylistic polish, and 100% preserved every code placeholder exactly.
Scored 1–5 per string by DeepSeek-V4-Pro, then re-scored by two more AI judges. Every language lands at 4.81 or higher; the spread between the best and weakest is just 0.19.
Every translation was scored by DeepSeek-V4-Pro, then re-scored from scratch by Kimi-K2.6 and GLM-5.2 — three AI judges from different model families than our June engine, so no model graded its own relatives. Across all 49 languages, the three agreed to within a 0.04 mean delta.
We don't bury the spread. Four languages land under 4.9 — Hindi, Danish, Bulgarian, Finnish — right there in the per-language table above. And where the three judges diverge most (GLM-5.2 and the others split a few tenths on a handful of lower-resource languages), we publish that too. Benchmarks that show their outliers are the ones worth trusting.
100 UI strings across 12 categories — buttons, error messages, plurals, and interpolated variables like {name} — translated by our production pipeline as it ran in June 2026. We’ve since changed models, so read these as June’s results.
The dataset, the verbatim judge prompts, and every translation with its per-string scores are open source. Re-judge our outputs with any LLM you like — including ones we didn't use — and compare against the published numbers.
See it on your strings
Don’t take a June number on faith. Your first 50 translations are free, no account required. Run a scan and judge today’s output yourself.
Get started