Turkmen Translator
← Journal
Published September 23, 2026· quality assurance, cat tools, turkmen, mtqe

COMET Doesn't Speak Turkmen, and My QA Panel Knows It

Automatic quality estimation is quietly reshaping how agencies triage translation — but the scores go blind on low-resource languages. Here's what actually catches errors in Turkmen work.

A PM sent me a batch last week with a note: "QE score is 0.82, should be a light edit." It was oil & gas Russian into Turkmen, and the score was decided by a model that has, near as I can tell, read almost nothing in my language. The 0.82 was a guess dressed as a measurement. I did a full edit anyway, because the alternative was shipping garbage with a confident number stapled to it.

This is the part of the QA conversation nobody wants to say out loud. Quality estimation — the automatic kind, the thing that spits out a number so the agency can route a job to "light" or "heavy" post-editing — works beautifully for the languages that already had it easy. English, Spanish, German, French. It works because those pairs have oceans of aligned text behind them. Turkmen has a puddle.

The score is only as honest as the training data

COMET, the metric a lot of these systems lean on, is trained on human quality judgments across specific language pairs. When your pair isn't well represented, the model still produces a number — it never refuses — but that number is closer to noise than signal. It can't tell you the machine confused a legal "shall" with a plain future tense, because it doesn't have the sense of Turkmen to know the difference matters.

Same with the agentic MTQE features showing up in Trados and Phrase and memoQ this year. Great tools. I use them. But watch what they're confident about. They flag number mismatches, tag breaks, doubled spaces, a date that flipped format. Real problems, worth catching. What they don't flag is a sentence that's grammatically clean and semantically wrong — which is exactly what modern MT produces in a low-resource pair. Fluent nonsense sails right past a QE score. It reads smooth. The model likes smooth.

So when an agency tells me the QE says light-touch, I hear: the automatic layer found no obvious mechanical faults. That's genuinely useful. It is not a statement about whether the translation is correct. Those are two different claims and the industry keeps merging them because merging them saves money.

What I actually keep switched on

I'll tell you my panel, because the useful ones and the theater ones are not the same.

On, always: number and date consistency, tag integrity, terminology against the client glossary, untranslated-segment detection, and inconsistency checks where the same source segment got two different translations. These are mechanical and Turkmen-blind, which is fine — they don't need to understand the language, they need to compare strings. That's honest work.

Off, mostly: the grammar and "style" checkers that were built with Latin-script European syntax in mind. Turkmen is agglutinative. One word can carry what English spreads across five, with suffixes stacking case and possession and tense. The checker sees a long token and panics. I've had a spellchecker underline a perfectly correct word because it couldn't parse the suffix chain, and I've watched it wave through an actual typo because the malformed word happened to sit in its short list. I killed those checks a long time ago. They cost me more attention than they save.

And the LLM-based "suggest a better translation" button — I treat it as a second opinion I never asked for. Sometimes it's right. Sometimes it rewrites a term I chose on purpose because the client's engineer in Ashgabat uses that exact word on the rig. The tool doesn't know that. I do.

Where I want the number, and where I want to be left alone

Here's my actual position for the PMs reading this. Send me the QE score. I want it. It tells me what the machine already touched and how it rated its own work, and that shapes how I read the file — I know where to be suspicious. That's real information.

Just don't let the score set my rate or my time. A 0.85 on English-German and a 0.85 on Russian-Turkmen are not the same object. One rests on millions of human judgments. The other rests on the model shrugging. If your pricing tiers treat them identically, you're paying accurately for the languages that don't need it and underpaying for the ones that do. The languages with thin data are precisely the ones where a human has to do the most, and the scoring system is precisely the one that can least tell you so.

There's a version of this that gets worse before it gets better. As more agencies wire QE directly into routing, the low-resource pairs get triaged by a metric that's structurally incapable of seeing their real errors. The clean-reading, wrong-meaning translation gets the green light. Someone in Ashgabat reads a contract clause that says the opposite of the English. And the audit trail shows an 0.82, so everyone upstream is covered.

Measure what the machine can measure. Tags, numbers, glossary hits — take them, they're free and they're right. But don't hand a language the tools never learned to a metric and call the output quality. Call it what it is: a mechanical pass, followed by me, doing the part the number can't.