Last month a PM sent me a Phrase project with 4,000 words showing in the count and about 900 words actually unlocked for me. The rest was greyed out. Not TM matches — I checked. Those segments had been machine-drafted, scored by a quality-estimation model, cleared the threshold, and gone straight to "approved." My job was the leftovers. The stuff the machine wasn't sure about.
The reasoning is clean on paper. Let the model grade its own output, set a confidence threshold, route only the shaky segments to a human. You pay for judgment where judgment is needed and nowhere else. Phrase, Smartcat, the whole crowd — they're all building this: AI draft, TM and terminology on top, a quality-estimation pass, automated routing by threshold, then targeted human editing. The CAT tool has quietly turned into the thing that decides how much of your language a person is allowed to touch.
And for Spanish or German, honestly, it mostly works. I've seen it. The QE models on high-resource pairs are calibrated against oceans of data. When they say a segment is a 92, it's usually a 92. The triage earns its keep.
The score for Turkmen is a coin flip in a suit
Here's the problem nobody upstream wants to hear. A quality-estimation model is only as honest as the data it was trained on. For English–Spanish that's a mountain. For English–Turkmen it's a puddle. The same architecture that gives you a trustworthy 92 in Spanish gives you a number in Turkmen that looks exactly as confident and means almost nothing.
So what gets waved through? Not the garbled segments — those trip the low-confidence flag and land on my desk, which is fine, I want them. It's the fluent ones. The lines that read smoothly, hit every grammatical marker, and are wrong. A case ending that points the sentence the wrong way. A term borrowed from Russian when the client's glossary wanted the Turkmen coinage. A negation dropped so cleanly the sentence still parses. The QE model sees fluency and rewards it. Fluency is precisely what a low-resource MT engine produces best and gets wrong most.
The threshold, in other words, filters out the errors I'd catch in two seconds and lets through the errors I actually get paid to catch. It's optimized backwards for my language and nobody set it that way on purpose. They just took the config that works for the twenty pairs that fund the platform and applied it to all of them.
The greyed-out segment is the one I want
When I raised this on that Phrase job, the PM's answer was reasonable and completely beside the point: "Those segments scored above 85, they're locked to keep the budget down." Nobody had asked whether an 85 in Turkmen is the same animal as an 85 in French. It isn't. It's a different model, a different data reality, a different reliability curve — reported on the same 0-to-100 scale so it can slot into one dashboard. That scale is a convenience for the PM, not a measurement of my language.
What I've started doing, and what I'd tell any agency running QE routing across a low-resource pair:
Unlock the whole file for the low-resource languages, or drop the threshold to near zero. If the QE model can't be trusted to grade Turkmen, don't let it grade Turkmen. Let me read every segment. Yes, that's more billable minutes. It's also the difference between a translation and a plausible-looking hazard.
Don't publish a per-segment QE score you can't defend per language. If you can't tell me what dataset calibrated the Turkmen estimator, the score is decoration. Decoration that's making routing decisions.
Track the leak. Every so often, pull a random sample of the auto-approved segments in a low-resource pair and have a human actually read them. If the error rate in the "confident" bucket is anywhere near the error rate in the "human-reviewed" bucket, your threshold isn't buying you quality, it's buying you a number on a report and a client who finds the mistakes later.
I'm not against the workflow. The CAT tool sitting between raw MT and the client, catching terminology drift and tone before anyone signs off — that's a good thing, and it's more necessary now than when it was just fuzzy matches and a spellcheck. The governance layer is real work and I'd rather it exist.
But governance means someone decides. Right now, for languages like mine, the deciding is being outsourced to a model that has never seen enough Turkmen to have an opinion worth acting on, wearing the same uniform as the models that have. The threshold looks like a quality decision. For Turkmen it's a budget decision in a quality decision's clothes.
Give me the greyed-out segments. That's where the interesting mistakes live.