There's a standard now, ISO 5060, meant for evaluating translation output — human, machine, post-edited, all of it — using error-based scoring. I read it the way I read most new standards: with one eyebrow up, wondering who's going to hand it to me as a rate justification.
And here's the thing. I'm not against it. Not even slightly. Most of my career, quality was whatever the last reviewer felt on a Tuesday. A number is better than a mood. But a number is also a story about what you decided to count, and that's where I want to argue.
What error-counting gets right
MQM — the framework most of this scoring leans on — breaks errors into categories. Accuracy, fluency, terminology, style, locale conventions. Each error gets a severity: minor, major, critical. You tally, you weight, you divide by word count, you get a score. Trados, memoQ, Phrase all bolt some version of this onto their QA panels now, and the LLM-assisted ones will even pre-flag candidates for you.
When it works, it works. The best thing about categorized errors isn't the total — it's the shape. If I get a Turkmen review back and 80% of the flags are terminology, that tells me the glossary was stale or nonexistent, not that I can't write. If they're all locale conventions — dates, decimal commas, the number formatting that changes when a document crosses from a Russian source into Turkmen — that's a settings problem, not a translation problem. The category is the diagnosis. The score is just the fever.
For oil and gas work especially, I want the critical bucket to exist and to mean something. A mistranslated tolerance or a flipped negation in a safety procedure is not a "major" style ding sitting next to a missing comma. Severity tiers force the reviewer to say out loud: this one could hurt someone. That's honest. I'll take it.
Where the number starts lying
The trouble begins the moment the score detaches from the reading. I've watched it happen. An agency sets a pass threshold — say, 98 on some normalized scale — and now the review stops being "is this good Turkmen?" and becomes "did we clear 98?" Those are not the same question. Not close.
Error-based scoring is brutal at catching the countable and blind to the thing that actually decides quality: whether a native reader trusts the text. You can score a 99 and still produce marketing copy that reads like it was assembled by committee from a phrasebook. No flaggable error. Every term correct. Dead on arrival anyway, because transcreation lives in choices that no rubric has a checkbox for. I've delivered near-perfect scores on copy I knew was flat, and I've had genuinely good work come back bleeding minor flags because the reviewer preferred their own synonyms.
And Turkmen is a small language. The evaluator pool is thin. When there's one reviewer, the "objective" MQM score is exactly as objective as that one person's preferences dressed up in a spreadsheet. The framework doesn't fix subjectivity. It hides it behind arithmetic. That's more dangerous than an honest opinion, because now you can't argue with it — it's a number.
How I'd actually use it
So here's my working position, for PMs who are about to roll ISO 5060 language into their contracts.
Use the categories as a conversation, not a verdict. Send me the annotated errors — the actual flags, with severity and category — not just "score: 94, revise." I wrote about this before: nobody reads past the score. The standard makes that worse if you let it, because now the score has an ISO number stapled to it and feels like physics. It isn't. It's a summary of someone's markup, and the markup is where the useful information lives.
Separate countable quality from felt quality, and pay for both honestly. Terminology, numbers, tags, locale formatting — score those hard, automate the checks, I want them caught. But don't run transcreation or brand voice through the same rubric and pretend the number means anything. Back-translation and error-scoring both die in the same place: the moment you treat a creative choice as a defect because it didn't match the source word-for-word.
Define severity before the project, not after. The single most argued-about thing in any LQA is whether a flag is major or minor. That one distinction swings the score more than the actual error count. Nail it down up front — for this content type, a terminology miss is major; for that one, minor. Otherwise the reviewer's mood sets the severity, and we're right back to Tuesday.
And keep a human who can override the number. The whole point of a standard is repeatability. But repeatable wrong is still wrong. If the score says 91 and the client's Turkmen distributor says the text sounds foreign, the distributor wins. Every time. She's the real evaluator. ISO 5060 just gave the rest of us a shared vocabulary for explaining to her why.
The measure is useful right up to the point where people forget it's a measure of a reading, and not the reading itself. Count the errors. Then go read the sentence.