Skip to the content
[ ]

LSPHIL

LSPHIL › Translation & interpreting › Why Machine Translation Makes Mistakes That Sound Perfect

Why Machine Translation Makes Mistakes That Sound Perfect

Fluent output is not evidence of accuracy. That principle sits at the heart of ISO 18587:2017, the international standard governing how human post-editors must handle machine translation output. The standard requires full post-editing to check accuracy and comprehensibility, improve readability, and correct errors—producing output comparable to human translation. The existence of this document, and of specialized post-editing guidelines like those developed for the U.S. government's BOLT program, reveals an uncomfortable truth: machine translation fails in ways that survive a casual read. The failures are systematic enough that buyers can predict them before committing budget.

English · 886 words

When the Source Hides What the Target Needs

ISO 18587:2017 applies only to content processed by machine translation systems, and its very existence signals where the risk concentrates. The standard is intended for translation service providers, their clients, and post-editors—stakeholders who have learned that raw output demands professional intervention. A translator doing post-editing compares source text with machine output and revises for intended purpose, a workflow described in the Wiley Encyclopedia of Applied Linguistics. This bilingual verification is the norm precisely because surface fluency misleads.

German teleprinter control panel at IMWWII.agr
German teleprinter control panel at IMWWII.agr. ArnoldReinhold · CC BY-SA 4.0 · Wikimedia Commons

The core problem is structural. Machine translation systems learn from parallel text, mapping source strings to target strings based on statistical or neural pattern matching. When the source language leaves meaning implicit—dropping pronouns, omitting temporal markers, or relying on verb aspect that English must render with explicit auxiliaries—the system guesses. The guess often produces grammatical English. The BOLT post-editing guidelines, developed by NIST for Arabic and Chinese to English translation, instruct editors to make machine output "have the correct meaning, use understandable English, and do so in as few edits as possible." The same guidelines require the edited version to "preserve the meaning of the reference translation and not add or omit information." The burden falls on humans precisely because the machine cannot recognize what it has added or lost.

Where Context Lives Outside the Sentence

Post-editing is defined as producing text comparable to human translation, while light post-editing merely aims for comprehensibility without human-translation quality. This distinction, documented in ACL Anthology research on post-editing levels, matters because it shows how far machine output can drift. A sentence that reads well in isolation may collapse when the reader discovers that "it" refers to a different antecedent three paragraphs back, or that a corporate "we" has silently shifted from the board to the subsidiary.

ISO-based post-editing usually requires checking machine output against source-language content. This requirement exists because meaning frequently depends on material outside the current sentence: prior discourse, document structure, or the controlled terminology defined in a separate glossary the translation system never saw. When terminology is fixed by client specification or regulatory requirement, the system invents plausible-sounding equivalents that violate binding definitions. The editor must catch these not by reading the target text for elegance, but by verifying against external reference.

Why Post-Editing Is a Different Job

The BOLT guidelines describe the goal as turning machine output into fluent English while preserving all meaning from the original Arabic or Chinese. This formulation exposes the task's asymmetry. Ordinary translation requires understanding source meaning and recreating it in target form. Post-editing requires diagnosing what the machine has distorted while it was pretending to do exactly that. The edited text must be faithful to the reference translation in both meaning and style—a dual requirement that adds cognitive load.

A translator working from scratch builds coherence sentence by sentence. A post-editor works backwards from fluent garbage, reconstructing what the source probably said by cross-referencing. The Wiley Encyclopedia source notes that post-editors compare source with output and revise for intended purpose. The intended purpose may be legal compliance, medical safety, or marketing persuasion—each demanding different tolerance for deviation. Post-editing is therefore not failed translation with cleanup, but a distinct competence with its own standards and training requirements.

The Expensive Mistake

The practical predictor of failure is not linguistic complexity but dependency on external structure. Content with controlled terminology, content requiring exact meaning preservation, and workflows demanding both fluency and fidelity all trigger the ISO 18587 process for reason. Buyers who skip this step because the sample paragraph "looks fine" discover the cost later: legal exposure, brand damage, or operational error propagated through fluent misdirection.

The BOLT guidelines' emphasis on minimizing edits while preserving meaning captures the economic pressure. Organizations want cheap speed, so they push for light post-editing. Light post-editing produces merely comprehensible text without human-translation quality. The compromise satisfies nobody who actually needs accuracy, yet persists because fluent output is easier to approve than accurate output is to verify.

ISO 18587:2017 requires full post-editing to produce output comparable to human translation. The standard's scope—translation service providers, their clients, post-editors—maps the chain of responsibility. Each link must understand that grammatical English is the default, not the achievement. The final check is never "does this read well?" but "does this match what the source actually said?" That check requires bilingual access, reference materials, and time that automated quality scores do not replace.

The coming procurement decision will not turn on whether machine translation has improved. It will turn on whether the buyer recognizes that improvement in fluency correlates weakly with improvement in adequacy. The systems get better at sounding right. The gap between sounding right and being right remains the domain where human post-editors earn their rates—and where unprepared buyers pay twice.