Translation quality can be measured, and enterprise buyers should expect to see the numbers. The most reliable approach combines three things: a human, error-based evaluation that uses a defined error list and severity levels; automated signals used for screening and monitoring; and operational metrics such as rework rate, on-time delivery, and terminology compliance. A vendor who cannot show you any of these is asking you to take quality on trust.

This guide explains the translation quality metrics that matter, how an error score is calculated, what automatic metrics can and cannot tell you, which metrics fit which content, and the questions to put to a vendor before you sign.

Why “we are ISO certified” is not a quality metric

Certifications are useful, but they describe a process, not a result. ISO 9001 shows that a quality management system exists. ISO 17100 sets requirements for translation service providers, including translator qualifications and revision by a second linguist. Both tell you that a provider works in a controlled way. Neither tells you how good the 40,000 words you received last month were.

That is the gap the international guidance on evaluating translation output, ISO 5060:2024, was written to address. It describes an analytic method: a trained evaluator compares the translation with the source, records each error by type and severity, applies penalty points, and produces an error score and a quality rating. It covers human translation, post-edited machine translation, and unedited machine translation. It is a guidance standard, which means it offers recommendations rather than requirements, and a company cannot be certified against it (ATC Certification).

In practice, ask for both: a certified process and measured output.

The three layers of translation quality measurement

No single number captures quality. Mature programs use three layers together.

Layer What it measures Best used for Main limitation
1. Human error-based evaluation Errors by type and severity in a sample of delivered text Acceptance decisions, vendor comparison, contractual thresholds Takes time and trained evaluators
2. Automatic metrics and quality estimation Statistical similarity or predicted quality Screening, engine comparison, trend monitoring Not reliable as the only judge of a final deliverable
3. Operational and business metrics Delivery, rework, consistency, cost Vendor management and process improvement Does not measure linguistic correctness directly

Layer 1: error-based evaluation (MQM and ISO 5060)

The most widely used approach is based on the Multidimensional Quality Metrics framework, usually shortened to MQM. ISO 5060 follows the same logic. An evaluator marks each error and assigns it a category and a severity.

Error categories. The standard’s typology groups errors into seven top-level categories:

  • Terminology
  • Accuracy
  • Linguistic conventions
  • Style
  • Locale conventions
  • Audience appropriateness
  • Design and markup

Severity levels. Each error is rated minor, major, or critical. The parties decide what counts as critical for their content. A single wrong number in a medical document may be critical, while an awkward phrasing in an internal memo may be minor (ISO 5060 overview, EU ATC).

Sampling. Evaluators rarely read everything. They review a defined sample, and the sampling method should be agreed in advance so that results are comparable over time.

Scoring. Each error earns penalty points based on its severity, and the total is divided by the volume evaluated, for example per 1,000 words. A threshold converts the score into a pass or fail or a quality rating.

A worked example

The weights and thresholds below are illustrative. In a real program, you and your vendor agree them for your content.

A 5,000-word sample of a product guide is evaluated. Minor errors count 1 point, major errors 5 points, and critical errors 25 points.

Severity Errors found Points each Total points
Minor 8 1 8
Major 2 5 10
Critical 0 25 0
Total     18

The error score is 18 points divided by 5 thousand words, which equals 3.6 points per 1,000 words. If your agreed threshold for this content is 5 points per 1,000 words with no critical errors, the delivery passes. If the same text contained a single critical error, it would fail regardless of the overall score, because critical errors should override the average.

This simple rule matters. An average score can hide a dangerous error, so always set a separate rule for critical errors.

Layer 2: automatic metrics and quality estimation

Automatic metrics are computed by software. Several types exist.

  • Reference-based overlap metrics such as BLEU, chrF and TER compare a machine translation with a human reference translation. They are fast and cheap, and they are mainly useful for comparing engines on the same test set.
  • Learned metrics such as COMET use models trained on human judgments, and they generally track human opinion better than simple overlap scores.
  • Quality estimation predicts the quality of a translation without needing a reference, which makes it useful for routing. Segments predicted to be good can pass with light review, and weak ones go to a human editor.

Use them for: choosing between engines, monitoring trends, deciding which content needs more human attention, and spotting outliers.

Do not use them for: the final acceptance of a legal, medical, or safety document, ranking individual linguists, or settling a dispute about a specific error. Research on using large language models to annotate errors in MQM style is active, and it is promising as a screening aid, but any such scoring should be calibrated against trained human evaluators before you rely on it.

Layer 3: operational metrics enterprise buyers should track

These metrics show how reliably a vendor performs over time.

  • On-time delivery rate. The share of deliveries that meet the agreed deadline.
  • First-pass acceptance rate. The share of deliveries accepted without any rework request.
  • Rework or revision rate. How often you need to send text back, and how long it takes to fix.
  • Post-delivery defects per 1,000 words. Errors your reviewers or customers find after delivery, which is the quality that actually reaches the market.
  • Terminology compliance. The share of approved glossary terms used correctly.
  • Consistency and translation memory leverage. How much approved text is reused, and whether repeated sentences are translated the same way.
  • Query response time. How quickly the vendor answers questions about the source.
  • Effective cost per accepted word. Price per word plus the cost of rework, which is often higher for cheap, error-prone work than the quote suggests.

A translation management platform can track many of these automatically. For example, a centralized system such as MarsCloud can report turnaround, volume, and quality checks by project, which makes trend analysis far easier than spreadsheets.

Match the metric to the content type

Quality means different things for different content. Weight the metrics to match the risk.

Content type Priority error categories Suggested approach
Marketing and brand content Style, audience appropriateness, locale conventions Evaluate with native reviewers who know the brand voice; tolerate
minor accuracy variation if tone is right
Technical documentation Terminology, accuracy Strict terminology compliance and consistency; see our
technical translation services
Software and UI text Design and markup, accuracy, length Check placeholders, tags and string length as well as language
Legal and official documents Accuracy, terminology, locale conventions Zero tolerance for critical errors; see our
official translation services
Medical and pharmaceutical Accuracy, terminology Zero critical errors, independent subject review; see
pharmaceutical translation services
Support and internal content Accuracy, clarity Lighter evaluation and faster turnaround

What a good vendor quality report contains

When you ask a vendor for a quality report, look for these elements:

  1. The sample size and how it was selected.
  2. Error counts by category and severity.
  3. The error score per 1,000 words and the threshold it was measured against.
  4. A separate flag for any critical errors.
  5. The qualifications of the evaluator, and whether the evaluator is independent of the translator.
  6. Corrective actions taken and the date they were completed.
  7. Trend data over several months, not just one delivery.

Questions to ask a translation vendor

  • Which error typology and severity levels do you use, and can we configure them?
  • How are samples selected, and what share of our volume do you evaluate?
  • Who performs the evaluation, and are they independent of the translator?
  • What is the pass threshold for each content type, and how are critical errors handled?
  • Can you share a sample quality report with client data removed?
  • How do you measure terminology compliance against our glossary?
  • What happens when we report an error after delivery, and how fast do you respond?
  • Do you measure the quality of post-edited machine translation separately from human translation?
  • How do you use automatic metrics or quality estimation, and do humans verify the results?
  • What are your on-time delivery and rework rates for the past six months?
  • How will you report quality trends to us each quarter?
  • Which certifications cover your process, and how do they relate to your quality measurement?

Common pitfalls to avoid

  • Relying on one score. A single number hides critical errors and different error types. Always report categories and severity.
  • Letting vendors grade their own work without oversight. Self-evaluation can be useful for internal improvement, but periodic independent checks keep it honest.
  • Using tiny or biased samples. If samples are chosen by the vendor, they may not represent the whole delivery.
  • Comparing scores across languages without care. Evaluators and difficulty differ by language, so compare trends within a language first.
  • Ignoring source quality. Ambiguous or inconsistent source text produces errors that are not the translator’s fault. Track source issues separately.
  • Measuring raw machine output and calling it the deliverable. Evaluate what the customer actually receives, after any human editing.

Build a quality program in 90 days

Days 1 to 30: define. Choose your content tiers, agree error categories and severity definitions, set thresholds for each tier, and build or refresh your glossary.

Days 31 to 60: measure. Run evaluations on a first set of deliveries, train your internal reviewers, and compare vendors on the same samples. Establish your baseline for rework and on-time delivery.

Days 61 to 90: improve. Review the results with each vendor, agree corrective actions, add automated checks where they are reliable, and schedule quarterly reviews. Share the numbers with stakeholders so quality becomes visible.

How CCJK approaches translation quality

CCJK runs projects through a translation, editing, and proofreading process supervised by project managers, with linguists who are native speakers and subject-matter specialists. Our processes are covered by ISO 9001:2015, ISO 27001:2013, ISO 17100:2015, and ISO 13485:2016 certifications, and we handle data in line with GDPR requirements. Clients can also add editing and proofreading as a standalone step, including light and full post-editing, and our document translation services include a final quality check before delivery.

If you are building or reviewing a quality program, talk to our team about the evaluation approach and reporting that suit your content and risk level.

Frequently asked questions

What is the best way to measure translation quality?

Use a combination. Human error-based evaluation gives the most trustworthy view of a delivery. Automatic metrics help with screening and trends, and operational metrics show reliability over time.

What is MQM in translation?

MQM stands for Multidimensional Quality Metrics. It is a framework for evaluating translation quality by classifying errors into categories and severity levels and then scoring them. ISO 5060 is based on the same approach.

Can I get ISO certified for translation quality evaluation?

ISO 5060 is a guidance standard, so there is no certification for it. Providers can certify their process under standards such as ISO 17100 and ISO 9001, and then apply an evaluation method like MQM or ISO 5060 to measure results.

Is BLEU a good measure of quality?

BLEU is a quick way to compare machine translation engines against a reference, but it is a weak guide to the quality of a final deliverable. Use it for engine testing, not for accepting or rejecting a delivery.

How many words should be sampled for evaluation?

It depends on volume and risk. Many programs evaluate a fixed sample, such as a few thousand words per delivery or per period, and increase the sample for high-risk content. Agree the method in advance.

Next step

If you want clear, measurable quality from your translation partner, contact CCJK to discuss your content, risk level, and reporting needs.

Sources and notes
  • ISO 5060:2024, Translation services: evaluation of translation output, general guidance. Summaries: ATC Certification and Slator.
  • Multidimensional Quality Metrics (MQM) framework documentation.
  • ISO 17100 and ISO 9001 for translation services and quality management requirements.
  • The scoring weights and thresholds in the worked example are illustrative only.