What a quality score measures
Established quality frameworks classify translation errors by type and severity, and produce a score for a sample of the work. That is useful information. It tells you how a supplier is performing, whether quality is improving over time, and where reviewers keep finding the same problems. It is a measurement of the work.
What the publish decision asks instead
The person approving a release is not asking how the supplier is performing. They are asking one question about one version: if we publish this, will readers understand the same obligations, claims, and conditions the source states? A translation can score well overall and still contain the single sentence that changes a refund window. Averages hide the passage that matters.
Severity is defined by consequence, not by category
An error taxonomy assigns severity by error type. A publish review assigns it by consequence. A missing comma in a marketing subheading and a missing condition in a refund clause can belong to the same category and carry completely different risk. Deciding which is which requires knowing what the content is for and who acts on it.
Use both, for different purposes
Keep quality measurement where it belongs: supplier management, sampling, and improvement over time. Add a review before publication for the content where a change in meaning creates a real consequence. The first produces a score. The second produces a decision and a record of the reasoning behind it.
Bring these questions to your next release.