Insights

Content Operations

Human approval points that keep multilingual content on-brand

Victory Harbour Operating Team · · 8 min read

Every content team has one person who can tell, in about four seconds, that something is off-brand. They usually cannot explain why, and that is the whole problem.

At low volume this works. The person reads everything, catches what is wrong, fixes it. The knowledge stays in their head, and nobody notices that it was never written down.

Then volume increases, or a second language is added, or that person goes on leave. Now the tacit knowledge is a single point of failure, and the standard reviewer instruction — “review everything” — stops being a control and becomes a queue.

Adding AI to content production makes this arrive faster. Production capacity goes up sharply while review capacity stays exactly where it was. The bottleneck does not disappear; it moves, and it moves onto the most experienced person you have.

Three different things wearing one name

“On-brand” is doing too much work as a phrase. Underneath it are three distinct requirements with different natures:

Facts. Specifications, prices, availability, names, dates. These are objectively checkable against a source of truth. A person is not required — a person is merely usual.

Compliance. Prohibited claims, regulated wording, mandatory disclosures, category rules per market. Also checkable, against a maintained list. The list is the hard part, not the checking.

Voice. Register, rhythm, how direct to be, what a brand does not say. This is genuinely tacit and genuinely human — but it is also the most stable of the three. Voice does not change weekly. Facts do.

Most review processes treat all three as one undifferentiated judgement performed by one person at one point. That is why they collapse: the human is spending their scarce attention on the two categories a machine handles better, and has none left for the one category only they can handle.

Separate them and the design becomes obvious:

Facts

Method: Automated check against product or knowledge source

Coverage: 100%, blocking

Compliance

Method: Automated check against maintained rule list per market

Coverage: 100%, blocking

Voice

Method: Human review

Coverage: Sampled — except where noted below

The multilingual multiplier

One error in a source text becomes five errors in five languages. This single arithmetic fact should determine the entire design.

It means the review effort belongs upstream. Checking the source text hard — before adaptation — is worth five times more than checking each translated version afterwards. Yet most workflows do the opposite: light review of the source, heavy review of each language, performed by native speakers who are expensive, in different time zones, and reviewing an error that should never have reached them.

The second implication: what happens after adaptation is not translation review, it is market fit review. The question is not “is this an accurate translation” — that is now largely solved. The question is “does this work here”: are the objections right, does the compliance wording match this jurisdiction, are these the terms people actually search for.

Those are different skills and they should be briefed differently. Asking a translator to check accuracy when you need a market judgement wastes the reviewer and misses the risk.

Where a person must stay in the path

Sampling is right for routine voice review. It is wrong in five places, and these should be blocking, 100%, permanently:

  • Anything with a price or a commercial term
  • Regulated claims — health, safety, environmental, financial
  • First content in a new market or language — until you have evidence of the failure modes there
  • First content in a new category or format
  • Anything a named individual will be quoted as saying

The general rule: a human stays in the path where an error is expensive, irreversible, or embarrassing in public. Everywhere else, sample — and audit the sample rate against the defect rate you actually observe.

The reviewer's job changes, and this is the valuable part

In a manual process the reviewer fixes things. They rewrite the awkward sentence and move on. The correction lives in the output and nowhere else.

In a governed workflow the reviewer's job is different: accept, reject, and state why. The rewrite is secondary. The reason is the product.

“Accept, reject, and state why. The rewrite is secondary. The reason is the product.”

Because rejection reasons, collected consistently, are the most valuable output the whole system produces. Within a few months they tell you exactly which categories fail, which market has the recurring compliance issue, which part of the brand voice the system has never learned. That is a prioritised improvement backlog, generated as a by-product of work you were doing anyway.

Almost nobody collects it. It requires exactly one thing: a structured reason field instead of a free-text comment box, and the discipline to use it.

This also changes what you should look for in a reviewer. In a manual process you want the best writer. In a governed one you want the best judge — someone who can articulate why something is wrong, consistently, in a form that improves the system.

What to measure

  • First-pass acceptance rate, tracked per language and per content type — a single average hides the one market that is failing
  • Time in review queue, not time spent reviewing. If items wait three days for four minutes of attention, your problem is scheduling, not capacity
  • Rejection reasons, by frequency — the improvement backlog
  • Escalation rate — how often sampled review escalates to full review. Rising means quality is drifting

In our own Operating Labs, batch production of 500 brand-constrained product images completed within one working day, and first-pass editorial acceptance for multilingual adaptations across EN, zh-Hant and DE reached 82%. Internal lab conditions, not client results. The 18% is the number worth designing around.

When not to do this

If your content's value is that a specific person wrote it — a founder's point of view, a partner's technical opinion, an analyst's judgement — do not put it through a production system. That content is not a throughput problem. Its scarcity is the point, and governing it will remove exactly what made it worth reading.

This system is for content that follows a pattern and needs to be consistent at volume. Most organisations have both kinds and should be honest about which is which.

Start with the workflow that matters

If brand consistency currently depends on one person having time, that is a risk worth examining before it becomes an incident. This is the discipline built into the AI Content Operations System.