Completions

AI governance for marketing teams

Most AI governance work produces a policy document that describes controls nobody built. The version that matters is five operational decisions, each expressed as a number, a rule, or a record — things a system enforces rather than things a slide asserts.

This page is the working version: what needs governing, the autonomy ladder and where operators stall on it, how thresholds get set, which regulations actually bite, what you should deliberately leave ungoverned, and the seven questions that separate a practice from a policy vendor.

Published August 22, 2026· 11-minute read

The five things that need governing

Each one has a failure mode you can recognise in your own operation today, whether or not you have deployed anything.

1

What may be said

The claims your brand is allowed to make, per vertical and per jurisdiction. This is not a tone-of-voice document — it is a set of checks an output can be scored against before it publishes. A claim that is fine in one state is an exposure in the next, and the model has no way to know which state it is writing for unless you tell it.

What it looks like when it is missing: A regulated claim reaches a customer surface and nobody can say which system produced it or which rule it broke.

2

Who approves what

Routing rules that send each output to auto-publish, to review, or to escalation, based on a score rather than a job title. The threshold is a number you can change on purpose.

What it looks like when it is missing: Everything queues to one person, the queue grows, and the team quietly starts approving in bulk without reading — which is worse than no gate at all.

3

How much autonomy each surface gets

Autonomy is per surface and per location, not per company. A review response at a mature location can publish unattended while the same agent at a newly acquired one stays in review, because the master record there is not trustworthy yet.

What it looks like when it is missing: One global setting, so the whole system runs at the speed of its least trustworthy corner.

4

What gets recorded

For every published output: the input, the model and version, the rules evaluated, the score, the routing decision, and who if anyone touched it. Recorded at the time, not reconstructed afterwards.

What it looks like when it is missing: A regulator, a franchisee, or an angry customer asks how something got published, and the honest answer is that nobody knows.

5

What happens when it is wrong

A defined path to pause a surface, correct what published, and feed the correction back so the same class of error is caught next time. The feedback half is the part almost everyone skips.

What it looks like when it is missing: Reviewers override the same mistake every week and the system never learns, so the human cost of running it never falls.

The autonomy ladder

Autonomy is set per surface and per location, so a single operation runs several of these at once. The useful question is never “what level are we at” but “which surfaces have earned the next one.”

  1. Level 0 — Draft only

    The agent produces, a human publishes everything. Correct for a brand-new surface, and a permanent state only if the surface is genuinely high-stakes.

  2. Level 1 — Review then publish

    Everything queues, but the queue is ordered by confidence so reviewers spend their attention where it matters. Most surfaces should sit here for the first month.

  3. Level 2 — Auto-publish above a threshold

    Outputs scoring above a set number publish unattended; the rest queue. This is where the labour saving actually appears, and where most operators stop moving because raising the threshold feels risky without evidence.

  4. Level 3 — Auto-publish with sampled review

    Everything publishes, and a random sample is reviewed after the fact to detect drift. Appropriate once you have months of scored history showing the gate holds.

  5. Level 4 — Autonomous with exception handling

    The system publishes and self-corrects, escalating only anomalies. Few marketing surfaces need this, and claiming it before Level 2 has run for a quarter is a sales position rather than an engineering one.

Almost every operation that has deployed anything is stuck between Level 1 and Level 2, because moving requires evidence that the gate holds, and nobody set up the scoring that would produce that evidence. That is the single most common finding, and it is a measurement problem rather than a risk problem.

Which rules actually bite

Buyer education, not legal advice — take the specifics to your own counsel. The point is that the obligations mostly predate AI and apply to the claim rather than the author.

EU AI Act — transparency obligations
If you operate in or market into the EU, AI-generated content aimed at the public carries disclosure obligations. The practical consequence is a per-output record of what was machine-generated, which you want anyway.
NIST AI Risk Management Framework
Voluntary in the United States, and the most useful available vocabulary for describing your controls to a board or an enterprise customer who asks. Map your gates to it once and reuse the mapping.
FTC advertising substantiation
Applies to a claim regardless of whether a model or a copywriter wrote it. An AI system that cannot show which claims it made, where, and on what basis is an substantiation problem waiting for a complaint.
Sector rules you already live under
Healthcare, financial services, alcohol, and state-level advertising and recording-consent rules do not change because the author is a model. The overlay has to encode them per jurisdiction, because your locations are not all in one.

What you should not govern

  • Internal drafts nobody outside the team will see. Governing these adds cost and removes the fast iteration that makes the system good.
  • Word choice within an already-approved claim. If the claim passed, arguing about synonyms is a taste dispute wearing a compliance badge.
  • Surfaces with no customer visibility and no regulatory exposure — internal summaries, backlog notes, drafts in a private queue.
  • Every output equally. Uniform scrutiny means the high-risk output gets the same thirty seconds as the low-risk one, which is how real problems slip through.

Over-governance fails in a way that looks responsible, which is why it survives far longer than under-governance. A team that reviews everything equally is a team that reads nothing carefully.

Seven questions for anyone selling you governance

A policy vendor answers these with frameworks and a document. A practice answers with numbers, a record, and a name. Ask us the same seven.

  1. 1Show me a governance system you built that is running unattended today, and tell me its auto-publish rate.
  2. 2What is the threshold, what units is it in, and who can change it?
  3. 3What exactly is written to the audit log for one published output? Show me a record.
  4. 4How does an override become a change to the system rather than a one-time correction?
  5. 5What did you deliberately choose not to govern, and why?
  6. 6What do I own at the end — the rules, the thresholds, the scored history?
  7. 7Which of these controls would you remove if I told you the queue was too slow?

The last one is the most revealing. Everyone has controls when the queue is empty; the question is which ones survive contact with a deadline.

Common questions

What is AI governance in a marketing context?
Five things, made operational rather than written into a policy document. What may be said — the claims allowed per vertical and jurisdiction, expressed as checks an output can be scored against. Who approves what — routing rules that send an output to auto-publish, review, or escalation based on a score rather than a job title. How much autonomy each surface gets, set per surface and per location rather than globally. What gets recorded — for every published output, the input, model and version, rules evaluated, score, routing decision, and any human touch. And what happens when it is wrong, including the feedback path that turns an override into a change to the system.
What are the levels of AI autonomy for marketing content?
Five, in practice. Level 0 draft only, where a human publishes everything. Level 1 review then publish, with the queue ordered by confidence so attention goes where it matters — most surfaces should sit here for a first month. Level 2 auto-publish above a threshold, which is where labour saving actually appears and where most operators stall because raising the threshold feels risky without evidence. Level 3 auto-publish with sampled post-hoc review to detect drift, appropriate once months of scored history show the gate holds. Level 4 autonomous with exception handling, which few marketing surfaces need and nobody should claim before Level 2 has run for a quarter.
How do you set an AI auto-publish threshold?
As a number in defined units, changeable on purpose by a named owner. Score every output on the dimensions that matter — brand fit, claim compliance, factual grounding against the master record — then set the auto-publish cut where your reviewed sample shows an acceptable error rate, and record why. The common failure is not a threshold set too high or too low; it is a threshold nobody can locate, in units nobody can explain, that therefore never moves. A gate that cannot be tuned becomes a queue, and a queue that grows gets approved in bulk without reading.
What should you not govern?
Internal drafts nobody outside the team will see, since governing them adds cost and removes the fast iteration that makes the system good. Word choice inside an already-approved claim, which is a taste dispute wearing a compliance badge. Surfaces with no customer visibility and no regulatory exposure. And crucially, not every output equally — uniform scrutiny means a high-risk output gets the same thirty seconds as a low-risk one, which is how real problems slip through. Over-governance fails in a way that looks responsible, which is why it survives longer than under-governance.
How do you tell an AI governance consultant from a policy vendor?
Ask to see a governance system running unattended today and its auto-publish rate. Ask what the threshold is, in what units, and who can change it. Ask what exactly is written to the audit log for one published output, and to see a record. Ask how an override becomes a change to the system rather than a one-time correction. Ask what they deliberately chose not to govern. A policy vendor answers these with frameworks and a document; a practice answers with numbers, a record, and a name.

How this shows up in practice

Governance is not a separate engagement here — it is the gate, the routing, and the audit trail built alongside the agents, because a control added afterwards is a control nobody designed the system around.

Ready to talk instead? Book the 30-minute consultation, or take the three-question diagnostic first. No email required for the diagnostic.

Part of what an AI strategy engagement should deliver.

Related: guardrails with override-learning · routing-decision audit trail · the governance decision router

Free — before you talk to anyone

The Scope and Sequence Sheet

Get the deliverable list for this engagement in writing — every artifact, the week it lands, and exactly what you own at the end.

  • The full deliverable list: code, prompts, configs, playbooks, vendor selections, and training material — all of it yours.
  • The week-by-week timeline, with the checkpoint that ends each phase.
  • The scoping questions worth answering before any call, so the first conversation starts at the real problem.

Opens on this page immediately. No attachment, no waiting on an email.