How to evaluate an AI marketing agency
Almost every marketing vendor now describes itself as AI-powered, and the phrase has stopped carrying information. The useful question is not whether a vendor uses AI. It is whether what they build keeps working after they leave.
We are one of the options on this page, and we are the third shape rather than the first. That is a disclosure, not a recommendation — the first shape is the right answer for a lot of companies, and the questions below work identically on us.
Published August 21, 2026· 10-minute read
Three shapes of help
Most disappointing engagements in this category are a shape mismatch rather than a quality problem. The vendor did what their shape does; it was not what the company needed.
An AI marketing agency
An agency that executes marketing work, using AI internally to do it faster or cheaper than a traditional agency would.
Right when: You have a defined volume of work — content, creative, campaign production — and you want it produced at lower cost per unit. You already know what should be made.
Wrong when: You need the system rebuilt rather than the output increased. An agency is measured on delivery, so it will deliver, whether or not delivery was the constraint.
The question that tells you: Ask what happens to the prompts, workflows, and fine-tuned assets when the contract ends. If they leave with the agency, you rented output.
An AI marketing consultant
An individual or small firm that assesses your operation and recommends what to build, sometimes staying to help build it.
Right when: You are early enough that the main risk is building the wrong thing, and a clear-eyed diagnosis is worth more than execution capacity.
Wrong when: You already know what to build. Paying for a second opinion you will not act on differently is the most common waste in this category.
The question that tells you: Ask whether they will still be there in ninety days. A recommendation whose author never has to run it is cheaper to write and worth proportionally less.
An embedded operator
A fractional executive who works inside your marketing function on an ongoing basis and is accountable for a number, operating AI systems as part of the job rather than as the product.
Right when: The problem is coordination across many surfaces or many locations, and it needs someone with decision rights rather than someone with a statement of work.
Wrong when: You have no marketing team to embed inside, or the work is a defined project with an end date.
The question that tells you: Ask how many days a week and what they can decide without asking. Vague answers to either mean this is advisory work wearing an operator label.
Five tells that separate capability from a wrapper
They cannot say what the system does when the model is wrong
Every real deployment has a gate, a threshold, and a human queue. A vendor without an answer here has not run anything at volume, because volume is where being wrong five percent of the time becomes visible.
The demo has no data in it
Ask them to run it on twenty of your real records, chosen by you, in the meeting. Capability and demo quality are almost uncorrelated, and this single request separates them faster than any reference call.
There is no source of truth in the architecture
Ask where the canonical record lives — the thing every AI output reads from before it writes. If the answer is that each tool holds its own copy, the outputs will contradict each other, and it will take a quarter for that to become obvious.
Nothing is emitted back
Ask what the system tells the rest of your stack after it acts. Systems that only consume are islands; each new one costs as much as the last. Systems that emit make the next one cheaper. This is the single clearest indicator of whether you are buying a tool or a capability.
The pricing is per seat, and the work is per location
Pricing that does not match the shape of your operation signals a product built for a different buyer. It is not disqualifying, but it predicts where the roadmap goes.
The fourth tell is the one worth pressing hardest. A system that acts without telling anything else what it did is invisible to the rest of your stack, so the next system you buy has to rebuild the same context from scratch. That is the mechanical reason some marketing operations get cheaper per capability added while others get more expensive with every tool.
Nine questions
Ask every one, of everyone, including us.
- 1What have you built that is still running unattended today, and how long has it been running?
- 2Run your system on twenty records I choose, live, in this meeting.
- 3Where does the canonical record live, and what happens when two systems disagree about a location?
- 4What does the system do when it is not confident? Show me the queue.
- 5What does it emit after it acts, and to where?
- 6What do I own at the end — code, prompts, configurations, the approved-and-rejected output history?
- 7Who specifically works on my account, and what else are they on?
- 8What is the first thing you would tell me not to automate?
- 9Show me an engagement where the honest recommendation was to do less.
The last two do the most work. A vendor who cannot name something you should not automate has not thought about failure modes, and a vendor with no engagement where the honest advice was to do less has either not run many or is not telling you about them.
Run a four-week paid pilot
No reference call resolves this as well as four weeks of real work. Structure it like this and the answer is unambiguous.
- 1
Pick one surface with high volume and low blast radius
Review responses, listing updates, or product descriptions. High volume so four weeks produces a real sample; low blast radius so a weak output is recoverable.
- 2
Give them your actual data, not a sanitised subset
Under an NDA if needed. Sanitised data hides exactly the messiness that determines whether the system works, which is why vendors are comfortable with it.
- 3
Define the acceptance test before they start
A specific number: the share of outputs publishable without human edit, or median latency, or error rate on a defined check. Agree it in writing, because it will be renegotiated at the end otherwise.
- 4
Pay for it
A free pilot is a sales activity and gets sales-quality effort. A paid pilot gets delivery-quality effort, and the price is a rounding error against a year of the wrong contract.
- 5
Keep everything produced
Write it into the pilot agreement. Even if you do not proceed, the prompts, the configuration, and the labelled outputs are worth more than the pilot cost and make the next evaluation faster.
Common questions
- What does an AI marketing agency do?
- An AI marketing agency executes marketing work — content, creative, campaign production, sometimes media buying — using AI internally to produce it faster or at lower cost than a traditional agency. The distinction that matters to a buyer is whether the AI capability transfers to you or stays with them. Most agencies in this category are measured on delivery, so they deliver, and the prompts, workflows, and tuned assets that made delivery cheap leave when the contract ends. Ask that question first, because it determines whether you are buying output or building capability.
- How do you tell a real AI capability from a wrapper?
- Five tells. They cannot say what the system does when the model is wrong — every real deployment has a gate, a threshold, and a human queue. The demo contains no data, so ask them to run it live on twenty records you choose. There is no source of truth in the architecture, meaning each tool holds its own copy of the location or customer record and outputs will eventually contradict each other. Nothing is emitted back to the rest of your stack, which makes every system an island where each new one costs as much as the last. And pricing is shaped for a different buyer than you, which predicts where the roadmap goes.
- Should you hire an AI marketing agency, a consultant, or an embedded operator?
- An agency when you have a defined volume of work and want lower cost per unit — you already know what should be made. A consultant when you are early enough that the main risk is building the wrong thing and diagnosis is worth more than execution capacity; the test is whether they will still be there in ninety days, since a recommendation whose author never has to run it is worth proportionally less. An embedded operator when the problem is coordination across many surfaces or locations and needs someone with decision rights rather than a statement of work — which requires an existing marketing team to embed inside.
- How do you run a pilot with an AI marketing vendor?
- Pick one surface with high volume and low blast radius, such as review responses or listing updates, so four weeks produces a real sample and a weak output is recoverable. Give them your actual data rather than a sanitised subset, because sanitised data hides the messiness that determines whether the system works. Define the acceptance test as a specific number before they start and agree it in writing. Pay for the pilot — a free pilot is a sales activity and gets sales-quality effort. And write into the agreement that you keep everything produced, since the prompts, configuration, and labelled outputs are worth more than the pilot cost even if you do not proceed.
- What questions should you ask an AI marketing agency?
- What have you built that is still running unattended today, and for how long. Run your system on twenty records I choose, live, in this meeting. Where does the canonical record live and what happens when two systems disagree. What does the system do when it is not confident — show me the queue. What does it emit after it acts, and to where. What do I own at the end, including the approved-and-rejected output history. Who specifically works on my account and what else are they on. What is the first thing you would tell me not to automate. And show me an engagement where the honest recommendation was to do less.
Where we fit
The third shape, for multi-unit franchises, multi-location retail, and DTC ecommerce. Every deliverable, the week it lands, and what you own at the end are already published, so you can hold them against anyone else on your list before a conversation happens.
Ready to talk instead? Book the 30-minute consultation, or take the three-question diagnostic first. No email required for the diagnostic.
If the third shape sounds right, what a fractional CMO is and how to hire one cover it without a sales conversation.