Criteria for evaluating agencies for AI visibility
The checklist above is the version you can run in a first call. The sections below cover what to ask and how to read the answer you get.
1. Relevant case studies
The most important criterion is verified case studies, ideally from your industry or one with a similar buying cycle. Ask each agency for their best examples of client brands appearing in your target AI search or chat platform, then check the prompts yourself while signed out. If they cannot produce this, they may be selling theory or too new to have results.
A case study worth reading tells you:
- Which prompts the brand appears in, and what the answer said before the work stated
- What the agency changed, at the level of specific pages, sources, and off-site assets
- How long the change took to show up in model responses
- Whether the visibility gain connected to pipeline, or stayed a vanity metric
Run the prompts in more than one model and more than one session. Answers vary between accounts and over time, so a single clean screenshot proves less than a pattern you can reproduce.
2. Whether the agency appears in AI search itself
An agency selling AI visibility should have some. Ask Google, ChatGPT or Perplexity for the best AI visibility agencies for B2B SaaS, or for whatever category describes you, and see which names come back.
When we built our list of top AI search agencies, we used AI visibility tracking data as the leaderboard, because an agency’s presence in AI answers is one of the few claims about its competence that you can check in a few seconds.
Treat it as a signal with limits though because agencies optimize themselves harder than they optimize clients, and category prompts about marketing services have thick training data compared with a niche industrial software category. An agency invisible in its own category is still a fair question to raise on the call.
3. How they explain the AI search ecosystem
You want a GEO agency that can explain where an AI answer comes from before it explains what it would sell you. There are two mechanisms at work, and they need different fixes. A model answering from memory reflects what its training corpus absorbed about your brand, which is slow to change and mostly shaped by third-party sources. A model answering with live retrieval pulls current pages and cites them, which is faster to influence and closer to conventional search work.
An agency that collapses those two into one story will sell you one tactic for both problems. Ask which of the two is costing you more right now, and expect an answer that references your category instead of a general principle.
The same question applies across platforms. Google AI Overviews and AI Mode draw on the search index in ways that reward strong technical and content foundations. LLMs with browsing behave differently again. A good answer to “which platform matters most for our buyers” is never “all of them equally.”
4. Entity disambiguation and brand consistency
For brand visibility, models care about entity presence, positioning consistency, external mentions, and source freshness. If you ask a potential agency about entity disambiguation and they do not immediately bring Wikipedia and Wikidata, knowledge graphs, schema markup, or the industry directories and review sites your category runs on, that is a red flag.
You may even give them a real problem to solve on the call. If LLMs confuse you with a similarly named company, or describe a product line you discontinued, ask what they would do about it. The strong answers describe fixing sources you do not own. The weak ones describe adding a page to your website or publishing more content only. We covered how ChatGPT builds and holds brand knowledge in more depth if you want a reference point for judging their reply.
5. Metrics, reporting, and who owns them
Before you sign, get the reporting spec in writing. AI visibility measurement answers a different question than SEO reporting does: what models say about you, which sources they cite when they say it, and how that shifts over time.

The metrics that belong in the report:
- Presence and share of voice across a defined prompt set, tracked per platform
- Citation frequency, plus which of your URLs and which third-party domains get cited
- Sentiment and factual accuracy of how models describe your positioning
- Competitor presence in the same prompts, so you can see relative movement
- AI-referred sessions in GA4 and what those sessions do after they land
- Pipeline contribution, once enough time has passed to attribute anything
Keyword rankings and organic sessions still belong in the report, since retrieval-based answers lean on search indexes and the technical foundations carry over. What should worry you is AI visibility reporting that is a rankings dashboard with a new cover page.
Ask how the report gets produced and how often. Expect a dedicated AI visibility tracking platform for prompt monitoring, mainstream SEO tools for the retrieval and technical side, and GA4 configured to separate AI referrals from the rest of your traffic. The tool names matter less than what the agency does with the output, so push on how a month of tracking data turns into a decision about what to change next.
Then ask who keeps the dashboard, the prompt library, and the baseline when the contract ends. If the monitoring lives entirely inside the agency’s account, your visibility history leaves when they do, and the next agency starts you from zero.
6. Content practices
Ask what they would stop or start publishing and how they would structure the content.
The formats that earn them are definitions and direct answers near the top of a page, comparison and alternatives content, documentation and support pages, and anything carrying first-party data a model cannot get elsewhere.

Volume-first content plans transfer poorly here. An agency that answers the question above by describing a publishing quota has not adjusted its content model, no matter what their deck says. The structural work that gets content cited is closer to editing and consolidation than to production.
For SaaS products, push hardest on documentation and support content, since those pages carry the specifics models reach for during evaluation. Ask how they would handle gated or authenticated docs, changelogs, and integration pages.
7. Technical execution
Ask what they would change in your stack in the first month, and listen for whether the answer is specific to how your site is built.
The questions worth asking are whether your CDN or WAF blocks the AI crawlers you want in, whether JavaScript rendering hides content from parsers, and whether structured data survives your CMS. Ask their position on llms.txt as well. There is no consensus on its value, so the answer matters less than whether they can defend one.
An agency working on your site should also know what changes when a page has to serve a reader and a parser at the same time.
8. Whether they will train your team
Ask whether the agency trains your team, and how they answer tells you what kind of relationship you are buying. Some agencies build a dependency on purpose. Others will hand over the frameworks or even train your marketing team on AI search.
True AI search expertise is scarce right now, and the teams who internalize it while it is scarce compound the advantage. Outsourcing the whole function means the knowledge stays with your vendor, and you renegotiate from a weaker position every year.
What good training looks like:
- Built on your data, your ICPs, and your competitors
- Includes executives, since AI visibility ownership crosses marketing, sales, and product
- Ends with a working monitoring dashboard your team runs, plus a documented start, stop, and continue plan
- Skips the fundamentals your team already has, which requires the agency to find that out first
We built our AI search training around that shape after enough conversations that started with an executive asking why the brand was missing from ChatGPT. It runs on a two-week data collection window before the session, so the working block reviews your live audit findings and ends with your roadmap. Total team time is under four hours.
Training and a retainer are not competing purchases. Plenty of companies do both, with the agency running execution while the in-house team learns to own measurement. If you want to check where your team currently sits before you decide, our readiness assessment scores the gap, and review the available GEO courses for teams that need a cheaper starting point.
Red flags to avoid in an AI visibility agency
Missing any of the criteria above is itself a warning. These are the specific signals that should end the conversation:
- No case studies showing growth you can verify by running the prompts yourself.
- They cannot explain where AI answers come from, or they describe memory and live retrieval as the same thing.
- No AI-specific methodology and/or deliverables. Overlap between SEO and AI search is real and the technical foundations carry over, so the red flag here is an identical scope of work, not an acknowledgement that the two connect.
- Ranking data is the primary evidence they offer for AI performance.
- They guarantee placements. Nobody controls which brands a model names in an answer, and anyone promising a fixed position is describing something they cannot deliver.
- No audit and no baseline before the engagement starts, which conveniently makes the first report unfalsifiable.