Knowing that ChatGPT mentions your company tells you almost nothing. A company that shows up in some answers could be leading its category or trailing it badly — the mention alone can’t tell you which. The only useful read is comparative: the same questions, asked the same way, across you and the three to five competitors a buyer would shortlist alongside you.
That’s the underlying problem with most conversations about AI visibility right now. Leaders check how they show up, feel either relieved or alarmed, and stop there. Neither reaction is justified without a baseline. When a buyer asks an AI assistant who to shortlist, the assistant is composing an answer from the whole competitive field. You should be reading it the same way.
A competitive AI search benchmark is a small, repeatable research exercise, not a software purchase. You can run the first pass in a week with a spreadsheet and a few hours of disciplined work.
Who belongs in the benchmark?
The three to five companies you actually lose deals to. Not the biggest names in the category, not the analyst-report leaders, not the aspirational peer set your board likes to reference. The companies your sales team hears about on discovery calls.
This matters because a benchmark against the wrong cohort produces the wrong fixes. If you compare yourself to a public company with a decade of press coverage, you’ll conclude the gap is unclosable and give up. If you compare yourself to the mid-market rival who wins your deals, you’ll find gaps you can close in a quarter.
Ask your head of sales for the loss reasons from the last two quarters. The names that appear repeatedly are your cohort. Cap it at five. Beyond that, the analysis gets noisy and nobody reads the report.
What questions should you ask?
The buyer’s own phrasing, at the shortlist stage. Not “what is [category]” — a buyer that early isn’t choosing anyone yet. You want the questions someone asks when they’re assembling a list: who are the best providers of X for a company like mine, how does A compare to B, what should I look for when evaluating X, who do companies our size typically use.
Start with the people closest to the buyer. Sales calls, onboarding interviews, the exact language prospects use when they describe their problem. Buyers rarely use your category label. They describe a situation.
Write ten to fifteen questions across a few intents: recommendation questions (“who should we consider for…”), comparison questions (“how does [you] compare to [competitor]…”), and evaluation questions (“what are the drawbacks of…”). Then stop. Consistency matters more than volume. Fifteen questions asked identically every quarter beats fifty questions rewritten every time, because the rewrite destroys your ability to see movement.
How do you run it honestly?
Answers vary run to run. That’s not a flaw in your method; it’s how these systems work. Ask the same question twice and you may get different companies, different orderings, different caveats. A single screenshot is an anecdote, and anecdotes make bad strategy.
So repeat each question — several runs per assistant, across ChatGPT, Gemini, and Claude — and read patterns, not instances. The signal you’re looking for isn’t “were we mentioned in this answer.” It’s “across all runs, how often does each company in the cohort appear, and in what position, and described how.” A competitor who appears in most runs across all three assistants is genuinely ahead of one who appears occasionally in one. A company that appears once, in one run, might be noise.
Log everything in a spreadsheet: question, assistant, date, which cohort companies appeared, how each was described, what sources were cited when the assistant surfaced any. Tedious, yes. But this log is the whole asset. It’s what makes the next quarter’s run a comparison instead of a fresh guess.
What should you compare beyond “were we mentioned”?
Mention frequency is the scoreboard. The signals underneath it explain the score, and those are what you can act on. For each company in the cohort, including yours, compare five things:
- Accuracy of description. When the assistant describes what each company does, is it right? Is it current? A competitor described crisply and correctly has fed the ecosystem clear language about itself. A company described vaguely, or with a positioning it abandoned two years ago, has a messaging problem before it has an AI problem — usually an entity clarity problem.
- Cited sources. When assistants point to sources, whose pages get cited — the company’s own site, or third parties? Which third parties? This tells you where the assistants are learning about your category, and where you’re absent.
- Page-level answers. Take your question set and read each company’s website against it. Does a specific page plainly answer “who is this for, what does it do, how is it different”? Many companies fail this test on their own homepage. The competitor who wins in AI answers usually has pages that answer buyer questions in direct language — the content structure AI assistants can cite.
- Third-party validation. Who else vouches for each company — industry publications, directories, review platforms, partner pages? Assistants compose answers partly from what the broader web says. A company that only talks about itself has one voice in the room. (More on authority and trust in AI search.)
- Technical readability. Can the site actually be read by a machine? If key content loads in ways a crawler can’t parse, or sits behind gates, the clearest positioning in the world never enters the conversation.
The pattern is usually humbling. Companies expect a mysterious algorithm gap and instead find a plain-language gap: the competitor gets recommended because its pages say clearly what it does, and yours don’t.
How do you turn the gap into fixes?
An AI assistant can only recommend what it can find, read, and understand. Every gap in your benchmark traces back to one of those three verbs — the same logic behind how AI assistants choose vendors.
So resist the giant roadmap. Pull three to five fixes from the benchmark that one team can ship in a quarter. Typical candidates: rewrite the two or three pages that should answer your highest-frequency buyer questions but don’t. Fix the description gap — if assistants describe you inaccurately, tighten the language on your own site until a stranger could restate your positioning in one sentence. Pursue presence on the two or three third-party sources the assistants cited most for your cohort. Resolve whatever technical issues keep pages from being readable.
Then stop and rerun before adding more. Small, sequenced, measured — that’s the feedback loop.
How often should you rerun it, and what goes upward?
Quarterly. Faster and you’re measuring noise; slower and you can’t connect fixes to movement. Same questions, same method, same cohort unless your loss data says the cohort changed.
Report upward on one page: mention frequency by company across assistants, quarter over quarter; the two or three most consequential description or citation gaps; the fixes shipped and what moved. Frame it honestly as an emerging channel you’re instrumenting early, not today’s number one source of pipeline. No one can guarantee a mention or a citation in these systems, and any report implying otherwise should make you skeptical of its author. What you can own is the decision to measure it rigorously before your competitors do.
Where does this break?
Small categories. If your niche has few published sources — little press, thin review coverage, sparse third-party writing — the assistants have little to compose answers from, and the answers get thin and unstable. Same question, wildly different responses, sometimes with companies that barely exist. In that situation, a benchmark mostly tells you the category itself is underdocumented, which is its own finding: the first company to publish clear, substantive answers to buyer questions has an open field. But treat quarter-over-quarter movement in a thin category with heavy skepticism.
One more boundary worth naming: this post covers the competitive comparison only. Measuring your own AI visibility over time — trend lines, coverage across your full question universe, share of AI voice, board reporting — is a separate discipline with its own method. Related work, different exercise. Don’t collapse the two into one spreadsheet. And if organic traffic is already sliding, sort out whether it’s AI or you before you benchmark anything.
What’s the real question underneath all this?
The benchmark is a diagnostic, and like most good diagnostics, it usually surfaces an underlying problem that predates AI: unclear positioning, pages that don’t answer buyer questions, a reputation footprint that’s thinner than the sales team assumes. The assistants didn’t create those gaps. They exposed them, in a format a CEO can finally see side by side with the competition.
That’s where the work stops being a research exercise and starts being a strategy question: which gaps matter most for the deals you’re losing, in what order, owned by whom. That’s the work I want to do alongside a leadership team — reading the benchmark against the revenue picture and turning it into a sequenced plan. If you’d rather start with a structured read of where your site stands against its competitors, the AI search audit does exactly that. If your buyers are already asking the assistants who to shortlist, the comparison is happening whether you’re measuring it or not. Better to be the company holding the spreadsheet.