A gap is the distance between the questions your buyers ask AI and the answers where AI names you. Finding it honestly is harder than it sounds.
- The prompt list decides the finding, so a list you brainstorm will flatter you
- Build it from customer language, demand data, query fan-out and live monitoring
- Four gap types: absent, described wrong, framed as an also-ran, named too late
- Most gaps are authority problems wearing a content costume
Your buyers are asking AI about your category today. Some of those answers name you. Most of them almost certainly don’t, and the difference between the two is the only thing worth measuring.
Finding that difference is straightforward work. Finding it without fooling yourself isn’t, because the method most teams use quietly guarantees a flattering result. That’s what this piece is about: how to build the test, what the results mean, and which gaps are worth your budget.
What counts as a gap in your AI search visibility?
A gap is the distance between the questions your buyers ask AI and the answers where AI names you. Ranking drops and visibility gaps measure different systems, and one doesn’t predict the other. You can hold page one on Google for your best keyword and still be missing from every answer a buyer reads before they run a single search.
That last part is what makes it urgent rather than interesting. Gartner surveyed 645 B2B buyers and found that 45% used generative AI during a purchase, mostly to gather information on vendors and products. The same survey found 69% then take what AI told them to a sales rep to check it. So the AI answer sets the agenda for the human conversation that follows, and decides who gets invited into it. That’s an influence channel, and it’s running whether or not you’re measuring it.
We’ve written separately about why your Google rankings don’t tell you what AI says, so we won’t repeat it here.
Why most AI visibility gap reports flatter you
Most gap reports test prompts the company picked itself, and companies pick prompts they expect to win. The output reads well and tells you nothing. The list decides the finding before a single prompt runs, so a list built from a brainstorm produces a picture that confirms what the room already believed.
Watch how it happens. Someone books an hour, gets marketing in a room, and asks everyone to write down the questions buyers ask. What comes out is the language the company uses about itself: its own category name, its own feature vocabulary, the comparisons it likes. They run those prompts, appear in a decent share of them, and conclude the gap is smaller than feared.
Of course they appear, and it’s a striking thing to watch a team celebrate. Those are the questions the company has been answering on its own website for three years.
If your prompt list came out of an internal workshop, treat every result from it as provisional. You have measured how well you answer your own questions.
The prompts that would have shown a real gap never made the list, because nobody in that room thinks in the words a buyer uses two months before they know what they need. Buyers don’t search your category name. They describe a problem, badly, in their own language, and let the model work out what they mean.
We build the list differently for exactly this reason, and it’s worth saying plainly: an audit is only as honest as its prompt list. Get the list wrong and everything downstream, every score, every chart, every recommendation, is a confident answer to the wrong question.
How to build a prompt list that tells you the truth
Build it from four sources rather than a brainstorm: the problems customers describe in their own words, search demand data showing which of those carry real volume, query fan-out to catch the sub-questions models generate on their own, and monitoring of what AI already says about your category.
None of these are exotic. What matters is that not one of them is you guessing.
The problems your customers describe
Start with recorded reality. Discovery calls, sales transcripts, support tickets, churn interviews, and the communities where your buyers complain in public. You’re looking for the sentence a person said before they knew what to call the thing they needed.
This is primary research, and it’s the richest source and the least used, because it takes an afternoon of reading rather than an hour of brainstorming. It’s also the only place you find phrasings that would never occur to your team. A buyer doesn’t ask for a customer data platform. They ask why their sales team and their product team have different numbers for the same account.
Which of those questions carry demand
Now check the volume. We tried skipping this step early on and it didn’t work: some of what you pull from calls is genuinely one person’s odd phrasing, and chasing it wastes the run.
Search demand data tells you which problem statements enough people share to be worth testing. You’re not building a keyword list here, so don’t treat it like one. You’re sorting a long list of real language into the parts that represent a pattern and the parts that represent one deal.
The sub-questions you’d never have picked
When someone asks a model a broad question, it breaks that question into narrower ones, answers each of those, and assembles the result. That decomposition is called query fan-out, and it’s where most of the surprises live.
Feed your real buyer questions to the models and capture the sub-queries they generate. You’ll find questions nobody on your team would have written down, because the model built them, not you. What comes back is its idea of what your category involves, and if that idea is wrong or incomplete, you’ve found something more useful than a visibility score. You’ve found how AI understands your market.
This surprised us the first time we ran it. It’s also the source no other gap report we’ve read even mentions, and it’s the one that changes the shape of the list most.
What AI already answers about your category
Last, look at what the models say right now, unprompted by you. Pay attention to which of your pages get pulled in, because the page types that get cited are rarely the ones teams expect. Who gets named on the broad category questions. Who gets recommended when the buyer narrows. What tone the answers carry about each vendor.
That gives you the competitive half of the picture, and it’s the part that gets attention internally. It’s much easier to fund work when you can show a founder that four competitors are named on the question their buyers ask most and the company isn’t one of them.
Somewhere between forty and eighty prompts, built this way, gives a stable read. The number matters far less than where it came from.
The four types of gap, and what each one means
Four types, and they call for different work. You’re absent from the answer entirely. You’re named but described wrongly. You’re named but framed as an also-ran. Or you’re named only once the conversation has narrowed, after the shortlist was already drawn.
| Gap type | What it looks like | What it usually is |
|---|---|---|
| Absent | AI names four vendors and none of them are you | An authority problem |
| Described wrong | You’re named, but the category, pricing or capability is off | A source problem |
| Framed poorly | You’re named last, with a caveat, or as the budget option | A third-party evidence problem |
| Named too late | You appear on narrow comparisons but not broad category questions | A presence problem, and the slowest to shift |
The fourth one gets missed constantly, and we’d look at it first.
AI conversations start broad and narrow down. A buyer asks a wide question, then keeps asking tighter ones until the model names a final three or four. If AI already knows you while the conversation is broad, you stay in contention as it narrows. If you only show up at the narrow end, you’re relying on being added late to a list that’s already been drawn, against criteria that were set without you.
That’s a different problem from the other three, and a more expensive one. It’s also the one where the broad questions are still open territory in most categories, because almost nobody is working them.
What a real gap looks like
Numbers make this concrete in a way that categories don’t. The pattern we see most often is a company scoring reasonably on branded and narrow questions, then dropping close to zero on the broad category questions where buyers start.
Run a properly built list and the shape usually looks like this. The company appears in a healthy share of the questions containing its own name. It appears occasionally on direct comparisons against one named competitor. And on the broad questions, the ones with no vendor named in them at all, it appears in low single digits while three or four incumbents split most of the answers between them.
That distribution tells you something a single score never would. The models know the company. They’ve filed it as a minor option, which needs different work from being unknown altogether.
We have written up what closing that gap looked like for one client over eight months. The comparison against competitors is the part that makes it land internally. A visibility score of 12% means nothing to a founder. The same finding phrased as four named competitors appearing on the question your buyers ask most, and you appearing on none of them, ends the debate about whether this matters.
A 30-minute call where we map what AI says about you and build a tailored roadmap.
What closes each type of gap
Absence is an authority problem wearing a content costume. Wrong descriptions are a source problem, so you fix what AI found. Poor framing sits on other people’s sites. And being named too late means you’re missing from the broad end of the conversation, which is the slowest of the four to shift and the one to start earliest.
Take those in turn.
If you’re absent, the instinct is to write more. Resist it. Absence often concentrates on one assistant before the others, and Perplexity is the one where it shows first. A site with authority behind it can publish a plain definition and get pulled into answers a better page on a no-name domain never reaches. The work that moves absence is the work that earns credibility: digital PR, links, brand mentions, reviews, genuine presence in the places your market already reads. Build the entity first. Publishing more into a domain nobody references is an act of faith.
If you’re described wrongly, go and find what AI is reading, and read why AI describes your product wrong for the longer version. Follow the citations on the answers that get you wrong. Nine times out of ten it’s a stale directory entry, an old comparison article, or a review page describing a version of your product you retired. Fix the source and the answer moves, often within weeks.
If you’re framed poorly, the problem is sitting on other people’s sites. Reviews, roundups and community threads carry more weight in these answers than anything you publish about yourself, which we’ve covered in how reviews and third-party mentions shape AI recommendations. Thin evidence there reads as a thin vendor.
And if you’re named too late, you need to be present in the category conversation before it narrows. That means being useful about the problem rather than about yourself. Early on, while the model is still working out who the players are and what matters, self-promotion gives it nothing to work with and gets skipped.
Whichever type you’re dealing with, decide what you’ll track before you start, so you can tell whether it worked. Share of voice, accuracy and sentiment move first, and they are the numbers our AI Search Optimization work is built to move. Pipeline follows, and we’ve set out how to connect AI search visibility to pipeline separately.
Frequently asked questions
Four questions come up every time we walk a team through this: how many prompts are enough, whether you need software to do it, how it differs from the SEO audit they already run, and how often to repeat it. Short answers below, and none of them require buying anything before you start.
How many prompts do you need before the picture is real?
Between forty and eighty, in our experience, and where they came from matters more than how many there are. Twenty prompts built from real customer language will tell you more than two hundred generated from your own feature list.
Can you do this without a monitoring tool?
For a first read, yes. Run the prompts by hand across the main assistants, record who gets named and what gets cited, and you’ll find the big gaps. Tooling earns its cost when you need to track the same set on a schedule and prove movement.
How is this different from an SEO audit?
An SEO audit asks whether your pages can rank. This asks whether AI names you in answers, which depends heavily on what other sites say about you. The disciplines are the same; the signals are different, and so is the measurement.
How often should you re-check?
Quarterly for the full set, monthly for a smaller tracking subset. We’ve covered how often to re-run it alongside the audit walkthrough.