TL;DR

A gap is the distance between the questions your buyers ask AI and the answers where AI names you. Finding it honestly is harder than it sounds.

  • The prompt list decides the finding, so a list you brainstorm will flatter you
  • Build it from customer language, demand data, query fan-out and live monitoring
  • Four gap types: absent, described wrong, framed as an also-ran, named too late
  • Most gaps are authority problems wearing a content costume

Your buyers are asking AI about your category today. Some of those answers name you. Most of them almost certainly don’t, and the difference between the two is the only thing worth measuring.

Finding that difference is straightforward work. Finding it without fooling yourself isn’t, because the method most teams use quietly guarantees a flattering result. That’s what this piece is about: how to build the test, what the results mean, and which gaps are worth your budget.

What counts as a gap in your AI search visibility?

A gap is the distance between the questions your buyers ask AI and the answers where AI names you. Ranking drops and visibility gaps measure different systems, and one doesn’t predict the other. You can hold page one on Google for your best keyword and still be missing from every answer a buyer reads before they run a single search.

That last part is what makes it urgent rather than interesting. Gartner surveyed 645 B2B buyers and found that 45% used generative AI during a purchase, mostly to gather information on vendors and products. The same survey found 69% then take what AI told them to a sales rep to check it. So the AI answer sets the agenda for the human conversation that follows, and decides who gets invited into it. That’s an influence channel, and it’s running whether or not you’re measuring it.

We’ve written separately about why your Google rankings don’t tell you what AI says, so we won’t repeat it here.

Why most AI visibility gap reports flatter you

Most gap reports test prompts the company picked itself, and companies pick prompts they expect to win. The output reads well and tells you nothing. The list decides the finding before a single prompt runs, so a list built from a brainstorm produces a picture that confirms what the room already believed.

Watch how it happens. Someone books an hour, gets marketing in a room, and asks everyone to write down the questions buyers ask. What comes out is the language the company uses about itself: its own category name, its own feature vocabulary, the comparisons it likes. They run those prompts, appear in a decent share of them, and conclude the gap is smaller than feared.

Of course they appear, and it’s a striking thing to watch a team celebrate. Those are the questions the company has been answering on its own website for three years.

Heads up

If your prompt list came out of an internal workshop, treat every result from it as provisional. You have measured how well you answer your own questions.

The prompts that would have shown a real gap never made the list, because nobody in that room thinks in the words a buyer uses two months before they know what they need. Buyers don’t search your category name. They describe a problem, badly, in their own language, and let the model work out what they mean.

We build the list differently for exactly this reason, and it’s worth saying plainly: an audit is only as honest as its prompt list. Get the list wrong and everything downstream, every score, every chart, every recommendation, is a confident answer to the wrong question.

How to build a prompt list that tells you the truth

Build it from four sources rather than a brainstorm: the problems customers describe in their own words, search demand data showing which of those carry real volume, query fan-out to catch the sub-questions models generate on their own, and monitoring of what AI already says about your category.

None of these are exotic. What matters is that not one of them is you guessing.

The problems your customers describe

Start with recorded reality. Discovery calls, sales transcripts, support tickets, churn interviews, and the communities where your buyers complain in public. You’re looking for the sentence a person said before they knew what to call the thing they needed.

This is primary research, and it’s the richest source and the least used, because it takes an afternoon of reading rather than an hour of brainstorming. It’s also the only place you find phrasings that would never occur to your team. A buyer doesn’t ask for a customer data platform. They ask why their sales team and their product team have different numbers for the same account.

Which of those questions carry demand

Now check the volume. We tried skipping this step early on and it didn’t work: some of what you pull from calls is genuinely one person’s odd phrasing, and chasing it wastes the run.

Search demand data tells you which problem statements enough people share to be worth testing. You’re not building a keyword list here, so don’t treat it like one. You’re sorting a long list of real language into the parts that represent a pattern and the parts that represent one deal.

The sub-questions you’d never have picked

When someone asks a model a broad question, it breaks that question into narrower ones, answers each of those, and assembles the result. That decomposition is called query fan-out, and it’s where most of the surprises live.

Feed your real buyer questions to the models and capture the sub-queries they generate. You’ll find questions nobody on your team would have written down, because the model built them, not you. What comes back is its idea of what your category involves, and if that idea is wrong or incomplete, you’ve found something more useful than a visibility score. You’ve found how AI understands your market.

This surprised us the first time we ran it. It’s also the source no other gap report we’ve read even mentions, and it’s the one that changes the shape of the list most.

What AI already answers about your category

Last, look at what the models say right now, unprompted by you. Pay attention to which of your pages get pulled in, because the page types that get cited are rarely the ones teams expect. Who gets named on the broad category questions. Who gets recommended when the buyer narrows. What tone the answers carry about each vendor.

That gives you the competitive half of the picture, and it’s the part that gets attention internally. It’s much easier to fund work when you can show a founder that four competitors are named on the question their buyers ask most and the company isn’t one of them.

Somewhere between forty and eighty prompts, built this way, gives a stable read. The number matters far less than where it came from.

The four types of gap, and what each one means

Four types, and they call for different work. You’re absent from the answer entirely. You’re named but described wrongly. You’re named but framed as an also-ran. Or you’re named only once the conversation has narrowed, after the shortlist was already drawn.

Gap type What it looks like What to investigate
Absent AI names four vendors and none of them are you Relevance, discoverability, and supporting evidence
Described wrong You’re named, but the category, pricing or capability is off Conflicting or outdated source information
Framed poorly You’re named last, with a caveat, or as the budget option Comparison criteria and third-party evidence
Named too late You appear on narrow comparisons but not broad category questions Coverage and evidence for broader buyer questions

The fourth one gets missed constantly, and we’d look at it first.

AI conversations start broad and narrow down. A buyer asks a wide question, then keeps asking tighter ones until the model names a final three or four. If AI already knows you while the conversation is broad, you stay in contention as it narrows. If you only show up at the narrow end, you’re relying on being added late to a list that’s already been drawn, against criteria that were set without you.

That’s a different problem from the other three, and a more expensive one. It’s also the one where the broad questions are still open territory in most categories, because almost nobody is working them.

A worked example: turn an AI visibility gap into a repair

A useful gap record connects an observed answer with a source you can inspect and an action you can take. The example below is illustrative, using a fictional SaaS vendor. It shows how we would investigate an inaccurate product description without treating one response as proof of a wider pattern.

Suppose a buyer asks, “Which customer data platforms support server-side events for a small product team?” The assistant recommends two vendors and says fictional vendor Northstar Data lacks that capability. Northstar’s current documentation says it supports it, while a cited comparison still describes an older product version.

Step What to record or do
Save the observation Record the exact prompt, date, platform, model if shown, answer text, and cited comparison URL. Label the finding as an inaccurate capability claim.
Verify the source Read the comparison and current product documentation. Identify the conflicting passage and confirm the feature’s availability with the product team.
Check whether it repeats Run the same question again in fresh conversations with consistent settings. Keep each result, including answers that describe the product correctly.
Choose the repair Make the existing integration documentation explicit about setup and availability. Request a factual correction from the comparison publisher, supplying the supporting documentation.
Connect the evidence Link the relevant product page to the integration documentation so readers can verify the capability. A second general blog post about the same feature is unnecessary.
Review the outcome Record when the source changes, then repeat the saved prompts. Track whether the incorrect claim recurs. A correction does not guarantee a recommendation.

Now consider a different result: the assistant describes Northstar accurately but never recommends it for the buyer’s requirements. Rewriting the same feature paragraph would not address the finding. We would investigate fit, comparison criteria, customer evidence, and the sources used to support the recommended vendors.

For an actual client engagement rather than this fictional example, our published SaaS case study explains the work and results over eight months.

A template for recording AI visibility gaps

Use one row per observed answer and keep the evidence attached to it. The fields below can be copied into a spreadsheet. They separate what happened from your explanation, proposed repair, and later result, so a suspected cause does not quietly become an established fact in your report.

Field What to enter
Run details Date, platform, model if visible, locale, search setting, and whether the conversation was fresh
Buyer question Exact prompt, source of the question, and buyer stage
Observed answer Saved answer or screenshot; brands recommended, mentioned, and cited recorded separately
Gap and evidence Absent, inaccurate, poorly framed, or late-stage only; quote the relevant passage and save cited URLs
Working explanation What may explain the finding, what supports it, and what still needs checking
Repair Existing page to update, missing page to create, or third-party source to correct; owner and completion date
Follow-up Recheck date, observed answer, source changes, and whether the same error or omission repeats

Prioritize a repeated factual error on a sales-critical question before a one-off omission on a broad topic. If you need the full research process, use our AI visibility audit walkthrough alongside this template.

Let's build your revenue engine.

A 30-minute call where we map what AI says about you and build a tailored roadmap.

Book a call →

What closes each type of gap

Absence is an authority problem wearing a content costume. Wrong descriptions are a source problem, so you fix what AI found. Poor framing sits on other people’s sites. And being named too late means you’re missing from the broad end of the conversation, which is the slowest of the four to shift and the one to start earliest.

Take those in turn.

If you’re absent, the instinct is to write more. Resist it. Absence often concentrates on one assistant before the others, and Perplexity is the one where it shows first. A site with authority behind it can publish a plain definition and get pulled into answers a better page on a no-name domain never reaches. The work that moves absence is the work that earns credibility: digital PR, links, brand mentions, reviews, genuine presence in the places your market already reads. Build the entity first. Publishing more into a domain nobody references is an act of faith.

If you’re described wrongly, go and find what AI is reading, and read why AI describes your product wrong for the longer version. Follow the citations on the answers that get you wrong. Nine times out of ten it’s a stale directory entry, an old comparison article, or a review page describing a version of your product you retired. Fix the source and the answer moves, often within weeks.

If you’re framed poorly, the problem is sitting on other people’s sites. Reviews, roundups and community threads carry more weight in these answers than anything you publish about yourself, which we’ve covered in how reviews and third-party mentions shape AI recommendations. Thin evidence there reads as a thin vendor.

And if you’re named too late, you need to be present in the category conversation before it narrows. That means being useful about the problem rather than about yourself. Early on, while the model is still working out who the players are and what matters, self-promotion gives it nothing to work with and gets skipped.

Whichever type you’re dealing with, decide what you’ll track before you start, so you can tell whether it worked. Share of voice, accuracy and sentiment move first, and they are the numbers our AI Search Optimization work is built to move. Pipeline follows, and we’ve set out how to connect AI search visibility to pipeline separately.

Frequently asked questions

Four questions come up every time we walk a team through this: how many prompts are enough, whether you need software to do it, how it differs from the SEO audit they already run, and how often to repeat it. Short answers below, and none of them require buying anything before you start.

How many prompts do you need before the picture is real?

Between forty and eighty, in our experience, and where they came from matters more than how many there are. Twenty prompts built from real customer language will tell you more than two hundred generated from your own feature list.

Can you do this without a monitoring tool?

For a first read, yes. Run the prompts by hand across the main assistants, record who gets named and what gets cited, and you’ll find the big gaps. Tooling earns its cost when you need to track the same set on a schedule and prove movement.

How is this different from an SEO audit?

An SEO audit asks whether your pages can rank. This asks whether AI names you in answers, which depends heavily on what other sites say about you. The disciplines are the same; the signals are different, and so is the measurement.

How often should you re-check?

Quarterly for the full set, monthly for a smaller tracking subset. We’ve covered how often to re-run it alongside the audit walkthrough.

Rafael De Jesus
Founder
Organic growth and AI search optimization specialist. Rafael has added seven figures in ARR to B2B SaaS and AI-native companies, built a 200K+ following, and generated 3 billion+ impressions across search and social. He writes about how to shape what AI says about your brand.