On this page
- What an audit is for
- Audit, check, or tracking?
- The four layers
- Layer 1: the answer audit
- Layer 2: the source audit
- Layer 3: the entity audit
- Layer 4: the site audit
- Where competitors fit
- Scoring without inventing a number
- What the deliverable should contain
- Three illustrative examples
- How often to re-run it
- What an AI visibility audit cannot tell you
- Running audits for clients
- Where our tooling fits
- Start with one layer
- Frequently asked questions
Most brands discover their AI visibility problem the same way: someone asks ChatGPT for the best option in their category, and a competitor comes back instead. That tells you there is a problem. It tells you nothing about why, and almost nothing about what to do on Monday morning.
An audit is the step between noticing and fixing. This guide sets out a four-layer framework you can run yourself, what to record at each layer, how to score it without inventing a number, and where an audit stops being able to help.
- An AI visibility audit is a diagnostic, not a measurement. Its job is to explain a result and produce a fix list, not to produce a score.
- Work in four layers: the answers engines give, the sources they cite, how the web describes your brand, and whether engines can retrieve your site.
- The source layer is where most real findings come from, and it is the layer most audits skip.
- Competitors are a lens applied to every layer, not a separate stage of the audit.
- Audit quarterly and track continuously in between. A full audit repeated monthly mostly measures noise.
An AI visibility audit is a structured, point-in-time review of how AI answer engines see your brand and why, ending in a prioritised list of fixes. It samples the answers, traces the sources behind them, checks how consistently the web describes you, and tests whether engines can retrieve your site.
What an audit is for
A free check answers one question: are you in this answer. That is a symptom reading. It is genuinely useful, it takes a minute, and it is where most people should start.
An audit answers a different question: why does the answer look like this, and what would change it. That means working backwards from the answer to the machinery that produced it. Engines do not invent brands. They assemble answers from sources they retrieved, about entities they recognise, described in language they found somewhere. Each of those is auditable, and each is fixable.
The distinction matters commercially too. Plenty of teams buy a tracking tool, watch a number sit still for a quarter, and conclude AI visibility is not workable. Usually the number sat still because nobody diagnosed the cause. Tracking tells you whether things changed. An audit tells you what to change. It is also the most common complaint about the tool category, and we cover why AI visibility tools stop at monitoring separately.
Audit, check, or tracking?
| Free check | Audit | Continuous tracking | |
|---|---|---|---|
| Question answered | Am I in this answer? | Why am I, or am I not? | Did anything change? |
| Scope | One prompt, a few engines | Fixed prompt set, plus sources, entity and site | Fixed prompt set, repeated |
| Effort | About a minute | About a day per brand | Setup, then ongoing |
| Output | A result | A ranked fix list | A trend |
| Best for | Finding out you have a problem | Deciding what to do about it | Knowing whether it worked |
| Fails at | Explaining anything | Detecting change | Explaining change |
The three are sequential, not competing. Check to find the problem, audit to understand it, track to confirm the fix landed.
The four layers
Each layer explains the one above it. Work top down, because a finding in the answer layer is only actionable once you have traced it into the layers beneath.
| Layer | What you audit | The question it answers |
|---|---|---|
| 1. Answers | What engines say about your category and your brand | Where do you stand right now? |
| 2. Sources | Which domains the answers cite | Where do those answers come from? |
| 3. Entity | How consistently the web describes your brand | Does the engine know what you are? |
| 4. Site | Whether engines can reach, render and extract your pages | Can you be retrieved at all? |
Layer 1: the answer audit
Start by fixing your prompt set, because a set you change between audits cannot be compared with itself. Ten to twenty prompts is enough for a first pass. Choose them the way a buyer types, not the way a keyword tool reports. Our guide on choosing prompts worth tracking covers the selection properly.
Split them deliberately into three groups, and never average the groups together:
Category prompts (“best pest control in Alpharetta”) test discovery. This is where growth comes from and where absence hurts.
Comparison prompts (“alternatives to [competitor]”) test whether you make the shortlist when an engine is explicitly assembling one. These are often the easiest answers to enter.
Brand prompts (“what is [your brand]?”) test accuracy rather than presence. You will nearly always appear, because the question contains your name. What you are reading is the description.
Then run every prompt on every engine you care about. Which engines to track goes into the trade-offs; for an audit, cover at least ChatGPT, Google AI Overviews and Perplexity, and add Gemini, Google AI Mode, Microsoft Copilot and Claude if your audience justifies them.
For each cell of that grid, record four things, not one:
- Named or not. The binary everyone records.
- Position. First in a list of five is a different outcome from fifth, and both count as “mentioned”.
- Description. The exact wording used about you. Copy it verbatim.
- Citations. Every source the answer showed. This feeds layer 2 and is the single most valuable thing you collect.
Run each prompt more than once if you can. AI answers are generated rather than retrieved, so the same prompt can return different brands on different runs. Two or three runs per cell turns a coin flip into a signal.
Layer 2: the source audit
This is the layer that separates an audit from a check, and it is the one most audits skip.
Take every citation you collected and tally the domains. You are looking for the handful that appear repeatedly behind the answers you lost. In most categories a small number of sources do most of the work: a couple of directories, one or two review platforms, a trade publication, a community thread, sometimes a single well-structured comparison page written by someone with no product in the race.
For each recurring source, answer three questions. Are you listed on it at all? If you are, is the entry current and complete? If you are not, can you be, and at what cost?
That produces the most direct fix list in the whole audit, because these are the pages the engine is already reading. A new post on your own blog competes for attention with everything else on the web. A corrected entry on a directory the engine cites in seven of ten answers is already inside the retrieval path. Our research on the sources AI engines trust covers which categories of source tend to carry weight.
One caveat worth stating: not every engine shows its sources, and those that do are not necessarily showing everything that influenced the answer. Citations are evidence, not a complete account. Treat a domain appearing in seven of ten answers as a strong signal and a domain appearing once as a note.
Layer 3: the entity audit
Engines assemble an understanding of your brand from everything they can find, then answer from that understanding. If the web describes you three different ways, you are a blurry entity, and blurry entities get left out of confident answers.
Check for consistency across every place your brand is described: your own site, Google Business Profile, LinkedIn, directories, review platforms, press coverage, partner and supplier pages. Specifically:
- Name. Exactly one form, including punctuation and any legal suffix.
- Category. What you would call yourself, matching what a buyer would search.
- Location and service area. Especially where you serve a region but list one address.
- What you actually do. The most common failure. Old positioning survives on the internet long after you have moved on.
- Size and stage signals. Enterprise-sounding language on an old press release will get you described as enterprise-only.
The uncomfortable finding here is usually not absence. It is being present and described from a source you would never have chosen: a three-year-old press page, a stale directory category, or a competitor’s comparison page written to make you look narrow. That description is now part of your positioning, repeated to buyers who never visit your site, and you cannot outrank it. You change it by changing the sources it is drawn from.
Layer 4: the site audit
The supply side. Before an engine can cite you, it has to be able to fetch and parse you.
Crawler access. Check robots.txt for the AI crawlers specifically, not just Googlebot. Many sites block them by accident, and blocks added at a CDN or WAF do not appear in the file you deployed, so check the live file on the live host. There are two distinct groups worth thinking about separately: crawlers that fetch a page live while someone waits for an answer, and crawlers that read pages into a model for training. Blocking the second is a legitimate business decision. Blocking the first removes you from answers.
Rendering. If your content only exists after JavaScript runs, assume some engines will not see it. View the raw HTML response and check the substance is in it.
Answer shape. Engines lift passages. Content written as one long unbroken argument is harder to lift than content where a clear question is followed immediately by a direct, self-contained answer. Look for pages where the answer to the page’s own question is buried in paragraph six.
Structure and freshness. Sensible headings, working schema, visible dates, no orphan pages. None of this is exotic, and none of it works if layers 2 and 3 are broken.
Our website audit automates most of this layer, but it is entirely doable by hand for a single site.
Where competitors fit
Competitive analysis is not a fifth layer. It is a lens you apply to the four you already have, and treating it as a separate stage produces a slide rather than a finding.
At the answer layer, the competitor question is: who is named where you are not, and how often. That gives you share of voice, and measuring share of voice in AI answers covers doing it defensibly.
At the source layer it becomes far more useful: which sources cite them and not you. That is a concrete, closeable gap rather than a general observation that they are winning.
At the entity layer: is their description sharper than yours? Frequently the brand that wins is not larger, it is just described more consistently.
At the site layer: what content shape do they have that you do not? Usually a page that answers the exact question, directly, in the first hundred words.
Scoring without inventing a number
An audit should produce a number, but only one you can define. The discipline is to state the denominator every time.
Visibility is best expressed over the grid you actually sampled: prompts multiplied by engines. Twelve prompts across five engines is sixty cells. Named in nine of them is 15 percent, and that sentence is defensible because every term in it is stated. A score with no visible denominator is a decoration.
Three rules keep an audit score honest:
Score the groups separately. Category, comparison and brand prompts measure different things. A brand scoring 100 percent on brand prompts and zero on category prompts has a discovery problem, and one blended number hides exactly the finding worth acting on.
Do not compare across different grids. Adding prompts widens the denominator, which usually lowers the score even if nothing got worse. If the prompt set changes, the series restarts.
Record the date and the engine versions if you can. Models update. An audit is a photograph, and an undated photograph is not evidence.
What the deliverable should contain
- The prompt set, written out in full, with the three groups labelled.
- The answer grid: every prompt against every engine, with named, position, description and citations.
- The score, per group, with the denominator stated and the date attached.
- The citation map: domains ranked by how often they appear behind answers, and whether you are present on each.
- The entity findings: every inconsistent or outdated description, with its source URL.
- The site findings: crawler access, rendering, answer shape, structure.
- The competitor read, expressed as gaps at the source and content layers rather than as a leaderboard.
- A ranked fix list: what to do, why it should move the result, roughly what it costs, and who owns it.
- The re-measurement plan: which prompts get tracked continuously, and when the next audit runs.
The last two items are what make it an audit rather than a report. A document that ends at findings has handed the hard part back to the reader.
Three illustrative examples
The answer layer misleads on its own. A services firm scores 2 out of 45 cells and assumes it needs a content programme. The source audit shows two directories behind most of the answers it lost, and the firm has no listing on either. The first fix is two listings and one corrected category, not a quarter of blog posts.
Good score, wrong description. A software company appears in most answers and treats the audit as a pass. Reading the recorded wording, two engines describe it as enterprise-only, traced to an old press page. Its growth segment is small teams. Nothing in the answer layer flagged this, because presence looked fine.
Invisible for a mechanical reason. A retailer is absent everywhere. Layers 2 and 3 look reasonable. The site audit finds the live robots.txt blocking live-fetch AI crawlers, added at the CDN months earlier and absent from the file in the repository. No amount of content would have fixed it.
How often to re-run it
Baseline once, then quarterly. In between, track a fixed prompt set continuously.
The reasoning is that the two activities have different clock speeds. Answers change week to week, so hand-auditing monthly mostly measures variance and burns the time you should be spending on fixes. The causes underneath change slowly: your entity data, your listings, your site structure. Quarterly is roughly the rate at which the diagnostic layers are worth re-examining.
Re-run early if something structural changes: a rebrand, a repositioning, a site migration, a new market, or a competitor visibly taking over your category answers.
What an AI visibility audit cannot tell you
Four limits, stated plainly, because an audit that hides them is selling confidence rather than findings.
It cannot tell you how often you were shown. There is no impression count for AI answers, and no engine publishes brand-level answer data. Every method available, including ours, is sampling.
It cannot tell you what any individual customer saw. Results vary by person, location, phrasing, session and model version.
It cannot promise a result. Nobody can make an engine recommend you. An audit identifies the variables you can influence; the engine still decides. Treat any tool or agency promising guaranteed AI recommendations as disqualified.
It cannot detect change on its own. One audit is one observation. The comparison is the point, and it needs a second reading taken the same way. This is the limit people discover only after spending a quarter on fixes with nothing to measure them against.
Running audits for clients
For agencies, the audit is usually the commercial entry point, and the four layers make a clean fixed-fee deliverable. Three things make the difference between an audit clients act on and one they file.
Lead with the citation map rather than the score. A prospect has usually never seen the list of domains shaping how AI describes their market, and it reframes the conversation from “your number is low” to “here is where the answers come from”.
Separate the fixes you can execute from the ones only the client can. Directory entries and content are usually yours. Legal name consistency, press pages and CDN rules usually are not.
Set the re-measurement expectation in the audit itself. Committing to a second reading on a stated date, using the same prompt set, is what turns a one-off document into a retained engagement, and it is also just honest measurement.
Where our tooling fits
You can run every layer of this by hand, and running it manually once is the fastest way to understand what the numbers mean. Tooling earns its place on repetition and comparability.
Visibility AI runs the answer grid across all seven engines on a fixed prompt set, records mentions, position, description and citations per cell, and aggregates the citation map for you. The website audit covers layer 4, and the brand profile holds the entity data from layer 3. If you want the fast version of layer 1 before committing to anything, the free Visibility Check runs a single question across three engines and shows the result on the page.
The honest framing: the tool removes the sampling and bookkeeping work. The judgement in layers 2 and 3, deciding which sources are worth pursuing and which descriptions actually matter, is still yours.
Start with one layer
If a full audit is more than you have time for, run layer 1 on five prompts and collect the citations. Then tally the domains. That single exercise, an hour at most, produces the finding that most often changes what a team does next: the answers about your category are being built from a handful of pages, and you are probably not on them.
Everything else in this framework is an expansion of that idea. Run the free Visibility Check to get a first reading, or start the grid by hand. The version you actually complete beats the thorough one you do not.
Frequently asked questions
What is an AI visibility audit?
An AI visibility audit is a structured, point-in-time review of how AI answer engines see your brand and why. It samples what engines say about your category, records which sources they cite, checks how the wider web describes you, and tests whether engines can retrieve your site. The output is a prioritised list of fixes.
How is an AI visibility audit different from an AI visibility check?
A check measures the symptom: are you named in this answer or not. An audit measures the causes underneath it. A check is one prompt and takes a minute. An audit samples a fixed prompt set, then works backwards through citations, your entity data and your site to explain the result and produce fixes.
How long does an AI visibility audit take?
A focused audit of one brand takes roughly a day of work: two to three hours sampling answers, two hours tracing citations and checking directory and review listings, and two hours on site retrieval and content. Doing it faster usually means skipping the source layer, which is the layer that explains the result.
How often should you run an AI visibility audit?
Once as a baseline, then quarterly. AI answers change far too often to audit monthly by hand, and a full audit repeated too frequently mostly measures noise. Between audits, track a fixed prompt set continuously so you can see movement without redoing the diagnostic work each time.
What should an AI visibility audit include?
Four things: the answer grid showing where you appear across prompts and engines, the citation map showing which domains the winning answers draw on, an entity check covering how consistently the web describes you, and a retrieval check on your own site. Plus a ranked fix list with owners.
Can I run an AI visibility audit myself?
Yes. Every layer can be done by hand: ask the engines directly, open the citations they show, search your brand name across directories and review sites, and check your robots.txt and rendering. It is slow rather than difficult. Tooling mainly saves repetition and makes the second audit comparable to the first.
What does an AI visibility audit cost?
Run yourself it costs time, roughly a day per brand. Agencies commonly sell one as a fixed-fee engagement. Paid tools reduce the sampling work and make repeat audits comparable, but no tool removes the judgement in the source and entity layers, which is where most of the actionable findings come from.
Why is my brand missing from AI answers?
Usually one of four reasons: the engines cite sources you are absent from, the web describes you inconsistently so you are a weak entity, your site cannot be retrieved or parsed usefully, or your content never answers the question a buyer actually asks. The audit exists to tell you which.