On this page
Measurement in this field is easy to do and easy to do uselessly. The difference is almost entirely in decisions you make before the first run.
This chapter is those decisions.
- Choosing questions buyers actually ask, not keywords.
- Why the prompt set must stay fixed, and what happens when it does not.
- Which engines, and why depth beats breadth.
- The four things worth recording per answer, and which one people skip.
Ask what buyers ask
“Best CRM for a two-person real estate team” is a question. “Enterprise CRM solutions provider” is a keyword, and nobody types it into an assistant.
Questions asked of AI are longer, more conversational and more situational than search queries. They carry context a search box never did: team size, industry, budget, the specific problem. If your tracked prompts are keywords with question marks bolted on, you are measuring against something nobody asks, and every number that follows is decorative.
A reasonable starting set covers the stages a buyer moves through. What is this category. What options exist. How do two options compare. Which is best for a situation like mine. You will lose different questions for different reasons, and that difference is itself diagnostic: losing “what is X” is a recognition problem, losing “best X for Y” is usually a source problem.
Choosing the prompts worth tracking goes through how to source these from real buyer language rather than inventing them at a desk.
Then leave the list alone
This is the mistake that quietly ruins more measurement programmes than any other, and it looks like diligence while it happens.
Every prompt you add or change breaks comparability with every previous run. Three months in you have a chart that appears to show movement, and some unknowable share of that movement is just the questions changing. The temptation is strong precisely because you learn as you go and keep thinking of better prompts.
Decide the set, write down why, and hold it. If you must add prompts, treat it as starting a second baseline rather than extending the first, and note the date. A tool that silently regenerates its prompt list between runs is not tracking anything, however good its dashboard looks.
This is also the first thing to ask any vendor: does the prompt set stay fixed, and can I see it?
Which engines
Cover where your buyers actually ask. For most businesses that means ChatGPT, Google’s AI surfaces and Perplexity at minimum, with Gemini, Copilot and Claude added when the audience justifies them.
Two things worth knowing before you optimise for engine counts.
Depth beats breadth. Five engines with the citations recorded per answer is worth considerably more than ten with only a mention count, because the citations are what tell you why you lost. A tool advertising a large engine list and no citation detail is selling the less useful half.
“Gemini” is not one surface. Google runs several AI answer surfaces that behave differently and routinely give different answers to the same question, and a single blended “Google” number hides which one you measured. The Gemini chapter of this problem covers the distinction, and which AI engines to track covers the trade-offs across the board.
What to record
Four things per answer. The first is obvious, the second is the one most tools skip, and the fourth is the one almost nobody captures.
Were you named in the answer text? Not somewhere in the citations, but in the sentence a buyer reads. Being used as a source and being recommended are different outcomes and should never be one number.
Which sources were cited? The most valuable field in the whole exercise. Aggregate it across a few dozen answers and a pattern appears fast: a small number of domains supply most of the answers in your category. That list is the closest thing to a work order this field produces, and it is the input to the entire next chapter.
Who was named instead? Competitor presence in the answers you lose is more informative than your own absence. It tells you the answer had room for a brand like yours and chose someone else. Share of voice turns this into a trackable number.
Did it retrieve at all? The presence or absence of citations is your proxy for which mode produced the answer, and as chapter 1 covered, that determines which problem you are looking at.
Cadence, and reading the results honestly
Weekly is enough for most businesses. Monthly is fine if you are moving slowly. Daily mostly measures noise.
Because answers vary between runs on their own, a single result is not a finding. A brand named in 3 of 20 answers this week and 3 of 20 next week is stable. One that moves from 3 to 9 over six weeks has genuinely changed. Reacting to individual runs will have you chasing variance.
The same caution applies to competitor claims and vendor demos built on a screenshot. It takes seconds to re-ask a question until you get the answer you wanted to show.
What measurement cannot tell you
Worth carrying into every conversation about this, because it is where overselling happens.
Nobody sees real user activity. No engine publishes brand-level answer data. Every number here comes from questions a tool asked itself. Useful, and not the same as observation.
Attribution stays partial. Someone who asks an engine about your category and visits you later often arrives with no referrer at all. Chapter 5 covers what can and cannot be recovered.
Correlation, not causation. Answers shift as models update and as the web changes underneath. A rise after your work is not proof your work caused it, and anyone claiming otherwise is overstating what the data supports.
How to track AI visibility walks through a practical setup, and the free visibility check will give you a first reading in about a minute if you would rather see numbers before building a process.
With a stable baseline in place, the interesting question becomes why you are absent from the answers you lose. That is the next chapter.
Frequently asked questions
How do I measure my brand's visibility in AI answers?
Run a fixed set of buyer questions across a fixed set of engines on a schedule, and record whether you were named, which sources were cited, and who was named instead. The fixed part matters more than the size: changing the questions between runs makes the comparison meaningless while appearing to add rigour.
How many prompts should I track?
Enough to cover the questions your buyers actually ask at each stage, which for most businesses is twenty to fifty rather than hundreds. A smaller set you keep stable and read carefully beats a large one you keep editing.
How often should I run a check?
Weekly suits most businesses, monthly if you are moving slowly. Daily mostly measures noise, because answers vary between runs even with nothing changed on your side.
Which engines should I track?
At minimum ChatGPT, Google's AI surfaces and Perplexity, because that is where most buyers ask. Add Gemini, Copilot and Claude if your audience justifies them. Depth beats breadth: five engines with citations recorded is worth more than ten with only a mention count.
What should I record for each answer?
Four things: whether you were named in the answer text, which sources were cited, which competitors appeared, and whether the engine retrieved at all. The citations are the most valuable and the most commonly skipped.
Why do my numbers move when I have not changed anything?
Because answers vary between runs, models update, and the web underneath changes. This is why a single check proves nothing and why you read trends across several runs of a stable prompt set rather than reacting to one result.