On this page
If you only track one AI engine, ChatGPT is the one, simply because more of your buyers are asking it questions than any other. That makes it the highest-value thing to measure, and also the easiest to measure badly.
The specific trap: ChatGPT does not produce answers one way. It produces them two ways, and the two fail for completely different reasons. A tracker that reports “you appear in 30 percent of answers” without separating them hands you a number you cannot act on.
- ChatGPT answers from model knowledge or from live retrieval, and the two need different fixes.
- When it searches, the citations tell you exactly which sources decide your category. That is the actionable half.
- When it does not search, being named depends on how consistently you are written about across the web. Slower, and not a switch you can flip.
- Answers vary between runs. A single check is close to meaningless; trends across a fixed prompt set are the real signal.
- Nobody can see inside other people's ChatGPT conversations. All honest tracking is sampling.
Two ways ChatGPT answers, two different problems
Ask ChatGPT “what is the best project management tool for a small agency” and it may search the web, read a few pages, and answer with citations. Ask it “what does Asana do” and it will likely answer from what the model already holds, with no sources at all.
Both are ChatGPT answers. They are not the same event.
Retrieval answers are the tractable ones. The model ran a search, retrieved some pages, and built an answer from them. If you are absent, it is because the pages it retrieved did not mention you. That is a specific, finite problem: you know which sources mattered because they are cited right there.
Memory answers come from the model’s training. If you are absent, it is because you were not well enough represented in the material the model learned from, or your identity was ambiguous enough that the model does not confidently associate you with your category. This is slower to change and it does not respond to a quick fix.
The practical consequence: when you look at a ChatGPT result where you lost, the first question is not “how do I rank better” but “did it search?” If it cited sources, you have a source problem. If it did not, you have a recognition problem.
What a ChatGPT visibility tracker should record
Four things, at minimum.
Whether you were named, in the answer text itself and not just somewhere in the citations. Being cited as a source and being recommended in the sentence a buyer reads are different outcomes, and the second is what you actually want.
Which sources the answer used. This is the most useful field on the whole record and the one thin tools skip. Over a few dozen answers a pattern emerges: a handful of domains supply most of the answers in your category. Those are your target list.
Who was named instead. Competitor presence in the answers you lose is more informative than your own absence. It tells you the answer had room for a brand like yours and chose someone else.
Whether it searched at all. The presence or absence of citations is your proxy for which mode produced the answer, and it determines which of the two problems you are looking at.
Choosing what to ask
A tracker is only as good as its prompt list, and this is where most people go wrong. Two rules matter more than the rest.
Ask what buyers ask, not what you want to rank for. “Best CRM for a two-person real estate team” is a real question. “Enterprise CRM solutions provider” is a keyword and nobody types it into ChatGPT. AI questions are longer, more conversational and more situational than search queries, and the whole exercise is worthless if you measure against questions nobody asks.
Fix the list and leave it alone. Every prompt you add or change breaks comparability with every previous run. It is tempting to keep adding prompts as you learn, and it quietly destroys your ability to say whether anything improved. Decide the set, then hold it still. We cover this in more depth in how to choose the prompts worth tracking.
A reasonable starting set covers the questions a buyer asks at each stage: what a category is, which options exist, how two options compare, and which is best for a specific situation. You will lose different questions for different reasons, which is itself diagnostic.
Why one check proves nothing
Ask ChatGPT the same question twice, ten minutes apart. You will often get different answers, sometimes different sources, occasionally a different set of brands entirely.
This is normal and it is not a flaw in your tracking. Generation is probabilistic, retrieval varies, and the web underneath moves. It has one hard implication: a single check cannot tell you anything reliable. Not about your visibility, not about a competitor’s, and definitely not about whether last month’s work paid off.
What works is the same fixed prompts, on a schedule, compared across several runs. A brand named in 3 of 20 answers this week and 3 of 20 next week is stable. A brand that moves from 3 to 9 over six weeks has genuinely changed. Anything read from a single run is noise wearing the costume of a metric.
This also means you should be suspicious of any competitor claim or vendor demo built on one screenshot. It is trivially easy to re-roll a question until you get the answer you wanted to show.
What to do with each result
You are absent and the answer cited sources. The most actionable case. Collect the cited domains across all the answers you lose and count them. A small number of sites will dominate. Those sites are the target list, and getting onto them is the work. For most local and service categories they are directories, roundups and review platforms; for software they skew toward comparison sites and community threads. More on this in the sources AI engines trust.
You are absent and the answer cited nothing. A recognition problem. Check the basics first, because they are cheap: is your name unambiguous, do your listings and site agree with each other, does your site state plainly what you do and who for. Then accept that the rest is a matter of being written about more consistently over time, which is slow and does not respond to a content sprint.
You are cited but not named. The engine read you and chose someone else for the sentence. This is where the substance of what those sources say about you matters, and where content genuinely is the lever.
You are named but described wrongly. More common than people expect, particularly after a rebrand, a move, or a merger. Worth catching, because a confidently wrong description does more damage than absence.
What tracking cannot tell you
Worth being direct about the limits, because vendors are often not.
Nobody can see other people’s conversations. There is no ChatGPT analytics feed, no impressions data, no query volume. Any tracker, ours included, works by asking questions itself and recording answers. That is sampling, and it is a genuinely useful method, but it is not measurement of real user activity and should not be sold as such.
You cannot attribute revenue cleanly. Someone who asks ChatGPT about your category and then visits your site later often arrives with no referrer at all. There are partial signals in AI referral traffic, and none of them add up to a clean attribution chain.
You cannot buy your way in. There is no ad slot inside an organic ChatGPT answer. You change the inputs and the model still decides.
Where ChatGPT fits with everything else
Track ChatGPT first if you track one thing. Then add breadth, because the answers differ meaningfully between engines and a brand can be strong in one and invisible in another.
Perplexity is the most legible engine, since it cites everything, which makes it the best place to learn which sources decide your category. Google’s AI surfaces reach people who never leave Search. We go through the trade-offs in which AI engines to track.
The important thing is that measurement is the start, not the deliverable. A tracker that gives you a percentage and stops has handed you the easy half of the job. What you needed was the source list underneath it, which is the thing you can actually act on. That is the argument in why AI visibility tools stop at monitoring, and it applies to ChatGPT more than anywhere, because the audience size makes it tempting to settle for a headline number.
Frequently asked questions
What is a ChatGPT visibility tracker?
A tool that runs a fixed set of buyer questions through ChatGPT on a schedule and records whether your brand is named, which sources the answer drew on, and which competitors appeared instead. The point is a comparable measurement over time, not a one-off screenshot, because a single answer tells you almost nothing.
Can you track brand mentions in ChatGPT?
Yes, by asking the same questions repeatedly and recording the answers. What you cannot do is see inside anyone else's ChatGPT conversations or get usage data from OpenAI. Every honest tracker works by sampling: it asks the questions your buyers plausibly ask and measures what comes back.
Why does ChatGPT mention my competitor and not me?
It depends which mode produced the answer. If ChatGPT searched the web, look at the sources it cited: your competitor is on them and you are not. If it answered from model knowledge without searching, your competitor is better represented across the material the model was trained on, which is a slower problem tied to how widely you are written about.
Does ChatGPT always search the web?
No, and this is the key thing to understand. It answers some questions from what the model already holds and searches for others, depending on the question and the settings. The two paths fail for different reasons and need different fixes, so a tracker that ignores the distinction gives you a number you cannot act on.
How often should I track ChatGPT visibility?
Weekly is enough for most businesses, monthly if you are moving slowly. Daily tracking mostly measures noise, because answers vary between runs even with no change on your side. What matters more than frequency is keeping the prompt set fixed so runs stay comparable.
Do ChatGPT answers change between runs?
Yes. The same question can produce a different answer minutes apart, with different sources and sometimes a different set of brands. This is why a single check is close to meaningless and why trends across a fixed prompt set over several runs are the only reliable read.
How do I get ChatGPT to mention my brand?
For the retrieval path, get onto the sources ChatGPT cites when it answers your category's questions, which you find by reading the citations on the answers you lose. For the memory path, it is a slower matter of being written about consistently and unambiguously across the web. Neither is a setting you can toggle.
Is tracking ChatGPT enough on its own?
It is the highest-value single engine for most businesses because of its audience size, but it is not the whole picture. Buyers also ask Google's AI surfaces, Perplexity, Gemini and Copilot, and the answers differ. Track ChatGPT first if you only track one, then add breadth.