On this page
This is where otherwise good programmes overclaim, and where credibility is either built or quietly lost. What follows is what can be proven, what can be partially shown, and what has to be presented as modelling.
- Re-measuring against a stable baseline, and the comparison that breaks it.
- The leading indicator that moves weeks before your score does.
- What AI referral traffic can and cannot show.
- Talking about ROI without inventing a number.
Re-measure like with like
Same prompts, same engines, compared across runs. This sounds trivial and is where most measurement falls apart, for the reason chapter 2 gave: the temptation to keep adding prompts is strong and it quietly destroys the comparison.
If you must expand, start a second baseline and date it. A chart built from a shifting prompt set shows movement that is partly just the questions changing, and you will not be able to tell which part.
Compare across several runs rather than reacting to one. Answers vary between runs on their own, so a single result is noise wearing the costume of a finding.
Watch the citations first
The most useful thing in this chapter.
When source work lands, new domains appear in your citation list before your mention rate moves, often by several weeks. The engine starts drawing on a source you have just joined, and only later does that translate into your brand being named.
If you only watch the headline percentage, you will conclude the work failed during exactly the period it is starting to work, and you will stop a quarter of effort one month early. Track which domains appear in the citations run over run and treat their arrival as the first evidence.
The second indicator worth separating: where you appear. Moving from cited-but-not-named to named in the answer text is real progress even when the percentage barely moves, and it is the step immediately before being recommended.
What referral traffic can tell you
Some assistants pass a referrer, so visits from them show up in analytics as referral traffic from the assistant’s own domain. That is genuinely useful and it is a floor, not a measure.
Three things it misses. Visitors who arrive with no referrer at all, which is common. People who read the answer and never clicked, which leaves no trace anywhere. And anyone who saw you named, remembered it, and searched for you directly later, which shows up as brand search rather than AI referral.
So the honest framing is that referral traffic proves a lower bound on the click-through half of the effect, and says nothing about the larger half where the answer did its job without a visit. Tracking AI referral traffic covers what is recoverable in GA4 and server logs.
Correlation, not causation
Answers shift on their own as models update and as the web changes underneath them. A rise following your work is not proof your work caused it.
This is not a reason to avoid claiming progress. It is a reason to state the standard you are using: a stable prompt set, compared over a reasonable window, with the work you did dated against it. That is a defensible claim. “Our visibility rose 40 percent because of our AEO programme” is not, and anyone senior will know it.
Being the person who says this first is usually better than being the person who gets asked.
Talking about ROI
Three tiers, and the discipline is keeping them separate.
Measured: visibility change against a stable baseline, which sources are cited, share of voice against named competitors. Solid.
Partial: referral traffic and the conversions attached to it. Real, and a floor.
Modelled: revenue attributed to AI visibility. This requires assumptions about the traffic you cannot see and the influence you cannot trace. You can build it, and you should label it as a model with its assumptions written down.
Presenting the third tier as if it were the first is the single fastest way to lose a stakeholder’s trust when they test one number. Proving the ROI of answer-engine optimization works through a version that survives scrutiny.
Choosing tooling, now that you know what to ask
Having read this far, the useful questions are obvious and most vendor conversations do not cover them.
Does the prompt set stay fixed, and can you see it? Are citations recorded per answer, or just a score? Are competitors tracked alongside you? And what does the product do after telling you that you are invisible?
That last one separates the category more than anything else, which is the subject of why AI visibility tools stop at monitoring, and the tools comparison groups the main options by who each is built for.
The whole thing, briefly
Answers are assembled from sources an engine retrieved, so being named depends on being in those sources. Measure with a fixed set of questions so the numbers mean something later. Read the citations under the answers you lose, because they say why. Fix identity first, sources second, content third. Then re-measure against the same baseline and describe what you find accurately.
None of that makes an engine recommend you. It changes what the engine has to work with, which is the only lever anyone actually has.
If you want to see where you currently stand, the free visibility check runs a live scan across the major engines in about a minute.
Frequently asked questions
How do I prove AI visibility work is working?
Re-run the same prompts on the same engines and compare like with like. Watch two things beyond the headline number: which sources are now cited, which moves first, and whether you are named in the answer text rather than only appearing in citations.
Can I track traffic from AI answers?
Partially. Some assistants pass a referrer you can see in analytics, and some send visitors who arrive with no referrer at all. An answer that satisfied someone without a click leaves no trace anywhere, so treat referral traffic as a floor rather than a measure.
What is the leading indicator for AI visibility?
The citation list. When source work lands, new domains appear in the citations before your mention rate moves, often by several weeks. If you only watch the headline percentage you will conclude the work failed while it is actually taking effect.
How do I calculate ROI on answer-engine optimization?
Honestly, and with stated assumptions. You can measure visibility change reliably, referral traffic partially, and revenue attribution barely. Build the case from what you can measure and be explicit about the modelled parts rather than presenting an estimate as a result.
How long before I should expect results?
Identity fixes can show within weeks. Source work usually takes a quarter, and its first visible sign is a change in which domains are cited rather than a change in your score. Set that expectation before starting, not after the first month looks flat.