Measurement

AI visibility tracking: how to measure whether you appear in AI answers

For two years the honest answer to "are we showing up in AI answers?" was that nobody could really tell you. That changed in 2026, partly. Google now reports some of it, and has already conceded the reporting is thin. Here is what you can actually measure today and how to assemble it into something you can report on.

Andrew Charon

Written by

SEO Consultant, Double Atari

Former General Mills · Technical SEO, GEO/AEO, and CMS strategy · Minneapolis-based consultant

No single source will tell you your AI visibility. You need four layers: Google's new generative AI performance report in Search Console for impressions inside AI Overviews and AI Mode, server logs for whether AI crawlers are reaching your pages at all, GA4 for the small trickle of referral traffic AI assistants actually send, and prompt sampling for whether you are named in answers. Each layer answers a question the others cannot. The gap nobody has closed is reliable share of voice, so treat any tool promising a single AI visibility score with suspicion.

What is AI visibility tracking?

AI visibility tracking is measuring whether your site appears in AI-generated answers, how often, and in what terms. It replaces the ranking report as the core measurement question, because a generated answer has no positions one through ten. You are either drawn on or you are not.

It breaks into four distinct questions, and confusing them is the most common reporting mistake I see:

  • Access. Can AI systems fetch your pages at all?
  • Inclusion. Do your URLs appear inside AI features?
  • Attribution. Does any traffic arrive from those appearances?
  • Representation. When a model discusses your category, are you named, and is the description accurate?

No tool answers all four. Build the stack accordingly.

What does Google's new generative AI report give you?

This is the biggest change in AI measurement in 2026 and plenty of people have not turned it on yet. Google announced dedicated Search Generative AI performance reports on June 3, 2026, in its Search Central announcement, and confirmed in the Search Console documentation that as of August 31, 2026 the insights reached all websites worldwide.

You find it in Search Console under Performance, then Search results, then the Generative AI entry in the report's left-hand menu. It shows how often URLs from your site appeared in generative AI features on Search, which currently means AI Overviews and AI Mode, with a companion view for Discover.

That is genuinely useful. For the first time you can separate the impressions you earn inside AI features from your ordinary search impressions, and watch that number move as you make changes.

What will the Search Console report not tell you?

Quite a lot, and Google has been unusually direct about it. Search Engine Journal reported in September 2026 that Google acknowledged the reporting is inadequate. Three limitations matter most for how you report:

  1. It is a filtered view, not additive. The generative AI numbers are already counted inside your regular Web search totals. Adding the two together double counts, and I fully expect to see that error in client decks this quarter.
  2. Position is block-level, not link-level. Every link inside an AI Overview inherits the position of the AI Overview itself, so average position tells you where the feature sat on the results page, not where you sat inside the feature. There is no link-level placement metric.
  3. It is interface only. The generative AI data is not exposed through the Search Analytics API or the BigQuery bulk export, which is why you cannot chart it automatically alongside the rest of your reporting yet. A practitioner re-verified this on a live property in August 2026. If you want it in a dashboard today, someone is exporting a CSV by hand.

It also only covers Google. Nothing here tells you anything about ChatGPT, Perplexity, or Claude.

How do you tell whether AI crawlers are reaching your pages?

Server or CDN access logs, and nothing else. This is the least fashionable layer and the one I would give up last, because it is the only place you can observe AI systems interacting with your site directly rather than infer it from a vendor's sample.

Filter your logs by user agent and answer three questions. Which AI crawlers are requesting pages? Which pages are they requesting? What status codes are they getting? A retrieval crawler receiving 200s on your priority pages is the cleanest evidence available that you are eligible to be cited. A retrieval crawler you have never seen is a diagnosis: you are not losing the AI visibility contest, you are not entered in it.

Two practical notes. Separate the training crawlers from the retrieval crawlers in your reporting, because they mean different things, and only the retrieval ones predict citations. And verify by IP where the operator publishes a verified range, since user agent strings are trivially spoofed and a meaningful share of traffic claiming to be an AI bot is not.

Can GA4 show you traffic from AI assistants?

Some of it, and the volumes will disappoint you. GA4 now has an AI Assistant default channel group, which buckets sessions arriving from recognized assistant referrers. It works, and it is worth checking.

Set expectations honestly though. On my own site, across a recent 90 day window, that channel accounted for two sessions out of 453. That is not a measurement failure, it is the nature of the surface: AI answers frequently resolve the user's question without a click. Which is precisely why impression-level and citation-level measurement matters more here than session counts, and why judging AI investment on referral sessions alone will lead you to abandon it right as it starts working.

Two things to get right in GA4. Watch engaged sessions rather than sessions, because assistant referrals tend to arrive with unusually specific intent and the engagement rate is the interesting signal. And filter out your own preview, staging, and proxy hostnames before you trust any of it, since a handful of internal sessions distorts a channel this small beyond usefulness.

How do you measure whether you are actually being cited?

By asking. Prompt sampling is the only direct read on representation, and a manual version costs nothing.

Build a fixed list of 20 to 40 questions a real prospect would ask, phrased the way people actually talk rather than as keywords. Include category questions, comparison questions, and explicit recommendation requests. Then run them on a schedule across the assistants your market uses, and record three things each time: whether you were mentioned, which competitors were, and which sources were cited.

The discipline that makes this worth doing is holding the prompt list fixed. Change the questions and you have thrown away your baseline. Also log the cited sources, not just the mentions, because that list is the most direct content and outreach roadmap you will ever get: those are the pages currently standing between you and the answer.

Two caveats to state in any report built on this. Answers are non-deterministic, so the same prompt on the same day can differ, which means you need repeat samples and you should report ranges rather than precise percentages. And personalization means your results are not necessarily a user's results.

What about the AI visibility tools?

The category is real and moving fast. As of 2026 the commonly compared options include Semrush's AI toolkit, Ahrefs Brand Radar, Profound, Peec AI, Otterly.AI, and Scrunch, with several roundups tracking a field that keeps reshuffling.

What they mostly do is automate the prompt sampling layer at a scale you cannot match by hand, then aggregate it into a visibility score. That automation is the legitimate value, and for a brand of any size it is worth paying for.

Be skeptical of the scores though. Every vendor samples a different prompt set through a different method, so their numbers are not comparable to each other and none of them is a measured share of voice. Treat a tool's score as an internal trend line you watch move, never as an absolute you report as market share. And do not let a tool substitute for the log layer, which no vendor can see on your behalf.

What does a measurement stack you can actually run look like?

QuestionSourceCadence
Can AI systems fetch us?Server or CDN logs, by user agentMonthly
Do we appear in Google AI features?Search Console generative AI reportMonthly
Does any traffic arrive?GA4, AI Assistant channel, engaged sessionsMonthly
Are we named in answers?Fixed prompt list, sampledMonthly
Who is being cited instead?Cited sources from the same samplingQuarterly
Is our entity described correctly?Direct prompts about your brandQuarterly

Capture a baseline for all six before you change anything. The most common reason AI visibility work cannot be defended later is that nobody recorded the starting point, and impressions in AI features are volatile enough that a single month tells you almost nothing.

What still cannot be measured?

Worth saying plainly, because the vendor messaging in this category tends to imply otherwise.

  • True share of voice. Nobody knows the real distribution of prompts in your category, so every share figure is an estimate built on a sampled prompt list.
  • Answers with no click. When a model resolves a question using your content and the user never visits, you get the influence and no measurable event. This is the structural blind spot of the whole discipline.
  • Link-level placement in AI Overviews. Not reported, as covered above.
  • Attribution to revenue. An assistant-influenced buyer who later arrives via direct or branded search looks like direct or branded search.

The reasonable posture is to measure the four layers you can, report them as directional rather than precise, and resist manufacturing a single number to make it feel tidier than it is. If you want the underlying mechanics of why these are the levers, see how language models decide what to cite. For the reporting workflow across all three platforms, see using GA4, Search Console, and Semrush together.

AI visibility tracking FAQ

What is AI visibility tracking?

AI visibility tracking is measuring whether your site appears in AI-generated answers, how often, and how you are described. It covers four separate questions: whether AI systems can access your pages, whether your URLs appear in AI features, whether any traffic results, and whether you are named accurately when a model discusses your category.

Does Google Search Console show AI search data?

Yes. Google announced generative AI performance reports on June 3, 2026, and rolled them out to all websites worldwide by August 31, 2026. You find them under Performance, then Search results, then Generative AI. The report shows impressions inside AI Overviews and AI Mode, with a separate view for Discover.

Can I get the Search Console AI data through the API?

Not currently. The generative AI data is not exposed through the Search Analytics API or the BigQuery bulk export, so it is interface only and has to be exported manually if you want it in a dashboard. The report is also a filtered view of data already inside your regular Web search totals, so never add the two together.

Why is my AI referral traffic in GA4 so low?

Because AI answers often resolve the question without a click. GA4's AI Assistant channel group captures only the sessions where someone followed a link through. Low numbers are expected and are not evidence the work is failing, which is why impression-level and citation-level measurement matter more than session counts on this surface.

How do I know if AI crawlers are visiting my site?

Server or CDN access logs are the only reliable source. Filter by user agent to see which AI crawlers request which pages and what status codes they receive, separate retrieval crawlers from training crawlers because only the former predicts citations, and verify by IP range where the operator publishes one, since user agent strings are easily spoofed.

Are AI visibility scores from tools accurate?

Treat them as internal trend lines rather than absolute measurements. Each vendor samples a different prompt set using a different method, so scores are not comparable across tools and none represents a true share of voice. They are genuinely useful for automating prompt sampling at scale and for watching direction over time.

Explore related Double Atari resources: analytics and reporting, GEO and AEO optimization, website audits, SEO consulting, and .

Andrew Charon

Written by

SEO Consultant, Double Atari

Former General Mills · Technical SEO, GEO/AEO, and CMS strategy · Minneapolis-based consultant

Last updated:

Want this measured properly?

Get an AI visibility baseline you can report on.

Double Atari sets up the full stack: crawler log analysis, the Search Console generative AI report, a clean GA4 view, and a fixed prompt list sampled on a schedule, with a documented baseline.