An AI-visibility score often arrives looking more certain than the underlying evidence. One number, a few bright arrows, perhaps a promise that the vendor can move it next month.
The first useful question is not whether the score is high. It is what was actually observed to produce it.
AI-search visibility can be measured in pieces: a platform may expose impressions or citations; analytics may show a known referral; a controlled prompt test may record a citation or an absence; a business may record a qualified inquiry. Those are different kinds of evidence with different denominators. They do not combine into a universal AI ranking.
Treat a vendor score as that vendor's calculation until its inputs, coverage, and missing-data rules are clear enough to inspect.
Start with an evidence ladder, not a dashboard
The useful order runs from the observation closest to the platform to the business decision farthest downstream. Each rung can answer a real question. None can honestly impersonate the whole ladder.
| Signal | What was actually observed | What it can support | What it cannot support | Source and date | Important denominator or availability limit | Appropriate next decision |
|---|---|---|---|---|---|---|
| Google Search Console generative-AI report, when available | Google-exposed generative-AI impressions in separate Search and Discover reports. Search can group by page, country, device, and date; Discover has no device dimension. | Whether the property received reported impressions in the named Google Search or Discover report and which reported pages or segments deserve inspection. | A citation count, a universal AI rank, visibility in every Google experience, or any result in another platform. | Google Search Console Search and Discover generative-AI performance reports and the June 3, 2026 Search Central announcement, rechecked July 21, 2026. | The reports are still rolling out; a property can lack access or enough impressions. Google says the Search feature list may change; device reporting is available for Search, not Discover. Chart and table totals can use different aggregation, and newest data can be preliminary. | Inspect the reported pages and their reader job. Do not turn an isolated impression change into a performance claim. |
| Bing Webmaster Tools AI Performance, when available | Total citations, average cited pages, sampled grounding queries, page-level citation activity, Citation Share, period comparison, intents, topics, and trends across Bing-supported AI surfaces. | That named URLs were cited, that aggregate citation activity changed within the report's stated scope, and what share of all citations Bing reports for a grounding query. | Ranking, authority, placement, page importance, the role of a URL in one answer, competitor visibility, traffic share, quality, or a cross-platform score. | Bing Webmaster Blog, AI Performance preview and June 16, 2026 expansion, rechecked July 22, 2026. | It remains a preview. Grounding queries are sampled. Bing calls Citation Share observational, not a ranking system or competitive scoreboard; it does not expose competitor domains, represent traffic share, or assign quality scores. | Check whether cited pages are accurate, current, and complete before deciding whether a page needs revision. |
| Known AI-domain referrals in site analytics | Sessions and on-site actions reported with a known referral source under the analytics property's processing rules. | That the property observed some referred visits from that disclosed source and what those measured visitors did on site. | A count of every AI-assisted visit, proof that an assistant caused an inquiry, or a complete acquisition picture. | Google Analytics traffic-source and direct-traffic documentation, rechecked July 21, 2026. | Missing referrer information can be processed as direct traffic; redirects, offline documents, ad blockers, and tagging gaps can remove source detail. There is no shared denominator for all assistant use. | Compare the known referral with landing-page behavior and separate capture and business outcomes before changing a budget or message. |
| A cited or mentioned URL | A named page was visibly cited or mentioned on a stated surface at a stated time. | That the specific page was referenced under those recorded conditions. | Stable frequency, authority, broad discoverability, referral volume, or commercial effect. | The platform report or public source URL, with observation date and method. | The collection method must say which surface, queries, dates, and pages were checked. A found example has no market-wide denominator by itself. | Verify accuracy. Correct or strengthen the page only when the reader need and evidence justify it. |
| Controlled prompt-panel observation | A fixed prompt, platform, date, location or session condition, and answer outcome were recorded. | A reproducible sample of what appeared under those test conditions, including an accuracy issue worth investigating. | A ranking, share of voice, market share, platform comparison, proof of discovery, proof of causation, or a complete baseline. | Versioned test record, with platform, prompt set, conditions, and observation date. | Report the full planned and completed denominator, including failed or untested observations. Wording, time, model behavior, location, and session conditions can all change the result. | Use the sample to form a hypothesis or perform an accuracy check. Re-test before asserting a trend. |
| Indexed or crawlable status | A named URL was accessible, eligible, crawled, or indexed according to a named system at a stated time. | Technical eligibility to investigate further. | A citation, a recommendation, a referral, a qualified inquiry, or a business outcome. | The named platform report or authorized technical check, with date. | This is a system-specific snapshot. Eligibility is necessary for some surfaces but does not establish selection by them. | Resolve a real technical blocker. For the retrievability requirements themselves, see the linked AI-search website guide. |
| Qualified inquiries or business outcomes | A business system recorded a qualified inquiry, won work, or another defined outcome under its own rules. | That a defined outcome occurred and can be compared with other business evidence. | That an AI system caused it, that one page received all credit, or that future outcomes will repeat. | The business's own defined record and date range; use sanitized aggregates when reporting internally. | State the outcome definition, all-inquiry denominator, attribution method, and excluded records. Most journeys have more than one influence. | Decide whether the outcome merits further investigation, not whether an AI-visibility score deserves credit. |
The ladder is deliberately inconvenient. That is its value.
A platform citation can be real without producing a referral. A referral can be real without becoming a qualified inquiry. An inquiry can become revenue without proving that one assistant, one prompt, or one page caused it. Putting all of those facts into one score hides the handoffs where the reasoning should happen.
Google's native report is evidence, not an AI rank
Google's generative-AI performance reports are a meaningful improvement over guessing when the property has access. The Search report covers supported Google Search generative-AI features, including AI Overviews and AI Mode; Google says it expects to update that list. A separate Discover report covers generative-AI features in Discover. The Search report can group impressions by page, country, device, and date; the Discover report does not provide a device dimension.
That report is still a limited rollout. A missing report can mean that the property has not been included yet or that it has not received enough impressions. It does not establish that a site is invisible to every AI-assisted search experience.
Google's normal Search Console methodology matters here too. In AI Mode, an external-link click counts as a click, standard impression rules apply, and position follows the same methodology as a Google Search results page. For an AI Overview, links share the Overview's one Search position. A number in that field is therefore a defined Google Search measurement, not a portable citation position or a stable answer ranking.
Use it to ask a narrow question: “Did Google report this page appearing in these supported features, and what should I inspect next?” That question is useful. “What is our AI rank?” asks the report to be something it does not claim to be.
Bing's citation data answers a different question
Bing's AI Performance view is more citation-oriented. Its preview documentation describes total citations, average cited pages, sampled grounding queries, page-level citation activity, and trends across supported AI experiences. A June 16, 2026 expansion added Citation Share, period comparison, intents, and topics. Bing defines Citation Share as the percentage of all citations for a grounding query, but frames it as an observational metric—not a ranking system or competitive scoreboard.
That gives a site owner a concrete inspection list: which URLs were cited, which query phrases appear in the sample, and whether the reported activity is changing. Bing is equally clear about the boundary. Citation counts do not indicate ranking, authority, placement, page importance, or a page's role in an individual answer.
That makes the report useful for editorial quality control. If a service page appears in the data, check whether its facts are current and whether the cited portion answers the reader's question cleanly. If it does not appear, do not leap from absence to a diagnosis before checking report availability, timeframe, coverage, and the page's technical eligibility.
The separate question of whether a site can be found and read by the relevant systems belongs in What AI Search and AI Assistants Need From a Local Service Website. This article is about what the later measurements can and cannot establish.
Referrals are useful evidence with blind spots
Analytics can show a known referral source when the information reaches the property and is processed into a traffic-source dimension. That is useful for understanding observed sessions, landing pages, and on-site behavior.
It is not a census of AI-assisted discovery. Google Analytics documents several paths by which source information can be unavailable or lost: missing tags, redirects, offline documents, and ad blockers can all contribute to direct or otherwise incomplete classification. A visit that arrived after an AI interaction may not preserve a usable referral; a visible referral does not prove that the assistant was the only reason the person came.
That is why analytics belongs beside the rest of the ladder. It can tell you about measured traffic. The article Why Google Analytics and Your Lead Count Don't Match explains why traffic, capture, handoff, qualification, and revenue should remain separate questions.
A prompt panel is an observatory, not a leaderboard
A controlled prompt panel can be useful when it behaves like a lab notebook.
Record the prompt, platform, date, locale or location condition, session condition where relevant, result classification, cited or mentioned URL, and the full denominator. Keep the wording stable unless the purpose is to test wording. Preserve failed and untested runs instead of quietly dropping them.
Then the panel can show a bounded observation: “Under these recorded conditions, this source appeared, did not appear, or was cited inaccurately.” That may be enough to investigate a content gap or correct a page.
It cannot tell you who “wins” AI search. The output can vary with time, model changes, prompt wording, location, personalization, and platform behavior. A panel with incomplete platform coverage is incomplete by definition. A score built from it needs the same warning label.
What to demand before trusting a vendor score
There is nothing inherently wrong with a vendor using a score for its own workflow. It becomes misleading when the score is sold as a shared market fact without a method.
Ask for these answers in writing:
- Inputs: Which signals feed the score—native account data, referral data, public citations, prompt observations, or something else?
- Platforms: Which products and surfaces are included? Are results aggregated across them or reported separately?
- Prompt protocol: What exact prompts were tested, who chose them, and were they fixed across periods?
- Collection conditions: What dates, locations, languages, devices, account states, and session conditions applied?
- Denominators: How many prompts, URLs, citations, or sessions were eligible? How many were actually observed?
- Missing-data rules: What happens when a platform report is unavailable, a prompt fails, a referrer is absent, or a result cannot be verified?
- Weighting: How does one citation, one referral, or one panel observation change the score? Can the calculation be reproduced?
- Trend comparability: Did the vendor change the panel, platforms, model, method, or weighting between periods?
- Claim boundary: Which decision can the score inform, and which outcomes does the vendor explicitly decline to promise?
If those answers are unavailable, the number may still be a private product metric. It is not evidence strong enough to buy a retainer, declare a visibility problem, or credit a business result.
Use the next action that fits the evidence
The most responsible next move is usually smaller than the sales pitch.
- A reported Google impression may justify inspecting the page and date range.
- A Bing citation may justify checking whether the cited page is accurate and useful.
- A known referral may justify reviewing the landing page and the defined on-site action.
- A prompt-panel anomaly may justify a reproducible re-test.
- A crawl or index failure may justify technical investigation.
- A qualified inquiry may justify asking the customer how they found the business, while admitting that memory and attribution are imperfect.
Those are real decisions. They keep a business from using a pretty score as an excuse to buy a vague promise.
Before accepting any AI-visibility report, write down the observation, its source and date, its denominator, what it supports, and what it cannot support. If the score cannot survive that exercise, it has not earned the authority of a ranking.
Sources
- Google Search Console, Generative AI performance report (Search), rechecked July 21, 2026.
- Google Search Console, Generative AI performance report (Discover), rechecked July 21, 2026.
- Google Search Central, Introducing Search Generative AI performance reports, June 3, 2026; rechecked July 21, 2026.
- Google Search Console, What are impressions, position, and clicks?, rechecked July 21, 2026.
- Microsoft Bing Webmaster Blog, Introducing AI Performance in Bing Webmaster Tools Public Preview, rechecked July 22, 2026.
- Microsoft Bing Search Blog, New AI Visibility Insights in Bing Webmaster Tools, June 16, 2026; rechecked July 22, 2026.
- Google Analytics, Campaigns and traffic sources, rechecked July 21, 2026.
- Google Analytics, Understand direct / none traffic, rechecked July 21, 2026.