Skip to main content

How to measure AI search visibility

A practical method for tracking brand mentions, recommendations, citations, referral visits, and outcomes across AI-generated answers.

22 min readWritten by Ampere

A question card branches to five paper symbols for an answer mention, recommendation, source citation, website visit, and completed inquiry.
Contents

Suppose someone asks an AI search product to recommend inventory software for a three-store retailer. Your company might appear by name in the answer. Its website might be linked as a source. The answer might recommend the product, describe it neutrally, or cite one of your guides while recommending somebody else. The person might click through, remember the name and return later, or never visit at all.

Each event answers a different question. Merge them into one metric and the report loses the distinctions that would make it useful.

AI search visibility is the observable presence of a brand, product, or source in AI-generated answers for a defined set of questions and conditions. A useful measurement practice separates five signals: mentions, recommendations, citations, referral visits, and business outcomes. It also labels what came from platform reporting, what came from your analytics, what you observed in a sample, and what remains unknown.

The method below gives a small team a repeatable baseline audit, a copyable tracking template, worked calculations, and a completed fictional example.

How AI search visibility builds on search visibility

Traditional search visibility asks whether pages from your site appear in search results, where they appear, how often people see them, and whether those appearances lead to clicks. AI search adds another surface between the question and the visit: a generated answer that may synthesize several sources and name entities that it does not link.

A generated answer creates several kinds of visibility. A brand can be present in the prose even when its site is absent from the source list. A useful article can be cited even when the answer does not recommend the company that published it. A linked source can receive no click. A person can encounter a brand in an answer, search for it later, and arrive through direct or branded search traffic that no analytics system can confidently connect to the earlier exposure.

SignalWhat you observedWhat it does not establish
Brand mentionThe answer named the brand or productEndorsement, a citation, or a visit
RecommendationThe answer presented the brand as a suitable option for the stated needThat the recommendation was seen, trusted, or acted on
CitationThe answer displayed a link or source reference to a URLA click, endorsement, or influence on a purchase
Referral visitYour analytics recorded a session arriving from an identifiable AI source or tagged linkEvery exposure, every later visit, or the visitor's full decision path
Conversion or inquiryA measured visitor completed a defined business actionThat one earlier AI answer caused the action

The rows can overlap, and any one can appear without the next. Report them separately, then study where they overlap.

Five paper cards labeled mention, recommend, cite, visit, and convert form a loose constellation connected by cords.
The five signals can overlap, split, or stop. Track citations, visits, and conversions independently, then examine where they coincide.
Three paper vignettes show a brand mention without a source, a citation that supports another option, and a recommendation supported by third-party pages.
Answer text and source lists need separate review. An answer may name a brand while citing third-party pages, or cite the brand’s page while recommending another company.

SEO, AEO, and GEO describe overlapping work

Search engine optimization (SEO) is the established work of making useful pages discoverable, understandable, and eligible to appear in search. Answer engine optimization (AEO) and generative engine optimization (GEO) are newer labels for work aimed at visibility in direct answers and generative search experiences.

The labels name useful areas of focus, although their boundaries remain unsettled. Google treats optimization for its generative search features as part of SEO and says its core search ranking and quality systems remain foundational. Its current guidance rejects the need for special AI schema, required content “chunking,” or a separate page for every possible fan-out query. The durable work is familiar: make pages crawlable, publish original and useful information, identify authorship and evidence, keep facts current, and make the page satisfying for the intended reader. See Google’s people-first content guidance and guide to generative AI features in Search.

Three technical terms help explain why measurement is different:

  • Retrieval is the system finding pages or other sources that may help answer a question.
  • Grounding is the system using retrieved material to support a generated answer. The terms are often used together because retrieved pages supply current evidence beyond what the model already contains.
  • Query fan-out is the system issuing several related searches to answer one user question. Those retrieval queries may be narrower than the question the person typed, so record them separately from the original prompt.

That last distinction matters in reporting. A grounding-query phrase in Bing Webmaster Tools describes retrieval activity associated with a citation; the user's original prompt remains unreported.

What the major platforms currently expose

Each AI search product exposes its own slice of publisher data. Interfaces, retrieval behavior, citation treatment, and first-party metrics differ, so the source should travel with every number.

Platform or experienceWhat affects the observationFirst-party measurement available to site owners
Google Search: AI Overviews and AI ModeAI Overviews appear within Search results; AI Mode supports follow-up questions and may fan one question into subtopics. Links shown inside these features follow Google Search impression and click rules.Search Console's Generative AI performance report reports link impressions from AI Overviews and AI Mode by page, country, device, and date. Google says it was rolled out worldwide on August 31, 2026. The ordinary Performance report counts clicks on external links in these features, but its generative report documentation is centered on impressions and does not expose answer text or a complete prompt log. Search Labs experiments are excluded.
Microsoft Copilot, AI summaries in Bing, and selected partner experiencesMicrosoft combines citation activity across several supported AI surfaces. The visible answer and cited pages may vary by experience.Bing Webmaster Tools AI Performance reports total citations, average cited pages, page-level citation activity, trends, and sampled grounding-query phrases. Microsoft describes the feature as a public preview and says its data is aggregated, refreshed daily with a delay, and not a complete log. Preview additions include intent, topic, comparison, and citation-share views.
ChatGPT SearchSearch may rewrite a question into one or more targeted queries. General location, memory, conversation context, and whether web search runs can affect the result. Answers may show inline citations and a Sources panel.OpenAI documents publisher crawl controls and automatically adds utm_source=chatgpt.com to referral URLs from ChatGPT search results. Its current public publisher documentation does not describe a Search Console-style impression or citation dashboard, so teams generally combine referral analytics with manual or vendor-run samples. See OpenAI’s search documentation and publisher FAQ.
PerplexityPerplexity answers are built around linked sources; search mode, selected model, location, account context, and follow-up history can change the source set.Perplexity documents PerplexityBot for search discovery and Perplexity-User for user-requested fetches in its crawler documentation. The public documentation reviewed for this guide does not describe a publisher visibility dashboard comparable with Google or Bing. Use visible citations, identifiable referrals, and clearly defined manual or vendor samples.
Gemini and other AI assistantsSome experiences search the web by default, some only when asked, and some answer largely from model knowledge. Model choice, interface, account state, and tool availability can all change the response.Keep data from a company's search product separate from its assistant app. Record the exact experience tested, then use its documented reporting if available; otherwise treat observations as a sample and referrals as partial outcome evidence.

Those boundaries tell you what each report can support. Search Console can tell you that links from your site were shown in supported Google generative search features. Bing can tell you that pages were visibly cited across supported Microsoft experiences. Answer text and a universal cross-platform visibility score remain outside both reports.

One question card connects to four different paper reporting stations representing impressions, referrals, citation trends, and visible source lists.
The same question can leave different evidence in different products. Attach the platform and experience to every metric so readers know exactly what each number measures.

Build the question set from customer needs

A baseline becomes useful when its questions represent decisions your audience actually makes. Start with evidence you already have: sales-call questions, support tickets, customer interviews, on-site search, Search Console queries, paid-search terms, community discussions, and the wording customers use in briefs.

Autocomplete suggestions and keyword-volume estimates can add ideas, and they supply different kinds of evidence. Autocomplete shows suggestions generated by a platform under particular conditions. A search-volume tool estimates search activity using its own data and method. An interview records what one person said. A prompt exported from an AI product records an observed question in that product. Keep the source labels to preserve those differences.

For a first audit, choose a set small enough that somebody will inspect every answer. Cover distinct jobs and skip dozens of near-duplicate phrasings:

Question typeExampleWhat it tests
Branded“What is Juniper Trail and who is it for?”Whether the product is recognized and described accurately when named
Category“What inventory planning software works for independent outdoor retailers?”Whether the brand surfaces without being supplied in the question
Solution“How can a three-store outdoor retailer reduce seasonal stockouts?”Whether the brand appears when the user starts with a problem rather than a product class
Comparison“How do retail planning tools compare with spreadsheets for seasonal buying?”Which criteria and sources shape the comparison
Recommendation“Recommend an inventory forecasting tool for a small outdoor retail chain.”Whether the brand is presented as a suitable option under stated constraints
Alternative“What are practical alternatives to enterprise retail planning software?”Whether the brand enters an adjacent competitive set

A branded question is a recognition and accuracy check. Supplying the name makes a mention easier, so report branded and non-branded results separately. Category and recommendation questions are more useful for discovery, while solution questions reveal whether your content connects the audience's problem to the category at all.

Interview, support, search, and autocomplete evidence cards feed a sorting point that separates one branded question from several discovery questions.
Question sources are evidence, too. Keep their labels, then separate brand-supplied recognition tests from the category and problem questions that test discovery.

Use the question set as a measurement instrument. Group misses by the underlying need, then decide whether one genuinely useful page can answer that need better. A prompt variation alone is a weak reason to create another page.

A repeatable baseline audit for a small team

The following method works with a spreadsheet, access to the products you want to observe, and your existing analytics. It creates a documented baseline for the chosen sample and conditions, ready to rerun later under a comparable setup. Statistical representation falls outside the scope of this audit.

1. Define the decision before collecting answers

Write down what the audit should help you decide. A good first objective is: “Find the audience questions where we are absent, described inaccurately, or cited through the wrong page, then choose two content or technical investigations.” This is more actionable than “increase AI visibility.”

2. Select questions and preserve their sources

Choose questions across branded, category, solution, comparison, recommendation, and alternative intent. Keep the original wording when it came from a real source. In the tracker, give each source its own label: “sales call,” “customer interview,” “Search Console,” “support ticket,” “autocomplete,” or “editorial hypothesis.” The labels preserve differences in evidentiary weight.

3. Fix and record the test conditions

For every run, capture:

  • platform and exact experience, such as “ChatGPT Search” rather than “GPT”
  • date and, if useful, approximate time
  • country or location setting, language, device, and account state
  • model or mode when the product exposes it
  • whether web search was enabled or visibly used
  • fresh conversation or continued conversation
  • relevant personalization, memory, or prior context
  • repeat number

Some variation will remain. Record enough context to see it. A fresh conversation reduces carryover from the preceding exchange; the answer can still change between runs.

4. Run independent repeats

Repeated runs show whether an observation recurs under your chosen conditions. Start each repeat in a fresh conversation unless conversation history is the thing you are testing. Count every question-platform-repeat combination as one answer observation.

Repeat counts and cadence are practical choices. More repeats cost more time and reveal more variability; fewer repeats are cheaper and give each appearance more influence over the result. Choose a method your team can sustain, document it, and keep it consistent when comparing periods.

Recent preprints reinforce the need for restraint. One study of commercial recommendations found large differences between paraphrases of the same buying intent, and another found that repeated identical questions continued to surface new brands and sources across many runs. Their findings apply to particular models, categories, and test designs, so they cannot set a benchmark for your brand. They support a narrower conclusion: one answer is weak evidence of a stable pattern. See the paraphrase-brittleness study and repeated-query study.

5. Code the answer, then preserve the evidence

For each observation, record whether the brand was mentioned, whether it was actually recommended for the stated need, and every cited URL from your domain. Save a screenshot or exported answer with a stable filename, such as 2026-09-15-chatgpt-q03-r1.png, and put that reference in the tracker.

Read the answer alongside the source chips. A page can be cited while the brand is absent from the prose. A brand can be mentioned because a third-party page described it. Count a recommendation when the answer presents the brand as suitable for the stated need; an appearance in a comparison table alone does not qualify.

6. Add first-party reports and analytics separately

Export the relevant Google and Bing publisher reports when available. In website analytics, inspect session source or referrer for identifiable AI referrals and connect those sessions to meaningful actions such as a pricing-page view, signup, demo request, or purchase. Keep this as a separate table from manual observations because the denominators are different.

Google Analytics distinguishes first-user acquisition, session acquisition, and event-level attribution. For this audit, Session source / medium is usually the clearest starting point for visits: it describes what initiated that session. A later key event may receive fractional or modeled credit under the property's attribution settings. Describe it as an attributed conversion under those settings; causal credit requires more evidence. See Google’s documentation on traffic-source scopes and the Traffic acquisition report.

7. Report findings at the level the evidence supports

Separate four classes:

  • Observed: a saved answer, a documented platform metric, or an analytics event you directly inspected
  • Sampled: a rate calculated from your defined question-platform-repeat set
  • Estimated: a vendor-modeled audience, impression, demand, or visibility figure
  • Unknown: unreported exposures, later unattributed visits, the user's private context, and causal influence on a purchase

This vocabulary keeps a small audit honest and still gives the team something useful to act on.

Four paper trays separate saved observations, a defined sample, projected estimates, and an envelope marked as unknown.
A useful report gives each claim an evidence class. The unknown tray preserves unanswered questions instead of filling them with estimates.

Copyable AI search visibility tracking template

Keep enough fields to reproduce and interpret an observation. Internal notes can sit elsewhere if they would make the shared view unwieldy.

FieldWhat to record
QuestionExact wording tested
Question sourceInterview, sales call, support, Search Console, autocomplete, observed prompt, or editorial hypothesis
Platform and experienceProduct plus mode, such as ChatGPT Search or Google AI Mode
DateDate of the answer observation
Relevant test conditionsCountry, language, fresh or continued conversation, search state, model/mode, and repeat number
Brand mentionYes or no, based on the answer text
RecommendationYes or no, using your documented definition
Cited URLExact URL from your domain, blank if none
Evidence referenceScreenshot, export, or saved-answer identifier

Download the CSV template and duplicate rows for every platform and repeat. A spreadsheet is enough until manual collection becomes the bottleneck.

Worked example: Juniper Trail’s first audit

The following business, questions, answers, and analytics are fictional and illustrative. They demonstrate the method with synthetic data; no real company or platform produced these results.

Juniper Trail is a fictional inventory-planning product for independent outdoor retailers. On September 15, 2026, its marketer tests eight questions in two experiences: ChatGPT Search and Perplexity Search. Each question is run twice in a fresh conversation, in US English, from Seattle, with no saved memory or prior conversation context. That creates:

8 questions × 2 platforms × 2 repeats = 32 answer observations

Eight question cards multiplied by two platform cards and two repeat tokens produce a folder containing 32 observation dots.
The denominator is the audit design: every question-platform-repeat combination counts once, including observations where the brand is absent.

The illustrative question set

IDQuestionSource used in the fictional audit
Q1What is Juniper Trail and who is it for?Support ticket
Q2How does Juniper Trail compare with planning tools built for large retail chains?Sales call
Q3What inventory planning software works for independent outdoor retailers?Customer interview
Q4How can a three-store outdoor retailer reduce seasonal stockouts?Customer interview
Q5How do retail planning tools compare with spreadsheets for seasonal buying?Search Console query theme
Q6Recommend an inventory forecasting tool for a small outdoor retail chain.Sales call
Q7How should an outdoor retailer plan seasonal buying across stores?Support ticket
Q8What are practical alternatives to enterprise retail planning software?Autocomplete suggestion

Illustrative results by question

Each row summarizes four answer observations: two platforms multiplied by two repeats.

IDMentionsRecommendationsCitations to junipertrail.exampleWhat the saved answers support
Q14 of 40 of 43 of 4The fictional product was consistently recognized when named; one answer mentioned it without citing its site
Q24 of 43 of 42 of 4The brand appeared in every branded comparison, but one answer described it without recommending it and two cited only third-party sources
Q31 of 41 of 41 of 4One non-branded category answer recommended and cited the product; three did not mention it
Q40 of 40 of 40 of 4No observed connection between the stockout problem and the fictional product
Q51 of 40 of 41 of 4One answer cited the product’s spreadsheet guide without recommending the product
Q61 of 41 of 40 of 4One recommendation named the product but linked only to third-party sources
Q70 of 40 of 40 of 4No observed presence for the seasonal-buying how-to question
Q80 of 40 of 40 of 4No observed inclusion in the alternatives set
Total11 of 325 of 327 of 32A baseline sample, not population coverage

The full 32-row illustrative dataset is included in the downloadable CSV after the blank template rows, so the totals can be checked observation by observation.

Mention rate and citation rate

For this audit, one answer observation counts once in the denominator. If an answer says “Juniper Trail” four times, it still contributes one mention. If it cites two Juniper Trail pages, it still contributes one cited answer to this rate; the individual URLs remain in the tracker for page-level analysis.

Sampled mention rate

answers that mention Juniper Trail ÷ all answer observations

11 ÷ 32 = 0.34375 = 34.4%

Interpretation: Juniper Trail appeared in 34.4% of this defined answer sample. Coverage across all real customer questions remains unknown.

Sampled citation rate

answers that cite at least one Juniper Trail URL ÷ all answer observations

7 ÷ 32 = 0.21875 = 21.9%

Interpretation: 21.9% of sampled answers displayed at least one URL from the fictional site. Clicks, endorsement, and recall use different denominators and remain unmeasured here.

The same data can be summarized at the question level: five of eight questions produced at least one mention, or 62.5%. This metric uses questions as its denominator. It shows breadth across needs while hiding inconsistency between platforms and repeats, so Juniper Trail should report it alongside the answer-level rate.

If the team later reports share of voice, it must define the competitive set and unit. One defensible sampled version is: Juniper Trail mentions ÷ mentions of all five predeclared brands across the same 32 answers. Changing the competitor list changes the denominator and therefore the score.

Illustrative referral and outcome data

For the 30 days ending September 15, the fictional company’s analytics show nine identifiable sessions from chatgpt.com or perplexity.ai. Four were engaged sessions, two reached the pricing page, and one submitted a demo request.

Those are observed website events in this fictional example. They belong to a separate dataset from the 32 manual answer observations. The analytics also miss people who saw the brand and returned later through search, copied an untagged URL, blocked analytics, or converted on another device. A careful report would say: “One demo request occurred in a session attributed to an identifiable AI referral.” Claiming that the sampled visibility generated the demo would overstate the evidence.

What the illustrative report can support

Here is the short report Juniper Trail could send to its founder.

Illustrative AI search visibility baseline — September 15, 2026

Juniper Trail appeared in 11 of 32 sampled answer observations (34.4%) and its site was cited in 7 (21.9%). Presence was concentrated in the two branded questions: eight of the 11 mentions occurred after the question supplied the name. Among six non-branded questions, the brand appeared in three of 24 observations. The clearest content gap is problem-led discovery: none of the four answers about reducing seasonal stockouts connected the need to Juniper Trail. One spreadsheet-comparison answer cited the company’s guide while recommending another product. That result separates source visibility from product recommendation.

In a separate 30-day analytics view, nine identifiable sessions arrived from ChatGPT or Perplexity and one produced a demo request. Those visits come from a separate dataset, leaving total AI exposure and causal impact on pipeline unknown.

Next: verify whether the stockout and seasonal-buying pages are crawlable and current; compare the cited spreadsheet guide with sources used in the missed answers; interview two customers about how they describe seasonal stockouts; then improve the strongest existing page if the gap is real. Repeat the same baseline conditions after the work has had time to be discovered, while logging platform or model changes.

The report points to an investigation. Before creating content, Juniper Trail should inspect whether a useful page already exists, whether the missed question belongs to its audience, and whether the answer would add something more helpful than a product pitch.

Use each signal for the decision it can support

Different measurements answer different questions:

  • Mentions and recommendations in a defined sample can reveal recognition, category inclusion, description errors, and needs where competitors recur but you do not.
  • Cited URLs can show which of your pages are being used as sources and which third-party pages frame the category. Inspect the actual passage and answer before deciding that the cited page is “winning.”
  • Google generative impressions can show whether links from your property were displayed in supported Google generative search features, broken down by page and selected dimensions.
  • Bing citation activity and grounding-query groups can show which pages and themes are associated with citations across supported Microsoft experiences. The report itself says the data is sampled and does not indicate ranking, authority, or a page's role in an answer.
  • Referral sessions and key events can show observable visits and downstream actions. Segment by landing page and intent, not just source, to learn whether visits are qualified.
  • Customer and sales evidence can tell you whether the audited questions resemble real buying work. It often improves the question set more than another hundred generated prompts.

A vendor score summarizes its sample. Citations record displayed sources, while referral analytics count identifiable sessions. A content edit needs stronger causal evidence than a before-and-after score. Model updates, demand shifts, indexing changes, location, language, personalization, search availability, and ordinary output variation remain plausible explanations.

Third-party tools can still be useful. Their methods simply need to travel with their numbers. Ahrefs, for example, defines mentions and citations at the answer level, derives its prompt set from search-backed questions, and models “impressions” using Google search-volume data. Semrush uses its own prompt collection and visibility formula. Profound’s current documentation says its main visibility score uses non-branded prompts. These are product-specific constructs with different scales. Review the vendor’s prompt source, platform coverage, counting rules, and formula before comparing periods. Compare scores only within the same documented methodology.

When manual checks are enough

Manual checks are enough when your question set is still being designed, the team serves one market and language, and somebody can inspect every answer. They are especially valuable at the start because reading the responses teaches you what the binary fields miss: a recommendation may be qualified, a mention may be inaccurate, and a citation may support only one sentence.

A paid visibility tool becomes useful when the collection work is crowding out analysis. Common triggers include monitoring several markets or languages, tracking a stable question set across multiple platforms, preserving answer history, comparing a defined competitor set, or needing alerts and exports for a regular reporting process.

A small manually reviewed answer stack and a larger paid-tool card sorter both connect to one human review sheet.
Manual work is strongest while the team is learning what to inspect. A paid tool earns its place when collection scale becomes the constraint; human review remains the decision point.

Question design still belongs to your team. Before subscribing to a tool, ask:

  1. Where do its prompts come from, and can you add questions grounded in your own audience evidence?
  2. Which exact platform experiences, models, locations, and languages does it test?
  3. How often does it run, how many repeats does it make, and can you inspect the underlying answers?
  4. How does it count a mention, recommendation, citation, position, impression, and share of voice?
  5. Does it preserve historical evidence when models or product interfaces change?
  6. Can you export the raw observations as well as the composite score?

Keep a small manual control sample even after adopting a tool. It gives you a way to notice changed definitions, mislabeled brands, broken prompts, or a score movement that disappears when you read the answers.

If the audit itself becomes a recurring operation, schedule the collection only after the questions, definitions, and review rules are stable. The guide to automating marketing with AI agents explains how to separate a repeatable trigger from the judgment that still needs a marketer.

Improve the evidence before chasing the score

A first audit will usually leave more unknowns than a traditional rankings report. The boundary shows where the team has evidence and where the market remains opaque.

Start with the gaps that connect to real audience needs. Check crawl and indexing access. Correct inaccurate entity information. Strengthen a useful existing page with original evidence, a clear answer, real examples, and current facts. Make important claims independently verifiable. Then repeat the same sample and compare the underlying answers, cited pages, first-party reports, referrals, and qualified actions.

A useful outcome is a report that shows a founder what was observed, sampled, estimated, and left unknown, with enough context to choose the next content decision.

Topics

  • SEO
  • Explainer
  • Template
  • Workflow

Share

XLinkedIn

From Ampere

Put this into practice with Ampere.

Ampere is your brand-aware AI marketing agent. It works from your saved brand memory to research, create, and package marketing work — you keep direction and approval.