Here's a situation I watch agencies walk into - and it's why AI search visibility has become the measurement question of 2026.
A client’s pages still rank for valuable searches. Organic traffic, however, is slipping. Then the agency asks ChatGPT a question a buyer might ask: “What are the best options in this category?” The answer names three competitors and leaves the client out.
That spot check raises a useful question, but it does not settle it. AI answers can vary by platform, prompt wording, and run. To tell a client whether they are consistently part of the answer, agencies need a repeatable way to measure AI search visibility.
The starting point is straightforward: define the buyer questions, check the AI surfaces that matter, record what each answer does, and compare results over time.
Measure mentions, citations, and recommendations separately
A brand can appear in an AI answer in three distinct ways:
- Mention: The answer names the brand.
- Citation: The answer links to the brand’s page as a source.
- Recommendation: The answer presents the brand as an option the buyer should consider.
These signals answer different questions. A product might be mentioned only to explain a limitation. An article might earn a citation while a competitor earns the recommendation. For a client selling a product or service, recommendation rate is often the clearest measure of shortlist presence; citation rate shows whether its content is being used as evidence.
StoryChief’s exposure, evidence, and impact framework extends this distinction to what an answer says about the brand and whether visibility produces qualified demand.

For client reporting, recommendation rate is the headline metric. Track all three, but lead with the one that maps to buying behavior.
Choose the surfaces your buyers use
ChatGPT, Google AI Overviews, Google AI Mode, Gemini, Microsoft Copilot, Perplexity, and Claude are all candidates for a tracking plan. The right selection depends on where a client’s buyers research their options. Results from different surfaces should remain visible separately: the same question can produce different answers, sources, and links.
For most B2B content programs, the six surfaces worth tracking are:
- ChatGPT - the largest conversational interface; increasingly browses the web for current queries
- Google AI Overviews - now appearing on a substantial share of informational queries
- Perplexity - answer engine with transparent citations; strong for research-intent queries
- Gemini - Google's assistant layer; behaves differently from AI Overviews
- Microsoft Copilot - Bing's AI layer; an underrated and often ignored referral source
- Claude - growing share of professional and analytical queries
Start with a manageable set and expand once the collection process is reliable. StoryChief’s AI Visibility Tracker currently covers ChatGPT, Google AI Overviews, Google AI Mode, Gemini, and Copilot. If Perplexity or Claude matters to the audience, check those separately.
Record the platform and, where relevant, the location, account state, and whether web access was enabled. These details make later comparisons easier to interpret.
Build prompts around buyer decisions
A keyword such as “project management software” identifies a topic. A prompt such as “What project management tool works well for a ten-person remote agency?” identifies a buyer and a decision.
Build an initial set of 10–20 prompts drawn from sales calls, site search, customer questions, and competitor comparisons. Include several types of intent:
- Discovery: “What are the best [category] tools for [use case]?”
- Comparison: “How does [client] compare with [competitor] for [specific need]?”
- Alternatives: “What are the alternatives to [market leader] for [buyer type]?”
Keep the core prompt set fixed from month to month. You can add wording variants to test how sensitive an answer is to phrasing, but label them separately so a changed question does not look like a changed result. StoryChief’s guide to measuring GEO offers more detail on selecting questions that reflect real buyer demand.

Repeat checks before calling a change a trend
One answer is an observation, not a stable visibility score. Research on repeated AI search queries found substantial variation in cited sources and warned that single-run estimates can appear more precise than they are. The study examined three platforms and three consumer product topics; it does not establish one universal number of runs for every agency prompt set.
Three runs per prompt and platform is a workable minimum for a manual pilot. It can reveal obvious variation, but it is not a statistical guarantee. If a result will drive a major client decision, collect more runs and show the range alongside the average.
For each run, log:
| Field | What to record |
|---|---|
| Context | Date, platform, prompt, location, and relevant settings |
| Brand presence | Whether the client was mentioned, cited, or recommended |
| Competitors | Which named competitors received those same outcomes |
| Sources | The cited URLs and their domains |
| Answer quality | Whether the client was described accurately and in the right category |
Save the answer or a link to it when the platform allows. That gives the team something concrete to review when a result looks surprising.
Start manually, then make tracking repeatable
A spreadsheet pilot teaches the team how to classify answers before a dashboard starts summarizing them.
- Agree on the buyer prompts, client brand, and genuine competitors.
- Run each prompt on the selected AI surfaces and record the results.
- Inspect the cited URLs. Separate the client’s pages from review sites, forums, publications, and competitors’ pages.
- Repeat on a consistent schedule before treating a change as a trend.
Once the prompt set is established, StoryChief’s AI Visibility Tracker can handle recurring checks across its supported surfaces. It suggests buyer questions using your company, audience, topics, and competitors; your team can edit those questions before tracking begins. Weekly results show brand mentions and website citations by provider, cited pages, and competitors appearing in the same answers. Open a question to inspect the response and its sources, then use a missed mention or citation to guide a content update. In StoryChief, start at Reports → AI Visibility.

Keep any Perplexity or Claude checks in the manual log or another tracking workflow. Maintain the same core prompts across measurement periods so the trend remains comparable.
For reporting, define each rate explicitly. Recommendation rate, for example, is the share of eligible runs in which the answer recommends the client. Report it by platform and buyer intent before calculating an overall figure. An overall average can hide a strong showing on research questions and weak presence when buyers ask for a shortlist.
Treat a combined visibility score as a starting point for investigation. Open the underlying answers, check whether a mention is accurate and useful, and report recommendation rate separately when shortlist presence is the client’s main concern.
Use the source mix to decide what to do next
Citation patterns suggest where to investigate, although a cited link does not prove precisely how an AI system selected its recommendations.
If answers repeatedly cite comparison sites, examine how the client appears on those sites and whether its positioning is current. If they cite forums, look for recurring customer questions the client has not answered clearly. If competitor pages provide the clearest comparisons, improve the client’s own explanations of fit, limits, and alternatives.
Llumo’s exploratory study of restaurant recommendations observed that some businesses recurred across runs even when their order and framing changed. That suggests a useful reporting question: Is the client regularly entering the shortlist? The study’s narrow category and uncontrolled variables mean agencies should test that pattern in their own markets rather than assume it applies everywhere.
Turn the findings into content work
Use the answers to identify a specific gap, make a change, and measure again. That might mean clarifying who a product is for, updating an outdated comparison, making important information easier to find in page text, or correcting inconsistent descriptions across relevant third-party sources.
StoryChief connects this step to the measurement workflow: teams can inspect a missed mention or citation, then plan, create, review, and publish content intended to address the gap. The next tracking cycle shows whether the answer changed.
Keep the technical work grounded in established search practice. Google’s guidance for AI features says its existing SEO best practices remain relevant and that there are no special technical requirements for inclusion in AI Overviews or AI Mode. It recommends crawlable pages, accessible text, helpful content, and accurate structured data where used. None of these changes guarantees a citation.
Review outcomes monthly. Record which pages or external profiles changed, then compare the next measurement window with the baseline. This makes the report a record of decisions and results rather than a collection of screenshots.
What to put in the client report
A useful one-slide report can show four things:
- Recommendation rate: How often the client appears as a suggested option on relevant prompts, split by intent and platform.
- Competitor comparison: How often the client and its closest alternatives enter the same shortlists.
- Citation source mix: Which domains support the answers and what action the team took in response.
- Business impact: Identifiable AI referrals, engaged visits, and conversions where available, with attribution limits stated.
Keep Google Search data alongside this report, but interpret it carefully. Google includes AI Overview and AI Mode activity within its broader Search Console reporting; referral analytics describes what identifiable visitors did after reaching the site.
Rankings still matter. They tell you something about search visibility, while repeated AI answer checks show whether a client is being named, used as a source, and offered as a choice. Start with 20 buyer questions and a spreadsheet. Once the questions and definitions are sound, use a tracker such as StoryChief’s to make the checks routine and connect the findings to content work.
After a few consistent measurement cycles, the team will have a much better answer when the client asks, “Are we showing up in AI search?”