To benchmark your AI visibility, fix a set of real buyer prompts, run each one several times in ChatGPT, Claude, Perplexity and Gemini under a written protocol, and score three numbers against named competitors: mention rate, citation rate and share of voice. Then repeat the identical run every month. The protocol is what turns screenshots into a benchmark.
This guide gives you the full method, a capture template, a scoring table you can copy, a worked example with the math shown, and rules for telling real change from noise.
Why a benchmark, not a spot check
Two facts make spot checks useless. First, AI answers change between runs. In Orbit Media's tracking of 13,184 citations, Claude kept only 30% of its cited URLs from one week to the next, Gemini kept 38% of its cited domains, and Perplexity was the most stable at 65%. Second, the engines disagree with each other. All four cited the same domain in only 1.7% of query and domain combinations.
So a single prompt in a single engine tells you almost nothing. A benchmark answers three questions you can act on: how often are we named, how often are we the source, and how do we compare with the companies we lose deals to. If you are just starting, read AEO and GEO for beginners first.
Step 1: Fix your competitor set
Pick five competitors and write them down before you run anything. Use the companies that show up in your CRM loss reasons and in sales calls, not the ones marketing admires.
- Two direct competitors you lose deals to most often
- One category leader buyers compare everyone against
- One adjacent alternative (a different approach to the same problem)
- One peer your size that seems to show up in AI answers already
Also list each brand's name variants (for example "Acme," "Acme Data," "AcmeData") so scoring catches all mentions. Keep this set fixed for at least two quarters. Changing it changes your share of voice denominator and breaks comparisons.
Step 2: Build the prompt set
Use 40 prompts for a single offering, and closer to 100 for a multi-product company. Source them from sales call transcripts, RFPs, discovery notes and support tickets. The detailed sourcing method is in how to build an AI visibility prompt set.
| Prompt family | Share of set | Example | What it tells you |
|---|---|---|---|
| Category | 30% | Best Snowflake implementation partner for healthcare | Whether you make the shortlist |
| Problem | 20% | How to migrate from Teradata without disrupting BI reporting | Whether you are cited as an expert before vendor selection |
| Comparison | 20% | [Competitor A] vs [Competitor B] for mid-market insurers | Whether you are included when buyers compare |
| Alternatives | 15% | Alternatives to [category leader] for smaller teams | Whether you are the named alternative |
| Validation | 15% | What does [your brand] do and how is it priced | Whether engines describe you accurately |
Tag each prompt with persona, funnel stage and a weight from 1 to 3. Weight 3 goes to prompts that map directly to your highest-value deals. Weighted scores come later; keep the raw scores too.
If Google matters in your market, add AI Overviews for question and comparison prompts. Seer Interactive found comparison queries triggered AI Overviews 95.4% of the time and question-format queries 85.9%, versus 8% for commercial queries.
Step 3: Write the capture protocol
This is the step most teams skip, and it is why their month-two numbers cannot be compared with month one. Write the protocol down and do not change it.
AI VISIBILITY BENCHMARK PROTOCOL v1.0
Owner: [name] Run window: first Tuesday to Thursday of each month
Engines: ChatGPT (search on), Claude (web search on), Perplexity, Gemini
Optional: Google AI Overviews for question and comparison prompts
Account: Logged out, or a clean account with memory and custom
instructions off. Never the marketing team's daily account.
Location: [country / VPN endpoint], language [en-US]
Runs: 3 runs per prompt per engine, each in a new chat
Prompt text: Exactly as written in prompt-set-v1.csv. No follow-ups.
Capture: Full answer text + list of cited URLs + screenshot
Scoring: Per answer, per brand, using scoring-guide-v1
Changes: Any change to engines, prompts or competitors creates v1.1
and is logged with the date and reason
Memory and personalization matter here, because they change what a given user sees. See ChatGPT memory and brand visibility and personalized AI search results.
Step 4: Run the engines and record every answer
Forty prompts, four engines and three runs is 480 answers. Record one row per answer and per tracked brand. That sounds heavy, but a shared sheet with dropdowns makes it fast, and a tracking tool can take over once the method is stable. The AI visibility tools comparison covers the main platforms.
benchmark-log.csv
run_date,engine,run_no,prompt_id,prompt_family,weight,brand,
mentioned(0/1),position(1st/2nd/..),cited_own_domain(0/1),
cited_urls,accuracy(ok/partial/wrong),sentiment(pos/neutral/neg),notes
2026-09-02,Claude,1,P07,category,3,YourBrand,1,2,0,
"clutch.co/...;competitor.com/case-study",ok,neutral,
"named after Competitor A"
Record cited URLs for every answer, not just yours. That column becomes your list of category gatekeepers, the third-party pages you need to appear on. How to act on it is in the off-site assets that get B2B brands cited.
Step 5: Score the three numbers
Mention rate = answers that name your brand / total answers
Citation rate = answers that cite your domain / total answers
Share of voice = your brand mentions / mentions of all tracked brands
Weighted mention rate = sum(weight x mentioned) / sum(weight)
Accuracy rate = answers describing you correctly / answers naming you
Keep mention and citation separate. They diverge a lot. In the same Orbit Media data, Claude mentioned tracked brands in 70% of answers but cited them in 40%, while ChatGPT's gap was narrower at 56% and 47%. A high mention rate with a low citation rate means engines know you but prefer other sources for evidence. A high citation rate with a low mention rate means your pages are used, but your brand is not being positioned as the answer.
For the scoring model behind share of voice, see how to calculate AI share of voice.
The scoring table template
Summarize the log in one table per month. This is the table to put in front of leadership.
| Engine | Answers | Mentions | Mention rate | Citations | Citation rate | Share of voice | Accuracy | Change vs last month |
|---|---|---|---|---|---|---|---|---|
| ChatGPT | 120 | |||||||
| Claude | 120 | |||||||
| Perplexity | 120 | |||||||
| Gemini | 120 | |||||||
| Total | 480 |
Add a second table with the same columns for each competitor, and a third that breaks your own scores down by prompt family. The family view is usually where the action items come from.
A worked example with the math
Illustrative example. A data consulting firm runs 40 prompts, three times, in four engines: 480 answers.
| Engine | Answers | Mentions | Mention rate | Citations | Citation rate |
|---|---|---|---|---|---|
| ChatGPT | 120 | 30 | 25.0% | 18 | 15.0% |
| Claude | 120 | 18 | 15.0% | 6 | 5.0% |
| Perplexity | 120 | 28 | 23.3% | 14 | 11.7% |
| Gemini | 120 | 20 | 16.7% | 5 | 4.2% |
| Total | 480 | 96 | 20.0% | 43 | 9.0% |
Across all answers, the five tracked brands collect 610 mentions. The firm's 96 mentions give it a share of voice of 96 / 610 = 15.7%. The category leader has 205 mentions, a 33.6% share.
By prompt family, the firm is named in 38% of validation answers but only 9% of category answers, and 6 of its 43 citations come from comparison prompts. The cited-URL column shows that Claude and Gemini repeatedly cite a Clutch category page and two listicles where the firm is absent.
Three actions follow directly from the numbers:
- Category prompts are the gap. Build industry-by-use-case pages with answer blocks, following the AEO checklist for a B2B website.
- Claude and Gemini citation rates are low and both lean on the same third-party pages. Complete the Clutch profile and request inclusion in the two listicles.
- ChatGPT already cites the firm's own pages at 15%. Refresh those pages first, since they are proven retrieval targets.
Telling real change from noise
Because answers vary, you need a rule before you celebrate or panic. Use the margin of error for a proportion as a quick guide:
95% margin of error = 1.96 x sqrt( p x (1 - p) / n )
Example: mention rate p = 0.20, answers n = 480
= 1.96 x sqrt(0.20 x 0.80 / 480)
= 1.96 x 0.0183
= about 0.036, or plus or minus 3.6 points
Per engine (n = 120): about plus or minus 7.2 points
So a move from 20% to 22% overall is inside the noise. A move from 20% to 27% is worth attention. For single engines, where each cell has only 120 answers, require either a larger move or the same direction for two consecutive months. This is a rough guide, because runs of the same prompt are not fully independent, but it stops most false alarms. More on this in LLM answer volatility and how often to track.
How to read the pattern
| Pattern | What it usually means | First fix |
|---|---|---|
| Low mention, low citation everywhere | Engines cannot reach you or do not know you | Crawler access and entity consistency |
| High validation, low category | Engines know you but do not see you as a category answer | Category and industry pages, listicle inclusion |
| High mention, low citation | Your reputation travels, your pages do not | Answer blocks and specifics on money pages |
| Strong in ChatGPT, weak in Claude | Own pages work; third-party footprint is thin | Review profiles and curated listicles |
| Strong in Perplexity only | Social and video footprint may be carrying you | Build primary pages and review profiles |
| Low accuracy | Old or conflicting information in sources | Correct sources; see fixing wrong AI information |
The monthly cadence
- Week 1: run the benchmark under the protocol.
- Week 1: score it, fill the tables, write a five-line summary: what moved, whether it is outside the noise, why you think it moved, what you will do, what you will stop.
- Weeks 2 to 4: ship the fixes the numbers point to.
- Quarterly: review the prompt set. Retire prompts buyers no longer ask, add new ones, and version the file so trend lines stay honest.
Tie it to revenue. Add an AI assistant option to your form's attribution field and compare benchmark trends with self-reported AI-sourced opportunities. AI search attribution and board reporting cover that step.
Common mistakes
- Running prompts from a personal account with memory on, then wondering why results look flattering.
- Changing the prompt set every month. Version it; do not edit it in place.
- Counting only your own brand. Without competitors there is no benchmark.
- Ignoring cited URLs. They are the most actionable column in the log.
- Reporting one blended number. Engines behave differently. Report by engine and by prompt family.
- Treating Google as solved by rankings. Google says AI Overviews and AI Mode have no special requirements beyond normal Search eligibility, but eligibility is not selection. Track them separately.
Next steps
Once you have a baseline, sequence the fixes with AEO in 90 days. For a deeper, scored audit of a single brand's perception, the AI visibility audit template adds a 100-point rubric. If you want us to build the prompt set and run the first benchmark against your real competitors, request a free audit from our AEO, GEO and SEO team.
