AEO & GEOUpdated Sep 11, 20268 min read

How to Benchmark Your AI Visibility: Prompt Set, Four Engines, Three Scores

Benchmark AI visibility step by step: build a prompt set, run ChatGPT, Claude, Perplexity and Gemini, score mention rate, citation rate and share of voice.

Short answerTo benchmark AI visibility, fix a set of 40 to 100 real buyer prompts, run each one three times in ChatGPT, Claude, Perplexity and Gemini under a written protocol, and score mention rate, citation rate and share of voice against five named competitors. Repeat the identical run monthly and treat small moves as noise.

To benchmark your AI visibility, fix a set of real buyer prompts, run each one several times in ChatGPT, Claude, Perplexity and Gemini under a written protocol, and score three numbers against named competitors: mention rate, citation rate and share of voice. Then repeat the identical run every month. The protocol is what turns screenshots into a benchmark.

This guide gives you the full method, a capture template, a scoring table you can copy, a worked example with the math shown, and rules for telling real change from noise.

Why a benchmark, not a spot check

Two facts make spot checks useless. First, AI answers change between runs. In Orbit Media's tracking of 13,184 citations, Claude kept only 30% of its cited URLs from one week to the next, Gemini kept 38% of its cited domains, and Perplexity was the most stable at 65%. Second, the engines disagree with each other. All four cited the same domain in only 1.7% of query and domain combinations.

So a single prompt in a single engine tells you almost nothing. A benchmark answers three questions you can act on: how often are we named, how often are we the source, and how do we compare with the companies we lose deals to. If you are just starting, read AEO and GEO for beginners first.

Step 1: Fix your competitor set

Pick five competitors and write them down before you run anything. Use the companies that show up in your CRM loss reasons and in sales calls, not the ones marketing admires.

  • Two direct competitors you lose deals to most often
  • One category leader buyers compare everyone against
  • One adjacent alternative (a different approach to the same problem)
  • One peer your size that seems to show up in AI answers already

Also list each brand's name variants (for example "Acme," "Acme Data," "AcmeData") so scoring catches all mentions. Keep this set fixed for at least two quarters. Changing it changes your share of voice denominator and breaks comparisons.

Step 2: Build the prompt set

Use 40 prompts for a single offering, and closer to 100 for a multi-product company. Source them from sales call transcripts, RFPs, discovery notes and support tickets. The detailed sourcing method is in how to build an AI visibility prompt set.

Prompt familyShare of setExampleWhat it tells you
Category30%Best Snowflake implementation partner for healthcareWhether you make the shortlist
Problem20%How to migrate from Teradata without disrupting BI reportingWhether you are cited as an expert before vendor selection
Comparison20%[Competitor A] vs [Competitor B] for mid-market insurersWhether you are included when buyers compare
Alternatives15%Alternatives to [category leader] for smaller teamsWhether you are the named alternative
Validation15%What does [your brand] do and how is it pricedWhether engines describe you accurately

Tag each prompt with persona, funnel stage and a weight from 1 to 3. Weight 3 goes to prompts that map directly to your highest-value deals. Weighted scores come later; keep the raw scores too.

If Google matters in your market, add AI Overviews for question and comparison prompts. Seer Interactive found comparison queries triggered AI Overviews 95.4% of the time and question-format queries 85.9%, versus 8% for commercial queries.

Step 3: Write the capture protocol

This is the step most teams skip, and it is why their month-two numbers cannot be compared with month one. Write the protocol down and do not change it.

AI VISIBILITY BENCHMARK PROTOCOL v1.0
Owner: [name]            Run window: first Tuesday to Thursday of each month

Engines:      ChatGPT (search on), Claude (web search on), Perplexity, Gemini
              Optional: Google AI Overviews for question and comparison prompts
Account:      Logged out, or a clean account with memory and custom
              instructions off. Never the marketing team's daily account.
Location:     [country / VPN endpoint], language [en-US]
Runs:         3 runs per prompt per engine, each in a new chat
Prompt text:  Exactly as written in prompt-set-v1.csv. No follow-ups.
Capture:      Full answer text + list of cited URLs + screenshot
Scoring:      Per answer, per brand, using scoring-guide-v1
Changes:      Any change to engines, prompts or competitors creates v1.1
              and is logged with the date and reason

Memory and personalization matter here, because they change what a given user sees. See ChatGPT memory and brand visibility and personalized AI search results.

Step 4: Run the engines and record every answer

Forty prompts, four engines and three runs is 480 answers. Record one row per answer and per tracked brand. That sounds heavy, but a shared sheet with dropdowns makes it fast, and a tracking tool can take over once the method is stable. The AI visibility tools comparison covers the main platforms.

benchmark-log.csv

run_date,engine,run_no,prompt_id,prompt_family,weight,brand,
mentioned(0/1),position(1st/2nd/..),cited_own_domain(0/1),
cited_urls,accuracy(ok/partial/wrong),sentiment(pos/neutral/neg),notes

2026-09-02,Claude,1,P07,category,3,YourBrand,1,2,0,
"clutch.co/...;competitor.com/case-study",ok,neutral,
"named after Competitor A"

Record cited URLs for every answer, not just yours. That column becomes your list of category gatekeepers, the third-party pages you need to appear on. How to act on it is in the off-site assets that get B2B brands cited.

Step 5: Score the three numbers

Mention rate   = answers that name your brand / total answers
Citation rate  = answers that cite your domain / total answers
Share of voice = your brand mentions / mentions of all tracked brands

Weighted mention rate = sum(weight x mentioned) / sum(weight)
Accuracy rate  = answers describing you correctly / answers naming you

Keep mention and citation separate. They diverge a lot. In the same Orbit Media data, Claude mentioned tracked brands in 70% of answers but cited them in 40%, while ChatGPT's gap was narrower at 56% and 47%. A high mention rate with a low citation rate means engines know you but prefer other sources for evidence. A high citation rate with a low mention rate means your pages are used, but your brand is not being positioned as the answer.

For the scoring model behind share of voice, see how to calculate AI share of voice.

The scoring table template

Summarize the log in one table per month. This is the table to put in front of leadership.

EngineAnswersMentionsMention rateCitationsCitation rateShare of voiceAccuracyChange vs last month
ChatGPT120
Claude120
Perplexity120
Gemini120
Total480

Add a second table with the same columns for each competitor, and a third that breaks your own scores down by prompt family. The family view is usually where the action items come from.

A worked example with the math

Illustrative example. A data consulting firm runs 40 prompts, three times, in four engines: 480 answers.

EngineAnswersMentionsMention rateCitationsCitation rate
ChatGPT1203025.0%1815.0%
Claude1201815.0%65.0%
Perplexity1202823.3%1411.7%
Gemini1202016.7%54.2%
Total4809620.0%439.0%

Across all answers, the five tracked brands collect 610 mentions. The firm's 96 mentions give it a share of voice of 96 / 610 = 15.7%. The category leader has 205 mentions, a 33.6% share.

By prompt family, the firm is named in 38% of validation answers but only 9% of category answers, and 6 of its 43 citations come from comparison prompts. The cited-URL column shows that Claude and Gemini repeatedly cite a Clutch category page and two listicles where the firm is absent.

Three actions follow directly from the numbers:

  1. Category prompts are the gap. Build industry-by-use-case pages with answer blocks, following the AEO checklist for a B2B website.
  2. Claude and Gemini citation rates are low and both lean on the same third-party pages. Complete the Clutch profile and request inclusion in the two listicles.
  3. ChatGPT already cites the firm's own pages at 15%. Refresh those pages first, since they are proven retrieval targets.

Telling real change from noise

Because answers vary, you need a rule before you celebrate or panic. Use the margin of error for a proportion as a quick guide:

95% margin of error = 1.96 x sqrt( p x (1 - p) / n )

Example: mention rate p = 0.20, answers n = 480
= 1.96 x sqrt(0.20 x 0.80 / 480)
= 1.96 x 0.0183
= about 0.036, or plus or minus 3.6 points

Per engine (n = 120): about plus or minus 7.2 points

So a move from 20% to 22% overall is inside the noise. A move from 20% to 27% is worth attention. For single engines, where each cell has only 120 answers, require either a larger move or the same direction for two consecutive months. This is a rough guide, because runs of the same prompt are not fully independent, but it stops most false alarms. More on this in LLM answer volatility and how often to track.

How to read the pattern

PatternWhat it usually meansFirst fix
Low mention, low citation everywhereEngines cannot reach you or do not know youCrawler access and entity consistency
High validation, low categoryEngines know you but do not see you as a category answerCategory and industry pages, listicle inclusion
High mention, low citationYour reputation travels, your pages do notAnswer blocks and specifics on money pages
Strong in ChatGPT, weak in ClaudeOwn pages work; third-party footprint is thinReview profiles and curated listicles
Strong in Perplexity onlySocial and video footprint may be carrying youBuild primary pages and review profiles
Low accuracyOld or conflicting information in sourcesCorrect sources; see fixing wrong AI information

The monthly cadence

  1. Week 1: run the benchmark under the protocol.
  2. Week 1: score it, fill the tables, write a five-line summary: what moved, whether it is outside the noise, why you think it moved, what you will do, what you will stop.
  3. Weeks 2 to 4: ship the fixes the numbers point to.
  4. Quarterly: review the prompt set. Retire prompts buyers no longer ask, add new ones, and version the file so trend lines stay honest.

Tie it to revenue. Add an AI assistant option to your form's attribution field and compare benchmark trends with self-reported AI-sourced opportunities. AI search attribution and board reporting cover that step.

Common mistakes

  • Running prompts from a personal account with memory on, then wondering why results look flattering.
  • Changing the prompt set every month. Version it; do not edit it in place.
  • Counting only your own brand. Without competitors there is no benchmark.
  • Ignoring cited URLs. They are the most actionable column in the log.
  • Reporting one blended number. Engines behave differently. Report by engine and by prompt family.
  • Treating Google as solved by rankings. Google says AI Overviews and AI Mode have no special requirements beyond normal Search eligibility, but eligibility is not selection. Track them separately.

Next steps

Once you have a baseline, sequence the fixes with AEO in 90 days. For a deeper, scored audit of a single brand's perception, the AI visibility audit template adds a 100-point rubric. If you want us to build the prompt set and run the first benchmark against your real competitors, request a free audit from our AEO, GEO and SEO team.

FAQ. Quick answers.

Still unsure? Ask us directly.

How many prompts do I need to benchmark AI visibility?

For a single B2B offering, 40 prompts run three times across four engines gives 480 answers, which is enough to see real movement. Larger portfolios often track 100 or more. What matters more than volume is that prompts come from real buyer language and stay fixed between runs, so month-over-month comparisons are valid.

What is the difference between mention rate and citation rate?

Mention rate is the share of AI answers that name your brand in the text. Citation rate is the share that link to your domain as a source. They diverge: in Orbit Media's study, Claude mentioned tracked brands in 70% of answers but cited them in 40%. Track both, because they need different fixes.

How do I calculate AI share of voice?

Count every mention of every brand in your competitor set across all answers. Divide your mentions by that total. If five tracked brands collect 610 mentions and you have 96, your share of voice is 15.7%. Calculate it overall and per engine, and keep the competitor set fixed so the denominator stays comparable.

Should I use a tool or benchmark manually?

Start manually for the first baseline so you understand the answers and can write good prompts. Once the prompt set is stable, a tracking tool saves time and adds run volume. Whatever you use, document the protocol: engine versions, logged-out or clean sessions, location, run count and dates. Changing tools resets comparability.

How often should I rerun the benchmark?

Monthly works for most B2B companies, on the same week each month. Weekly tracking of a smaller core set helps during active campaigns or after a major site change. Because answers vary between runs, only treat a change as real when it exceeds your margin of error or repeats for two consecutive months.

Turn this into pipeline. We can run it with you.

Tell us the revenue number and the market. We will come back with the stages that matter most for you, and the ones you can skip.

  • 20 minutes with a senior operator, not an SDR
  • Bring your revenue target and markets; we bring the pipeline math
  • Slots across US, Canada, India, Singapore and GCC time zones

Prefer email? growth@lemniscategrowth.com

Pick a 20-minute slotStraight to a senior operator. No SDR screen.