AEO & GEOUpdated Jul 25, 20269 min read

How to Track Your Brand's Visibility in ChatGPT, Claude, Gemini and Perplexity

A practical method to track brand visibility in AI assistants including ChatGPT, Claude, Gemini and Perplexity, covering prompt panels, metrics and cadence.

How do you track brand visibility in AI assistants?

Tracking brand visibility in AI means running a fixed panel of buyer prompts across ChatGPT, Claude, Gemini and Perplexity on a set schedule, then recording whether your brand is mentioned, where it sits in the answer, and which sources the assistant cited. There is no rank position to look up. The measurement has to be manufactured deliberately rather than pulled from a report.

The output is a small set of numbers you can trend. The most useful are mention rate, the share of panel prompts where your brand appears at all, citation rate, the share where one of your own URLs is linked as a source, and share of voice against a named competitor set. Everything else is diagnostic detail sitting underneath those three headline figures.

The discipline is closer to survey research than to rank tracking. Answers vary between runs even with identical prompts, so a single observation tells you almost nothing about your position. What tells you something is the same panel, run repeatedly on a fixed cadence, with the variance measured and reported rather than quietly ignored.

Two constraints shape the whole exercise. The assistants publish no visibility reporting of their own, so every number you hold is one you generated. And responses are shaped by session state, memory and geography, so results are comparable only when those conditions are held constant. A program that ignores either constraint produces figures that look precise and mean very little once someone senior starts asking how they were derived.

Why traditional rank tracking does not transfer

Rank tracking assumes a stable ordered list of results for a query that many people type identically. AI assistants break all three assumptions at once. There is no list, the answer is synthesized prose, and the prompts buyers actually type are longer, more conversational and far more varied than keywords. Two people asking the same underlying question in different words can receive materially different brand recommendations.

Responses are also non-deterministic. Running the same prompt five minutes apart on the same platform can produce different brands, different orderings and different citations. Practitioners generally see enough run-to-run drift that a single measurement should be treated as one sample rather than a reading. Three to five repetitions per prompt per platform is the usual minimum before a number is worth putting in front of anyone.

Personalization and account state add another layer of noise. Memory, prior conversation, logged-in status, geography and subscription tier all influence output. Measurement should therefore run from clean sessions with memory disabled wherever the platform allows it, and the session conditions should be written down so that results stay comparable from one month to the next.

None of this makes measurement impossible, but it does change what a good number looks like. In classic rank tracking a position is a fact. In AI visibility a mention rate is an estimate with a spread around it, and the honest way to present it is as a proportion across repeated runs with that spread shown. Teams that report it as a single hard number invite challenges they will not be able to answer.

The prompt panel is the decision everything else depends on

The prompt panel determines whether your measurement means anything at all. A workable panel contains between forty and one hundred prompts, written in the language buyers actually use, and covering the full funnel from category definition through vendor comparison to implementation, integration and pricing questions. Panels smaller than about thirty prompts are too noisy to trend reliably across months.

Build the panel from real inputs rather than keyword tools. Sales call recordings, support tickets, RFP questions and the questions asked live in demos are the strongest sources, because they capture phrasing that keyword research systematically misses. Mix unbranded category prompts, competitor-branded prompts, and prompts naming your own brand directly, since each of the three measures a genuinely different thing.

Then freeze it. The entire value of the panel comes from comparability over time, so resist the urge to rewrite prompts each quarter because a new campaign launched. Add a small tranche of new prompts when the category genuinely shifts, but keep the original core untouched and report on it separately so that the long trend line stays intact and readable.

Document the panel like a research instrument. Record why each prompt was included, which funnel stage it maps to, and which competitor it is designed to test you against. That documentation is what allows a new analyst to run the panel two years later and produce comparable output. It is also the difference between a measurement program and a recurring task that drifts quietly until nobody trusts the trend anymore.

The Five-Column Visibility Ledger

The Five-Column Visibility Ledger is a recording format that keeps every run of the panel comparable. Column one is the prompt, verbatim. Column two is the platform and the run date. Column three is presence, a simple yes or no on whether the brand was named. Column four is position, meaning whether the brand led the answer, sat mid-list or was mentioned only in passing. Column five is provenance, listing which domains were cited.

Provenance is the column most teams skip and the one that produces the most action. It tells you which third-party pages the assistants are reading to form an opinion about your category. When a review site, a community thread or an analyst page shows up repeatedly across platforms, that page is effectively part of your positioning whether or not you have ever influenced what it says.

Position deserves more nuance than a binary. Being the first vendor named in a recommendation carries far more commercial weight than being the fourth item in a list, and the gap is wide enough that teams tracking mention rate alone can report improvement while their actual influence on buying decisions declines. Recording position as a separate field prevents that particular reporting failure.

Which metrics actually mean something

Four metrics carry most of the signal. Mention rate answers whether you are in the conversation. Position-weighted share of voice answers whether you are winning it. Citation rate answers whether your own content is being used as evidence rather than someone else's description of you. Sentiment and factual accuracy answer whether what the assistant says about you is correct, which is the metric with the most direct revenue consequence.

Accuracy failures are common and badly under-monitored. Assistants routinely describe products with stale pricing, discontinued features or competitor capabilities attributed to the wrong vendor. Enterprise teams that audit for factual errors typically find problems in ten to twenty percent of brand-specific responses at the start of a program. Each error is correctable, and most trace back to an outdated source page the model keeps reading.

Resist composite scores. Vendors and internal teams both like a single visibility number because it fits on a slide, but blending mention rate, sentiment and citations into one index hides which lever actually moved. Keep the metrics separate on the dashboard and let the written narrative do the synthesis, so that a change in the number always has a traceable cause.

One further metric earns its place once a program matures: how the assistant frames the category itself. Assistants usually define a category before recommending vendors, and that definition determines which vendors qualify to be named. If the prevailing definition excludes the way your product actually works, no amount of mention optimization will fix it. Tracking the definitional language across runs surfaces the problem while it is still a content problem rather than a positioning crisis.

How often to measure and how to handle variance

Monthly is the right default cadence for most enterprise programs, with weekly runs reserved for a small subset of high-value prompts during an active campaign. Daily measurement mostly captures noise. The models themselves change on release cycles measured in months, and content changes take weeks to propagate through retrieval, so a monthly rhythm matches the timescale of the underlying system rather than fighting it.

Handle variance by repeating and averaging rather than by pretending it does not exist. Run each prompt three to five times per platform per cycle and report mention rate as a proportion rather than a binary. Record the model version wherever the platform exposes it, because a version change is the single most common explanation for a sudden swing that has nothing whatsoever to do with your work.

Report the four platforms separately. Perplexity leans heavily on live retrieval and tends to move fastest in response to content changes. ChatGPT and Claude blend retrieval with parametric knowledge in ways that differ by prompt type. Gemini interacts with Google's index on its own terms. Averaging the four into one figure produces a number that describes nothing in particular and hides where the movement came from.

Turning measurement into a decision

A measurement program earns its budget when each metric maps to an owner and an action. A low mention rate on unbranded category prompts is a content coverage and third-party presence problem. A low citation rate alongside healthy mentions means your content is not retrievable, and the fix is technical and structural. Accuracy errors are a source-correction problem, handled by finding and updating the specific pages the models are reading.

Report upward in commercial terms. Executives do not need mention rate broken out by prompt. They need to know what share of the buying questions in the category currently return your brand, how that has moved since last quarter, and what the movement is worth against known deal sizes. Tag every prompt to a funnel stage and report the bottom-funnel subset on its own line.

Tooling helps but is not the starting point. A spreadsheet running a frozen panel monthly beats an expensive platform running an unexamined default panel. Lemniscate Growth publishes free AEO and GEO checkers inside The GrowthGPT for teams that want a first baseline before committing to a vendor, and the sequence matters: establish the panel, prove the trend is readable, then buy automation for the parts that are genuinely tedious.

Set expectations on timelines before the first report goes out. Technical and structural fixes on your own pages can move citation rate within weeks. Coverage gaps that require new content move on publishing and indexing timelines, usually one to two quarters. Brand-level presence that depends on what third parties write about you moves slowest of all. A program judged on thirty-day results tends to be cancelled before the slower levers have had any chance to register.

FAQ. Quick answers.

Still unsure? Ask us directly.

Can Google Search Console or standard analytics measure AI visibility for us?

No. Search Console reports on Google Search surfaces and does not cover third-party assistants, and analytics only sees users who clicked through. Since most AI answers resolve without a click, analytics undercounts influence severely. Referral data from assistant domains is a useful supporting signal, but the primary measurement still has to come from running prompts and recording answers.

Should we run measurement from logged-in accounts or anonymous sessions?

Anonymous or memory-disabled sessions, for comparability. Logged-in accounts accumulate history that biases responses toward brands the account has discussed before, which flatters your own numbers. Run the core panel from clean sessions, and if you want to understand personalization effects, run a separate small panel from typical buyer-like accounts and report the two streams independently rather than blending them.

What mention rate should a mid-market B2B brand expect in its own category?

It varies widely by category concentration. In categories with three or four dominant vendors, challengers often start in the range of five to fifteen percent on unbranded prompts. In fragmented categories, twenty to thirty percent is achievable sooner. The absolute number matters less than the trend and than your position relative to the specific competitors you lose deals to.

How do we tell whether a change in the numbers came from our work or from a model update?

Log model versions and dates alongside every run, and watch whether competitors moved at the same time. A shift affecting every brand in the panel simultaneously is almost always platform-side. A shift confined to your brand, on prompts related to pages you recently changed, is more likely attributable to your work. Neither is provable, so record both readings.

Which team should own AI visibility measurement inside an enterprise?

Usually organic search or content, because the remediation work is mostly theirs, with product marketing owning accuracy and messaging corrections. What matters more than the reporting line is that one person owns the panel and the cadence. Programs split across three teams tend to produce inconsistent panels, which destroys comparability and makes the trend line useless within two quarters.

Turn this into pipeline. We can run it with you.

Tell us the revenue number and the market. We will come back with the stages that matter most for you, and the ones you can skip.

  • 20 minutes with a senior operator, not an SDR
  • Bring your revenue target and markets; we bring the pipeline math
  • Slots across US, Canada, India, Singapore and GCC time zones

Prefer email? growth@lemniscategrowth.com

Pick a 20-minute slotStraight to a senior operator. No SDR screen.