
To measure AI visibility, ask the AI engines your customers use the questions they actually ask, then record whether your business is mentioned, cited, or recommended in the answers. AI visibility is the degree to which ChatGPT, Claude, Perplexity, Gemini, and Google AI Overviews surface your brand, and you can measure it with a simple, repeatable testing process.
Start With the Questions Your Customers Ask
Your measurement is only as good as your prompt set. Build a list of 20 to 50 questions that map to real buying journeys: best-of queries, comparisons, local searches, and problem statements your product or service solves. Pull them from sales calls, support tickets, review sites, and the search terms already bringing people to your website.
Group the prompts by intent and topic so you can spot patterns later. A plumber might separate emergency prompts from renovation prompts; a software company might separate category prompts from alternative-to-competitor prompts. Once you start measuring, keep the list fixed, because changing prompts midstream makes your trend data meaningless.
Worked Example: Building a Prompt Set for a Regional Accounting Firm
Imagine a 12-person accounting firm in Denver that serves small businesses and dental practices. Its prompt set starts with category questions: best small business accountant in Denver, top CPA firms for dental practices, who should handle bookkeeping for a new restaurant. Next come comparison and alternative questions: the firm's name versus two local rivals, and alternatives to the national chains its prospects consider first. Then problem questions: how much does outsourced bookkeeping cost, when should a small business switch from DIY accounting, and do I need a CPA or a bookkeeper.
That gives the firm roughly 30 prompts spread across four intents. Each prompt maps to a real conversation its partners have had with prospects, which is the test every prompt should pass: if no customer would ever ask it, it does not belong in your measurement set.
Test Across Multiple AI Engines
No two engines answer the same way. ChatGPT may recommend you while Perplexity ignores you, and Google AI Overviews may cite a directory that lists you rather than your own site. Run your full prompt set on every engine your customers plausibly use, and treat each one as its own channel with its own score.
Always test in fresh sessions with personalization and memory turned off, so the results reflect what a brand-new customer would see rather than what the engine has already learned about you.
It is also worth understanding why the engines disagree. Each one pulls from different sources, refreshes them on different schedules, and blends live retrieval with training data in different proportions, so your visibility can be strong in one engine and absent in another. Our guide on how AI engines decide which businesses to recommend explains the mechanics behind those differences.
Record Every Result the Same Way
Consistency is what turns spot checks into measurement. For every prompt and engine combination, log the same fields:
- Mentioned: was your brand named anywhere in the answer, yes or no
- Position: first recommendation, mid-list, or a passing reference
- Citation: did the answer link to your site or to a third-party page about you
- Sentiment: positive, neutral, or negative framing of your brand
- Competitors: which other brands appeared in the same answer
A dated spreadsheet works fine at small scale. What matters is that every test is captured the same way, so month-over-month comparisons are honest and defensible.
Turn Raw Results Into a Score
Roll your logs up into a handful of numbers you can chart: mention rate, which is the share of prompts where you appear; citation rate; average position; and a simple sentiment score. Many teams then combine these into one composite AI visibility score, so leadership can see the direction of travel at a glance while marketers work the underlying details.
Worked Example: Scoring a Month of Results
Suppose you run 40 prompts on three engines, for 120 total answers. Your brand appears in 30 of them, so your mention rate is 25 percent. Of those 30 appearances, 12 include a link to your site or to a profile about you, a citation rate of 40 percent among your mentions. You are the first recommendation 6 times, mid-list 18 times, and a passing reference 6 times. Sentiment is positive in 21 answers, neutral in 8, and negative in 1.
Next month you run the identical set. Mention rate rises to 30 percent and first-position appearances double, while everything else holds steady. That is a real, defensible improvement you can tie to the listings cleanup and comparison pages you shipped in between, which is exactly the cause-and-effect story a folder of ad hoc screenshots can never tell.
Repeat on a Fixed Cadence
AI answers shift as models update, sources change, and competitors publish new content. Measure monthly as your baseline, and move to weekly checks in competitive categories or during an active optimization push. Judge trends, not snapshots: one missing mention is noise, while three consecutive months of decline is a signal worth acting on.
A Step-by-Step Monthly Measurement Routine
Once the pieces above are in place, the whole process compresses into a repeatable routine you can finish in an afternoon:
- Open a clean session on each engine with memory and personalization switched off
- Run your fixed prompt set, one prompt per chat, and paste each answer into your log
- Score every answer for mention, position, citation, sentiment, and competitors named
- Roll the scores up into mention rate, citation rate, average position, and share of voice
- Compare each number against last month and flag any movement that repeats for a second month
- Pick one or two gaps to work on, and note the fix you are attempting next to the metric it should move
- Book the next run on the calendar before you close the spreadsheet
The last two steps are the ones teams skip, and they are the ones that turn measurement into improvement. A score without an attached experiment is trivia; a score with a named fix and a re-test date is a strategy.
Common Mistakes That Undermine Your Measurement
Most broken AI visibility programs fail in the setup, not the intent. Watch for these traps:
- Testing from a logged-in, personalized account: your history skews answers toward brands you already engage with, including your own
- Changing the prompt set every month: new questions reset your baseline and turn your trend lines into noise
- Treating one answer as the truth: engines vary between runs, so single checks overstate both wins and losses
- Testing only one engine: customers are spread across several, and strength on one says nothing about the others
- Testing only branded prompts: asking about your own name measures reputation, not discovery; unbranded category prompts are where new customers are won
- Measuring without acting: a scorecard that never changes what you publish or fix is overhead, not strategy
Common Objections, Answered
We are too small to worry about this
Small businesses are often the biggest beneficiaries, because AI answers flatten the advantage of large ad budgets. When a customer asks for the best provider nearby, the engine recommends whoever the evidence supports, and a small firm with clear information and strong reviews can beat a big one. Measuring is how you find out whether you are that firm. Once you know where you stand, our guide on how to improve AI visibility for your business covers what to do next.
AI answers change too often for measurement to mean anything
Variability is an argument for measurement, not against it. Because single answers fluctuate, the only way to know your real position is to sample repeatedly and read the trend. Rates calculated across dozens of prompts and multiple runs are stable enough to act on, even when any individual answer is not.
We already rank well on Google, so we must be fine
Google rankings and AI recommendations are related but not interchangeable, and businesses regularly discover they are strong in one channel and weak in the other. The engines draw on a wider evidence base than a ranking algorithm does, and they compress ten results into a two or three brand answer. The only way to know how you fare in that compression is to test the AI channel on its own terms, which is exactly what this process does.
Adjacent Terms Worth Knowing
A shared vocabulary keeps reports honest. These are the terms you will meet across this topic, and they all sit inside the bigger picture drawn in our complete guide to AI visibility for businesses.
- Prompt set: the fixed list of customer questions you test on every engine, every cycle
- Mention rate: the share of tested prompts where your brand appears in the answer
- Citation: a linked source in an answer that points to your site or to a page about you
- AI share of voice: your portion of all brand mentions across the answers in your category
- Sentiment score: a consistent rating of how positively answers describe your brand
- Answer engine: any AI system that responds with a synthesized answer rather than a list of links, such as ChatGPT, Perplexity, or Google AI Overviews
How the Pieces Fit Together
Measurement is one discipline with several connected parts, and each part deserves more depth than a single article can give. The numbers you log each month roll up into the AI visibility metrics and KPIs every business should track, and the competitive version of those numbers is AI share of voice, which tells you how much of the category conversation you own rather than simply how often you appear.
When you need a comprehensive one-time picture rather than a monthly pulse, run a full AI visibility audit. And to connect all of this activity to money, learn how to measure the ROI of AI visibility efforts. Together these form a loop: audit to establish the baseline, measure monthly to track movement, benchmark to stay honest about competitors, and report ROI to keep the work funded.
Frequently Asked Questions
Can I measure AI visibility manually?
Yes, at small scale. A fixed prompt list, a spreadsheet, and a monthly hour of testing will give you a genuine baseline. Manual measurement becomes painful once you track multiple engines, several competitors, and dozens of prompts, which is the point where automated platforms earn their keep.
What is a good AI visibility score?
There is no universal benchmark, because categories differ enormously. The practical standard is your own trend line plus your closest competitors: if your mention rate is rising and you appear more often than the rivals you actually lose deals to, you are winning.
How often should I re-measure?
Monthly is the right default for most businesses. Weekly makes sense in fast-moving or highly competitive categories, and immediately after a major model release, since new model versions can reshuffle which brands get recommended.
Do I need special tools to measure AI visibility?
No. The process in this guide runs on the engines' own free interfaces plus a spreadsheet. Tools earn their place when scale arrives: more engines, more prompts, more competitors, and more months of history than a manual routine can sustain without errors creeping in.
Should I test branded prompts, unbranded prompts, or both?
Both, in separate groups. Unbranded category prompts measure discovery, meaning whether you appear when buyers do not yet know your name. Branded prompts measure reputation, meaning what engines say when someone checks you out directly. The two move independently and call for different fixes, so blending them into one number hides more than it reveals.
Who in the company should own AI visibility measurement?
Whoever owns organic growth is the natural owner, usually the marketing lead. What matters more than the title is that one named person runs the same process on the same cadence, because measurement that belongs to everyone belongs to no one.
Related reading
- How to Check if ChatGPT Recommends Your Business
- Benchmarking Your AI Visibility Against Competitors
- Tools for Tracking AI Search Visibility
- How to Track Brand Citations in AI Answers
- Measuring Sentiment: What AI Engines Say About Your Brand
Start Measuring Before You Start Optimizing
You cannot improve what you have never measured, and most businesses guess wrong about how AI engines present them. Run your first prompt set this week, log the results, and put a repeat date on the calendar. Or skip the spreadsheet entirely: GrowBiz10x tracks your mentions, citations, and sentiment across the major AI engines automatically. Get your AI visibility score with GrowBiz10x and know exactly where you stand.
Want to see how your business performs across AI platforms?
Get a free AI visibility audit and personalized insights for your brand.
About the Author
GrowBiz10x Team
AEO SpecialistWe share actionable insights on AI visibility, content strategy, and digital growth to help businesses get discovered and grow faster.
Related Posts
Categories
- Commercial / Agency / GrowBiz10x20
- Industry-Specific10
- Measuring AI Visibility10
- Practical AEO/GEO10
- SEO vs AEO vs GEO10
- Brand Signals & Authority10
- How AI Understands & Recommends Businesses10
- The Search Shift10
- Pillar: AI Visibility10
Get the latest insights in your inbox.
Subscribe to our newsletter and never miss an update.
