You type your business name into ChatGPT, ask whether it would recommend you, and there you are in the answer. Relief. The AI knows you, it says good things, and you close the tab reassured that your visibility is fine. That single check felt like proof. It was almost nothing of the sort.
Tracking your presence in AI answers is a sensible thing to want. The problem is that the quick, casual way most people do it produces a picture that is not just incomplete, it is actively misleading. It tells you comforting things that are not true and hides uncomfortable things that are. Businesses make real decisions on the back of these checks, and the decisions are only as good as the tracking behind them.
This guide explains why casual AI brand tracking misleads you, the specific traps that catch people out, and what reliable tracking looks like instead. None of it is complicated. It mostly comes down to understanding one thing about AI answers that changes how you should measure them.
What brand tracking means in AI search
Brand tracking in AI search is the practice of checking how your business appears when people ask AI tools the questions that lead to a sale. Done properly, it tells you where you show up, how you are described, whether you are recommended or merely mentioned, and how that changes over time.
That is a genuinely useful thing to measure. It is the AI-era version of watching your rankings, and it should inform where you spend effort. The trouble is not the idea of tracking, it is the method. A ranking on Google is relatively stable. You can check it today and expect a similar answer tomorrow. AI answers do not behave that way, and treating them as if they do is where the trouble starts.
If you want the formal version of what good measurement looks like, our guide to what an LLM Visibility Score is and how it is measured sets it out. This post is about why the casual version goes wrong, so you understand what you are trying to avoid before you build something better.
How this differs from tracking Google rankings
It helps to see exactly how AI tracking differs from the rank tracking you may already know, because the habits that work for one actively mislead you in the other.
A Google ranking is a position on a page, and it is stable enough to check once and trust for a while. You search, you see position four, and position four is a fact you can act on. Rank trackers exist because that number is consistent enough to be worth recording daily.
AI presence is not a position, it is an outcome that is regenerated each time someone asks. There is no single slot to occupy and no fixed number to read. Two people asking the same question can get different answers, and the same person asking twice can too. That is why the tools and instincts built for rank tracking fail here. They assume a stability that AI answers do not have.
The practical lesson is to bring the discipline of rank tracking, a fixed method repeated over time, without the assumption that any single reading is the truth. You measure the pattern, not the position.
The one thing that breaks casual tracking: AI answers vary
Here is the fact that undoes most quick checks. Ask an AI tool the same question twice and you can get two different answers. Ask it a third time and you might get a third. The businesses named, the order they appear in, and whether you are included at all can shift from one run to the next, even when nothing about your business has changed.

Ask the same question three times and the answer moves. A single check sees only one version.
This is not a fault you can avoid, it is how these systems work. They generate answers rather than looking them up, and a degree of variation is built in. For tracking, the consequence is serious. A single check catches one version of a shifting picture, and you have no way of knowing whether you caught a good day or a bad one. You might appear in that run and be absent in the next two, or the reverse. Either way, the single check told you far less than it seemed to.
Once you accept that answers vary, the whole approach to tracking has to change. You stop trying to capture the answer, because there is no single answer, and you start measuring the pattern across many answers instead. That shift is the difference between tracking that misleads and tracking that informs.
The specific ways your tracking misleads you
Variance is the root cause, but it shows up in several distinct traps. Most casual tracking falls into at least one of them, and often several at once.

Each of these habits quietly turns a check into a false reassurance.
You check once
The most common trap, and the most damaging. A single run captures one version of a variable answer. Whatever it shows, good or bad, it is not reliable on its own. You need several runs of the same prompt to see the real pattern.
You test one prompt
People rarely ask a question just one way. They phrase it differently, add context, and come at it from different angles. Checking a single prompt tells you about that one phrasing and nothing about the dozens of others your customers actually use. Reliable tracking rests on a proper set of prompts, which is why finding the prompts your customers ask comes first.
You are signed in
This one is quietly disastrous. If you are logged in and have been researching your own business, the tool personalises its answers to your history and cheerfully shows you your own name. It feels like a strong result. It is a mirror. Always track in a fresh, logged-out session with personalisation and memory turned off, so you see what a stranger sees.
You check one platform
Appearing well in ChatGPT tells you little about Gemini, Perplexity or Copilot. Businesses frequently show up strongly on one and are absent on another, and that difference is important information. Checking a single platform hides it. This is also why a proper visibility audit spans several platforms rather than one.
You mistake a mention for a citation
Seeing your name in an answer is not the same as the AI using your website as the source. The first is a mention, the second is a citation, and they mean very different things. Counting mentions as citations makes your position look stronger than it is. Our guide to the AI citation audit explains how to tell them apart, and why the difference matters.
You keep the run you liked
When answers vary, it is human nature to remember the flattering one. If you run a prompt three times and quote the run where you appeared, you have not tracked anything, you have cherry-picked. Record every run, including the ones that leave you out, because those are the ones telling you the truth.
What reliable tracking looks like
The fix is not complicated. It follows directly from understanding variance. Reliable tracking swaps a single lucky look for a consistent, repeatable process.

Reliable tracking is the same questions, repeated, across tools, logged out, measured on a schedule.
A fixed set of prompts. Use the same list of real buyer questions every time, so your results are comparable across checks. Changing the prompts each time makes the numbers meaningless.
Several runs of each prompt. Run each question two or three times and record how often you appear, rather than whether you appeared once. The pattern is the measurement, not any single run.
Across the main platforms. Check ChatGPT, Gemini, Perplexity and Copilot, and record each separately, so you can see where you are strong and where you are missing.
Logged out. Always in a fresh, non-personalised session, so the answers reflect what real prospects see rather than your own history.
Tracked over time. Repeat the same process on a regular schedule and keep the results. The trend is what matters, and it only appears when you measure the same way repeatedly.
Do these five things and your tracking stops flattering you and starts informing you. It takes more effort than a single check, but not much more once the prompt list exists, and the difference in what it tells you is enormous.
The metrics that actually matter
Once you are tracking properly, a handful of metrics carry almost all the value. These are the numbers worth watching, and they are far more honest than a yes-or-no glance.
Put these on a simple dashboard and revisit them each time you track. None of them can be read from a single check, which is rather the point. They only exist once you measure across many runs, and that is exactly why they are trustworthy where a one-off glance is not.

A few honest metrics, tracked over time, beat any number of one-off checks.
Share of voice. Across your prompts and runs, what portion of answers include you? This single figure captures your real presence far better than any individual check.
Citation rate. How often is your website used as the cited source, not just mentioned? This tells you whether the AI relies on you or merely knows you exist.
Prompt coverage. How many of your important questions do you appear for at all? Gaps here show you exactly where you are missing.
Consistency. When you appear, do you appear reliably across runs, or only occasionally? A business that shows up in one run of three has a fragile presence, and that fragility is a signal in itself. Together these feed the kind of figure described in our LLM Visibility Score guide, which rolls them into one number you can watch move.
How to build a simple tracking sheet
You can set up reliable tracking in a spreadsheet in under an hour. Here is a straightforward way to do it.
List your prompts down the side. Use your fixed set of real buyer questions, the same ones every time, so results stay comparable.
Add a column for each platform. ChatGPT, Gemini, Perplexity and Copilot, kept separate so you can see where you are strong and weak.
Record runs, not a single result. For each prompt and platform, run it a few times and note how often you appeared, such as two of three.
Mark cited versus mentioned. A quick tag in each cell showing whether you were the source or just named keeps the two from blurring together.
Date each pass and keep the old ones. A new tab per month lets you see the trend rather than overwriting your history.
Once the sheet exists, each pass is mostly repetition, which makes it quick to maintain. You can extend the same sheet to cover competitors, turning it into the basis for an AI competitor analysis as well, so one piece of work feeds several.
A worked example: a West Sussex windows and doors company
A windows and doors company in West Sussex is confident about its AI visibility. The owner has checked ChatGPT a few times, seen the business named, and moved on reassured. Then they run a proper tracking pass, and the picture changes.
They take fifteen real questions, run each three times across three platforms, all logged out. The results are sobering. On ChatGPT, where the owner had been checking, the business appears in most runs, which is why it felt fine. On Gemini and Perplexity it barely appears at all. Across everything, it shows up in about a third of runs, and it is cited as a source almost never. The earlier confidence rested entirely on checking the one platform, signed in, that happened to know them.
More usefully, the tracking shows where the presence is fragile. On several important questions the business appears in one run and vanishes in the next two, which means it is on the edge of the answer rather than secure in it. Those are the prompts to shore up first. None of this was visible from the casual checks. It took a proper, repeated process to see it, and once seen, it turned a vague sense of comfort into a clear list of priorities.
The owner’s first reaction was to argue with the tracking, because it contradicted what they had seen. But once they re-ran the prompts themselves, logged out and across platforms, the pattern held. That is often the moment proper tracking earns its place. Not when it confirms what you hoped, but when it calmly shows you something you would rather not have found, and turns out to be right.
How much tracking is enough
It is possible to overdo this. Because AI answers move day to day, tracking obsessively produces a stream of noise that tells you nothing and eats your time. The goal is a rhythm that catches real change without drowning you in random variation.
For most businesses, a full pass each month or quarter is plenty. Run your fixed prompt list across the platforms, log the numbers, and compare with last time. Between passes, a quick weekly glance at your five or six most important prompts is enough to spot any sudden drop worth investigating, as long as you remember that a single bad run is not yet a trend.
The effort is front-loaded. Building the prompt list and the sheet takes an hour or two the first time. After that, a monthly pass is a short, repeatable task that one person can own. That modest, steady investment gives you a far truer picture than hours of frantic, one-off checking ever would.
Common mistakes
Beyond the traps already covered, a few habits weaken even well-intentioned tracking.
Changing the prompt list each time. If the questions move, the results are not comparable. Fix the list and keep it stable so you can measure change.
Tracking only your own name. Without competitor context, your numbers float free. Appearing in a third of answers means one thing if rivals appear in all of them and another if they appear in none.
Ignoring why the numbers move. A metric that drops is a prompt to investigate, not just a figure to log. Pair your tracking with a look at the trust signals and content behind the change.
Measuring too often. Answers vary day to day, so checking constantly produces noise. A steady monthly or quarterly rhythm shows the real trend without the jitter.
Tracking without acting. The point of measurement is to guide work. If your tracking never leads to changes in content or presence, it is a hobby rather than a tool. Use it to decide what to fix, drawing on how businesses get recommended by AI.
Frequently asked questions
Why do AI answers change when I ask the same thing twice?
Because AI tools generate answers rather than retrieving a fixed result, some variation is built in. The businesses named and their order can shift from run to run even when nothing about them has changed. This is normal, and it is the main reason a single check is unreliable.
How many times should I run each prompt?
Two or three runs of each prompt is usually enough to see the pattern. What matters is that you record how often you appear across the runs, rather than treating any single run as the answer.
Do I need a paid tool to track properly?
No. You can track reliably with free accounts and a spreadsheet, as long as you use a fixed prompt list, several runs, multiple platforms, logged-out sessions, and a regular schedule. Tools help at scale and add consistency, but the method matters more than the software.
Why does being signed in matter so much?
AI tools personalise answers to your history. If you have been researching your own business while signed in, the tool is more likely to show you your own name, which flatters your result. Always track logged out, with personalisation off, to see what a real prospect sees.
How often should I track?
A monthly or quarterly rhythm suits most businesses. Answers vary day to day, so checking too often produces noise. Regular, consistent checks reveal the trend, which is what you actually want to know.
Track it properly, decide with confidence
The reassuring glance at ChatGPT is not tracking, it is a coin toss you happened to win. Real tracking accepts that AI answers vary and measures the pattern instead of the moment. It takes a little more effort, but it replaces false comfort with an honest picture you can actually act on.
Start by building a fixed prompt list and running it properly just once, logged out, several runs, across platforms. Compare what you find with the casual checks you have been relying on. The gap between the two is usually the moment the value of proper tracking becomes obvious.
The businesses that get this right are not the ones checking most often. They are the ones checking consistently, believing what the numbers show, and acting on it. That is a low bar to clear, and clearing it puts you ahead of most of your market, which is still glancing at ChatGPT once in a while and calling it tracking.
If you would rather have your AI visibility tracked and interpreted for you, with the metrics that matter reported clearly, that is part of our AI optimisation services. Book a discovery call and we will show you where you really stand in AI answers, not just where a single check suggested.





