Almost every agency conversation I have at the moment is around being ‘found’ across LLMs. It’s a fair question, but unfortunately, one that it’s impossible to answer.
If a client asks, “How do we know if we’re showing up in ChatGPT?” I’ll be honest and tell them, “Well, you can’t, with any real degree of certainty”.
But then, a ‘GEO expert’ sends them an email saying “we can help you show up on ChatGPT” with a dashboard tracking 25 prompts that has a graph going up, and so, within a week, we’re looking at a score or percentage – but nobody can actually tell us how this data is useful in any way, shape, or form.
It’s costing agencies and clients real money to track a metric that’s mostly guesswork dressed up in a fancy UI. And the time we spend trying to work with this data is time being taken away from activities that actually make a difference (*cough actual SEO*).
Are AI visibility tools accurate?
AI visibility tools offer a rough snapshot of whether a brand appears in generative engines, but they are not fully accurate or absolute. Because large language models randomise, personalise, and dynamically change answers based on context, these third-party trackers only measure a small, simulated sample of queries rather than what people are actually prompting.
What can AI visibility tools actually do?
To be fair to the category, because it’s easy just to say it’s all a load of tripe and onions: AI visibility trackers do one genuinely useful thing.
You give them a list of prompts, they run those exact prompts against ChatGPT, Perplexity, Gemini, Copilot, Google AI Overviews (AIOs) and so on, on a schedule, and they log whether your brand got mentioned, where you sat in the answer, what got cited, and how you stack up against named competitors.
The number goes up or down. It tells you something changed. But it doesn’t tell you why, and it definitely doesn’t tell you how often real humans are typing anything remotely close to the prompts you fed it. It’s a best guess.
What can’t they do?
They can’t see real user prompts
Every tool in this space works from a prompt list you or they invent. Nobody has access to the actual query logs of ChatGPT, Gemini or Perplexity. You are testing your best guess at what people ask, not what people actually ask. For me, this is the most obvious flaw.
The auto-generated prompts are often just wrong
A lot of tools will scan your site or category and generate a starter prompt set for you if you don’t want to write your own. Reviewers have caught these auto-generated lists producing prompts with nothing to do with what the business actually does, missing smaller or newer competitors entirely, and skewing towards whatever the tool’s crawler happened to associate with your niche rather than what buyers actually ask. If you’re not checking and editing that list by hand (if that functionality even exists) before it starts scoring you, you’re being benchmarked against a fictional version of your own business.
LLM answers aren’t stable
Ask the same model the same question twice, and you can get different citations, different phrasing, sometimes a different brand entirely. Run it daily, and you’re measuring noise as much as signal.
Coverage is patchy and paywalled by engine
Most tools give you three or four engines on the entry plan and charge extra for the rest. Claude, Perplexity, Gemini, and Google AI Mode/AIOs are routinely add-ons rather than included.
A mention isn’t a customer
None of these tools can currently draw a reliable line from “you got cited in an AI answer” to “someone clicked through and converted.” The attribution simply doesn’t exist yet in any dependable form. If someone clicked and converted from an LLM, the best place to see this is in Google Analytics (GA4), adding a dimension to see the landing page that sent them.
Hallucinations happen
Just because a tool tells you that your brand has shown up doesn’t mean the information the LLM provided is correct.
Prompt volume is expensive
Track a serious number of prompts across several engines, several markets, and you’re into four figures a month before you’ve optimised a single page. And as highlighted above, because it isn’t based on real user data – IT IS PROBABLY WRONG.
Popular AI visibility tracking tools
Peec AI
Peec is well regarded for being clean, fast to set up and genuinely built for this one job. Reviewers consistently flag it as the best entry point if visibility tracking is literally all you want: a straightforward percentage, average position, sentiment, and source-level detail on what the AI is citing, exportable to Looker Studio or via API.
But that’s where it ends. It’ll tell you a competitor shows up in 62% of buyer prompts and you show up in 8%. Okay, so what? If you have a couple of massive competitors, that’s probably coming as no surprise.
You’re also capped at three engines per plan and a fairly small prompt allowance before the cost climbs, which for an agency running this across several client accounts adds up fast.
Otterly.ai
Otterly.ai is generally seen as the friendliest, cheapest credible way into this category, starting around $29 a month with unlimited seats, which matters for agencies juggling multiple client logins. It covers ChatGPT, Google AI Overviews, Perplexity and Copilot on the base tier, with Gemini and Claude sitting behind paid add-ons.
It turns a manual, spreadsheet-based process of manually checking prompts into something automated and comparable over time, which is a real improvement.
But even the reviewers championing it admit there’s currently no reliable correlation between an AI brand mention and actual traffic landing on your site. So effectively, you’re still just automating the measurement of something whose real-world impact still can’t be proven.
Semrush AI Visibility Toolkit
The advantage here is obvious if you already run Semrush: it sits next to the SEO data you already trust, so brand visibility, prompt research, and competitor comparison all live in one place instead of yet another login.
Engine coverage is thin; ChatGPT and Google’s Gemini, AI Overviews, and AI Mode are covered, but Perplexity and Claude aren’t.
Regional and language coverage is limited too, and the pricing stacks up fast: $99 a month buys one domain, 25 prompts, and every extra domain, seat or 50-prompt block is billed on top. Smaller brands with less search volume also get flagged repeatedly for sparse, occasionally erratic data.
The tool’s own auto-generated prompt suggestions produce irrelevant prompts, making the measurement of the AI Visibility score problematic.
Profound
Profound sits at the premium end and is worth including because it does one thing differently from the rest of this list: it claims to monitor real, live consumer-facing AI interfaces rather than just running your prompt list through an API, which in theory gets it closer to real usage patterns than pure prompt-based tools. It’s also the most feature-rich, with agent-based content workflows layered on top of the tracking.
The trade-off is cost and complexity. Starter is $99 a month for ChatGPT only with a low prompt cap. Growth jumps to around $332 to $399 a month billed annually for three engines and 100 prompts, and full engine coverage is Enterprise-only with custom pricing that reviewers put anywhere from $2,000 to $5,000+ a month once you factor in seats and platforms.
Other brands worth knowing
New entrants show up most weeks, but a few others come up repeatedly in comparison lists:
- Scalenut, which pairs visibility tracking with content creation and optimisation in one workflow rather than leaving you with a dashboard and nothing to do about it
- AthenaHQ, which covers eight engines with action-oriented workflows
- RankScale, generally rated as the widest engine coverage for your money.
However, no matter which tool you choose, they can’t solve the fundamental problem. None of these tools are bad software. They do what they say on the tin, but the tin itself is a proxy for something nobody can actually observe directly.
So what can you actually do?
Well, you can write out thousands of realistic prompts covering every stage of the buying journey, in every phrasing a real person might plausibly use, and you run them through each LLM in batches (or one by one for AIOs) or via an API/agent based workflow across every major model.
Even then, you’ve only tested the prompts you thought of. Real conversations with AI assistants are exactly that, conversational. People don’t type neat keyword-style queries; they follow up, rephrase, and ask “what about for a small business” three messages deep into a thread you’ll never see.
While you can check your own chatbot data (if you have one) and speak to the sales team or your customer service department to identify pain points (again, if you have one) to help give you some steer on what people are actually asking, no amount of prompt volume closes that gap completely.
You can build an AI visibility tool in-house, and plenty of agencies and some clients already are. Set up a workflow in something like n8n, feed it your prompt list, call ChatGPT, Claude, Gemini and Perplexity via APIs on a schedule, log the responses and citations to a spreadsheet or database, and it’s done. It’s not that hard to build, and it gives you full control over the prompt set, which sidesteps the auto-generated prompt problem above entirely.
But be prepared to absolutely burn through tokens if you’re running this properly. A few hundred prompts, run regularly, across four or five models, and multiplied across several client accounts adds up fast. Plus, it’s knowing what to do with that data once you have it, and working out if you can glean any genuinely useful information from it.
The personalisation problem
There’s a layer underneath all of this that makes the “we tested 1,000 prompts” approach even shakier than it already sounds, and it’s one most AI visibility conversations skip entirely: personalisation.
We’ve lived with this in Google for years. Signed-in search results have been tailored by location, search history, and account signals for well over a decade, and every SEO knows to caveat rankings data with “results will vary by user.” That’s annoying, but at least it’s a known, bounded problem. When talking about ranking data with clients – especially at keyword level – it’s important to explain the data should be used as a barometer rather than something to be treated as gospel. And at least we have some indication of approximate search volume (either through keyword tools or Google Search Console data queries) to marry it up with. Still not perfect, but better than nothing.
AI chat tools make the personalisation problem worse. ChatGPT now uses saved memory and prior conversation context to actively rewrite the prompt before it even runs a search, so “where should I eat” can silently become “vegan restaurants in a specific city” for its fan-out query based on what it already knows about the person asking.
Two people typing exactly the same words into ChatGPT can get citations to entirely different brands, and no analytics tool can see into another user’s memory profile to explain why.
That’s the bit that makes this harder than logged-in Google ever was. Google’s personalisation still runs against a query you typed. ChatGPT’s personalisation can change the query itself before anything gets searched, based on a memory profile you can’t see and a rewrite you’ll never read.
When a visibility tool runs your test prompt list, it’s running it logged out, with no memory, no history, no persona. It is not what most real users are getting, and there’s currently no way for any tracking tool, however good, to close that gap. Some providers are starting to segment logged-in versus logged-out and fast-path versus considered responses in their tracking to at least flag the volatility, which is a step in the right direction, but it’s a flag on the uncertainty, not a fix for it.
What actually gives you a useful steer, alongside a tool if you’ve got budget for one, is your own first-party data:
Google Search Console (GSC)
Filter for long, conversational, question-style queries with multiple words. These often read nothing like traditional keyword searches and may indicate that they’ve come from AI Mode.
Bing Webmaster Tools (BWT)
Similarly worth digging through for grounding data and queries that hint at Copilot-sourced traffic rather than a straight Bing search.
Here’s something else interesting that you can find in this data. If you filter your GSC query data and see things like “ok”, “yes”, “yes, continue” or “yes go on” showing up with their own impressions and clicks, that’s not junk data as you may expect.
When someone follows up on an AI Mode answer, Google logs the follow-up itself as a brand new query, and John Mueller has confirmed this behaviour directly. Spot your page cited alongside one of those, and you’ve got about as close to proof of an actual AI conversation happening where your brand was mentioned as you’re going to get from any tool on the market right now, free or paid.
But like all things Google (and it’s a properly annoying one), that same “yes go on” data is quietly inflating your impression counts. Every one of those follow-up non-queries gets logged as its own fresh query with its own impression, and sometimes its own click, sitting right alongside your actual keyword data in the performance report. So your overall impressions figure is padded out with strings like “yes” that tell you nothing about search demand and everything about someone mid-conversation with a chatbot.
It means the topline number you’d normally report to a client, ‘impressions are up’, isn’t as clean any more. Some of that growth might just be Google logging conversational filler as if it were a keyword. It’s worth filtering those rows out before you present anything, not just mining them for the odd useful signal if you see this cropping up a lot.
The bit I actually want clients to hear (even though it’s a hard conversation)
No tool, however polished the dashboard, has a Google Search Console equivalent for ChatGPT, Perplexity, or Claude. Nobody outside those companies can see the real query logs. Every visibility percentage you’re being sold is built on a synthetic prompt list, run through a model that won’t give the same answer twice, on engines that mostly won’t tell you when someone actually visited your site off the back of it.
We’ve been here before. Google pulled organic keyword data out of Analytics back in 2011, and by 2013 it was gone for the vast majority of traffic, replaced with the infamous “(not provided)”. That’s nearly fifteen years of the industry accepting that some data simply isn’t ours to have, and building workflows around directional signals instead of pretending we still had the full picture. AI visibility is the same problem, just newer and with worse tooling.
None of that means don’t bother at all. Track what you can, use GSC and BWT, run a sensible tool if the budget’s there and you understand exactly what it can and can’t prove, and keep an eye out for the “yes go on” rows because for Google at least, they’re the closest thing to showing you’ve been cited we have. But don’t forget to strip them out of your headline impressions number where they don’t belong.
Just don’t sell a client a single percentage figure and let them think it’s an accurate measure of their AI visibility, because it isn’t, and every reviewer who’s actually stress-tested these platforms ends up saying some version of the same thing.
We’re not there yet, and pretending otherwise costs everyone money for a number that doesn’t hold up – and that you shouldn’t base an entire digital marketing strategy on.