What Is an AI Visibility Audit? (And What a Proper One Should Include)

An AI visibility audit is a structured test of how AI assistants like ChatGPT, Claude, Gemini and Perplexity represent your company: whether you appear when buyers ask relevant questions, how you rank against competitors, whether the claims made about you are accurate, and which sources are shaping the answers. A proper audit tests multiple engines with a meaningful number of prompts, scores the results systematically, and ends with a diagnosis and a prioritised fix list, not just screenshots.

That second sentence is the important one, because "AI visibility audit" currently means anything from a genuine research exercise to someone typing your name into ChatGPT and exporting the chat. This post sets out what the real thing involves, so you can judge any audit on offer, including mine.

Why this needs auditing at all

Half of B2B software buyers now start their research in an AI chatbot rather than a search engine, and AI recommendations have overtaken review sites as the biggest influence on vendor shortlists. Those conversations happen in private. No analytics tool shows you the answer your prospect just read, which means companies are routinely misdescribed, out-positioned or simply absent in front of active buyers, and nothing in their dashboard ever says so.

An audit exists to make that invisible layer visible: to answer, with evidence, the question "what does AI actually tell buyers about us?"

What a proper audit should include

Seven things. If an audit you're considering skips several of these, you're buying a screenshot deck.

1. Multiple engines, not just ChatGPT

ChatGPT is the biggest, but it's nowhere near the whole picture, and the engines disagree with each other constantly. In the audits I run, it's completely normal for a company to rank well on one engine and barely exist on another, because each engine reads different sources and weighs them differently. A single-engine audit isn't a smaller version of the answer. It can be a different answer entirely. Four engines (ChatGPT, Claude, Gemini, Perplexity) is a sensible baseline.

2. Prompts that mirror real buying questions, at meaningful volume

Testing your brand name proves nothing; AI engines will happily describe any company when asked directly. The revealing prompts are the ones that don't mention you: "best [category] for [customer type]", "alternatives to [competitor]", "X vs Y for [use case]", "we're scaling, what should we look at". A serious audit runs dozens of prompts across several categories of buying question. Mine runs 45 prompts across six categories, per engine, which is the sort of volume where patterns become trustworthy rather than anecdotal.

3. Clean, unpersonalised sessions

Any test run from a logged-in account is contaminated: the engine knows the tester, remembers previous chats, and skews accordingly. It's the single most common flaw in DIY checks, and it flatters you. Proper audits query the engines via API or clean sessions, so the answers are what a stranger sees.

4. Scoring, not vibes

Every response should be scored on at least three dimensions: presence (were you mentioned at all?), position (first name or grudging afterthought?), and accuracy (is what it said about you true?). Accuracy is the one thin audits skip, and it's often where the expensive findings live. In a recent audit I ran on a multi-billion-dollar software company, the engines mentioned them constantly and got the facts wrong repeatedly, including one product that doesn't exist. Presence without accuracy is its own kind of invisibility.

5. The source trail

When an engine cites its sources, those citations are the closest thing AI visibility has to a "why". If your competitor keeps winning a prompt, the sources explain what content is doing the winning: a comparison article, a review platform, a stale listicle from 2022. An audit that doesn't examine sources can tell you that you're losing, but not what to do about it.

6. Competitor benchmarking

Your score in isolation is half a finding. Present in 40% of answers: is that bad? It depends entirely on whether your nearest rival is at 15% or 85%. A proper audit runs the same prompt set against your key competitors so every number has context.

7. A diagnosis and a prioritised fix list

This is the difference between detection and value. The deliverable should end with why the results look the way they do (which gaps in your content, facts and third-party presence are driving them) and what to fix first for the biggest movement. If the final page of the report is a score with no roadmap, you've paid for a thermometer.

Red flags in thin audits

The compressed version, for anyone comparing providers: one engine only. A handful of prompts, or prompts that just test your brand name. No mention of clean sessions or methodology. Single runs presented as fact, when AI answers vary between runs. No accuracy checking. No sources examined. No competitors. A deliverable that's screenshots with commentary. Any of these alone is survivable; three or more and you're buying reassurance, not research.

DIY, tools, or done-for-you?

Honest answer: you can do a useful version of this yourself, and I've published the full 20-minute method [link: DIY check post]. It won't have the volume, the API-clean sessions or the benchmarking, but it will tell you whether you have a problem, and I'd genuinely rather you ran it than took anyone's word, including mine.

Monitoring tools sit in the middle: good at tracking presence over time, weaker on accuracy, diagnosis and what-to-fix. And a done-for-you audit makes sense when you want the full evidenced picture with a roadmap attached, or when someone senior needs convincing with more than a spreadsheet you built yourself.

What mine includes, for transparency

Since this post will be read by people comparing options, here's mine plainly: 45 prompts across six buying-question categories, run across ChatGPT, Claude, Gemini and Perplexity (180 scored responses), via clean API sessions, with every response scored for presence, position and accuracy, source analysis, competitor benchmarking, and a written report ending in a prioritised fix roadmap. It costs £1,750. You can see what the findings look like in practice in my teardown of a multi-billion-dollar software company [link: teardown #1], which is the same methodology applied to a company that isn't a client.

And if you just want to know whether you have a problem before spending anything: the free mini audit gives you a score in about a minute [link: mini audit].

FAQ

What does an AI visibility audit check? It checks how AI assistants represent your company across four dimensions: presence (whether you're mentioned when buyers ask relevant questions), position (how you rank against competitors in recommendations), accuracy (whether claims made about you are true), and sources (which content is shaping the answers). Results are scored across multiple engines and benchmarked against competitors.

How much does an AI visibility audit cost? Prices vary widely with depth. DIY costs nothing but your time. Monitoring tools typically run from tens to hundreds of pounds per month. Done-for-you audits generally range from several hundred to several thousand pounds depending on scope; mine is £1,750 for a four-engine, 180-response audit with a full diagnosis and fix roadmap.

How often should you audit your AI visibility? A full audit once or twice a year, with a lightweight monthly check in between, is sensible for most B2B SaaS companies. AI answers shift as models update and new content gets indexed, so the trend across checks matters more than any single result.

Next
Next

Do B2B Buyers Actually Use ChatGPT to Find Software? Here's the Data.