You want to know where you stand. You don't want to buy anything to find out, and you shouldn't have to. Before spending a dollar on a tool or a consultant, you can build a genuinely accurate picture of your AI visibility in an afternoon, using nothing but the AI engines themselves and a spreadsheet.
The cost of skipping this step is that every later decision gets made on a guess instead of a baseline. Teams that jump straight into "publishing more content" or "fixing schema" without first checking where they actually stand often spend weeks fixing problems they don't have while ignoring the one that's actually holding them back. A few hours of manual audit work now prevents that, and it costs nothing but time.
Here's the real method, the same one a paid tool runs under the hood, done by hand.
Step 1: Build Your Prompt List (30 minutes)
Write down 15 to 20 questions your actual buyers would type into an AI engine. Not keywords, questions, phrased the way a real person searching for a solution would phrase them, including their context: company size, budget, use case, or the competitor they're already comparing you against.
Pull these from real sources if you have them: sales call notes, support tickets, the specific language customers use in your reviews. If you don't have those, write the list from the buyer's actual decision journey instead of your own product's feature list: "best [category] for a [size] company," "[your brand] vs [top competitor]," "is [your brand] worth it," "alternatives to [category leader]." A prompt list built from genuine buyer language is the single biggest factor in whether the rest of this audit produces something useful or something misleading.
Aim for a spread across the buyer's journey, not twenty variations on the same question. A reasonable split for twenty prompts: five broad discovery questions ("best tools for X"), five direct comparison questions naming you and a specific competitor, five decision-stage questions ("is X worth it," "X pricing," "X vs Y for [specific use case]"), and five category or definitional questions where your brand could plausibly be mentioned as an example even though it isn't the direct subject. That spread surfaces different problems: discovery prompts test whether you're in the running at all, comparison prompts test how fairly you're positioned against named competitors, and definitional prompts test whether you're recognized as a credible example of your category in the first place.
It's tempting to treat prompt list creation as a five-minute formality before the "real" audit work begins. Resist that. A rushed, generic prompt list is the single most common reason a DIY audit produces a result that doesn't match how the brand actually performs when a real customer asks. If you can, show the draft list to someone in sales or support before running it, and ask whether it sounds like the questions they actually hear. Adjust before moving to step two.
Step 2: Run the Prompts Across Engines (60 to 90 minutes)
Open ChatGPT, Perplexity, Google Gemini, and Google AI Mode in separate tabs. Run every prompt through each one, and for each result record four things in a simple spreadsheet: whether you appear at all, where you rank relative to competitors if a list is given, what the engine specifically says about you, and whether anything it says is factually wrong.
This is the tedious part, and it's also the part no paid tool can shortcut meaningfully better than you doing it by hand for a first pass. Twenty prompts across four engines is eighty checks. Budget the time and don't rush it. The value of this audit lives entirely in the specific wording each engine uses, not just a binary yes or no on whether you showed up.
Note patterns as you go, not just individual results. Are you consistently missing from one engine but present in the others? Are you present but consistently ranked below the same one or two competitors? Is the description of your product accurate on some engines and outdated or wrong on others? Patterns matter more than any single data point.
Here's what one filled-in row actually looks like, so the format is concrete rather than abstract. Prompt: "best [category] tool for a 20-person agency." Engine: Perplexity. Appeared: yes, third position. Competitors ahead: two, both larger, established players. Description accuracy: mostly accurate, though it described your pricing tier one level higher than it actually is. That last detail, a specific, checkable inaccuracy, is worth more to you than the position number, because it's something you can trace back to a source (probably an outdated third-party listing) and correct directly.
A note on judging "accurate" fairly: don't mark a description wrong just because it's not flattering, or because it emphasizes a feature you'd rather it didn't. Mark it wrong only when it states something factually incorrect: wrong pricing, wrong category, wrong ownership, a feature you don't actually have. Conflating "unflattering" with "inaccurate" is the most common way this audit produces a skewed, overly defensive result.
Step 3: Audit Your Structural Presence (45 minutes)
Separately from the prompt run, check the foundational elements that make you legible to these systems in the first place.
Check whether you have a Wikidata entry, and if you do, whether the information on it is current. Check your Organization and Product schema markup on your core pages, using a free schema validator (there's no need to pay for this check). Check your G2, Capterra, or Trustpilot profile, if relevant to your category, for completeness: correct category, current logo, description that actually matches how you'd describe yourself today. Check whether your site's About page and homepage state clearly, in plain language near the top, what you do and who you do it for. This sounds basic, but a surprising number of sites bury this under a headline that's more clever than clear, and clarity here is exactly what an AI engine needs to categorize you correctly.
Step 4: Check for Entity Confusion (30 minutes)
This step alone justifies the afternoon for some brands. Search your brand name alongside "vs," "review," "who owns," and "is [brand] the same as" across the engines from step 2. Look specifically for the AI confusing you with a similarly named company, citing an acquisition or leadership change that isn't accurate, or describing your category incorrectly.
Entity confusion is more common than most founders expect, particularly for brands with generic names or names that overlap with another company in an unrelated industry. A common shape it takes: your brand shares a name with an unrelated company in a different country or a different industry entirely, and an AI engine occasionally blends facts from both when asked a broad question, citing the wrong headquarters, the wrong founding date, or a completely unrelated product line. Another common shape: an old acquisition or rebrand that a training snapshot picked up correctly at the time but that's since become outdated, so the engine keeps repeating a fact that was true two years ago and isn't anymore.
If you find any of this, note it separately from everything else in this audit. It's usually the highest-priority fix that comes out of the whole exercise, because until it's resolved, other improvements are competing against a fundamentally confused starting point, and no amount of new content fixes a problem rooted in the engine misidentifying who you even are.
Step 5: Score Yourself Honestly (30 minutes)
With the spreadsheet filled in, calculate two simple numbers per engine: the percentage of your prompt list where you appeared at all, and the percentage where the description of you was accurate. These two numbers, run side by side, tell you more than a single composite score would. A brand that appears frequently but is often described wrong has a different, more urgent problem than a brand that's accurately described but simply absent.
Do the same calculation for your two or three closest competitors, using the same prompt list. Your numbers mean little in isolation. Knowing you appear in 40 percent of prompts while your top competitor appears in 75 percent tells you something real. Knowing you appear in 40 percent with no comparison point tells you almost nothing.
Common Mistakes That Skew the Results
A handful of habits quietly undermine an otherwise well-run audit, and they're worth checking for before you trust the numbers.
Running the prompts only once, right after signing in with an established account. Some AI engines personalize results based on account history and prior conversations. Where possible, run your audit prompts in a fresh or logged-out session, or at least be aware that a heavily personalized account may show you a rosier or stranger picture than a new user would actually see.
Writing prompts that are secretly about you. "What makes [your brand] the best choice for X" is not a neutral prompt, it's a leading one, and the response will reflect that framing rather than reflect how the engine would answer an open, comparative question. Every prompt in your list should be one a genuinely undecided buyer would ask, with no hint baked in about which answer you're hoping for.
Stopping at the first sentence of a long AI response. These engines often bury the most useful detail, a specific caveat, a comparison point, a factual error, several sentences into a longer answer. Read the full response for each prompt, not just the headline recommendation, or you'll miss exactly the kind of specific, actionable finding this audit is supposed to surface.
Auditing once and never again. A single snapshot tells you where you stand today. It says nothing about direction. Even a rough repeat of the short version a month later turns a static number into a trend, and a trend is what actually tells you whether anything you're doing is working.
What This Audit Tells You, and What It Doesn't
Done properly, this afternoon produces a genuinely useful baseline: where you stand today, per engine, with specific examples of what's working and what's broken, and a comparison point against real competitors.
It has a real ceiling, though, and it's worth naming honestly. A manual audit like this works well up to roughly twenty prompts across four engines, which is already eighty individual checks and a full afternoon of focused work. Push much beyond that, say fifty prompts across six engines checked weekly, and the manual approach stops scaling. At that point you're not choosing between free and paid because the free version is inferior, you're choosing because tracking that volume by hand every week becomes a part-time job on its own, and that's the actual point where a paid tracking tool starts to earn its cost: not because it does something fundamentally different, but because it automates a task that's no longer reasonable to do by hand.
There's a second honest limitation worth stating. This audit tells you where you stand right now. It doesn't tell you why, and it doesn't fix anything by itself. Once you have the numbers, the actual remediation work, fixing entity confusion, closing schema gaps, building comparison content, is a separate project that this audit only points you toward. Treat the afternoon as diagnosis, not treatment, and you'll use the results correctly.
Turning This Into a Repeatable Check
Once you've run this the first time, the fastest version to repeat monthly is a scaled-down check: your top eight to ten prompts, run across the two engines that matter most for your category, checked in under an hour. That's enough to catch a meaningful regression or improvement without repeating the full afternoon every time. Save the full version for a quarterly deep check, and use the short version to keep a rough pulse on things in between.
Keep every audit's raw spreadsheet, not just the summary numbers. Six months from now, being able to pull up exactly what an engine said about you in a specific prompt back in month one is worth more than any trend line, especially if you're ever explaining to someone else on the team why a particular fix mattered.