MONITORING · 5 ENGINES
GEOscanAI

AI VISIBILITY GUIDE

Building Your Prompt Set

How to build a prompt set from real buyer language instead of a guess, and why the quality of that list determines whether anything else you track means anything.

Look at the list of "buyer questions" your team now tracks across AI engines every week. Best project management tool for startups. Top CRM for small business. Somebody wrote that list down in a meeting, and nobody in the room had actually seen a real customer type those words. It came from a whiteboard, informed guesses about the category, and whatever competitor's tracked-keyword list somebody glanced at once.

That's not a minor detail. It's the foundation everything else in this discipline stands on, and if it's wrong, every number built on top of it is wrong in the same direction, consistently, in a way that feels like real data because it comes with a real percentage attached to it. A visibility score built on invented prompts tells you how you rank against questions nobody asks, which is a precise measurement of something that doesn't matter.

Here's the actual method for building a prompt set from real signal instead of a guess, along with the honest limits of what that method can achieve.

Why This Matters More Than Almost Anything Else in GEO

Every downstream decision, what content to prioritize, which engines to focus on, whether a given month's numbers count as progress, inherits whatever bias is baked into the prompt list. A list skewed toward your own product's feature names instead of the buyer's problem language will make you look artificially strong (you'll naturally rank well for terms close to your own vocabulary) while missing the actual questions costing you deals. A list that's too narrow, ten near-identical variations on one query, will make normal week-to-week noise look like a meaningful trend. Getting this step right is worth more time than almost any single piece of content you'll produce this quarter.

Step 1: Mine Real Buyer Language

Start with sources where actual prospects used their own words, not your team's words. Sales call recordings and transcripts are the richest source, if you have them: search specifically for how a prospect described their problem in their own first sentence, before the salesperson reframed it into product language. Support tickets are almost as good, particularly tickets from prospects still evaluating rather than existing customers with a specific bug. Your own G2, Capterra, or Trustpilot reviews are a third strong source, since reviewers frequently describe the exact use case or comparison that led them to you, in language a real buyer would recognize.

If you don't have access to any of these, the fallback is still workable: read the reviews of your two or three closest competitors on the same platforms. Competitor reviews, especially critical ones, are full of exactly the comparative language ("I switched from X because...") that maps directly onto real AI search prompts.

Pull twenty to thirty raw phrases from this research before trying to shape them into anything. Resist editing at this stage. The goal is volume and authenticity first, structure second.

Step 2: Structure the Set Across the Buyer Journey

A prompt set that's all discovery questions or all comparison questions gives you a lopsided, misleading picture. Structure your list to mirror how a real buyer actually moves through a decision, and build roughly even coverage across four categories.

Discovery prompts, broad category questions where you're one of many possible answers: "best [category] for [context]," "tools that help with [problem]." These test whether you're in the conversation at all.

Comparison prompts, naming you directly against a specific competitor: "[your brand] vs [competitor]," "[competitor] alternatives." These test how fairly and accurately you're positioned when a buyer is actively deciding between options.

Decision prompts, closer to purchase, often including qualifiers: "is [your brand] worth it," "[your brand] pricing," "[your brand] for [specific use case or company size]." These test whether the engine can answer the practical, close-to-conversion questions accurately.

Category or definitional prompts, where you might not be the direct subject but could plausibly be cited as an example: "what is [category term]," "how does [category] work." Appearing here signals a deeper kind of authority than simply ranking in a comparison list.

For a working set of twenty to thirty prompts, aim for something close to an even split across these four types. An imbalanced set doesn't just misrepresent your standing, it also biases which fixes look like they're working, since improving comparison-stage visibility won't move a discovery-stage number and vice versa. Teams that skip this structure tend to gravitate toward comparison prompts specifically, because they're the easiest to write and the most flattering to check, and end up with a set that looks rigorous but silently ignores whether the brand is even part of the conversation at the discovery stage in the first place.

Step 3: Write Each Prompt the Way a Person Actually Talks

Once you know the categories, write each individual prompt in natural, specific language, not the compressed, keyword-shaped phrasing that belongs in a search engine optimization brief. "CRM small business" is a keyword. "What's a good CRM for a 15-person sales team that doesn't need a ton of setup" is a prompt. The second version produces a meaningfully different, more realistic AI response, because the engine has actual context to reason with: team size, complexity tolerance, implicit budget sensitivity.

Include the specific qualifiers your real buyers care about: team size, industry, budget tier, technical sophistication, geography if relevant. These details are exactly what differentiate a generic answer from one that actually reflects how the engine handles your real market segment, and they're usually the details a quickly assembled, meeting-built prompt list leaves out entirely.

A Worked Example: Turning a Bad Prompt Into a Good One

Abstract advice about "natural language" is easy to agree with and hard to apply consistently, so here's an actual before and after.

Bad prompt, straight from a meeting: "best marketing automation software." This is a keyword phrase wearing a question mark. It has no context, no company size, no use case, and it produces a generic AI response that reads like a top-ten listicle rather than an actual recommendation.

Better prompt, pulled from an actual support ticket where a prospect described their situation: "what marketing automation tool works well for a 5-person team that doesn't have a dedicated ops person to manage it." Notice everything this version carries that the first one doesn't: team size, a specific pain point (no dedicated ops person), and an implicit requirement (ease of setup and maintenance). An AI engine answering this second prompt has to reason about simplicity and team capacity, not just list the five biggest names in the category, and the answer it gives will look completely different from the answer to the first version.

The gap between these two prompts is the entire point of this guide. The first version measures whether you show up in a generic popularity contest. The second measures whether you show up for the actual, specific situation your buyers are in, which is a far more useful and far more actionable thing to know.

What to Do When You Don't Have Sales Calls or Reviews Yet

Early-stage teams sometimes don't have a deep well of transcripts or reviews to mine, and that's worth addressing directly rather than skipping the step entirely.

If you're pre-revenue or very early, use your own discovery calls, even a handful, and pay close attention to the exact words a prospect used to describe their problem in the first two minutes, before your pitch reframed it. If you have even five or six of these, you'll usually find enough recurring phrasing to build a first-pass prompt set. Failing that, look at the questions your prospects ask in your own demo or trial onboarding flow, if you track those. And if none of that exists yet, the fallback is reading competitor reviews closely, since even a company with no reviews of its own operates in a category where competitors almost certainly do, and critical reviews in particular tend to spell out the exact comparison and decision language you're looking for.

Whatever the source, treat an early prompt set as a first draft you'll revise within a quarter once you have real usage data of your own, not as a permanent list.

Step 4: Decide How Many Prompts You Actually Need

More isn't automatically better. A list of eighty prompts sounds thorough and is usually unmanageable to track consistently, which means half of it gets checked sporadically and the resulting data is unreliable in a different way than a too-short list is.

For a small team checking manually or with light tooling, fifteen to twenty-five well-built prompts, evenly split across the four journey stages, is a realistic, sustainable set. For a team with dedicated tracking infrastructure, forty to sixty is workable, provided the list still maintains that same structural balance rather than growing through twenty near-duplicate variations of the same three questions. Quality and structure matter more than raw count at almost every scale.

Step 5: Revise on a Real Cadence, Not a Whim

A prompt set isn't a one-time deliverable. Buyer language shifts as your category evolves, as competitors reposition, and as your own product changes. Revisit the full set every quarter: re-mine your sales calls and reviews from the last three months, check whether any prompts have gone stale (referencing a product or competitor that's no longer relevant), and replace anything that's stopped reflecting how buyers actually talk.

Avoid revising it more often than that outside of a real trigger, like a major competitor entering the market or a significant product repositioning. Changing the prompt set every few weeks makes it impossible to build a clean trend line, since you'd be comparing your visibility against a moving target rather than a stable baseline.

The Honest Limitation

Here's what needs to be said plainly, because it's easy to forget once you have a clean spreadsheet that looks authoritative: no AI engine publishes its actual query logs. Nobody outside these companies knows the real, complete distribution of what people type when they're evaluating your category. Every prompt set, no matter how carefully built from sales calls and reviews, is an informed approximation, not ground truth.

That doesn't make the exercise worthless. A prompt set built from real buyer language is dramatically more accurate than one built from a meeting, and it's the best approximation available without access to data no outside company has. But treat every number that comes out of tracking it as a hypothesis about your visibility, tested against your best guess at real buyer language, not as a precise measurement of your actual market position. When you present these numbers to anyone else on your team, say this part out loud too. A tracking dashboard that looks precise to two decimal places is still built on an approximation underneath, and pretending otherwise sets up an unrealistic expectation the data was never built to support.

Putting It Into Practice This Week

If your current prompt list was built in a meeting and never revisited, that's the single highest-leverage fix available to you right now, ahead of any content work. Block two hours this week: one hour mining real language from calls, tickets, or reviews, one hour restructuring the list across the four journey stages. Everything you track from that point forward will mean something closer to what you think it means.

One last practical note: once the new set is built, don't discard the old one silently. Keep a record of what changed and why, especially if you're presenting visibility numbers to anyone else on the team. A score that appears to drop because you swapped in a more honest, harder prompt set looks like a regression if nobody remembers the methodology changed. A short note, "revised prompt set on [date], numbers before and after aren't directly comparable," saves a confusing conversation three months from now when someone pulls up the historical chart and asks what went wrong.

Start tracking your AI visibility.

Track your AI visibility daily across ChatGPT, Claude, Gemini, Perplexity, and Tavily, and see exactly what to fix next.