- Build prompts from jobs-to-be-done, not keywords — buyers ask assistants full questions with context, constraints and use cases, not two-word queries.
- Cover the funnel deliberately: category discovery, comparison, validation and objection prompts measure different things and move at different speeds.
- Never seed your own brand into unbranded prompts — "is X the best…?" inflates mention rate by construction and measures nothing.
- A workable core suite is 20–50 prompts with paraphrase variants, frozen and versioned like code — because every edit breaks comparability with your history.
Every AI visibility number is downstream of one artifact: the list of prompts you test. Sample size, confidence intervals, engine coverage — all of it quantifies uncertainty within the suite. None of it fixes a suite that asks the wrong questions.
And most suites do. They're written in ten minutes by transplanting the SEO keyword list ("best CRM software", "CRM pricing") into a chat window, or — worse — by typing the questions the brand team wishes buyers asked. The resulting metrics are precise measurements of an irrelevant quantity.
This is the methodology we use to build suites that measure something real: how the questions get written, how coverage gets structured, and how the suite survives contact with time.
The suite is the measurement
A prompt suite plays the role a question wording plays in survey research: it defines the population of situations you're estimating. "How visible are we in AI answers?" is unanswerable in general — visible in response to what? A brand can have a 60% mention rate on comparison questions and 5% on category-discovery questions. Both numbers are true; which one you report depends entirely on suite composition.
That cuts two ways. It means a sloppy suite silently biases every metric downstream. It also means a well-designed suite is a strategic document: it encodes your actual bet about which buyer conversations matter. Treat it with the seriousness of an annual plan, not a brainstorm.
The stakes justify the care. With Gartner projecting a ~25% drop in traditional search volume by 2026 as buyers shift to assistants,2 and users clicking through far less when AI summaries appear,3 these synthesized answers increasingly are the shortlist. The suite is your window into how that shortlist gets built.
Write prompts the way buyers actually ask
People don't talk to assistants the way they typed into Google. Search trained us to compress ("crm small business"); chat invites us to describe. Real assistant queries are longer, contextual, and framed as jobs-to-be-done: a situation, a constraint, a desired outcome.
Compare:
- Keyword thinking: "best project management software"
- How a buyer asks: "I run a 12-person design agency and we're drowning in client feedback spread across email and Slack. What tool would help us manage projects without a big setup effort?"
The second prompt retrieves different sources, triggers different reasoning, and produces different shortlists — often dramatically so. If your suite is built from the first style, you're benchmarking a conversation your buyers aren't having.
Where to mine real phrasings: sales-call transcripts and discovery-call notes ("what were you looking for when you found us?"), support and onboarding questions, community threads where your category gets discussed, and win/loss interviews. You want the awkward, specific, constraint-laden sentences real humans produce — budget caps, team sizes, integration requirements, migration fears. Ten prompts lifted from actual buyer conversations beat fifty invented in a conference room.
Cover the funnel, deliberately
Buyers ask assistants different kinds of questions at different stages, and each kind stresses a different part of your visibility. A balanced suite covers four layers:
- Category discovery — the buyer doesn't know the players. "What tools exist for X?" This is where mention rate is hardest to earn and most valuable: you're competing to be on the shortlist at all.
- Comparison — the buyer has names and wants a verdict. "How does A compare to B for a team like mine?" Here recommendation rate and the accuracy of what's said about you dominate.
- Validation — the buyer is nearly decided and stress-testing. "Is A reliable? What do users complain about?" Sentiment and accuracy are the metrics; a single recurring criticism surfacing here can quietly kill deals.
- Objection — the buyer probes dealbreakers. "Does A work with Salesforce? Is it GDPR-compliant?" Binary factual visibility: does the engine know, and is it right?
| Funnel stage | Example prompt | Metric to watch |
|---|---|---|
| Discovery | "What are the best options for automating invoice processing at a mid-size company?" | Mention rate, share of voice |
| Comparison | "Compare the leading invoice automation tools for a 200-person company on price and ERP integrations." | Recommendation rate, accuracy |
| Validation | "What are common complaints about [shortlisted tool]? Is support responsive?" | Sentiment, accuracy |
| Objection | "Does [shortlisted tool] support NetSuite and EU data residency?" | Accuracy |
Paraphrase clusters: one intent, several phrasings
Engines are sensitive to wording in ways buyers aren't. "Best PM tool for agencies", "what project software do creative agencies use" and "recommend something to manage client projects at a design studio" express one intent — and can yield noticeably different brand lists. If you test only one phrasing, you've measured that phrasing, not the intent.
The fix is to structure the suite as intent clusters: each core intent carries 3–5 paraphrase variants, and you report metrics at the cluster level (the average across variants). This does two useful things. It stops any single lucky phrasing from dominating the metric, and it lets you see phrasing sensitivity itself — a cluster where variants disagree wildly is a cluster where your visibility is fragile, resting on a few retrieved pages rather than broad corroboration. The original GEO research made a related design choice, evaluating tactics across ~10,000 diverse queries precisely because effects vary so much query to query.1
Vary systematically, not randomly: formality ("optimal" vs. "decent"), specificity (with and without team size, budget, industry), and framing (question vs. described situation). Keep the underlying job identical within a cluster.
The brand-leading trap
The single most common way teams corrupt their own benchmark: putting the brand in the prompt. "Is Acme the best invoicing tool?" all but forces the answer to discuss Acme — the engine is completing your framing, not consulting its model of the category. Congratulations on your 98% mention rate; it measures the prompt, not the market.
Brand-leading is often subtler than naming yourself. Prompts that embed your tagline, your exact positioning language, or a feature combination only you offer ("tools with offline-first sync and per-seat pricing under $10") are rigged in the same way — they describe you, then ask the engine to name you.
- "What's a good invoicing tool for a freelance studio that bills in multiple currencies?" — unbranded, real constraints
- "Compare Acme and Billify for a 5-person agency" — branded, but symmetric: fair comparison framing
- "What do users dislike about Acme?" — branded on purpose, in the validation layer, to audit sentiment
- Keep branded and unbranded prompts in separate layers with separate metrics
- "Is Acme the best invoicing tool?" — forces the mention, inflates the metric by construction
- "Why is Acme better than Billify?" — presumes the conclusion; you'll measure sycophancy
- "Best tool with [your unique feature combo]" — a description of yourself wearing a fake mustache
- Mixing branded prompts into the unbranded average to make the topline look healthier
Branded prompts absolutely belong in the suite — in the validation and objection layers, where buyers genuinely name brands. The rule isn't "never mention yourself"; it's "never let branded prompts masquerade as discovery, and never blend the two layers into one number".
MentionBeat generates an unbranded, funnel-covering prompt suite for your category, runs it repeatedly across ChatGPT, Claude, Gemini and Perplexity, and reports mention rate and share of voice with confidence intervals — the methodology in this post, automated.
Get a free visibility reportLocale and persona: the dimensions teams forget
Assistants tailor answers to context, and buyers supply context whether you test for it or not.
Locale. "Best payroll software" gets a different shortlist asked in German, or asked in English with "for a UK company" appended. If you sell in five markets, your visibility is five different numbers — often wildly different, because your third-party corroboration (reviews, forums, press) is unevenly distributed across languages. Test at least your top two or three markets in local language and local framing.
Persona. "As a CTO evaluating…" versus "as a small-business owner with no IT staff…" shifts both retrieval and tone. If your positioning targets a persona, include prompts voiced from it — and from adjacent personas you don't target, so you can see whether engines route the wrong buyers toward you (an accuracy problem masquerading as a visibility win).
Add dimensions frugally: every locale × persona combination multiplies run costs. Pick the two or three cells that carry real revenue and instrument those properly, rather than spreading samples thin across a grid of hypotheticals.
Size, freezing and versioning
How big should the suite be? Big enough to cover intents, small enough to sample deeply. In practice: 20–50 core intents, each with 3–5 paraphrase variants, run repeatedly per engine. Under 20, single-intent quirks dominate the topline; past 50, most teams end up unable to afford the repeated runs that make any of it statistically meaningful — and repeated runs matter more than marginal intents.
Then treat the suite like code:
- Freeze it. Between measurement waves, nothing changes. Comparability is the entire value of a tracking benchmark.
- Version it. When change is warranted — a new product line, a shifted ICP, a phrasing you learned from sales calls — ship suite v2 explicitly. Run v1 and v2 in parallel for one wave to bridge the histories, then retire v1.
- Log the context. Record engine and model versions alongside each wave. When numbers jump, the first question is always "did we change, or did the engine?" — and it should be answerable from the log, not from memory.
- Review on a calendar, not on impulse. Quarterly is right for most categories. The temptation to tweak prompts mid-wave when numbers disappoint is exactly how benchmarks die.
A governance tip: the person who owns the visibility KPI should not be the only author of the suite. Have sales or research contribute the buyer phrasings and someone outside the KPI's reporting line sign off on changes — the same separation you'd want between a test and its grader.
The pre-flight checklist
- Every prompt traceable to a real buyer phrasing or job-to-be-done
- All four funnel layers covered, discovery weighted heaviest
- Each core intent carries 3–5 paraphrase variants, reported at cluster level
- Zero brand names or self-describing feature combos in discovery prompts
- Branded validation/objection prompts kept in a separate layer with separate metrics
- Top locales and personas represented — and no more than you can sample deeply
- Suite frozen, versioned, and engine/model versions logged per wave
- Change control: scheduled reviews, sign-off from outside the KPI owner
Frequently asked questions
Yes — identical prompts across ChatGPT, Claude, Gemini and Perplexity, so differences in your metrics reflect the engines, not the questions. The temptation to tailor prompts per engine ("Perplexity users write shorter queries") reintroduces a confound. If you believe usage styles differ by platform, encode that as paraphrase variants run everywhere.
Audit inbound signals quarterly: new objections in sales calls, new comparison targets in win/loss notes, new question patterns in support. If a theme shows up repeatedly in real conversations but has no prompt cluster, that's a coverage gap — queue it for the next suite version rather than patching mid-wave.
As a drafting step, absolutely — LLMs are excellent at generating paraphrase variants once you've defined the intents. As the source of the intents themselves, it's risky: you'll get plausible generic questions, not your buyers' actual constraints and vocabulary. Ground the intents in real conversations; automate the variation.
Sources & further reading
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., Deshpande, A. — "GEO: Generative Engine Optimization", KDD 2024 / arXiv:2311.09735.
- Gartner — "Gartner Predicts Search Engine Volume Will Drop 25% by 2026, Due to AI Chatbots and Other Virtual Agents", February 2024.
- Pew Research Center — "Google users are less likely to click on links when an AI summary appears in the results", July 2025.