Blog/Measurement
Measurement

Designing a prompt suite: how to benchmark your brand across AI assistants

Sampling rigor gets the attention, but the prompts you sample decide what your metrics actually mean. Garbage prompts in, garbage mention rate out — no confidence interval can save a suite built on questions no buyer would ever ask.

ML Maya Lindqvist · Head of Research May 5, 2026 11 min read
Key takeaways
  • Build prompts from jobs-to-be-done, not keywords — buyers ask assistants full questions with context, constraints and use cases, not two-word queries.
  • Cover the funnel deliberately: category discovery, comparison, validation and objection prompts measure different things and move at different speeds.
  • Never seed your own brand into unbranded prompts — "is X the best…?" inflates mention rate by construction and measures nothing.
  • A workable core suite is 20–50 prompts with paraphrase variants, frozen and versioned like code — because every edit breaks comparability with your history.

Every AI visibility number is downstream of one artifact: the list of prompts you test. Sample size, confidence intervals, engine coverage — all of it quantifies uncertainty within the suite. None of it fixes a suite that asks the wrong questions.

And most suites do. They're written in ten minutes by transplanting the SEO keyword list ("best CRM software", "CRM pricing") into a chat window, or — worse — by typing the questions the brand team wishes buyers asked. The resulting metrics are precise measurements of an irrelevant quantity.

This is the methodology we use to build suites that measure something real: how the questions get written, how coverage gets structured, and how the suite survives contact with time.

The suite is the measurement

A prompt suite plays the role a question wording plays in survey research: it defines the population of situations you're estimating. "How visible are we in AI answers?" is unanswerable in general — visible in response to what? A brand can have a 60% mention rate on comparison questions and 5% on category-discovery questions. Both numbers are true; which one you report depends entirely on suite composition.

That cuts two ways. It means a sloppy suite silently biases every metric downstream. It also means a well-designed suite is a strategic document: it encodes your actual bet about which buyer conversations matter. Treat it with the seriousness of an annual plan, not a brainstorm.

The stakes justify the care. With Gartner projecting a ~25% drop in traditional search volume by 2026 as buyers shift to assistants,2 and users clicking through far less when AI summaries appear,3 these synthesized answers increasingly are the shortlist. The suite is your window into how that shortlist gets built.

Write prompts the way buyers actually ask

People don't talk to assistants the way they typed into Google. Search trained us to compress ("crm small business"); chat invites us to describe. Real assistant queries are longer, contextual, and framed as jobs-to-be-done: a situation, a constraint, a desired outcome.

Compare:

The second prompt retrieves different sources, triggers different reasoning, and produces different shortlists — often dramatically so. If your suite is built from the first style, you're benchmarking a conversation your buyers aren't having.

Where to mine real phrasings: sales-call transcripts and discovery-call notes ("what were you looking for when you found us?"), support and onboarding questions, community threads where your category gets discussed, and win/loss interviews. You want the awkward, specific, constraint-laden sentences real humans produce — budget caps, team sizes, integration requirements, migration fears. Ten prompts lifted from actual buyer conversations beat fifty invented in a conference room.

Cover the funnel, deliberately

Buyers ask assistants different kinds of questions at different stages, and each kind stresses a different part of your visibility. A balanced suite covers four layers:

  1. Category discovery — the buyer doesn't know the players. "What tools exist for X?" This is where mention rate is hardest to earn and most valuable: you're competing to be on the shortlist at all.
  2. Comparison — the buyer has names and wants a verdict. "How does A compare to B for a team like mine?" Here recommendation rate and the accuracy of what's said about you dominate.
  3. Validation — the buyer is nearly decided and stress-testing. "Is A reliable? What do users complain about?" Sentiment and accuracy are the metrics; a single recurring criticism surfacing here can quietly kill deals.
  4. Objection — the buyer probes dealbreakers. "Does A work with Salesforce? Is it GDPR-compliant?" Binary factual visibility: does the engine know, and is it right?
Discovery · 16 Comparison · 12 Validation · 8 Obj. · 4 A typical 40-prompt core suite weights discovery heaviest — it's where shortlists form and where new visibility is won.
Illustrative allocation of a 40-prompt core suite across funnel stages. Weight discovery heaviest unless your problem is demonstrably late-funnel.
Funnel stageExample promptMetric to watch
Discovery"What are the best options for automating invoice processing at a mid-size company?"Mention rate, share of voice
Comparison"Compare the leading invoice automation tools for a 200-person company on price and ERP integrations."Recommendation rate, accuracy
Validation"What are common complaints about [shortlisted tool]? Is support responsive?"Sentiment, accuracy
Objection"Does [shortlisted tool] support NetSuite and EU data residency?"Accuracy

Paraphrase clusters: one intent, several phrasings

Engines are sensitive to wording in ways buyers aren't. "Best PM tool for agencies", "what project software do creative agencies use" and "recommend something to manage client projects at a design studio" express one intent — and can yield noticeably different brand lists. If you test only one phrasing, you've measured that phrasing, not the intent.

The fix is to structure the suite as intent clusters: each core intent carries 3–5 paraphrase variants, and you report metrics at the cluster level (the average across variants). This does two useful things. It stops any single lucky phrasing from dominating the metric, and it lets you see phrasing sensitivity itself — a cluster where variants disagree wildly is a cluster where your visibility is fragile, resting on a few retrieved pages rather than broad corroboration. The original GEO research made a related design choice, evaluating tactics across ~10,000 diverse queries precisely because effects vary so much query to query.1

Vary systematically, not randomly: formality ("optimal" vs. "decent"), specificity (with and without team size, budget, industry), and framing (question vs. described situation). Keep the underlying job identical within a cluster.

The brand-leading trap

The single most common way teams corrupt their own benchmark: putting the brand in the prompt. "Is Acme the best invoicing tool?" all but forces the answer to discuss Acme — the engine is completing your framing, not consulting its model of the category. Congratulations on your 98% mention rate; it measures the prompt, not the market.

Brand-leading is often subtler than naming yourself. Prompts that embed your tagline, your exact positioning language, or a feature combination only you offer ("tools with offline-first sync and per-seat pricing under $10") are rigged in the same way — they describe you, then ask the engine to name you.

✓ Do
  • "What's a good invoicing tool for a freelance studio that bills in multiple currencies?" — unbranded, real constraints
  • "Compare Acme and Billify for a 5-person agency" — branded, but symmetric: fair comparison framing
  • "What do users dislike about Acme?" — branded on purpose, in the validation layer, to audit sentiment
  • Keep branded and unbranded prompts in separate layers with separate metrics
✕ Don't
  • "Is Acme the best invoicing tool?" — forces the mention, inflates the metric by construction
  • "Why is Acme better than Billify?" — presumes the conclusion; you'll measure sycophancy
  • "Best tool with [your unique feature combo]" — a description of yourself wearing a fake mustache
  • Mixing branded prompts into the unbranded average to make the topline look healthier

Branded prompts absolutely belong in the suite — in the validation and objection layers, where buyers genuinely name brands. The rule isn't "never mention yourself"; it's "never let branded prompts masquerade as discovery, and never blend the two layers into one number".

See your suite in action
A clean benchmark, without building the machinery

MentionBeat generates an unbranded, funnel-covering prompt suite for your category, runs it repeatedly across ChatGPT, Claude, Gemini and Perplexity, and reports mention rate and share of voice with confidence intervals — the methodology in this post, automated.

Get a free visibility report

Locale and persona: the dimensions teams forget

Assistants tailor answers to context, and buyers supply context whether you test for it or not.

Locale. "Best payroll software" gets a different shortlist asked in German, or asked in English with "for a UK company" appended. If you sell in five markets, your visibility is five different numbers — often wildly different, because your third-party corroboration (reviews, forums, press) is unevenly distributed across languages. Test at least your top two or three markets in local language and local framing.

Persona. "As a CTO evaluating…" versus "as a small-business owner with no IT staff…" shifts both retrieval and tone. If your positioning targets a persona, include prompts voiced from it — and from adjacent personas you don't target, so you can see whether engines route the wrong buyers toward you (an accuracy problem masquerading as a visibility win).

Add dimensions frugally: every locale × persona combination multiplies run costs. Pick the two or three cells that carry real revenue and instrument those properly, rather than spreading samples thin across a grid of hypotheticals.

Size, freezing and versioning

How big should the suite be? Big enough to cover intents, small enough to sample deeply. In practice: 20–50 core intents, each with 3–5 paraphrase variants, run repeatedly per engine. Under 20, single-intent quirks dominate the topline; past 50, most teams end up unable to afford the repeated runs that make any of it statistically meaningful — and repeated runs matter more than marginal intents.

Then treat the suite like code:

🧭

A governance tip: the person who owns the visibility KPI should not be the only author of the suite. Have sales or research contribute the buyer phrasings and someone outside the KPI's reporting line sign off on changes — the same separation you'd want between a test and its grader.

The pre-flight checklist

Frequently asked questions

Yes — identical prompts across ChatGPT, Claude, Gemini and Perplexity, so differences in your metrics reflect the engines, not the questions. The temptation to tailor prompts per engine ("Perplexity users write shorter queries") reintroduces a confound. If you believe usage styles differ by platform, encode that as paraphrase variants run everywhere.

Audit inbound signals quarterly: new objections in sales calls, new comparison targets in win/loss notes, new question patterns in support. If a theme shows up repeatedly in real conversations but has no prompt cluster, that's a coverage gap — queue it for the next suite version rather than patching mid-wave.

As a drafting step, absolutely — LLMs are excellent at generating paraphrase variants once you've defined the intents. As the source of the intents themselves, it's risky: you'll get plausible generic questions, not your buyers' actual constraints and vocabulary. Ground the intents in real conversations; automate the variation.

Sources & further reading

  1. Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., Deshpande, A. — "GEO: Generative Engine Optimization", KDD 2024 / arXiv:2311.09735.
  2. Gartner — "Gartner Predicts Search Engine Volume Will Drop 25% by 2026, Due to AI Chatbots and Other Virtual Agents", February 2024.
  3. Pew Research Center — "Google users are less likely to click on links when an AI summary appears in the results", July 2025.
Share
ML
Maya Lindqvist

Head of Research at MentionBeat. Maya leads the measurement methodology behind MentionBeat's visibility metrics — prompt-suite design, sampling, and confidence intervals — and writes about how generative engines choose what to say.

Know where you stand in AI answers

MentionBeat samples real buyer prompts across ChatGPT, Claude, Gemini and Perplexity — and turns them into metrics you can act on.

Get your free visibility report
No credit card. Results in about a minute.