- Baseline: mentioned in 12% of answers (95% CI 10–14%) across a 40-prompt suite on four engines; share of voice 9%.
- Diagnosis found three blockers: thin spec pages, PDF-locked data, and near-zero third-party corroboration.
- The 90-day plan: evidence-first page rewrites, FAQ/Product schema, six third-party placements, and an llms.txt file.
- Outcome at day 90: 41% mention rate (95% CI 38–44%) and share of voice up from 9% to 23% — with the biggest gains on retrieval-heavy engines.
About this case study: this is an illustrative composite drawn from the MentionBeat playbook — a worked example of the method, not a named customer's audited results. Details are anonymized and numbers are representative of what the approach produces; treat them as a realistic model to plan against, not a promise.
The subject — call them "Nordvind Instruments" — is a European manufacturer of humidity, temperature and CO₂ measurement instruments for HVAC systems, data centers and pharmaceutical cleanrooms. Mid-market, roughly 30 years old, well regarded by the engineers who already know them, invisible to everyone else. Classic B2B: long sales cycles, technical buyers, a website last redesigned around trade-show priorities.
Their trigger moment will sound familiar. A sales engineer lost a data-center deal and asked the prospect why. The answer: "Our facilities team shortlisted three vendors with ChatGPT. You weren't in it." That sentence funded the project.
Day 0–10: measure before touching anything
The first instinct — rewrite the website immediately — is the wrong one. Without a baseline, you can never attribute improvement to the work. So the first ten days went to measurement design:
- A 40-prompt suite mirroring real buyer language, built from sales-call transcripts and support tickets. Four intent groups: category discovery ("best humidity sensors for data centers"), use-case fit ("CO₂ monitoring for pharma cleanrooms GMP"), comparisons ("Nordvind vs [competitor] transmitters"), and problem-first prompts ("condensation false readings in cold aisle containment").
- Four engines: ChatGPT, Claude, Gemini and Perplexity, hit via API and UI sampling.
- 25 runs per prompt per engine spread across two weeks — 4,000 sampled answers — because single runs of stochastic systems are anecdotes, not data.
- Three metrics: mention rate, share of voice against seven tracked competitors, and factual accuracy of any claims made about Nordvind.
The baseline came back: mentioned in 12% of answers (95% CI 10–14%), share of voice 9%, versus 34% for the category leader. Worse: in the answers that did mention them, two recurring factual errors — an obsolete product name from a 2021 rebrand, and a humidity accuracy spec from a discontinued line.
Day 10–20: the diagnosis
With the baseline in hand, the team audited why engines had so little to work with. Three blockers explained most of it:
- Thin spec pages. Product pages averaged 140 words of visible text: a hero image, three adjectives ("precise, reliable, robust") and a "Download datasheet" button. Nothing liftable, nothing evidence-bearing.
- PDF-locked data. Every number that mattered — accuracy classes, drift specs, calibration intervals, operating ranges — lived exclusively in PDF datasheets. The company's best content was effectively invisible to retrieval pipelines chunking HTML.
- Zero third-party corroboration. Outside their own domain, Nordvind barely existed in text: no comparison-site presence, no forum footprint, one stale Wikipedia mention. Engines synthesizing "best of" answers had no independent source that had ever put Nordvind on a list. The GEO research says engines reward citations and corroborated claims1 — Nordvind offered neither.
Day 20–30: the 90-day plan
The plan matched one workstream to each blocker, plus technical hygiene:
| Workstream | Scope | Blocker addressed |
|---|---|---|
| Page rewrites | 12 money pages rebuilt on the evidence-first pattern: 40-word liftable opener, HTML spec tables, number + source + method claims, honest comparison sections | Thin pages |
| Data liberation | Every datasheet spec republished as on-page HTML tables; PDFs kept as supplements | PDF lock-in |
| Corroboration | Six third-party placements: two trade-publication technical articles, two industry-directory listings with full spec data, one HVAC engineering forum AMA, one comparison-site profile | Zero corroboration |
| Technical | Product + FAQPage schema on all rewritten pages,2 llms.txt manifest,3 robots.txt audit confirming GPTBot, ClaudeBot and friends weren't blocked4 | Hygiene |
Deliberately excluded: a blog-volume push ("ten posts a month" was proposed and cut), anything aimed at engines' training runs (too slow for a 90-day window), and any tactic that couldn't plausibly move the 40-prompt suite. Focus is the discipline here — the suite defines what "winning" means, so the work is whatever moves the suite.
Resourcing, since everyone asks: the whole program ran on roughly 1.5 people. A product marketer owned the rewrites at two to three pages a week (each one needing an hour with a product engineer to source real numbers for the evidence pattern), the web developer spent about a week total on schema, llms.txt and freeing the datasheets, and an outside writer handled the two trade-publication articles. No agency, no new headcount — the scarce input wasn't budget, it was engineering time to verify claims. That's typical: evidence-first content is bottlenecked on facts, not words.
Day 30–75: execution notes
Three details from the execution phase that mattered more than expected:
The openers did heavy lifting. Each rewritten page leads with a self-contained answer paragraph — e.g. "The HMT-340 is a duct-mount humidity transmitter for data-center cold-aisle monitoring. It measures 0–100% RH with ±1% accuracy, drifts less than 0.5% per year, and ships with an ISO 17025-traceable calibration certificate." In later sampling, engines quoted these paragraphs nearly verbatim more often than any other passage on the site.
Comparison honesty was the hardest sell internally. The rewritten comparison pages name competitors and concede specific wins ("choose [X] if you need wireless mesh; choose the HMT line for calibration traceability"). Legal and sales pushed back for two weeks. The pages went live anyway — and in the day-90 sample, comparison-type prompts showed the largest single improvement, because Nordvind's page was frequently the most balanced source retrieved.
Placements were chosen for retrievability, not prestige. The six third-party placements weren't press releases; they were technical content on domains that already ranked for category queries — the pages engines actually pull when synthesizing "best humidity transmitter" answers. Two were bylined engineering articles on trade publications, with spec tables included, corroborating the same numbers the site now published.
The accuracy fixes traveled furthest. Republishing current product names and specs in crawlable HTML didn't just lift mentions — it corrected the two recurring factual errors. By day 90, the obsolete product name had disappeared from sampled browsing-enabled answers entirely, though it still surfaced occasionally in offline (non-browsing) responses, which only a future training run will fix.
The midpoint check earned its keep. At day 60, a reduced sampling run (10 runs per cell instead of 25) showed mention rate at roughly 26% — clearly up, but with two prompt groups barely moving: the problem-first prompts ("condensation false readings…") and everything on Gemini. The diagnosis took an afternoon: the troubleshooting content answering problem-first queries was still buried in a support portal behind a session wall. The team pulled the six most-asked troubleshooting guides into public, indexable pages in week nine — a fix that wouldn't have happened until after the project ended if measurement had only run at the bookends.
Day 76–90: the remeasure
Same 40 prompts, same four engines, same 25-runs-per-cell protocol — the comparison is only valid because the instrument didn't change. Headline numbers:
By engine, before and after — the pattern is the story:
Two readings of that chart matter. First, the gains concentrate where retrieval dominates: Perplexity and search-enabled ChatGPT answers responded fastest and hardest, because the project shipped exactly what their pipelines consume — crawlable, quotable, corroborated pages. Second, the smallest gain (Claude, in configurations leaning on parametric knowledge) is the honest reminder that 90 days cannot rewrite a model's memory. That's next year's training runs, fed by the corroboration that's now accumulating.
And to be equally honest about uncertainty: with these sample sizes the confidence intervals are wide enough that the per-engine ordering could shuffle on a re-run. What's not in doubt is the aggregate movement — 12% (CI 10–14) to 41% (CI 38–44) doesn't overlap by any reading.
One number deliberately left off the headline: recommendation rate — how often answers didn't just name Nordvind but positively suggested it — moved from 5% to 19%. It lagged mention rate the whole quarter, which is the expected shape: engines start naming you as soon as your pages become retrievable, but they start recommending you when corroborating sources agree you belong on the shortlist. If the two rates ever converge from above — mentioned everywhere, recommended nowhere — that's a positioning problem, not a visibility problem, and no amount of publishing fixes it.
Everything here began with one measurement: how often AI assistants mention you today. MentionBeat runs that baseline across ChatGPT, Claude, Gemini and Perplexity — prompts, sampling and confidence intervals included.
Get a free visibility reportWhat generalizes (and what doesn't)
The transferable lessons from this composite:
- Measure first, always. The baseline cost ten days and made every later decision defensible — including killing the blog-volume idea.
- Free your data from PDFs. For spec-driven B2B, this is routinely the single highest-leverage fix. Engines can't quote what they can't parse.
- Corroboration multiplies content. The page rewrites alone drove the day-60 midpoint (26%); the jump to 41% came after third-party placements repeated the same claims elsewhere.
- Honest comparisons win comparative prompts. The pages sales feared most performed best.
- Accuracy is a metric, not a hope. Tracking what answers said surfaced two errors that mention-rate tracking alone would have missed.
What doesn't generalize: the magnitude. Nordvind operated in a niche with thin source material, where a dozen good pages and six placements can visibly shift what engines retrieve. In a crowded consumer category, the same 90 days of work faces thousands of competing sources — expect a slower grind and smaller steps. Which is exactly why you run your own baseline instead of borrowing this one.
Frequently asked questions
It's representative of what's achievable in a low-competition B2B niche with severe, fixable blockers — thin pages, PDF-locked specs, zero corroboration. Brands starting from a healthier baseline, or competing in crowded categories, should expect smaller and slower movement. The method transfers; the magnitude depends on your starting point.
LLM answers are stochastic — the same prompt names different brands run to run. At 25 runs per prompt per engine, a 40-prompt suite yields 1,000 samples per engine, tightening the confidence interval on a mention rate to a few points. Fewer runs means wider intervals, and a "gain" that might just be noise.
In this composite, the sequencing suggests page rewrites plus data liberation produced the first jump (baseline to day-60 midpoint) and corroboration produced the second. But the honest answer is that they compound: corroboration works by repeating claims the pages now make quotably. Running one without the other buys you less than half the result.
Sources & further reading
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., Deshpande, A. — "GEO: Generative Engine Optimization", KDD 2024 / arXiv:2311.09735.
- Schema.org — Schema.org vocabulary (Product, FAQPage types).
- Answer.AI — "The /llms.txt file" proposal.
- OpenAI — "Overview of OpenAI crawlers"; Anthropic — ClaudeBot documentation.