Generative Engine Optimization · Answer Engine Optimization

Case studies of successful LLM-visibility improvement

A curated, credibility-rated evidence base of organizations that measurably improved how often they are cited, mentioned, and recommended in ChatGPT, Perplexity, Gemini, Claude & Google AI Overviews — filtered for what actually transfers to a niche B2B instrument maker. Companion to the LLM Visibility Master Playbook.

Citation / mention rate AI referral traffic & leads Share of voice B2B & technical focus Credibility-rated
Version 1.0
Compiled June 2026
Owner Digital / Product Marketing
Pairs with Vaisala-LLM-Visibility-Playbook.html
Read this first — most "case studies" in this space are marketing

GEO is a young field. The overwhelming majority of published case studies are self-reported by the agency or tool vendor that sold the service, with no independent audit, no control group, and survivorship bias (nobody publishes the campaign that flopped). They are still useful — but as evidence of what teams tried, not proof of effect size. To keep this honest, every case below carries a credibility tier.

Credibility tiers used in this document

Tier 1 · Rigorous Controlled / peer-reviewed / independent dataset. Trust the effect size.
Tier 2 · Independent-ish Third-party analytics or analyst data, not seller-funded. Directionally reliable.
Tier 3 · Vendor-reported Self-published by the agency/tool. Take the method, distrust the numbers.

Outcome metrics in focus (per your brief): citation / mention rate, AI referral traffic & leads, and share of voice vs. competitors.

1The one controlled study — start every argument here

Everything else is anecdote; this is the closest thing to evidence. If you cite one source internally, cite this one.

Academic · Princeton · Georgia Tech · Allen Institute for AI · IIT Delhi

“GEO: Generative Engine Optimization” (Aggarwal et al., KDD ’24)

Tier 1 · Rigorous

A controlled study (randomized + quasi-experimental designs) over GEO-Bench — ~10,000 diverse queries across domains, measuring which on-page content changes increase the probability of being cited in an AI-generated answer. It is the paper that coined the term “GEO.”

up to 40%
visibility lift from combined GEO methods
+41%
from adding statistics
~+30%
from citing sources
+28%
from quotations / fluency

What moved the needle (and what didn't)

Tactic testedEffect on AI visibilityRead-across
Add cited statistics (concrete numbers)strongest (~+41%)Specs, field-hours, standards numbers in every section
Cite authoritative sources~+30%; up to ~+115% for lower-ranked pagesThe underdog lever — biggest for niche/low-authority topics
Add direct quotations from authorities~+28%Quote IEC / WMO / peer-reviewed sources verbatim
Improve fluency / clarity~+15–28%Plain, answer-first prose
Keyword stuffingno gain (can hurt)Don't. Write for intent.

Across methods, the paper also reports organic-style visibility gains in the ~12–26% range, with efficacy varying by domain — i.e. there is no universal recipe; technical domains need technical evidence.

Maps to your playbook: validates Pyramid Layer 3 (extractability) and Layer 4 (corroboration), plus the house-style rules “numbers with units inline” and “define before you describe.” The “+115% for lower-ranked pages” finding is the single best argument for why a niche instrument like WindCube has more to gain from GEO than a famous brand does.

Sources: arXiv 2311.09735 · ACM SIGKDD KDD’24 · Search Engine Land summary

2Closest analogs — B2B, industrial & technical

The most transferable cases: technical buyers, long sales cycles, authority-driven verticals. All are vendor-reported, so weigh the method over the headline number.

Industrial chemicals · technical B2B

Chemours — authority-led citation in a highly technical vertical

Tier 3 · Vendor-reported
82–84%
AI citation rate across target query set
$90M+
pipeline attributed to AI-assisted discovery

What they did: deep E-E-A-T — content authored by named, recognized industry experts; technical papers that cite patents and peer-reviewed research; and an authoritative backlink profile from industry-specific publications. This is the closest published analog to a scientific-instrument manufacturer.

Why it matters for Vaisala: it shows the winning posture in a technical vertical is primary technical evidence and named expertise, not marketing copy — exactly your playbook's Layer 4 corroboration bet. (Treat the $90M as illustrative; the method is the takeaway.)

Source: DigitalAgencyNetwork — GEO case-study roundup.

B2B financing platform ($400M+ backed) · anonymized

The most fully documented before / after

Tier 3 · Vendor-reported
MetricBeforeAfterChangeWindow
AI referral sessions4386~+100%Apr–Sep 2025
AI Overview keywords75311~+315%same
Page-1 organic keywords2975~+159%3–4 mo
Qualified leads+300%multi-month

What they did (maps almost 1:1 to your checklist): answer-first TL;DR blocks, FAQ schema, llms.txt, Q&A-structured content, named expert author bios, consistent monthly publishing, plus earned authority (TechCrunch / G2 / Crunchbase) and a Wikipedia page.

Honest read

Baselines are tiny — “+100%” is 43→86 sessions. Dramatic percentages, small absolute volume. This is typical of early-stage GEO results and exactly why §6 matters.

Source: Concurate — B2B financing GEO case study.

B2B SaaS · GEO agency

“CITABLE” 90-day program

Tier 3
24%
citation rate in 90 days (from ~0)
47
AI-referred leads, ~2.8× organic conversion

Useful as a 90-day program template. Numbers are agency-reported.

Source: Discovered Labs — B2B SaaS, 3× citations in 90 days.

Technical publisher

IEEE Spectrum

Tier 2

Deep, long-form technical content drove a surge in ChatGPT referrals (94% of its AI mix), with referrals still climbing month-over-month in 2025. Closest content-strategy analog to a technical authority brand.

Read-across: long, genuinely useful technical explainers — your hub-and-spoke “how it works” pages — are themselves citation magnets.

Source: Digiday · RebelMouse.

3Cross-sector wins with transferable lessons

Different industries, but the levers transfer. The recurring pattern — definitional/answer-first content, entity clarity, structured data, third-party authority — is the same spine as your playbook.

Brand / sectorReported resultPrimary leverTier
HubSpot
SaaS
Cited in AI Overviews for 3,000+ marketing queriesYears of definitional, entity-first “What is X” contentT2
via Semrush
Go Fish Digital
Agency (self-test)
+43% AI referral traffic · +83% AI conversionsPrompt-mapping → 5–8 cornerstone assets built for fact-density + external authorityT3
Go Fish Digital
Auto-insurance brand
Insurance
+447% AI Overview mentions (6 mo)Structured content, entity clarity, quotable insightsT3
DigitalAgencyNetwork
LS Building Products
Building materials
+540% AI Overview mentions · +67% organicRebuilt content architecture around AI-friendly structureT3
DigitalAgencyNetwork
Farringdons
Creative / web
+140% AI traffic · +62% AI mentionsLLM-optimized content + entity reinforcementT3
DigitalAgencyNetwork
The signal across all of them

Strip out the vendor numbers and the same four levers remain in every winning case: (1) answer-first / definitional content, (2) entity clarity, (3) structured data (FAQ / schema), (4) third-party authority. That convergence — across rigorous and vendor sources alike — is the part worth trusting.

4The business case — why visibility is worth the work

Aggregate, mostly third-party data on AI referral traffic and conversion quality. Use these to justify the program; treat precise figures as directional.

Conversion quality T1/T2

AI-referred visitors convert far above organic. A peer-reviewed Marketing Science study plus analytics vendors put ChatGPT at ~16%, Perplexity ~10–11%, vs Google organic ~1.8%; AI-chatbot arrivals were ~38% more likely to purchase in retail.

Sources: Marketing Science (INFORMS) · ALM Corp · Digiday.

Traffic growth T2

Outbound ChatGPT referrals to the web grew ~206% in 2025; ChatGPT referrals up ~52% YoY (Sep–Nov 2025). AI traffic overall rose ~7× from early-2024 to mid-2025.

Sources: Digiday · TechCrunch · Superlines.

Where the traffic is T2

In B2B referrals (early 2026): ChatGPT ~63%, Claude ~18%, Gemini ~11%, Perplexity ~7%. ChatGPT still dominates volume, but optimizing for it alone now covers a third less of the landscape than a year ago — multi-engine matters.

Sources: Goodie (117K+ B2B leads) · Goodie 2026 report.

Share of voice is platform-specific — measure per engine

The same brand and query set produces wildly different share-of-voice by engine (one documented set: Perplexity 28–38%, Gemini 12–20%, ChatGPT 10–16%, Claude 3–7%). This is because engines cite different sources — Perplexity leans heavily on Reddit; ChatGPT favors earned media and Wikipedia.

Maps to your playbook: confirms §8 — run a versioned prompt suite across all major engines and report SOV per engine, never as a single blended number.

5Findings & learnings — what the evidence establishes

The synthesis. Each finding is a claim the research supports; each row shows the evidence behind it (with its tier), a confidence rating, and the learning — what to actually do. Confidence reflects the strength of the underlying evidence, not how often the claim is repeated.

High backed by the controlled / peer-reviewed study, or by converging evidence across tiers Medium independent analytics, or a clear pattern across several cases Tentative rests mainly on vendor-reported cases — direction only
FindingEvidenceConf.Learning — what to do
F1 · Structure and evidence-density cause citations — keywords don't. Princeton controlled study: +41% from statistics, ~+30% from citing sources, +28% from quotations; keyword stuffing gave no gain. T1 High Put concrete numbers, specs and standards in every section; quote authorities verbatim; never optimise for keywords.
F2 · The lower your current authority, the more GEO helps. Same study: citing sources lifted lower-ranked pages by up to ~115% — far above already-authoritative pages. T1 High A niche instrument has the most to gain. Seed canonical, corroborated facts first — whoever does owns the answer. Prioritise flagship products.
F3 · In technical B2B, authority/E-E-A-T is the winning lever. Chemours: 82–84% citation via named experts + papers citing patents/peer review T3; IEEE Spectrum surge on deep technical content T2; consistent with the study's corroboration finding T1. Medium Lead with primary technical evidence and named expert authorship, not marketing copy. Your IEC/WMO/paper assets are the raw material.
F4 · One repeatable on-page recipe recurs in every win. B2B financing (answer-first/TL;DR, FAQ schema, llms.txt, expert bios) T3; HubSpot definitional content → 3,000+ AI Overview queries T2; multiple vendor cases T3; all consistent with F1 T1. Medium Apply the playbook must-haves: definition-first opening, key-facts block, FAQ schema, HTML spec tables, entity/sameAs wiring.
F5 · AI-referred visitors convert far above organic. Peer-reviewed Marketing Science study + analytics: ChatGPT ~16%, Perplexity ~10–11% vs Google organic ~1.8%; AI arrivals ~38% likelier to buy. T1 T2 High Even small visibility gains are high-value. Use this to justify the investment — quality, not just volume.
F6 · ChatGPT leads volume, but the landscape is fragmenting. B2B referrals early 2026: ChatGPT ~63%, Claude ~18%, Gemini ~11%, Perplexity ~7%; ChatGPT-only now covers a third less than a year ago. T2 Medium Optimise multi-engine. Claude's B2B share is now too large to ignore; don't single-platform.
F7 · Share of voice is platform-specific. Same brand + query set varies 5–10× by engine (e.g. Perplexity 28–38% vs Claude 3–7%), because engines cite different source types. T2 T3 Medium Measure and report SOV per engine, never as one blended number. Tailor corroboration to where each engine looks.
F8 · Most published numbers are inflated by tiny baselines & survivorship. Methodological: e.g. “+100%” = 43→86 sessions; only the controlled study isolates cause from effect. T1 High Trust the direction, not the decimal. Demand absolute numbers before quoting any case externally.
F9 · AI influence is partly invisible to analytics. AI often shapes a purchase without sending a click; referral counts understate, pipeline-attribution claims overstate. T2 Medium Don't judge success on referral clicks alone. Track citation/mention rate & spec-accuracy directly via the prompt suite.

The four universal levers (the “so-what” of F1–F4)

  • Be extractable — answer-first, definition-led, structured HTML, FAQ schema
  • Be dense with evidence — concrete statistics, specs, standards in every section (the #1 lever in the controlled study)
  • Be corroborated — independent papers, standards bodies, named expert authors, Wikipedia/Wikidata
  • Be measured per engine — versioned prompt suite, SOV tracked per platform, re-run on a cadence

The single most important learning (F2)

The controlled study's standout result — citing sources lifts lower-ranked pages by up to ~115%, far more than authoritative ones — flips the usual disadvantage. For a niche instrument with thin search volume and sparse model knowledge, whoever establishes the canonical, corroborated facts first effectively owns the answer. A durable, winnable position, not a race against a famous incumbent.

Playbook tie-in: quantitative backing for prioritising Layer 2 (entity) + Layer 4 (corroboration) on flagship products first.
What we can say with confidence vs. what we can't

Confident (Tier-1 backed): structure + cited statistics + quotations increase AI citation; the effect is largest for low-authority pages; AI traffic converts far better than organic. Directional only: the specific uplift percentages from agency cases, exact per-engine share figures, and dollar-pipeline claims — these tell you what to try, not how much you'll get.

6How to read these numbers — the honest caveats

Apply this filter to every case study you encounter (including the ones above).

  1. Selection bias. Nobody publishes the campaign that flopped. Every percentage is a best case.
  2. Tiny baselines. “+315%” often means “75→311 keywords.” Always ask for absolute numbers.
  3. Confounded with ordinary SEO. Most “AI Overview” wins are partly just good SEO, since AI Overviews lean on the existing search index.
  4. Attribution is hard. AI assistants often influence a purchase without sending a click (“dark” influence) — so referral counts understate impact while pipeline-attribution claims overstate precision.
  5. Only the Princeton study isolates cause and effect. Use it as the backbone; use the rest as illustrations of what teams tried.
  6. Figures churn quarterly. Per your playbook's own note: optimize for the direction (structure helps, corroboration helps, entities help), not the decimal.

7What this means for Vaisala

The evidence is unusually well-aligned with the strategy already in your playbook. Three takeaways to carry into the WindCube work and the portfolio rollout.

1 · You're the underdog — that's good

The controlled evidence says niche, lower-authority pages gain most from citing sources and adding statistics. Vaisala's deep technical evidence base (IEC, WMO, papers, field data) is exactly the raw material that wins here.

2 · Evidence > copy

The closest analog (Chemours, technical B2B) won on named expertise + primary research + standards, not marketing language. Lead with specs, standards, and cited numbers — the SSOT discipline you already enforce.

3 · Measure per engine, prove direction

SOV differs 5–10× between Perplexity and Claude for the same query. Stand up the versioned prompt suite (§8 of the playbook), baseline now, and report per-engine citation rate, SOV, and spec-accuracy — not a blended figure.

Bottom line

No public case study is a perfect proxy for a scientific-instrument maker, and the rigorous evidence is thin — but it all points one way, and that way is the strategy you've already written down: be the clearest, most evidence-dense, most corroborated true source about your product, and measure it per engine. The case studies don't change the plan; they justify the investment.

§Sources & references

Grouped by credibility tier. Vendor-reported figures (Tier 3) should be cited internally with that caveat attached.

Tier 1 — rigorous / peer-reviewed

Tier 2 — independent analytics / analyst data

Tier 3 — vendor / agency self-reported (method > numbers)

Compiled June 2026 via live web research. Figures reflect what each source reported at time of access; AI-visibility metrics shift quarterly — re-verify before quoting externally. Optimize for the direction, not the decimal.