Flagship Research Paper Jun 12, 2026 · 28 min read

We Analysed One Million AI Responses

A methodology for measuring how AI assistants recommend brands, sources and entities — and what a representative benchmark corpus reveals about the mechanics of AI visibility.

KernelX Labs Research · Benchmark Series KBC-1M · Version 1.2

Download PDF

Section 01Executive Summary

Generative AI assistants have become a primary discovery surface for products, services and information. Yet the mechanics that determine which brands they recommend, which sources they cite, and how consistently they do either remain poorly instrumented. This paper introduces a reproducible methodology for measuring those mechanics at scale.

We describe the construction and analysis of the KernelX Benchmark Corpus (KBC-1M): a representative benchmark dataset of approximately one million assistant responses, generated from roughly 500,000 unique prompts spanning 70 industries, four frontier model families and 15 country/locale configurations. All quantitative results in this paper are derived from this benchmark corpus and should be read as representative benchmark findings that demonstrate the methodology, not as externally audited market measurements.

~1,000,000assistant responses in the benchmark corpus
~500,000unique prompts across 12 intent classes
70industry verticals segmented
4frontier model families evaluated
15country / locale configurations

Key findings

  • Official brand websites appeared in approximately 80–82% of responses that contained at least one citation — the single most common citation destination in the corpus.
  • Community sources matter disproportionately: Reddit was the most frequently referenced community domain, with the strongest influence in consumer software and electronics prompts.
  • Wikipedia and structured encyclopedic sources remained among the strongest entity-establishment signals: brands with complete, consistent entity coverage were recommended materially more often than functionally similar brands without it.
  • Brands that publish structured comparison content received substantially more recommendations in "best X for Y" prompts than brands relying on generic landing pages.
  • Original research and data-led content was cited meaningfully more often than standard blog content of equivalent topical relevance.
  • Model agreement is partial: the four model families agreed on the top recommendation in roughly half of commercial prompts, implying that visibility must be measured per model, not in aggregate.
Research Note

Framing. Throughout this report, quantities are expressed as ranges (for example, "approximately 80–82%") rather than false-precision point estimates. The corpus is a controlled benchmark built to make AI-visibility measurement reproducible; it is not a census of production assistant traffic.

Key takeaways

  • AI recommendation behaviour is measurable, repeatable and materially different across model families.
  • Citation behaviour concentrates on official sites, community platforms and encyclopedic sources — in that order.

Business implications

  • Brands can treat AI visibility as an instrumentable channel with its own metrics, baselines and levers.

Recommendations

  • Adopt a per-model measurement discipline before investing in optimisation tactics.

Section 02Introduction & Research Questions

Search behaviour is migrating from ranked lists of links to synthesised answers. When an assistant answers "what is the best CRM for a 10-person startup?" it performs three operations that classical search never fully collapsed into one step: it retrieves candidate information, it selects a small set of entities to name, and it frames those entities with sentiment and justification. The selection step — which brands get named at all — is the core object of study in Generative Engine Optimisation (GEO).

Existing SEO instrumentation does not transfer cleanly. Rank positions, impressions and click curves have no direct analogue inside a generated paragraph. A measurement programme for AI visibility therefore needs new primitives: mention detection, citation extraction, recommendation scoring, consistency measurement and hallucination detection — each defined precisely enough to be reproduced.

Research questions

  1. RQ1 — Selection: Which observable properties of a brand correlate with being recommended by AI assistants?
  2. RQ2 — Citation: When assistants cite sources, which source classes do they prefer, and does this differ by model family?
  3. RQ3 — Consistency: How stable are recommendations across repeated sampling, phrasing variants and locales?
  4. RQ4 — Agreement: How often do different model families converge on the same recommendations?
  5. RQ5 — Failure: At what rate do assistants produce non-existent entities, wrong attributes or unverifiable claims in commercial answers?
Key Observation

The unit of competition in AI search is the entity, not the page. Pages are evidence; entities are what get recommended. Measurement systems that stay page-centric systematically mis-model the channel.

Key takeaways

  • GEO requires new measurement primitives; classical SEO metrics do not transfer.

Business implications

  • Teams should reframe reporting from "rankings" to "share of recommendation" per prompt set.

Recommendations

  • Define a fixed prompt panel per category before measuring anything else.

Section 03Methodology: The AI Recommendation Pipeline

We model every assistant answer as the output of a five-stage pipeline. The pipeline is a descriptive framework — it does not assume access to any vendor's internals; it decomposes the observable answer into stages that can each be measured independently.

Figure 1 — Framework
The AI Recommendation Pipeline
Five-stage flow diagram: Prompt Interpretation → Retrieval & Grounding → Entity Selection → Framing & Justification → Citation Assembly. Each stage annotated with the metrics this methodology attaches to it.

The visual communicates where each measurement primitive attaches: intent classification at stage 1, source-class analysis at stage 2, recommendation scoring at stage 3, sentiment framing at stage 4 and citation extraction at stage 5.

Measurement primitives

PrimitiveDefinitionOutput
Mention detectionNamed-entity recognition tuned for brand and product names, resolved against a curated entity registryEntity set per response
Recommendation scoringPosition-weighted scoring of entities in explicit recommendation contexts ("we recommend", ranked lists, superlatives)Score per entity per response
Citation extractionParsing of inline citations, reference lists and link annotations into a normalised source taxonomySource-class distribution
Consistency measurementRepeated sampling (n≈5 per prompt) plus paraphrase variants; agreement computed with a Jaccard-style overlap indexStability index 0–1
Hallucination detectionEntity resolution failures + attribute checks against the registry (pricing tiers, feature existence, company facts)Failure rate per model per category

Table 1 — The five measurement primitives used throughout the KernelX research programme.

The Visibility Score

To make results comparable across categories we aggregate primitives into a single bounded score. Weights below are the benchmark defaults; the sensitivity analysis in the appendix shows rank orderings are robust to ±10-point weight shifts.

VisibilityScore(brand, panel) = 100 × Σᵢ wᵢ · fᵢ f₁ recommendation share (w₁ = 0.40) f₂ citation share (w₂ = 0.20) f₃ first-mention share (w₃ = 0.15) f₄ consistency index (w₄ = 0.15) f₅ sentiment-weighted framing (w₅ = 0.10) RecommendationConfidence(brand, prompt) = agreement(models) × stability(samples) × specificity(context)
Important Limitation

Every composite score embeds editorial choices. We publish the weights, the components and the sensitivity analysis so that any team can recompute the score under different assumptions. A score whose construction is hidden is a marketing number, not a measurement.

Key takeaways

  • Answers decompose into five measurable stages; each primitive is independently reproducible.

Business implications

  • Vendors and in-house teams can adopt the same primitives and compare results.

Recommendations

  • Publish weighting schemes whenever composite visibility scores are reported.

Section 04Dataset Construction

The corpus was constructed in four passes: prompt taxonomy design, prompt generation, response collection and cleaning. The design goal was coverage with controlled variation — every industry receives the same intent mix, so cross-industry comparisons are like-for-like.

Prompt taxonomy

Intent classExample patternShare of corpus
Best-of selection"best {category} for {segment}"~18%
Direct comparison"{brand A} vs {brand B} for {use case}"~14%
Alternatives"alternatives to {brand}"~11%
Purchase guidance"which {category} should I buy if {constraint}"~11%
How-to with tooling"how do I {task} — what tools do I need"~10%
Category explanation"what is {category} and who are the main providers"~9%
Pricing & value"is {brand} worth it / pricing comparison"~8%
Reputation"is {brand} reliable / what do people say about {brand}"~7%
Integration fit"does {brand} work with {ecosystem}"~5%
Local / regionallocale-conditioned variants of the above~4%
Regulated advicefinance / health phrasing with constraint language~2%
Adversarial phrasingnegations, typos, mixed languages~1%

Table 2 — Twelve intent classes; every industry panel receives an identical intent mixture.

Collection & cleaning

  • Sampling: each prompt issued to each model family under default consumer settings, with repeated sampling for the consistency subset.
  • Locale control: 15 country/language configurations applied to the locale-sensitive intent classes.
  • Cleaning: deduplication of near-identical responses, removal of refusals and malformed outputs (~2–3% of raw collection), normalisation of citation formats across model families.
  • Entity registry: a curated registry (~30,000 brands and products) providing canonical names, aliases and verifiable attributes for resolution and hallucination checks.
Figure 2 — Dataset
Corpus composition by industry group
Software & SaaS — 26%Consumer & retail — 22%Finance & insurance — 14%Health & wellness — 12%Travel & local — 11%Other verticals — 15%

The visual communicates the balance of the benchmark corpus across major industry groups; no single group exceeds ~26% of prompts.

Key takeaways

  • Identical intent mixtures per industry make cross-industry findings comparable.

Business implications

  • A brand's panel can be reconstructed with modest effort: taxonomy + registry + sampling discipline.

Recommendations

  • Treat the prompt panel as versioned infrastructure; changing it invalidates trend lines.

Section 05Findings I: Citation Behaviour

Across the corpus, roughly 60–65% of commercial responses contained at least one identifiable citation or source attribution. Within that cited subset, source classes distribute as follows.

Figure 3 — Citations
Share of cited responses referencing each source class
Official brand sites80–82%
Community (Reddit et al.)37–40%
Encyclopedic (Wikipedia)32–35%
Review platforms28–31%
News & trade media23–26%
Independent blogs16–18%
Academic / research8–10%

The visual communicates that official sites dominate citation behaviour, but community and encyclopedic sources form a strong second tier. Categories are not mutually exclusive; a response can cite several classes.

Key Observation

Original research earns citations out of proportion to its volume. Data-led pages (benchmarks, surveys, indices) made up a small fraction of available brand content in the registry, yet appeared in the cited set at 2–3× the rate of standard blog content on equivalent topics.

Source trust ordering

We summarise citation preferences in a Citation Trust Framework: assistants behave as if sources are ordered by verifiability and institutional accountability — official documentation first, then community consensus, then editorial content. The ordering was stable across model families even where absolute rates differed.

Trust tierSource classesObserved behaviour
Tier 1 — CanonicalOfficial sites, documentation, structured dataCited by default when available; anchors factual claims
Tier 2 — ConsensusReddit, Stack Exchange, review platforms, WikipediaUsed to justify subjective judgements ("users report…")
Tier 3 — EditorialNews, trade media, expert blogsUsed for context, comparisons and recency
Tier 4 — PromotionalGeneric marketing pages, thin affiliate contentRarely cited; occasionally paraphrased without attribution

Table 3 — The Citation Trust Framework: a descriptive ordering of source classes by observed citation preference.

Key takeaways

  • Official sites anchor factual claims; community sources anchor subjective judgements.

Business implications

  • A brand's documentation quality is now a distribution channel, not just a support asset.

Recommendations

  • Invest in canonical, structured, factual pages before investing in more editorial volume.

Section 06Findings II: Recommendation Patterns & Model Agreement

Recommendation behaviour concentrated sharply: in a typical category, the top three recommended brands captured roughly 55–65% of all recommendation weight, while the long tail of remaining brands shared the rest. Concentration was highest in mature categories (CRM, cloud storage) and lowest in emerging ones (AI developer tools), where model families disagreed more.

Figure 4 — Agreement
Cross-model agreement on the top recommendation, by prompt intent
~71%
Category explain
~54%
Best-of
~49%
Comparison
~43%
Alternatives
~41%
Purchase advice

The visual communicates that model families converge on explanatory prompts but diverge substantially on advisory ones — the prompts with the highest commercial intent are the least consistent.

What correlates with being recommended

SignalCorrelation strengthNotes
Entity completeness (registry coverage, consistent naming)StrongStrongest single differentiator between comparable brands
Structured comparison contentStrongEffect concentrated in best-of and comparison intents
Third-party citation footprintStrongBreadth of independent sources matters more than volume on owned channels
Original research & data assetsModerate–strongOutsized citation rate; slower to influence recommendation share
Community presence (Reddit, forums)ModerateCategory-dependent; highest in consumer software/electronics
Domain authority (classic SEO)ModerateCorrelates, but with notable exceptions in both directions
Publishing frequency aloneWeakVolume without structure showed little measurable effect

Table 4 — Observed correlations between brand-side signals and recommendation share in the benchmark corpus. Correlational, not causal.

Practical Implication

Because agreement on advisory prompts sits near coin-flip levels, a brand can be dominant in one assistant and nearly invisible in another while its team believes "AI visibility" is a single number. Per-model dashboards are a requirement, not a refinement.

Consistency and hallucination

Repeated sampling produced stable top recommendations in roughly 70–80% of cases, dropping for long-tail brands. Hallucination-class failures — non-existent products, discontinued tiers, invented attributes — appeared in approximately 3–6% of commercial responses depending on model family and category, concentrated in fast-moving categories where training data ages quickly.

Key takeaways

  • Recommendation weight is concentrated; advisory prompts show the lowest cross-model agreement.

Business implications

  • Aggregate visibility numbers hide model-level risk; per-model measurement is essential.

Recommendations

  • Track a stability index alongside share metrics; volatile visibility is fragile visibility.

Section 07The Entity Authority Model & AI Visibility Funnel

Synthesising the findings, we propose two working frameworks for practitioners.

Entity Authority Model

A brand's probability of recommendation is modelled as a function of four layers, evaluated bottom-up: Identity (is the entity unambiguous?), Evidence (is there canonical, verifiable content?), Consensus (do independent sources corroborate it?), and Salience (is it associated with the specific prompt context?). Failures at lower layers cap the value of investment at higher ones — community advocacy cannot compensate for an ambiguous entity.

Figure 5 — Framework
Entity Authority Model (four-layer pyramid)
Pyramid diagram: Identity → Evidence → Consensus → Salience, annotated with the signals measured at each layer and typical failure modes.

The visual communicates the dependency ordering: each layer is necessary for the layers above it to convert into recommendations.

AI Visibility Funnel

Funnel stageMetricBenchmark median (all industries)
RecognisedEntity resolution rate~90–95% for established brands; far lower for sub-brands
RetrievedAppears anywhere in relevant answers~35–45%
RecommendedNamed in recommendation context~12–18%
PreferredFirst-mentioned recommendation~5–8%
TrustedCited as a source itself~3–6%

Table 5 — The AI Visibility Funnel with representative benchmark medians. Individual categories vary widely; the funnel shape is the durable insight.

Best Practice

Diagnose before optimising: locate the funnel stage where a brand leaks most. Retrieval problems are content problems; recommendation problems are consensus problems; preference problems are differentiation problems. The treatments are different.

Key takeaways

  • Authority is layered; each layer gates the next.
  • The funnel identifies where visibility is actually lost.

Business implications

  • GEO budgets can be allocated against a specific failing layer instead of spread thinly.

Recommendations

  • Run the funnel per model family and per intent class; fix the deepest leak first.

Section 08Limitations

  • Benchmark, not census. The corpus samples controlled prompt panels, not live user traffic. Real prompt distributions differ by product and audience.
  • Temporal validity. Model families update frequently; findings describe the collection window and decay with time. Trend tracking (see the quarterly series) matters more than any snapshot.
  • Correlation only. Observed relationships between brand signals and recommendation share are correlational. Controlled interventions are the subject of ongoing work.
  • Registry coverage. Hallucination detection is bounded by registry completeness; failures outside registered attributes go uncounted.
  • English-market weighting. Locale coverage is broad but English-language markets remain over-represented relative to global usage.
Important Limitation

No external audit of this corpus has yet been performed. We publish the methodology at full depth precisely so that results can be independently reproduced and challenged.

Key takeaways

  • Findings are reproducible benchmark results with explicit boundaries.

Business implications

  • Consumers of this research should weight the methodology, not the headline numbers.

Recommendations

  • Re-run panels quarterly; treat deltas as more informative than levels.

Section 09Conclusion: The Future of AI Visibility

Search engines, assistants, agents and recommendation systems are converging into a single discovery fabric. The same entity graph that determines whether a chatbot names a brand will shortly determine whether an autonomous agent shortlists it, negotiates with it, or transacts with it. In that world, visibility stops being a marketing metric and becomes an interface: the machine-readable representation of what a company is, evidenced by what independent sources say about it.

Three shifts follow. First, authority compounds: entity-level trust accrues slowly and transfers across surfaces, favouring brands that invest early in canonical evidence and independent consensus. Second, measurement becomes continuous: as models update weekly, visibility monitoring converges on observability practices familiar from infrastructure engineering. Third, the content economy re-prices: material that machines can verify and cite — research, data, documentation, structured comparisons — appreciates; interchangeable editorial volume depreciates.

The methodology in this paper is offered as a foundation for that discipline: a shared vocabulary of primitives, a reproducible corpus design, and frameworks that turn measurements into decisions. Subsequent reports in this series — the quarterly State of GEO and the comparative study GPT vs Claude vs Gemini — apply this framework longitudinally and across model families.

Key takeaways

  • Discovery surfaces are converging on shared entity infrastructure.

Business implications

  • Early authority investment compounds across future surfaces, including agentic commerce.

Recommendations

  • Build the measurement discipline now; optimisation without instrumentation is guesswork.

Section 10Appendix: Glossary, References & Methodology Notes

Glossary

TermDefinition
GEOGenerative Engine Optimisation: improving how AI systems recommend and cite a brand
Recommendation shareA brand's weighted share of recommendation contexts within a prompt panel
Stability indexOverlap of recommended entity sets across repeated samples of the same prompt
Entity registryCurated database of canonical brand names, aliases and verifiable attributes
Citation shareShare of cited responses in which a brand's owned properties appear as sources
Hallucination rateShare of responses containing unresolvable entities or attribute errors

Methodology notes

  • Prompt panels are versioned; this paper reports against panel v1.2.
  • Recommendation scoring weights first mentions at 1.0, subsequent mentions decaying harmonically.
  • Sensitivity analysis: Visibility Score rank orderings unchanged for the top decile under ±10-point weight perturbations.
  • Refusal and safety-template responses are excluded from denominator counts.

References

  1. KernelX Labs Research. The State of GEO — Q3 2026. Quarterly series applying the KBC framework longitudinally.
  2. KernelX Labs Research. GPT vs Claude vs Gemini: How Recommendations Differ. Comparative study on the KBC-50K commercial subset.
  3. KernelX Labs Research. The AI Visibility Index 2026. Illustrative industry rankings computed with the Visibility Score.
  4. Public model documentation and system cards of the evaluated assistant families (versions as of the collection window).