What SEOs believe, and what 12,778 AI answers actually did

A September 2026 survey asked 100+ SEO experts to rank Google's ranking factors. It produced 13,665 data points and never mentions AI search, AEO, GEO, LLM citation or Schema / structured data. We had been measuring the AI answer surface for 164 days. Here is where belief and measurement disagree.

The survey is not wrong for being a survey. It measures what practitioners believe an engine does, which is a real and useful thing to know. It is being read for something else: what engines actually do. Those are different kinds of claim, and only one of them can be checked against data.

Three collisions

Expert consensus

29.4%

named behavioral click signals a top-tier ranking factor. User satisfaction polled under 20 percent.

What we measured

0.061%

of the time, ChatGPT returns the identical set of sources twice.

Measured across 24,633 answer pairs. A surface that unstable does not hold still long enough for click signal to accumulate against a ranking.

Collision 1The factor assumes a stable ranking. On this surface we did not find one.

Expert consensus

Absent

Schema and structured data appear nowhere in the survey's ranked factors.

Matched-control experiment

−4.6%

change in AI Overview citations after adding JSON-LD.

Ahrefs, 1,885 pages against 4,000 matched controls. Schema earns its place through entity resolution and rich results, not citation lift.

Collision 2Consensus and experiment agree. This is the one row where they do.

Expert consensus

0 mentions

AI search, AEO, GEO and LLM citation appear nowhere across all 13,665 data points.

What we measured

12,778

answers across five engines over 164 days.

Across that window the surface the survey does not mention cited 8,577 distinct domains in 109,674 citation events.

Collision 3The largest gap is the one nobody was asked about.

Which engine you ask outweighs most factors on the list

Every answer in our corpus was a real query put to a real engine, repeated over months. The question we can answer that a survey cannot is whether an engine returns the same sources twice.

0.00.10.20.30.40.5Claude0.417Perplexity0.357DeepSeek0.333Gemini0.231ChatGPT0.143
Median Jaccard overlap between repeat runs of one prompt. 1.0 would mean an engine returns the identical sources every time. Whiskers are 95% confidence intervals from a cluster bootstrap. On this measure Claude's citation sets overlap 2.9 times more than ChatGPT's; on the stricter test of returning a wholly identical set, Claude does so 27 times as often.

No ranking-factors survey has a row for which engine the question was asked in. On this evidence it moves the answer more than most rows that do exist.

The finding with money attached

If an engine rarely repeats itself, a single measurement of it is not a measurement. That is quantifiable, and it sets a floor on how many times a prompt must run before a change in AI visibility means anything.

95% confidence interval half-width for one prompt's mean Jaccard, by number of runs. Pooled within-prompt SD 0.1491, from 878 groups with three or more runs.
Runs95% CIVerdict
1n/ano variance estimate at all
2±0.2067too noisy to act on
3±0.1688too noisy to act on
4±0.1462directional only
6±0.1193directional only
8±0.1033directional only
9±0.0974usable for large moves
12±0.0844usable for large moves
14±0.0781usable for large moves

Below 9 runs per prompt, the confidence interval is wider than ±0.10. A tool reporting that your AI visibility improved this month, on fewer runs than that, is reporting noise. Most tools in this category run a prompt once.

Where the numbers came from

Measured, our corpus, recomputed 2026-09-17

Relayed, third party, linked at source

Withdrawn, published by us, does not reproduce

Seven domains entered tracking on different dates. boatlaw.com (69 mentions over 7 days) sits below the reporting floor and is excluded from pooled conclusions. Controlled to the six domains present throughout, median Jaccard is flat across April and May.

The full dataset ships as JSON under CC BY 4.0. Re-use it, check it, or disagree with it in public.