Bottom line: early exploratory work provided some support for using website structure as a technical diagnostic. It did not calibrate the current 44-rule score, establish reliable prediction of live AI citations, or show that changing a score causes traffic, enquiries or revenue.
Two questions must stay separate
A structural comparison asks whether selected well-built and poorly-built pages receive different technical scores. A live-citation study asks whether named answer engines cite a site for a frozen set of prompts at a stated time. They use different outcomes and must not be presented as the same validation.
What the early work can support
The earlier exploratory comparisons suggest that the old score detected some differences in selected website structure. A separate citation comparison found only a weak association. The samples, score versions and analysis records are not sufficient to treat those results as calibrated performance evidence for the current score.
What remains unknown
It remains unknown whether a given technical score predicts Google or Bing rankings, live mentions or citations in ChatGPT, Claude, Gemini or Perplexity, qualified traffic, enquiries, conversion or revenue. Correlation would not by itself establish causality.
How the current score is tested
The current 44-rule scanner is checked against a frozen local classification and security matrix, including positive, negative, unavailable, malformed and adversarial evidence. That supports repeatable rule behaviour. It does not establish real-world business impact or live-engine selection.
Evidence needed for a stronger claim
A future study would need a preregistered sample and query set, frozen engine and score versions, repeated observations, independent adjudication, uncertainty intervals, preserved failures and a reproducible analysis. Any causal claim would additionally require a controlled intervention or another credible causal design.
Read the current scoring boundary
The Methodology page explains what each score measures, how unknown evidence is handled, and what the result cannot prove.
Read the Methodology