Zero Data Leakage Protocol Enacted: GroupKFold by Query & Domain • Own Domains Strictly Excluded • Honest Empirical Results
c
cited.
Key Research Takeaway • Honest Failure Clause

Classic Search Rank Explains >70% of AI Citations.

We built an empirical research pipeline evaluating 50 buyer-intent queries and 1,500 candidate pages across Google Gemini Search Grounding, Perplexity Sonar, and Google AI Overviews. Controlling strictly for classic organic SERP rank, on-page semantic features yield only a marginal +0.018 PR-AUC gain. Removing rank collapses model prediction to near-random noise.

Rank-Only Baseline
0.4400
PR-AUC [0.402 - 0.484]
ROC-AUC 0.7189 • Prec@1: 50%
Full LightGBM
0.4136
PR-AUC [0.381 - 0.461]
35 features (text+embed+tech)
Logistic Regression
0.4583
PR-AUC [0.422 - 0.503]
Highest precision baseline (+0.018)
Ablation: No Rank
0.3171
Random baseline: 0.3079
Collapses without rank (-28%)
Counselrise Product Feature

Client AI Visibility Audits

Real automated Share of Voice (SoV) diagnostics, competitor citation matrices, and PDF export.

AI Share of Voice
Citation frequency in vertical
Brand Citations
Verified source links
Competitor Citations
Across same query set
Total Engine Citations
Perplexity + Gemini + AIO

Engine Mention & Citation Rate

Top Competing Citation Domains

Domain Citations Engines

Prescribed Engineering Interventions

Research Artifacts

Figures F1 through F5

High-resolution publication figures generated from 1,500 evaluated candidate pages with bootstrap confidence intervals.

Figure 1: Rank Confounder
Enlarge
Figure 1

Rank Confounder Decay Curve

P(cited | rank) drops exponentially. Over 70% of citations originate from organic ranks 1–3.

PR-AUC: 0.4400 (Rank alone)
Figure 2: Precision-Recall Curves
Enlarge
Figure 2

Precision-Recall Curves

Full model tracks the classic rank baseline. Logistic Regression achieves 0.4583 (+0.018 linear boost).

5-fold GroupKFold by Query
Figure 3: SHAP Feature Importance
Enlarge
Figure 3

SHAP Feature Importance

Classic rank accounts for >60% of total SHAP impact. On-page features contribute in the tail.

TreeSHAP Explainer
Figure 4: Engine Overlap
Enlarge
Figure 4

Cross-Engine Overlap

Pairwise Jaccard similarity is ~0.24, showing answer engines favor distinct source subsets despite shared rank bias.

Mean Jaccard: 0.238
Figure 5: Brand Share of Voice
Enlarge
Figure 5

Brand Share of Voice

Client citation share in aesthetic clinics, urology, and hospitality verticals across all 3 generative engines.

Counselrise Client Metric
Ablation Study (Table 1)

Empirical Benchmark & Ablation Matrix

Strict 5-fold GroupKFold by Query ID (1,350 training rows, 27.8% positive rate). Bootstrap 95% Confidence Intervals.

Model / Specification PR-AUC [95% CI] ROC-AUC [95% CI] P@1 MRR Brier Confounder Status / Note
Random Baseline 0.3079 [0.277, 0.340] 0.5236 [0.492, 0.554] 20.0% 0.5202 0.3186 Class prevalence floor
Rank-Only Baseline (1/rank) 0.4400 [0.402, 0.484] 0.7189 [0.689, 0.748] 50.0% 0.7042 0.2324 The Primary Confounder
Top 3 Organic Heuristic 0.3557 [0.322, 0.391] 0.6114 [0.585, 0.637] 74.0% 0.8202 0.2115 High precision, poor recall in tail
Logistic Regression (Full) 0.4583 [0.422, 0.503] 0.7121 [0.684, 0.740] 60.0% 0.7437 0.2156 Highest overall PR-AUC (+0.018)
LightGBM (Full 35 features) 0.4136 [0.381, 0.461] 0.6765 [0.648, 0.706] 50.0% 0.6877 0.2188 Non-linear interactions overfit noise
Ablation: All Except Rank 0.3171 [0.293, 0.352] 0.5719 [0.547, 0.602] 28.0% 0.5182 0.2381 Performance collapses to noise
Domain Grouped Split (LightGBM) 0.4118 0.6719 48.0% 0.6811 0.2201 AUC drop = 0.0046 (Zero domain memorization)
Open Data Release

Dataset & Query Explorer

50 buyer-intent queries and 1,500 labeled observations across 3 commercial verticals.

ID Query Text Vertical Intent Locale

Sample Labeled Records (100 Rows Preview)

Q-ID Rank Cited? URL Max Cosine Sentence Len Flesch RE Schema?
Publications & Documentation

Research Paper & Datasheet