Selective Wikipedia concept routing

Outcome

The extension now uses a compact, offline-trained model to recognise broad concepts in short searches which do not confidently name one Wikipedia entity. It can infer a model domain, nation, historical period or subject role and give matching results a small ordering nudge. For example, Luftwaffe ace can supply German evidence and Kriegsmarine boat can supply German and Second World War evidence even though those labels are absent from the query.

This is deliberately a ranking-only feature. Rules results still come first, dense calibration still decides which semantic results are eligible, and the inferred concepts do not alter the text embedded for the query. The largest possible concept adjustment is 0.02 on the existing cosine score.

Why the model predicts concepts rather than exact themes

An initial experiment attempted to route a query directly to one of the 5,214 latent Wikipedia graph themes. The themes are intentionally fine-grained (median size: five subjects), so short descriptions often identify several equally plausible communities. The best selective operating point reached only 90.83% precision at 2.12% coverage. That missed the 95% promotion gate and was rejected.

The promoted model instead predicts 45 interpretable labels across four families:

  • 9 model domains;
  • 12 nations;
  • 7 historical periods; and
  • 17 subject roles.

These labels are broad enough to transfer across themes, but specific enough to help order a semantic tail. No binary vector index is shipped: the preceding wiki-binary-embedding-evaluation.md found that sign-only retrieval missed its quality gate. Likewise, the pseudo-labelling pass weakened worst-label precision, so this router uses the selective first stage and abstention principle without promoting snowballed labels.

Training and leakage controls

The model is a selective sparse one-vs-rest classifier, analogous to the high-precision, abstaining stage of a hybrid binary snowball workflow. It is trained entirely from the pinned English Wikipedia snapshot:

  1. source-derived domains and facets provide weak labels; previously propagated facets are excluded from both training and gold evaluation labels;
  2. informative unigrams and adjacent bigrams are counted from titles, redirects, descriptions and keywords;
  3. a feature-label association needs repeated support, at least 88% raw precision, a 70% Wilson lower confidence bound and at least 1.5 times the label’s base rate;
  4. each label retains no more than 120 features; and
  5. whole graph themes, rather than individual entities, are assigned to the 70/20/10 train, calibration and evaluation partitions. No fine-grained theme can leak across a split boundary.

Per-family score and margin thresholds are chosen on the calibration partition for at least 98% precision. The model abstains when neither threshold is met. A small non-negative pairwise model then learns how much each label family should contribute to ranking from 20,632 calibration pairs; the weights are tested on a separate 10,104 pair holdout.

Held-out evidence

Measure Result
Evaluation entities 3,053
Short evaluation queries 9,083
Accepted concept predictions 6,011
Prediction precision 97.55%
Query coverage 55.62%
Indirect prediction precision 97.43%
Indirect query coverage 50.50%
Pairwise ranking accuracy 80.20%
Learned mean ranking margin 0.3168
Equal-weight mean ranking margin 0.1716
Label family Precision Query coverage
Domain 97.81% 48.31%
Nation 96.75% 15.59%
Period 98.15% 0.59%
Role 97.39% 1.68%

An indirect prediction is one where the query does not contain the canonical predicted label. This prevents the headline result being explained solely by trivial examples such as German predicting German.

The shipped artefact contains 2,292 features and 2,650 weighted postings across 45 labels. It adds 72,257 bytes to the minified knowledge pack. On the production pack, sparse concept inference measured 0.0041 ms median and 0.0110 ms p95 in the Node runtime benchmark.

Runtime safeguards

  • The model and Wikipedia pack are local; searches and favourites are not transmitted anywhere.
  • An exact, confident entity link still takes precedence over concept routing.
  • Concept routing cannot add a result to the semantic tail and cannot reorder anything above the rules tier.
  • Concept evidence can add at most 0.02; the combined knowledge adjustment remains capped at 0.04.
  • The query embedding text is unchanged, so existing favourite vectors and the semantic model tag remain valid.
  • Pack auditing checks label indices, sparse postings, thresholds, normalised ranking weights, evaluation evidence and the adjustment ceiling.
  • DevTools verbose search logs show inferred labels and their triggering terms, while the ordinary interface stays unchanged.

Limitations

The probes are short queries generated from held-out Wikipedia subjects, not logged searches or human relevance judgements; the extension deliberately collects no such telemetry. The model is therefore promoted only as a small ordering prior behind dense admission, not as a replacement for semantic search. Strict calibration also makes period and role predictions uncommon. A missed concept simply falls back to the existing behaviour, while a wrong concept cannot introduce a result and is limited by both the 0.02 concept cap and the 0.04 combined cap.

Reproducing the model

python -m tools.wiki_pipeline train-query-router `
  --pack data\wiki-entities.json `
  --output data\wiki-entities.json `
  --report docs\wiki-concept-router-report.json

The complete machine-readable split, calibration, per-family and ranking evidence is in wiki-concept-router-report.json.