Wikipedia-aware semantic search diversification

Outcome

Broad concept searches now expose a slightly wider range of relevant subjects when several semantic results are near-tied. The extension uses the existing unsupervised Wikipedia graph themes to reduce repetition in the semantic tail. For example, a broad search is less likely to spend most of its eight result slots on several boxings of one subject or on a single tightly related aircraft family.

This changes ordering only. Rules matches remain first, dense calibration still decides which favourites are eligible, and the strongest semantic result is always fixed in first place. There is no new control, label or stored preference in the manager; verbose DevTools logs show the movement and its evidence.

Runtime method

The promoted model is a conservative maximal-marginal-relevance (MMR) pass over at most 24 already accepted semantic candidates. It runs only when the selective query router has recognised a broad domain, nation, period or role concept. Exact Wikipedia entity searches and unclassified queries retain their previous order.

After preserving the leader, each position compares only candidates no more than 0.01 below the best remaining base score. A compact redundancy score is then applied:

  • another boxing linked to the same Wikipedia subject: 1.00;
  • another subject in the same latent graph theme: 0.85; and
  • a direct compact-graph neighbour: 0.30.

The resulting penalty cannot exceed 0.01. A candidate outside the score band is never considered, however diverse it might be. This makes the model a bounded tie-break rather than a general recommendation system.

Offline calibration and holdout

The evaluator creates natural broad queries from domain/facet combinations observed on at least three subjects in the pinned enwiki-20260801 snapshot. It routes them through the production selective concept model, embeds them with the extension’s pinned MiniLM revision, and ranks the complete 30,000-subject pack using its verified dense vector cache. Complete query text is deterministically split by SHA-256 parity: 143 queries tune 36 conservative parameter combinations and 177 remain untouched for final evaluation.

Holdout measure Result
Leading-result preservation 100%
Mean nDCG@8 0.996922
Mean base-score retention 0.999733
Mean query-concept retention 1.000468
Additional distinct themes in the top eight 0.322034
Reduction in same-theme result pairs 0.013923
Largest observed score sacrifice 0.008446
Evaluator/runtime ranking mismatches 0

A separate duplicate-boxing stress test adds two close-scoring copies of each leading subject. It retains the leader on every query, adds 0.175 distinct subjects to the top eight on average, and reduces same-subject pairs by 0.0115. This is intentionally modest: the relevance band is allowed to prevent diversification when the next distinct subject is materially weaker.

The runtime pass measured 0.0787 ms median and 0.1523 ms p95 over a 24-candidate synthetic worst case. Its compact configuration adds 678 bytes to the knowledge pack; it adds no per-favourite storage and performs no network request.

Safeguards and limitations

  • Rules results and dense semantic admission are unchanged.
  • The first semantic result is immutable.
  • The hard score-loss band and the diversity penalty are independently capped at 0.01.
  • Diversification requires at least four candidates and a selective concept-query route.
  • The offline evaluator scans a 256-result dense shortlist. Across all 320 routed queries, the strongest possible unseen knowledge adjustment still remained at least 0.019857 below the final 24-candidate cut-off.
  • Candidates without a confident Wikipedia match receive no penalty; ordinary semantic ranking remains their only evidence.
  • Pack auditing checks model provenance, scope, weights, caps, holdout quality, duplicate stress evidence and exact evaluator/runtime parity.
  • The holdout uses Wikipedia-derived broad queries rather than private search telemetry. This is why the feature is kept behind dense admission and constrained to near-ties.

Reproducing the evaluation

The dense vector cache, model files and Node runtime remain beside the local snapshot rather than in Git:

node tools\wiki_pipeline\evaluate-search-diversity.mjs `
  --pack data\wiki-entities.json `
  --vectors <external-work-directory>\wiki-vectors.f32 `
  --model-root <offline-model-directory> `
  --runtime-module <offline-runtime-directory>\node_modules\@huggingface\transformers\dist\transformers.node.mjs `
  --output docs\wiki-search-diversity-report.json `
  --output-pack data\wiki-entities.json `
  --query-limit 320

The complete machine-readable calibration grid, holdout, stress-test and promotion gates are in wiki-search-diversity-report.json.