Binary embedding and snowball evaluation

Decision

Do not replace the extension’s graph themes or dense favourite search with binary-only clustering. Keep a 384-bit shortlist followed by dense reranking as an offline scale-up option, but do not add its 1.44 MB index to the current extension. Do not ship the prototype domain classifier until it passes a holdout made from sparse, realistic KingKit text.

This is an evidence-gated decision rather than a rejection of either source idea. The Hybrid Binary Snowball paper describes selective one-vs-rest classification and iterative coding, not an unsupervised clustering algorithm. The useful analogue here is to seed labels from trusted Wikipedia evidence, abstain on weak predictions, then admit only high-margin pseudo-labels. The binary embedding technique is a compelling compact candidate generator: each sign becomes one bit and Hamming distance replaces a float dot product.

Method

  • The complete 30,000-subject production pack was embedded with the extension’s pinned Xenova/all-MiniLM-L6-v2 revision, q8 model and Transformers.js 3.7.6 runtime.
  • Input used Wikipedia title, description, observed/inferred facets and keywords. The domain label under test was deliberately excluded from the embedding text.
  • Sign quantisation reduced each 384-float vector to 384 bits (48 bytes).
  • Ranking quality used 65,698 unique compact-graph pairs against deterministic same-domain, different-theme negatives, plus 256 deterministic full-corpus nearest-neighbour scans.
  • Nine sufficiently represented single-domain labels were split into seed (50%), calibration (20%), pseudo-label pool (15%) and untouched test (15%) sets. Positive-versus-rest majority-bit prototypes could abstain; the calibration threshold targeted at least 95% precision. One pseudo-labelling pass tested the snowball effect.
  • Model artefacts and the official Node runtime are hash-checked before inference. Cached dense vectors and source snapshots remain outside the repository.
  • The later sparse query-router and bounded search-diversifier additions change only pack-level metadata. The machine report retains the hash of the pack actually evaluated, records the current pack hash separately, and regression-checks the unchanged per-entity embedding-corpus identity.

Results

Measure Dense Binary Interpretation
Storage for 30,000 subjects 46.08 MB 1.44 MB Binary is exactly 32 times smaller.
Graph-pair discrimination 86.73% 84.91% A real but modest 1.83-point loss.
Graph-neighbour recall at 10 33.79% 29.56% Binary retains 87.48% of dense recall.
Same-theme purity at 10 13.91% 11.52% Binary-only clustering is less coherent.
Median 30,000-item scan 9.80 ms 2.69 ms Faster, but dense is already cheap at this scale.

Binary and dense agreed on only 59.77% of first neighbours and shared 57.97% of their top tens. Binary works much better as a coarse stage: its top 64 preserved 91.13% of the dense top ten; top 128 preserved 95.63%.

The label-blind classifier reached 95.38% micro precision, 96.11% macro precision and 81.83% coverage on the untouched Wikipedia test split; its weakest label still reached 94.12%. After one snowball pass it reached 95.19% micro precision, 95.71% macro precision and 82.35% coverage - a coverage gain of only 0.52 percentage points, with the weakest label falling to 93.33%. The pseudo-label pool passed the aggregate 95% gate at 95.32%, although its weakest label reached only 93.94%. That validates selective prototypes inside the Wikipedia distribution, but not transfer to KingKit’s much sparser product text or repeated self-training.

Promotion gates

  • Binary-only clustering: failed. Its neighbourhood fidelity is not high enough to displace the existing graph-derived themes or dense scoring.
  • Binary shortlist plus dense reranking: passed as an experiment. It is worth retaining if a future pack grows far beyond 30,000 subjects or an offline build stage needs cheaper all-pairs candidate generation.
  • HBS-style labelling: provisionally passed in-domain, blocked for runtime. A sparse product-text holdout is required. Iterative snowballing should also stop after a round unless it adds materially more coverage without eroding the precision gate.

The machine-readable figures and release decision are in wiki-binary-embedding-report.json.