Research
What we can show, what we cannot yet, and the measurements that told us which is which. Published including the ones that came back empty.
Validity
Does the crowd peak land on the headliner’s own record?
Moments are ranked by measured response; names are joined on time afterwards and never feed the score. So the two are independent, and agreement between them is corroboration rather than a circular check.
- 12/24named top-3 moments are the headliner’s own record
- 5/8named #1 moments
- ~16of the ~800 pairwise human labels the evaluation gate wants
Suggestive, not validated. The label count is the reason.
Null result
Cross-set matching of unreleased IDs finds nothing
Audio fingerprints should recognise the same unreleased record turning up in two different sets. The mechanism works and is switched on. At a threshold where precision is defensible it reports zero clusters on this corpus, and we publish that rather than lowering the bar.
- 0cross-set clusters at the honest threshold
- 169 / 176 / 194score of one correct and two wrong matches — score does not separate them
- 1true match found before the threshold rose past it
Fixed by window coverage and a different corpus shape, not by a lower threshold.
Bottleneck
Segmentation limits everything downstream
Where a track starts and ends is currently decided by spectral change plus a minimum length — and a crossfade plays two records at once, which is exactly where spectral change detection fails. Recognition is the sharper boundary detector.
- 30/193segments are one track split in two
- 63/193had recognition probes disagree with each other
- 2.9–3.5 minmedian segment in every set — the floor is doing the segmenting
This, not better names, is the real argument for a paid recognition API.