Skip to main content

Expanded vocal benchmark and edition policy

The September 12 pilot now contains 13 analyzed recordings: the original six references plus Titanium, Levels, One Kiss, Don't You Worry Child, Get Lucky, This Is What You Came For and Poker Face. Disclosure — Latch and a standalone original/radio Don't Start Now were not found by filename and are skipped at the operator's request.

Recordings before scores

Use standalone album/original, radio and extended recordings for the primary benchmark. Radio edits give compact vocal contrasts; original/extended mixes retain longer introductions, exits and musical phrases. Keep DJ edits, mashups and continuous-mix excerpts in a separate stress cohort. Never transfer timestamps between editions.

The existing Don't Start Now recording is a VDJ JD edit, retained as an edit stress test. Levels, Don't You Worry Child and Get Lucky are labelled radio edits in both filenames and embedded metadata. The other additions have matching track/artist tags. These checks do not establish unmixed provenance: listening review remains necessary. Body Language comes from a compilation tagged Defected In The House Miami 2017; verify whether that particular recording has overlapping adjacent tracks before accepting it as an instrumental control.

The versioned aidj-platform/tools/audio-analysis/benchmark-set.json records the exact filenames, canonical PCM hashes, track IDs, durations, edition notes and intended test cases. Original music files remain unchanged. No example timings from a different recording are treated as ground truth.

New recording results

These are duration-weighted raw music voice scores, not vocal fractions or measured accuracy. High means score ≥0.8; low means score below 0.2. Unknown coverage remains explicit, including short tails without model frames.

RecordingDurationMean voice scoreHigh-score secondsLow-score seconds
Titanium245.07 s0.86421222
Levels — Radio Edit199.93 s0.683643.93
One Kiss214.87 s0.8731622.87
Don't You Worry Child — Radio Edit212.89 s0.9171862.89
Get Lucky — Radio Edit248.44 s0.8101700
This Is What You Came For222.19 s0.9072002
Poker Face237.23 s0.9202044

Titanium has extended high-score spans, and Levels has much more intermediate-score time. Listening labels must establish whether these spans follow the actual vocals. The instrumental/texture controls still receive substantial scores: Aerodynamic averages about 0.515 and Glue 0.522. This prevents interpreting a model score as a calibrated probability of vocal presence or absence.

Reusable evaluation

The benchmark exporter selects reports matching the current canonical PCM and native-analysis revision. It produces two-second scores, contiguous diagnostic score bands, a standalone HTML viewer, and an empty listening-annotation template. The template binds annotations to the exact PCM hash and requires an annotator for nonempty labels.

Presence labels (present, absent, unknown) are separate from dominance (dominant, background, absent, unknown). Sampled/chopped vocals can be present without dominating. Ambiguous vocal-like textures should remain unknown until adjudicated. The current model does not predict dominance, so dominance accuracy remains null.

The evaluator integrates model/label overlap in seconds and reports TP/FP/TN/FN, precision, recall, specificity, F1 and annotated coverage. Unknown labels, missing scores and unlabelled intervals do not become negative examples. It rejects wrong-recording hashes, invalid intervals, overlaps and contradictory labels. Structure labels can be recorded for listening review, but semantic boundary scoring is not implemented yet.

Keep whole recordings held out from threshold tuning. Report results separately for lead, sampled, dense-arrangement and instrumental cases, and for standalone versus edited recordings. Adjacent windows of a song are not independent test recordings. Approximately 4.815 seconds of model context plus two-second output windows limits boundary precision.

The local artifacts are aidj-platform/data/benchmark-vocal-expanded.{json,html} and benchmark-vocal-labels.json. No human listening labels have been invented. Seven benchmark-metric tests, actual Mongo exports, wrong-PCM rejection and desktop/mobile HTML checks were exercised. Full-library processing remains paused; importing these selected files does not qualify them for automatic transitions.