Evaluate timing, audio, and intervention
Metrics support listening tests; they do not replace musical judgment. No scores or performance claims exist yet. Compare runs with the same track pool, starting conditions, analysis versions, and evaluation definitions.
| Metric | Proposed definition | Required evidence |
|---|---|---|
| Beat error | Absolute phase offset in ms at eligible transition anchors; report distribution and threshold failure rate | Valid aligned beatgrids and executed positions |
| Phrase error | Distance in beats from an annotated phrase boundary at transition start | Phrase annotations with confidence |
| Clipping | Fraction of rendered samples above configured ceiling, plus peak dBFS | Pre-limiter audio measurement |
| Bass overlap | Time both contributing decks exceed a declared low-band energy threshold | Per-deck post-channel audio/stems |
| Silence | Master duration below a configured loudness threshold, excluding intentional silence | Rendered master and scenario labels |
| Energy continuity | Loudness/energy change across declared transition windows | Windowing method and calibrated signal features |
| Harmonic compatibility | Compatibility under a named key model; unknown keys excluded | Key estimate, confidence, model version |
| Tempo change | Rate and magnitude of applied tempo changes | Executed transport events |
| Transition duration | Frames from annotated audible overlap start to completion | Explicit transition-window rule |
| Human takeover | Count and rate per AI-controlled minute | Ownership journal in live sessions |
Store units, thresholds, eligible denominator, exclusions, and missing-data counts with every result. A heuristic such as bass overlap must not be labelled a perceptual defect without listening validation. Human takeovers cannot be inferred from AI-vs-AI offline runs.
Experiment protocol
Select representative fixtures, pin input assets/configuration, run paired scenarios across seeds, report distributions and failures, then perform blind listening comparisons. Separate tuning material from held-out evaluation sessions. Save run IDs and recordings with results.
Batch runs and faster-than-real-time rendering follow reliable single-run reproduction. Report wall-clock throughput, audio duration, machine configuration, and underruns separately; 5×/20×/100× remain targets to measure.