Skip to main content

Evaluate timing, audio, and intervention

Metrics support listening tests; they do not replace musical judgment. No scores or performance claims exist yet. Compare runs with the same track pool, starting conditions, analysis versions, and evaluation definitions.

MetricProposed definitionRequired evidence
Beat errorAbsolute phase offset in ms at eligible transition anchors; report distribution and threshold failure rateValid aligned beatgrids and executed positions
Phrase errorDistance in beats from an annotated phrase boundary at transition startPhrase annotations with confidence
ClippingFraction of rendered samples above configured ceiling, plus peak dBFSPre-limiter audio measurement
Bass overlapTime both contributing decks exceed a declared low-band energy thresholdPer-deck post-channel audio/stems
SilenceMaster duration below a configured loudness threshold, excluding intentional silenceRendered master and scenario labels
Energy continuityLoudness/energy change across declared transition windowsWindowing method and calibrated signal features
Harmonic compatibilityCompatibility under a named key model; unknown keys excludedKey estimate, confidence, model version
Tempo changeRate and magnitude of applied tempo changesExecuted transport events
Transition durationFrames from annotated audible overlap start to completionExplicit transition-window rule
Human takeoverCount and rate per AI-controlled minuteOwnership journal in live sessions

Store units, thresholds, eligible denominator, exclusions, and missing-data counts with every result. A heuristic such as bass overlap must not be labelled a perceptual defect without listening validation. Human takeovers cannot be inferred from AI-vs-AI offline runs.

Experiment protocol

Select representative fixtures, pin input assets/configuration, run paired scenarios across seeds, report distributions and failures, then perform blind listening comparisons. Separate tuning material from held-out evaluation sessions. Save run IDs and recordings with results.

Batch runs and faster-than-real-time rendering follow reliable single-run reproduction. Report wall-clock throughput, audio duration, machine configuration, and underruns separately; 5×/20×/100× remain targets to measure.