Methodology
The protocol, not a percentage.
This is how transcription quality is measured before anything on this site claims it: what the test set is, how it was degraded, how notes are scored — and what the measurement cannot tell you about your recording.
Ground truth that is exact
Comparing a transcription with a recording has an alignment problem: the ground-truth score and the audio rarely agree on exactly when a note began. We removed the problem rather than arguing with it — the test audio is rendered from the reference score, so the ground truth is exact to the millisecond and a miss is a real miss.
The material is public-domain piano MIDI from Mutopia, rendered with FluidSynth and the FluidR3 GM soundfont at 44.1 kHz, cut to 30-second windows. The trade is honest: a sampled-GM render is not a living-room recording, so a real human performance — Edward Neeman's public-domain recording of the Chopin étude — is checked alongside, qualitatively, to catch synthetic-only artefacts.
Bach, Invention No. 1, BWV 772
Sparse two-voice counterpoint — the easy case
220 notes / 30 s
Bach, WTC I Prelude No. 1, BWV 846
Steady broken-chord polyphony, no pedal
120 notes / 30 s
Satie, Gymnopédie No. 1
Slow, heavily pedalled
50 notes / 30 s
Chopin, Étude Op. 10 No. 12
Fast, dense, virtuosic left-hand runs
528 notes / 30 s
Six conditions
Each piece is degraded in the ways real uploads arrive degraded — 24 scored files in total. Noise levels are solved to a target signal-to-noise ratio and verified within 0.01 dB.
Clean
Baseline render
Phone
A phone on the music stand (band-limited, compressed)
Reverb
A live room (three discrete echoes, decaying)
Noise, 20 dB SNR
A quiet room (pink noise at a solved gain)
Noise, 10 dB SNR
A noisy room
MP3 at 64 kbps
A heavily compressed upload
How a note is scored
Notes are matched with mir_eval, the standard research scorer: a detected note counts only if its pitch is within 50 cents of the reference and its onset within 50 ms. A stricter variant also requires the note to end correctly — the difference between "found the notes" and "engraved the rhythm".
Every number is measured on CPU with pinned tool versions, and the run is reproducible end to end: same MIDI sources, same ffmpeg filters, same scorer. Measurements taken this way tuned the detection thresholds the product ships with — measured first, shipped second.
What the measurement said
Findings, not scores
The shapes of the failures are publishable. The bare F1 values are not — measured on four pieces, they would pretend to describe your recording, and they cannot.
The piece matters more than the channel
Across phone, reverb, noise and compression, the model's score moved far less than it moved across pieces. A fast dense étude on clean audio is harder than a sparse invention through a phone filter. Recording-quality advice is still correct — but what is being played dominates it.
Overtone ghosts are the model's worst habit
A general-purpose pitch detector splits piano partials and invents octave doublings even on clean audio. An invented note is a wrong note you have to find; a missing note is a gap you fill in. That asymmetry is why the shipped thresholds are stricter than the model's defaults.
Noise, counter-intuitively, can help
Adding pink noise buried the model's weakest spurious detections — precision rose. We took the hint deliberately instead: the preview runs at onset/frame thresholds of 0.7/0.4 rather than the 0.5/0.3 defaults, a measured improvement at no recall cost.
Reverb, not pedal, is what smears note endings
On a genuinely pedalled piece the detector truncated sustained notes; it was room reverb that stretched them. The transcription report's warnings follow the measurement, not the intuition.
What this does not tell you
- 01Four pieces, one soundfont, one renderer. The protocol ranks models reliably; it does not characterise a stranger's recording.
- 02No amateur playing. Every test file is a metronomic render of a masterwork — users record themselves, with wrong notes and uneven tempo, and timing jitter is exactly what the benchmark cannot see.
- 03A real-recording check exists (a public-domain human performance, scored qualitatively), but there is no hand-corrected ground truth for real audio yet.
- 04Nothing here measures whether the score reads well. Note-level scoring says nothing about engraving — that is the notation pipeline's own test suite.
That gap — between a benchmark and your piano, your phone, your room — is exactly why the first 30 seconds of every transcription are free, in your browser, on your audio. How the pipeline works · Uncorrected example transcriptions
Questions, answered
A percentage measured on four chosen pieces would not survive contact with a recording we did not choose. We publish the protocol — the pieces, the degradations, the scorer — because a number without it would be marketing, not evidence.
Each detected note is compared with the reference score: a correct detection must match pitch within 50 cents and start within 50 ms. F1 combines precision (no invented notes) and recall (no missed notes). We also measure a stricter variant that requires the note to end at the right time.
The scored set is — deliberately. Rendering public-domain MIDI with a known soundfont makes the ground truth exact to the millisecond, so a miss is a real miss and never an alignment error. A public-domain human recording is checked separately, without scoring, to catch synthetic-only artefacts.
That is the honest limit of any benchmark, ours included — which is why the first 30 seconds of every transcription are free, in your browser, on your own audio. The preview is the test that cannot lie to you.
The only test that matters.
Thirty seconds of your own recording, free, in your browser — the benchmark that cannot lie to you.