Skip to content
Piano to Sheet Music

Methodology

The protocol, not a percentage.

This is how transcription quality is measured before anything on this site claims it: what the test set is, how it was degraded, how notes are scored — and what the measurement cannot tell you about your recording.

Ground truth that is exact

Comparing a transcription with a recording has an alignment problem: the ground-truth score and the audio rarely agree on exactly when a note began. We removed the problem rather than arguing with it — the test audio is rendered from the reference score, so the ground truth is exact to the millisecond and a miss is a real miss.

The material is public-domain piano MIDI from Mutopia, rendered with FluidSynth and the FluidR3 GM soundfont at 44.1 kHz, cut to 30-second windows. The trade is honest: a sampled-GM render is not a living-room recording, so a real human performance — Edward Neeman's public-domain recording of the Chopin étude — is checked alongside, qualitatively, to catch synthetic-only artefacts.

Bach, Invention No. 1, BWV 772

Sparse two-voice counterpoint — the easy case

220 notes / 30 s

Bach, WTC I Prelude No. 1, BWV 846

Steady broken-chord polyphony, no pedal

120 notes / 30 s

Satie, Gymnopédie No. 1

Slow, heavily pedalled

50 notes / 30 s

Chopin, Étude Op. 10 No. 12

Fast, dense, virtuosic left-hand runs

528 notes / 30 s

Six conditions

Each piece is degraded in the ways real uploads arrive degraded — 24 scored files in total. Noise levels are solved to a target signal-to-noise ratio and verified within 0.01 dB.

Clean

Baseline render

Phone

A phone on the music stand (band-limited, compressed)

Reverb

A live room (three discrete echoes, decaying)

Noise, 20 dB SNR

A quiet room (pink noise at a solved gain)

Noise, 10 dB SNR

A noisy room

MP3 at 64 kbps

A heavily compressed upload

How a note is scored

Notes are matched with mir_eval, the standard research scorer: a detected note counts only if its pitch is within 50 cents of the reference and its onset within 50 ms. A stricter variant also requires the note to end correctly — the difference between "found the notes" and "engraved the rhythm".

Every number is measured on CPU with pinned tool versions, and the run is reproducible end to end: same MIDI sources, same ffmpeg filters, same scorer. Measurements taken this way tuned the detection thresholds the product ships with — measured first, shipped second.

What the measurement said

Findings, not scores

The shapes of the failures are publishable. The bare F1 values are not — measured on four pieces, they would pretend to describe your recording, and they cannot.

The piece matters more than the channel

Across phone, reverb, noise and compression, the model's score moved far less than it moved across pieces. A fast dense étude on clean audio is harder than a sparse invention through a phone filter. Recording-quality advice is still correct — but what is being played dominates it.

Overtone ghosts are the model's worst habit

A general-purpose pitch detector splits piano partials and invents octave doublings even on clean audio. An invented note is a wrong note you have to find; a missing note is a gap you fill in. That asymmetry is why the shipped thresholds are stricter than the model's defaults.

Noise, counter-intuitively, can help

Adding pink noise buried the model's weakest spurious detections — precision rose. We took the hint deliberately instead: the preview runs at onset/frame thresholds of 0.7/0.4 rather than the 0.5/0.3 defaults, a measured improvement at no recall cost.

Reverb, not pedal, is what smears note endings

On a genuinely pedalled piece the detector truncated sustained notes; it was room reverb that stretched them. The transcription report's warnings follow the measurement, not the intuition.

What this does not tell you

  • 01Four pieces, one soundfont, one renderer. The protocol ranks models reliably; it does not characterise a stranger's recording.
  • 02No amateur playing. Every test file is a metronomic render of a masterwork — users record themselves, with wrong notes and uneven tempo, and timing jitter is exactly what the benchmark cannot see.
  • 03A real-recording check exists (a public-domain human performance, scored qualitatively), but there is no hand-corrected ground truth for real audio yet.
  • 04Nothing here measures whether the score reads well. Note-level scoring says nothing about engraving — that is the notation pipeline's own test suite.

That gap — between a benchmark and your piano, your phone, your room — is exactly why the first 30 seconds of every transcription are free, in your browser, on your audio. How the pipeline works · Uncorrected example transcriptions

Questions, answered

A percentage measured on four chosen pieces would not survive contact with a recording we did not choose. We publish the protocol — the pieces, the degradations, the scorer — because a number without it would be marketing, not evidence.

The only test that matters.

Thirty seconds of your own recording, free, in your browser — the benchmark that cannot lie to you.

Transcribe a recording