Key takeaways
The three points that matter most
PDF to MusicXML recognition benchmark
20 of 20 routing expectations matched
The checked-in 2026-06-30 regression report correctly kept clean, difficult, and unusable repository samples in their expected candidate/final lanes.
PDF to MusicXML recognition benchmark
This is not note accuracy
The regression result measures workflow classification, not aligned pitch, rhythm, key-signature, or lyric correctness.
PDF to MusicXML recognition benchmark
Four percentages remain unpublished
We will not infer note, rhythm, key, or lyric accuracy until a versioned gold set and alignment report exist.
Explanation 01
Evidence currently available
The repository contains a 20-file regression report with 20 matched expected outcomes: 11 clean/final, 6 difficult/draft, and 3 unusable/non-final cases. It also contains checksum-pinned real-engine fixtures for multi-page, rotation, crop, degraded, corrupted, timeout, and cancellation scenarios, plus a recorded Audiveris 5.10.2 runtime qualification dated 2026-08-17.
Those artifacts prove that explicit workflow and engine gates exist. They do not justify a claim that a given percentage of notes or rhythms is correct. The downloadable JSON records this distinction in machine-readable form.
- Candidate-routing regression: measured and versioned
- Runtime qualification: measured and versioned
- Pitch accuracy: not published
- Rhythm error rate: not published
- Key-signature retention: not published
- Lyric retention: not published
Explanation 02
Gold-set protocol for publishable accuracy
A future recognition percentage must identify the exact engine version, corpus license, scan matrix, alignment algorithm, sample count, and exclusions. Results must be reported separately by scan condition and notation complexity rather than averaged into a promotional number.
- Create human-verified MusicXML references and immutable source PDFs.
- Run clean 300 dpi, lower-resolution, rotated, cropped, shadowed, lyric, piano, and dense-voice conditions.
- Align parts, measures, voices, and events before scoring pitch or duration.
- Report note precision/recall, duration edit rate, measure structure, key/time retention, lyrics, and unsupported symbols separately.
- Publish failures, engine version, checksums, scripts, and reviewer disagreements.
Explanation 03
How to read this report
A blank metric means evidence is insufficient, not that performance is zero. A successful job means an artifact was produced; only aligned comparison with a human-reviewed reference measures recognition correctness. Users should still test a representative page and inspect the candidate.
Step-by-step
Follow this order
Download the evidence JSON
Review the scope, source files, dates, counts, and explicit non-claims.
Download the public PDF and MusicXML
Inspect the open sample pair without an account.
Run the product on a permitted score
Test a representative page and retain the source beside the candidate.
Compare measure by measure
Review structure, rhythm, voices, pitch, key, lyrics, and repeats.
Do not generalize one sample
Report results by source condition and notation type rather than as a universal percentage.
FAQ
Frequently asked questions
Why not publish 99% accuracy?
Because no aligned, versioned gold-set evidence supports such a universal claim across scan qualities and notation types.
Does 20/20 mean every note was correct?
No. It means each regression sample followed its expected final/draft/non-final workflow outcome.
When will note and rhythm percentages appear?
Only after the gold-set runner produces an auditable report with aligned references, engine version, checksums, and reviewer-approved scoring.
Verification sources
Primary and upstream sources
These links verify formats, software behavior, and public limitations; they do not imply commercial endorsement of ScoreTransposer.
Continue reading