No such thing as “Close Enough”
Ask any ensemble director what "close enough" sounds like in a triad, and the honest answer is that it doesn't exist. A chord that's 90% in tune still makes an audience wince. Music doesn't grade on a curve the way most fields do. A typo in an email gets skimmed past. A rounding error in a spreadsheet rarely changes the bottom line. A wrong note gets heard, immediately, by anyone who has spent years training an ear to catch it.
So when I hear "AI can read sheet music now," my first question isn't whether it works. My question is how close to 100%, because in this field, anything less isn't a shortcut. It's a demo.
This past weekend, MusEdLab ran the actual numbers on that question, using the most literal, foundational task there is: can an AI system look at a page of sheet music and correctly transcribe what's on it. Here's what we found, plainly, no spin.
We tested two dedicated tools built specifically to read sheet music, homr and audiveris, against 685 pages of real, professional repertoire. (BTW: huge thank you to Juan Carlos Martinez Sevilla from the Universidad de Alicante for allowing us to use his SMB dataset) Researchers score this kind of thing using a metric called Normalized Edit Distance: essentially, how much correction it would take to turn the AI's version into the correct one. Zero means a perfect match. The stronger of the two tools, homr, averaged around 0.38 across the full set, and climbed to 0.45 on the harder pieces we tested separately. Audiveris averaged worse, around 0.49, and failed to produce a usable result at all on roughly 1 in 6 pages. Neither number is close to zero. Neither tool is "basically perfect" yet, even on clean, professionally engraved pages.
Then we tried something more telling. We skipped the dedicated tool entirely and asked Claude, the general AI model that powers the rest of MusEdLab, to just read the page and transcribe it directly, the way a lot of "AI reads music now" claims imply. It was dramatically worse: an average error score of 0.925, more than double homr's, about as close to unusable as this metric gets. It also failed to produce a usable transcription at all on about 1 in 6 pages. When it went wrong, it most often lost track of where one measure ended and the next began, the kind of exact structural bookkeeping a general AI model isn't built to hold onto over a dense page.
The number that matters most isn't a research metric. It's what shows up in front of a director. We tested MusEdLab's real rehearsal-prep tool, the one that reads a score and hands back teaching notes, against ground truth on a set of foundational facts: key signature, time signature, measure count, instrumentation. Key signature was correctly identified less than half the time, 46.2%, in our test set. That number deserves a caveat: the pilot set included a few pieces without a conventional fixed key, which drags the average down in a way that isn't purely a failure of the AI, so treat the exact figure as directional, not gospel. But even generously read, it's nowhere near a number you'd trust without checking. Overall accuracy across those four foundational fields, in the tool as it actually runs in production today, landed at 60.3%.
One more, harder to explain away: when the tool cited a specific measure number in its notes on a challenging passage, that measure number didn't actually exist in the piece about 1 in 14 times. That's a fabricated fact stated with total confidence, the kind a director without the score in front of them would have no way to catch.
The right response to numbers like these is honesty about exactly where the line sits today. MusEdLab would rather publish that line than pretend it isn't there. This is a snapshot of where the tool stands as of this weekend, and we intend to keep moving that number.
The instinct a lot of music educators carry, that this technology threatens to replace a trained ear before it has earned the right to, is one our own research treats as a legitimate concern, not something to be talked out of. Given where these numbers actually sit, that instinct is the correct read right now.
That's exactly why every tool we build treats a director's judgment as the final check, not an optional one. AI can draft a starting point, flag a passage worth a closer look, or save an hour of manual data entry. It cannot yet be trusted to get the foundational facts right on its own, and until the numbers say otherwise, we're not going to pretend it can.
Closing that gap is the actual work now. Everything above was tested against professional concert repertoire, chosen because it's a well-established research benchmark, not because it looks like what's actually on a teacher's stand. Our next benchmark uses real K-12 material instead: elementary general music, beginning band and orchestra arrangements, middle and high school choir octavos, across every grade level we serve. That's the material this has to work on to matter, and it's the only dataset that will tell us how much closer we actually are to the number that counts.
Comments
No comments yet — be the first to share your thoughts.