Someone actually made a searchable database of all the music used to train AI models

Someone actually made a searchable database of all the music used to train AI models

7 0 0

The Atlantic’s Alex Reisner did something that’s been sorely needed: he dug up four major music datasets used to train AI models and made them fully searchable for the public. No more guessing what’s in the black box.

Two of these datasets are absolutely massive. We’re talking 12 million tracks and 9 million tracks respectively. The other two are smaller but still pack over 100,000 songs each. That’s a lot of music being fed into algorithms, and according to Reisner, these sets have been downloaded thousands of times.

It’s impossible to know exactly who has used them, but Google and Stability AI have both confirmed in research papers that they’ve pulled from these datasets. That’s not surprising—Stability has been all-in on generative audio, and Google’s music AI projects have been quietly chugging along.

Some of the sources, like the Free Music Archive dataset, are free to stream for personal use but have ambiguous licensing when it comes to training commercial AI. That’s the kind of gray area that keeps lawyers busy and artists angry. I’ve been saying for a while that the lack of transparency in training data is the real scandal here, not the tech itself. This database doesn’t solve the legal mess, but it at least gives creators a way to see if their work was used.

A smiling computer surrounded by music notes.

The database is searchable by artist, track, and dataset. So if you’re a musician wondering whether your stuff ended up in a training set, you can actually check now. That’s a huge step forward from the usual “trust us” approach from AI companies.

I’d love to see more of this kind of transparency across the board—image, text, video. But for now, at least music has a starting point. Read the full story at The Verge if you want the nitty-gritty on how Reisner pieced this together.

Comments (0)

Be the first to comment!