How did 3 datasets quietly scrape millions of songs for AI training?
Mapping the unlicensed music AI datasets
STVDIO maps the future of music with deep visual reports. Helping 4,900+ artists, music industry professionals and founders navigate music industry shifts with weekly visual maps.
Supported by Secretly Distribution: For more than 25 years, Secretly Distribution (SD) has been enabling independents to challenge the mainstream. SD is not only a full-service global physical and digital music distributor, but it also provides hands-on marketing, financial and technological support to some of the most exciting record labels and artists across geographies and genres. By supporting businesses who are self-reliant by choice, not just necessity, SD champions the musicians and creative works that drive culture forward. Listen to the latest SD releases on the Secretly Weekly playlist.
Mapping the AI music datasets
Last week the Atlantic revealed several unlicensed datasets used to train some of the large AI music models.
Most of us in the music industry instinctively knew this was happening. We knew that music was being fed into the AI machine without permission. But I didn’t fully appreciate the scale and the detail of what was happening.
If you’ve only read the headlines, I encourage you to keep reading and look at exactly how these datasets were built and fed into mass-extraction song-rippers.
Like many artists, I found my own music in two of these datasets, and even though I’m excited for AI there’s something unsettling about reading through them and the cold disregard for music and copyright.
Today we’ll go through them in detail.
Note: not all AI music companies use these datasets. There are plenty of brilliant ethical AI music generation tools that have licensing agreements and fair compensation for artists. I’ll list some of them at the end.
First … how does it work?
It’s important to understand the datasets do not actually contain music. Here’s how it works:
A giant spreadsheets of metadata: artist names, song names, lyrics and corresponding links to the song on either YouTube or Spotify.
AI companies and developers then feed these spreadsheets into mass-extraction software to rip the audio.
The extracted audio is used to train the models.
There’s an important legal difference between each one … which we’ll get into later. First let’s go through the datasets and who made them.
1. LAION DISCO 12M 🟠
Made by: LAION
A broad dataset of 12 million links to music on YouTube.
Released for “academic research”
But widely downloaded by commercial users, then fed into audio extraction software.
LAION DISCO is the largest dataset uncovered by the Atlantic. It’s a huge document of metadata: 12 million artist names, song titles and links to YouTube.
This LAION DISCO dataset is designed to capture music at scale, by scraping YouTube’s “related artists” algorithm.
Who made it?
LAION is a German non-profit research organisation. They believe in open source datasets to ensure that AI isn’t controlled by a few big companies.
LAION defends the dataset by saying it’s only a list of links for “research purposes” … not to be used for “creating end products.”
Yet at the same time, it’s fairly clear they knew how this dataset would be used:
2. Sleeping DISCO 9M 🟠
Made by: Sleeping AI
Contains links to 9 million songs on YouTube.
Lyrics scraped from Genius.
Hyper-curated to capture “popular and well known songs”.
Sleeping DISCO 9M is similar to the first dataset, in that it contains metadata and links to YouTube. But this one is optimised to capture the most popular songs and relevant artists.
In other words, it’s designed so that AI music models can mimic and recreate “world-renowned” artists and their sounds. It also scraped lyrics from Genius.com so that AI models can learn from the lyrics of famous songs.
Who made it?
Sleeping AI. They are a non-profit research lab, self described “hackers who provide datasets.” Like LAION, the group said the dataset is for research purposes and should not be used for commercial products.
However, there are plenty of points in their paper where they clearly acknowledge how their dataset would be used:
1. To rip YouTube audio:
2. To build commercial models:
Notably: Team members from SleepingAI also contributed to LAION so these two teams have close links. The teams are fully public on their websites, on the published papers and on LinkedIn. Since Atlantic published their article last week, Sleeping AI has shut down their website.
3. The “Private Pointer” 🔴
Made by: Unknown / anonymous
A smaller list of 100,000 Spotify and YouTube links to popular songs.
Created anonymously and circulated among developer forums
This dataset is smaller but in some ways more egregious than the previous two. There’s no “research organisation” behind it or ethical pretence of open source data. It’s purely a black-market training document passed around specifically to build unlicensed AI music models.
The dataset contains Spotify and YouTube links that can be fed into a mass-extraction software to rip the audio.
Who made it?
We don’t know. It’s an anonymous dataset which may have numerous contributors as it was passed around through forums and developer channels.
4. The Free Music Archive 🟢
Made by: WFMU
A creative commons dataset of ~100,000 songs.
Used by Google and Stability to train its music models.
This one should not be viewed in the same way as the others. It’s a creative commons dataset of royalty-free music built with clean intentions. Artists have submitted their music and made it available to stream, share and use under a CC license.
However, AI labs have then downloaded and used this “royalty-free” library to train its models. We don’t yet know if this counts as a breach of terms.
Who made it?
A New Jersey radio station called WFMU. They started it in 2009 to create a library of legal, royalty-free, downloadable music for podcasters, filmakers and music fans. The downloads come with limits on commercial use and artists are supposed to be credited.
How is this allowed?
Organisations like SleepingAI and LAION are open and public about these datasets. They publish their scraping methodology in academic papers, with names attached.
So, how is allowed?
Both organisations claim the datasets are for “research” purposes.
They’re not ripping the songs directly. They’re just aggregating metadata that’s already public (song names and links).
They include disclaimers to stress that their data shouldn’t be used in products or “commercial purposes.”
They lean on EU regulations that allow scraping and copying data for research purposes.
They’ve created a layer of plausible deniability, which shifts the blame onto AI developers or companies that actually rip the audio and train the models.
The argument: making a list of links is not illegal, but ripping the music might be.
Reminder: fully licensed AI music does exist!
It’s important to say: not all AI music is pirated.
There are many, many great tools that have not used these unlicensed databases. Some have cut deals with independent organisations or major labels. Some have built their own models using music contributed by artists. Some have properly licensed royalty-free libraries, opt-ins, and compensate artists when their work is used.
ElevenLabs - Fully licensed training model via independent label and publisher agreements.
Mozartai.com - Fully licensed AI music generation built using ElevenLabs’ model.
Ace-Step - Uses a combination of licensing agreements and copyright-free music libraries.
Stable Audio - Licensing agreements with Warner, Universal (+ the Free Music Archive library).
We recently mapped a full list of AI music tools and mapped the current licensing deals if you want to explore further.
Thanks for reading
If you found this useful, please do forward it to your colleagues or share with your friends. If you have questions or thoughts, reach out to me on LinkedIn or by email at benjamin@stvdio.io. See you next week with another music industry report.










