AI-Assisted Royalty Dispute Identification and Documentation
Machine learning now identifies royalties owed but never paid at a scale humans cannot match.

The first thing to understand is that disputes don't announce themselves. A missing PRO registration throws no error. A wrong split doesn't trigger an alert. It just produces silence, and money that should flow somewhere simply doesn't. Absent a system actively hunting for that silence, nobody knows it's there.
The volume problem makes this structurally intractable. Platforms like YouTube, Spotify, and Apple Music generate billions of lines of streaming data, and the discrepancies that create disputes live inside all of it: wrong song titles, alternative titles that don't cross-reference cleanly, missing rightsholder names, bad split percentages. Scanning that systematically isn't a task any human team can perform, regardless of how large or well-resourced they are. I've sat across from rights teams with real budgets and real intent, and watched them catch what they could reach and miss everything else. That's not a criticism. It's the ceiling.
Then there's international fragmentation, which compounds everything in ways people underestimate until they're in it. Rights registered with one collecting society don't automatically map to another. Dozens of CMOs operate with their own identifiers and their own intake logic. A work that's correctly registered in one jurisdiction can be completely unmatched somewhere else. The documentation burden multiplies across borders until it becomes genuinely punishing.
The AI training data litigation illustrates how severe this gets at scale. Labels used audio fingerprinting to identify recordings inside AI training datasets, and an original complaint against one platform covered 560 works. When the same evidentiary process was scaled using AI tools, that number expanded to 61,026 recordings. The Recording Industry Association of America (RIAA) filed these suits against Suno and Udio in June 2024. The number of works involved isn't a variable in the problem. It is the problem.
Without automated tooling, building an evidentiary trail for a single dispute means pulling statement data manually, cross-referencing registrations across CMOs, assembling timelines by hand. For a catalog of any real size, that is not a workflow. It's a department you probably can't staff into existence even if you wanted to.
How AI systems actually identify royalty disputes across streaming data
Machine learning models do one thing here that no human team can replicate: they scan at scale and surface anomalies a reviewer would never reach. That capability gap isn't incremental. It's categorical.
Muserk, operating under the Blue Matter brand, built a patented AI system that searches billions of lines of streaming data across major platforms specifically to find royalties owed but not received. Their reported results include over $100 million recovered for songwriters and artists worldwide, with customers seeing royalties double in their first year. The company currently represents around 22 million copyrights. Their founder's stated position is that manual handling is no longer viable regardless of artist size or resources. That is a claim made by the company based on its operational experience, not an independently verified finding.
BMG, working with Google Cloud, built StreamSight for digital royalty forecasting and anomaly detection. The system flags missing sales periods, unexpected market-level fluctuations, and mismatches between reported revenue and rights ownership. Each of those flags previously required hours of manual review per instance. At the volume these platforms generate, that math breaks almost immediately without automation. Most people, even inside the industry, have not fully internalized just how fast it breaks.
Worth being specific about what detection actually means in practice, because the word gets used loosely. A useful system doesn't tell you "something is off." It tells you "this work is missing its ASCAP registration and the split shows 100% to publisher when the underlying agreement shows 70/30." That's actionable. Vague flags are just more noise, and you don't need more noise.
Turning a detected dispute into a documented claim
Detection without documentation produces nothing. A flagged discrepancy still needs a paper trail before any society or platform will act: registration proof, split history, identifier cross-references, submissions formatted to that specific CMO's requirements. Finding the problem and resolving the problem are two different workflows, and building the bridge between them is where a lot of systems quietly fail.
Claimy illustrates what the full loop looks like when it's actually closing. The system identifies missing rights and ownership gaps, prepares claims, generates documentation automatically, follows up with societies without manual intervention, and continuously verifies catalogs across societies for accurate splits, identifiers, and cross-CMO alignment. Direct registration relationships with 17 of the largest CMOs across Europe, North America, India, Australia, and South Africa mean submissions aren't generic templates. They're formatted to each society's actual intake requirements.
What that looks like on a per-claim basis is a structured packet: the specific data points each CMO requires, an audit trail showing the discrepancy and its source, and timeline records establishing when the error originated. That last element matters more than you'd expect, because retroactive pay eligibility is often where the meaningful money sits. Without the timeline, you're leaving it on the table.
Speed matters more than you might realize, too. Societies process disputes on fixed windows, and a late or incomplete submission loses its place in queue. Another full processing cycle can add weeks or months. A documentation packet assembled by AI is complete on submission, rather than still being pulled together as the window closes.
This same logic is beginning to spread beyond music. Book publishing workflows are starting to use AI-assisted tools to flag discrepancies in royalty statements and reduce the reconciliation burden on finance and rights teams. The underlying problem — statements that don't match, splits that don't align, periods that go unreported — is structurally identical across any rights-intensive business. That's either reassuring or alarming depending on where you sit.
The legal and regulatory context shaping what AI-assisted documentation must produce
The legal environment has moved fast enough that documentation standards have already changed, whether people have noticed or not. The RIAA launched lawsuits against AI music platforms Suno and Udio in June 2024. A lawsuit against Google was filed in March 2026. Statutory damages in copyright infringement cases can reach $150,000 per work under 17 U.S.C. § 504(c)(2). The expansion from 560 to 61,026 recordings in the Suno case isn't just an impressive number; it's a demonstration of why catalog-wide evidentiary assembly is only tractable with AI.
The fair use question in those cases is still moving through the courts, with summary judgment expected around April 2027. The outcome will shape how AI training licensing works across the industry for years. Early signals suggest the market isn't waiting for the ruling. Universal Music Group settled with Udio in October 2025. All three major labels licensed with Klay in November 2025. Sweden's STIM launched an AI music license in September 2025 for legal model training, with royalties flowing back to creators. The industry is already pricing in the direction this goes.
Platform enforcement is running in parallel, and some of the numbers are worth sitting with. More than half of tracks uploaded to Deezer daily are now AI-generated, roughly 90,000 songs, up from about 10% at the start of 2025, according to Deezer's public statements. Deezer launched AI detection technology in June 2025 specifically to filter fraudulent AI-generated tracks. Starting July 15, 2026, TIDAL stops royalty payments on any track identified as 100% AI-generated, making it the first major platform to move from disclosure requirements to outright demonetization. Michael Smith was charged by the U.S. Department of Justice with using AI-generated songs and bot streams to collect more than $8 million from major platforms. Apple Music flagged roughly two billion fraudulent streams in 2025, according to the company's reported figures. Every fraudulent stream claiming a share of a fixed payout pool thins every legitimate artist's share. That's not abstract.
The Copyright Office's January 2025 ruling added another layer: AI-generated work can be copyrighted when it embodies meaningful human authorship. That ruling creates a new documentation requirement for human creators to establish authorship provenance, which is itself a systematic tracking problem. The compliance infrastructure this requires hasn't been fully worked out yet.
The cumulative implication is that dispute documentation isn't administrative anymore. It's evidentiary. The standard for what constitutes adequate documentation is rising, and it will keep rising as litigation works through the courts.
Why the scale of this work rules out traditional staffing approaches
Adding headcount doesn't change the underlying math when the volume of data exceeds what any team can process in time for disputes to be actionable. The ceiling is structural.
Walk through what your team actually does for each dispute: pull streaming data across multiple platforms, cross-reference registration records across CMOs, identify the discrepancy type with enough specificity to support a claim, build a documentation packet to each society's submission requirements, then track and follow up through processing windows that typically run six to twelve weeks, longer for complex international or multi-split disputes. That's the workflow for one dispute. A catalog of meaningful size has not dozens of these situations but potentially thousands, sitting unresolved because no one has the cycles to touch them.
Each traditional alternative hits the same ceiling differently. A royalty administrator is slow to ramp, expensive, and still working with manual tooling that can't cover billions of data points. A consultant analyzes the problem and hands back a report, which is useful context but doesn't build a system or own an outcome. Off-the-shelf software flags some anomalies but typically won't generate documentation, follow up with societies, or adapt to catalog-specific patterns over time.
The alternative that actually scales is an embedded AI engineer: someone who builds and operates AI systems inside your workflow, knows your catalog, understands the CMO relationships, and iterates as the system learns. Job postings for forward-deployed engineers grew 729% year over year between April 2025 and April 2026, according to Indeed data reported by Business Insider, from 643 postings to 5,330. That growth reflects something real about what organizations are learning when they try to solve operationally complex problems through other means.
A 2025 MIT study found that roughly 95% of enterprise generative AI pilots showed no measurable impact on profit and loss. Researchers traced the failure to flawed integration, not weak models. A system built on the wrong integration model — disconnected from real catalog data, misconfigured for the relevant CMOs, lacking a feedback loop from outcomes back into detection — will produce activity without revenue. Model quality is not the variable. Deployment is.
What a working AI royalty dispute system looks like when it's operational
A properly built system is not a one-time audit. It runs continuously against live streaming data and registration records, finding discrepancies as they emerge rather than accumulating a backlog between review cycles. Disputes have time windows, and a system that batches its work loses money on timing alone.
On an ongoing basis: it monitors streaming data across platforms for unreported or anomalous activity, checks registrations across CMOs for identifier mismatches and missing splits, flags new discrepancies as they appear, generates and submits documentation without manual assembly, and tracks dispute status through processing windows with automatic follow-up. The workflow doesn't stop between submissions. That continuity is the point.
The output is revenue, not just data. Dispute identification that doesn't reach resolution is overhead. The 422% cumulative royalty growth over three years reported by Muserk customers is the relevant measure, not the number of discrepancies surfaced. One number means something. The other is just activity.
There's a compounding effect that builds quietly over time, and it's the part that's hardest to explain until someone has watched it happen. As the system processes more disputes and learns catalog-specific patterns, detection accuracy improves. The value grows without proportional increases in cost, which is a fundamentally different economics than staffing, where cost scales with volume. Past a certain catalog threshold, the embedded system model doesn't just outperform the department model. It's not a close comparison.
Building it correctly requires a few things you can't skip: integration with your actual catalog data and existing systems, configuration for the CMOs and platforms relevant to your catalog, documentation templates matched to each society's submission requirements, and a feedback loop from dispute outcomes back into the detection model. A system missing any of those elements will produce diminishing returns faster than you'd expect, usually within the first few months. I've seen it happen more than once.
That's what Genius Bar is built around: embed an AI engineer, define one measurable outcome in recovered royalties, ship a working system in weeks, and let the compounding do its work. No six-month hiring cycle. No consultant who hands back a report and exits. The engineer stays in the workflow, inside the actual business, iterating against real data until the number moves. The $424 million sitting unmatched in the system didn't accumulate because the industry lacked intent to fix it. It accumulated because fixing it at scale was structurally impossible without the right tooling, and for a long time, that tooling didn't exist.


