A group of authors filed a lawsuit against Meta, alleging the unlawful use of copyrighted material in developing its Llama 1 and Llama 2 large language models....
Meta has acknowledged using parts of the Books3 dataset but argued that its use of copyrighted works to train LLMs did not require "consent, credit, or compensation." The company refutes claims of infringing the plaintiffs' "alleged" copyrights, contending that any unauthorized copies of copyrighted works in Books3 should be considered fair use.
Furthermore, Meta is disputing the validity of maintaining the legal action as a Class Action lawsuit, refusing to provide any monetary "relief" to the suing authors or others involved in the Books3 controversy. The dataset, which includes copyrighted material sourced from the pirate site Bibliotik, was targeted in 2023 by the Danish anti-piracy group Rights Alliance, demanding that digital archiving of the Books3 dataset should be banned and is using DMCA notices to enforce those takedowns.
What sort of crack are they on that they think unauthorized use of an entire work for commercial gain is fair use? I think copywrite laws are ridiculous but that is a pretty low bar they are trying to set.
They should have to pay for their usage or retrain the model without it. Going to guess they would prefer to pay up.
Training an AI does not involve copying anything so why would you think that fair use is even a factor here? It's outside of copyright altogether. You can't copyright concepts.
Downloading pirated books to your computer does involve copyright violation, sure, but it's a violation by the uploader. And look at what community we're in, are we going to get all high and mighty about that?
Actually it does. It involves making use of a copy that is not the original. Fair use is about experiencing media for sake of dialog (criticism or parody) or for edification. That means someone is reading the book or watching the movie, or using it for transformative art or science.
Remember when media companies tried to sue switch manufacturers because their routers held copies of packets in RAM and argued they needed licensing for that?
Training an AI can end up leaving copies of copyrightable segments of the originals, look up sample recover attacks. If it had worked as advertised then it would be transformative derivative works with fair use protection, but in reality it often doesn't work that way
Remember when piracy communities thought that the media companies were wrong to sue switch manufacturers because of that?
It baffles me that there's such an anti-AI sentiment going around that it would cause even folks here to go "you know, maybe those litigious copyright cartels had the right idea after all."
We should be cheering that we've got Meta on the side of fair use for once.
look up sample recover attacks.
Look up "overfitting." It's a flaw in generative AI training that modern AI trainers have done a great deal to resolve, and even in the cases of overfitting it's not all of the training data that gets "memorized." Only the stuff that got hammered into the AI thousands of times in error.
Yes, but should big companies with business models designed to be exploitative be allowed to act hypocritically?
My problem isn't with ML as such, or with learning over such large sets of works, etc, but these companies are designing their services specifically to push the people who's works they rely on out of work.
The irony of overfitting is that both having numerous copies of common works is a problem AND removing the duplicates would be a problem. They need an understanding of what's representative for language, etc, but the training algorithms can't learn that on their own and it's not feasible go have humans teach it that and also the training algorithm can't effectively detect duplicates and "tune down" their influence to stop replicating them exactly. Also, trying to do that latter thing algorithmically will ALSO break things as it would break its understanding of stuff like standard legalese and boilerplate language, etc.
The current generation of generative ML doesn't do what it says on the box, AND the companies running them deserve to get screwed over.
And yes I understand the risk of screwing up fair use, which is why my suggestion is not to hinder learning, but to require the companies to track copyright status of samples and inform ends users of licensing status when the system detects a sample is substantially replicated in the output. This will not hurt anybody training on public domain or fairly licensed works, nor hurt anybody who tracks authorship when crawling for samples, and will also not hurt anybody who has designed their ML system to be sufficiently transformative that it never replicates copyrighted samples. It just hurts exploitative companies.
There actually isn't a downside to de-duplicating data sets, overfitting is simply a flaw. Generative models aren't supposed to "memorize" stuff - if you really want a copy of an existing picture there are far easier and more reliable ways to accomplish that than giant GPU server farms. These models don't derive any benefit from drilling on the same subset of data over and over. It makes them less creative.
I want to normalize the notion that copyright isn't an all-powerful fundamental law of physics like so many people seem to assume these days, and if I can get big companies like Meta to throw their resources behind me in that argument then all the better.
Humans learn a lot through repetition, no reason to believe that LLMs wouldn't benefit from reinforcement of higher quality information. Especially because seeing the same information in different contexts helps mapping the links between the different contexts and helps dispel incorrect assumptions. But like I said, the only viable method they have for this kind of emphasis at scale is incidental replication of more popular works in its samples. And when something is duplicated too much it overfits instead.
They need to fundamentally change big parts of how learning happens and how the algorithm learns to fix this conflict. In particular it will need a lot more "introspective" training stages to refine what it has learned, and pretty much nobody does anything even slightly similar on large models because they don't know how, and it would be insanely expensive anyway.
Especially because seeing the same information in different contexts helps mapping the links between the different contexts and helps dispel incorrect assumptions.
Yes, but this is exactly the point of deduplication - you don't want identical inputs, you want variety. If you want the AI to understand the concept of cats you don't keep showing it the same picture of a cat over and over, all that tells it is that you want exactly that picture. You show it a whole bunch of different pictures whose only commonality is that there's a cat in it, and then the AI can figure out what "cat" means.
They need to fundamentally change big parts of how learning happens and how the algorithm learns to fix this conflict.
Cranky enough to demand satisfaction (in the courts if not the dueling field), but no one in the company will think their own ire warrants empathy for those from whom they pirate.