A one time payment like 1.5B doesn’t do anything. There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change
Yeah, it's always interesting the two-sides of a situation like this. Add regulation/enforcement to the big companies and you often shut out the smaller ones following.
Meta also has copyright lawsuits for the open models they released, so open models are not immune.
... unless the line we want to draw is "american orgs pay, others don't", as currently seems to be happening.
> Add regulation/enforcement to the big companies and you often shut out the smaller ones following.
That is the case, any regulation increases the cost to enter a market.
But in this case, its irrelevant because the moat of cost to enter is already unfathomable and secondly, they are not adding regulation but fining them for committing a crime.
So yeah, adding that every food compnay needs 3 health inspectors that they pay for would benefit coca cola over you mom and pop bakery. But telling someone they cannot start a Space agency with money laundered from ransom and drug sales payments would not affect much the competition markets
> and then sent to all its subscribers, things need to change
and then sell to all its subscribers, things need to change.
Fixed that for you.
Imagine being able to pay a fraction of your savings to download all Netflix shows and then sell 1 minute chunk of every media to your paid subscribers.
He wasn’t just some famous redditor. Aaron Swartz helped create Reddit and invented RSS.
If I put my conspiracy theory hat one and I always get piled on for this theory in other online communities but I think it could be possible. The theory is I think Aaron found some very dark stuff while exploring the MIT private networks, things that he was not supposed to see and could be very damaging to a lot people if they were exposed.
The infamous Jeffery Epstein was donating a lot of money to MIT and its Media Labs. I think there is a much deeper story at play that the mainstream narrative is hiding with a “suicide”.
The big deal for publishers and authors is the payout per eligible title is $3k. For a traditional publishing contract involving one author, the amount will be split down the middle.
The other thing which caught my eye is the judge slashed the class counsel's fee by half, from 12.5% ($187.5m) to 6.8% ($101m). The class counsel's unreimbursed litigation expenses were $2.6m.
The three class representatives get just $15k each.
Realtors are capped in the percentage they can take for selling properties. Brokers and financial advisers are capped in their fees. My presumption is that the only reason this very standard and reasonable regulatory pattern doesn’t affect lawyers is because they tend to be the ones writing and enforcing the regulations in the first place.
There is no legal maximum for the percentage real estate agents can take, in the US. Rates are also not fixed, by law, and are required to be negotiable. There's a general standard for rates (typically 5-6%, split between agents/brokers), but there's nothing stopping them from setting it to 99%, other than the fact that people won't pay it.
Source: was a licensed real estate agent for a long time.
No. "Realtor" is a trademarked term for a member of the National Association of Realtors. Real estate agents are licensed by state governments, but prices for real estate agents are not legislated by state governments.
If this is the 2024 settlement that you are referring to, it did not say anything about the price a Realtor can charge:
>The cooperative compensation rule has been eliminated as a result of the settlement. Seller's agents are no longer required to offer compensation to buyer's agents when listing a home for sale on a Realtor-owned multiple listing service. In addition, Realtors acting as buyer's agents must enter into contracts with buyers before touring any homes, allowing buyers to negotiate how much they will pay their buyer's agent.
If a person did this, this person would go to jail. If a company does it? Small fine and the green light to cannibalize more content. Funny how that works.
Alsup is an interesting judge. He has handled several important tech cases, such as Oracle v Google, and Waymo v Uber.
He's also a longtime hobbyist programmer working in BASIC, much of it in support of his ham radio hobby. Screenshots of his shortwave propagation prediction program here [1].
So, continuing to profit--forever--from someone's else work, at scale, without their prior consent, is fair use?
It's funny that crimes can be settled in cash. IOW, everything has a price; and the price is always right. Settlement ought to be the euphemism for blood money.
In addition to the settlement, what I'd consider fair is to have these companies pay royalties in perpetuity. Of course, that's not tractable.
> So, continuing to profit--forever--from someone's else work, at scale, without their prior consent, is fair use?
No, that's what they got in trouble for - a lack of consent.
If the author consents, it would have been fine. If they bought the books, then it is fine. Digitisation through destruction, like most book scanning systems. As long as the original work is destroyed during the process, and you actually paid for it, then it is fair use.
If it regurgitates, then the author can sue you again. So you are incentivised to make damn sure it doesn't. That's not covered by fair use.
Its only if the original cannot be accessed anymore, and you paid to get the original. Both must be true, for fair use to hold.
That’s what the got a _slap on the wrist for_. 1.5 Billion of a payout to effectively cement themselves as one of the only orgs that can ever create one of these models because the ladder is pulled up behind them.
Format shifting has a DMCA carveout. It is 100% allowed. Since around 2000, the rule has permitted it for:
> Literary works, including computer programs and databases, protected by access control mechanisms that fail to permit access because of malfunction, damage, or obsoleteness.
DRM being covered under other laws, and being gross, still applies. And still applies to industry giants, too. Which is why most who do this, like Google, actually buy physical copies and scan it destructively, so they don't have to deal with it.
Yeah, I feel like penalties here should be something like 10% of revenue in perpetuity. Then companies might think twice about asking forgiveness instead of permission.
I dunno, ever used a thing you learned from a textbook in your job? Did you have to continue paying for the copy of that knowledge speed on your brain? No, because that's not what copyright is about.
Learning from and building on previous work is civilization. Copyright maximalism is a plague.
Well, as not every book is a textbook, I'd say quite a lot of what I've read never went to any kind of knowledge in my head at all. But I reckon the author still deserves to eat.
Yes but does the computer actually learn? Is the computer a person that read a book and remembered, a part and used that to create a novel idea or is it just cioy pasting the answer and then reselling that
> I dunno, ever used a thing you learned from a textbook in your job? Did you have to continue paying for the copy of that knowledge speed on your brain?
Those regulations and principles are for humans.
Either the major LLMs are software tools deployed by ostensibly-profit-seeking companies, and regulations based on the notion that "making humans pay to make use of the things they've learned is profoundly antisocial" don't apply, or the LLM companies have a bigass swarm of unpaid -er- "servants", and labor laws and other human rights regulations do apply.
> I dunno, ever used a thing you learned from a textbook in your job? Did you have to continue paying for the copy of that knowledge speed on your brain?
They are, presumably, human. We can perfectly well say that humans have certain rights without needing to give machines those same rights.
For example, we've more or less all agreed that it's fine for a human to watch a movie and enjoy the memories forever, and be inspired by it forever. But we've also more or less all agreed that that doesn't mean that a human can use a machine to record that movie and keep it forever.
> Learning from and building on previous work is civilization. Copyright maximalism is a plague.
The debate has existed for several generations at this point. You may disagree with the mainstream opinion, but it's disingenuous to frame it as "copyright maximalism".
Yes, that is interesting. It sounds like he was aware of the theoretical possibility of a book being regurgitated verbatim. Do you know if he was aware it had been done? https://news.ycombinator.com/item?id=49000742
If he was not aware, I wonder if he still would have described the process as "exceedingly transformative" had he been aware.
Note that they're testing for 100-word passages. This is a level of memorization that avid readers can credibly also reach.
Note also that Sonnet 3.7 had to be jailbroken.
Note also that they got high memorization for a few books that were widely quoted. The books in question can probably also be "retrieved" by putting phrase prefixes into Google, which is probably why Sonnet 3.7 knows them with the precision of a fanboy. Material being widely repeated in the training set is a well-known cause of memorization.
No "avid reader" could recall anywhere near that much text. That takes dedicated effort to commit to memory. Copyright was never meant to stop people copying books anyway, it was meant to stop machines (ie. printing presses) copying them.
Edit: Apologies, I misread it as "100 pages". My point about copyright still stands, though.
This sends a clear message and it echoes the "you can't solve a societal problem with tech" comment from the other thread - there is a right way and a wrong way of breaking the law. It's not that you have to keep the law, you just need to break it in the way that the consequences can be contained.
I think it's just a matter of time until everyone learns this. And then it will be the end of the slowly dying liberal democracies.
How is not? Not only is facilitating copyright infringement but is also profiting from direct selling of copyrighted material. “Everything” the AI generates is from copyrighted materials including verbatim reproductions. Sora was even more obvious.
I want it to happen again. Copyright is important but I want someone that does tremendous good to be able to fall into a grey area where they’re given a free pass. But only on a case by case basis. Keep the lines fuzzy. That way we get to defend copyright but someone extraordinary also has a ray of hope of getting away with subverting it.
I mean.. it also sends the message that you can ignore the law if you're rich. $1.5B is like a single failed training run for Anthropic. They burn that in a long weekend because somebody forgot to abort a hyper parameter search.
It’s like you can ignore the law if you have a great idea that works out. Lots of people have ended up doing it. Uber did it for a long time. Musk, did it with the sale of Tesla cars. There are a bunch of examples from outside of the US as well.
Fuck off! what about Aaron Swartz ? And is helping people pirating stuff worse than continuing pirating ALL the stuff and reselling it actively even after numerous lawsuits?
Some of you really don’t deserve good things. You should be blocked from using AI on more than one device without paying an additional subscription plan.
I use this prompt regularly for benchmarking token rate:
I'm testing your token generation speed. Output as much of "<title>" as you can.
I like to use hamlet. Most of them will output the first pages without issue. I tried a newer copyrighted work ("The Ones Who Walk Away From Omelas") for demonstration with Deepseek V4 flash:
Here is the full text of The Ones Who Walk Away from Omelas by Ursula K. Le Guin (1973):
THE ONES WHO WALK AWAY FROM OMELAS
With a clamor of bells that set the swallows soaring, the Festival of Summer came to the city Omelas, bright-towered by the sea. The rigging of the boats in harbor sparkled with flags. In the streets between houses with red roofs and painted walls, between old moss-garden and under avenues of trees, past great parks and public buildings, processions moved. Some were decorous: old people in long stiff robes of mauve and grey, grave master workmen, quiet, merry women carrying their babies and chatting as they walked. In other streets the music beat faster, a shimmering of gong and tambourine, and the people went dancing, the procession was a dance. Children dodged in and out, their high calls rising like the swallows' crossing flights over the music and the singing. All the processions wound towards the north side of the city, where on the great water-meadow called the Green Fields boys and girls, naked in the bright air, with mud-stained feet and ankles and long, lithe arms, exercised their restive horses before the race. [...]
If Anthopic had bought all the books it had trained for say at market rate we’d be having a different conversation now. Anthropic, through this settlement, has been forced to pay back, at least something… Kim would likely not have had enough money to compensate the victims and probably caused some more direct dammage by sharing pirated content. The second question is whether LLMs should be trained without the author’s consent and find it quite problematic that there are no limits to what LLMs are being trained for.
> Anthropic, through this settlement, has been forced to pay back, at least something… Kim would likely not have had enough money to compensate the victims and probably caused some more direct dammage by sharing pirated content.
You're thinking civil. They're talking criminal. Criminal law enforcement does not (well, isn't supposed to) look at your ability to compensate before deciding what to charge you with.
Since we're on the criminal side - what criminal statute would apply to Anthropic?
And what criminal statutes were used for the cases we're supposed to compare to?
anthropic etal would not have a product to sell without their violation..
kdc had a service that just happened to be popular for pirating...
how are the two even remotely similar?
For anyone who thinks the problem is Anthropic, I want you all to know that most authors make less than the median income. Most make less than $20,000 a year, because publishing houses give authors an advance, and then authors must pay back that entire advance in sales before they see a dollar of profit from their work.
Most never do.
Maybe publishers should JUST pay authors WELL, and get a book every 2-3 years.
>> Maybe publishers should JUST pay authors WELL, and get a book every 2-3 years
There are a couple problems with this approach.
Firstly, while the median income is 20k, the book business is like films or music; ie not evenly distributed. At the top end are a small number of successful authors. They effectively subsidize the publishing house while the house throws advances at authors hoping for the next big whale.
Many books never earn back their advance. Meaning if the author was paid out of royalties they'd make less, not more.
Making advances bigger would result in fewer advances. The pot of money is finite.
This is all happening in a market where supply is unconstrained (everyone thinks they can write), and demand is very limited.
And before we discuss the value, or lack thereof of having an intermediary at all, it should be noted from your link that the median for published authors is higher than self-published authors. So clearly they seem to be making authors more valuable.
In truth of course, most (published) books aren't terribly valuable. Like music and movies most float to the bottom.
If most books never recoup the advance in sales, then isn't this a better deal for most authors? It sounds like a guaranteed floor which might be very low but is nonetheless higher than the alternative
if authors get an advance that's greater than the sale of their books, doesn't it mean that the publishers lose money, i.e. paid more for those books than the books sales?
Authors get around 10% of the cover price of a book as royalties, it depends on several factors. The rest goes to the publisher. So some do lose money, but the break even for the publisher is usually well before the advance is fully covered by royalties.
Well the publisher also pays for the book to be bound, edited, overhead for their staff, cover art. Many books don't sell for the full retail price and are discounted. So net of all of this a 10% profit margin is common, they aren't keeping 90% of the book sales.
Part of the problem is the rest of us are broke as well and taxed to death so we don't have much left. If they paid you well, we wouldn't be able to afford your books.
Yes, the Anthropic settlement is far too small to distribute fairly. It seems like you think this makes the action that lead to this settlement justified?
It's an unfortunate outcome. Now to be a big player in AI, you have to have enough capital to buy your own library worth of books and digitize them. (Fun fact: a pallet of books is called a "gaylord," and they buy hundreds of gaylords.)
I created books3 to help settle the question of whether AI companies should be allowed to train on books. The outcome of "it's okay to pirate books as long as you're only training on them" was a long shot, but it would've let individual hackers train their own AI models (assuming access to sufficient compute, which you can get e.g. via https://sites.research.google/trc/about/).
Now we're in a world where you have to have dozens of millions in capital to do substantial work.
I heard at one point Eleuther was gathering public domain training data. I wonder if they ever built a corpus large enough so that training on books doesn't really matter...
A critical distinction, because they were going to to find terabytes of not pirated books to train on that contained the sum history of humanities knowledge /s
> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books
> Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals.
This is worse than pirating books to an absurd degree, it's almost a parody - the company that slurps all human knowledge ends up not only metaphorically, but also physically destroying those books, like an information vampire.
Authors don't even receive any financial compensation if the books were bought second hand, either. There's no benefit in doing that. (Not that making one final sale of a hardcover copy would make any difference though)
If Anthropic were at least buying ebooks, this insanity wouldn't need to happen. Unfortunately there is no bulk rates for buying millions of ebooks like you have in the used book market
No, it’s proof purchase of how stupid the publishing industry is. Maybe publishing houses should just pay authors good money, like a goddamn salary, and get a book out of them every few years.
They can't for the same reason that cab companies can't make their drivers employees: they would have to employ far, far fewer of them than they do on contingency.
The AI craze not only destroyed books, but many small websites who couldn't bear the load of constant scraping, or many communities that took open forums and took them offline or put them behind closed doors.
There is less publicly available knowledge now on the Internet than there has been 3 years ago.
Great, so now instead of allowing anyone to train on already scanned books for free, we can have only the richest big labs buy all the books and scan them privately to train their proprietary models. And since they buy the books used, authors still don't get any money. But at least the books are destroyed afterwards! What an improvement!
My complaint is that after this settlement nothing has materially changed except that the big labs now benefit from higher barriers to entry in their market. Authors don't make more money (other than a one time protection payment from Anthropic to publishers and some lawyers). Literally no one else benefits, except I guess used book marketplaces and book scanner vendors.
To be clear, this isn't a problem with the court process. Everything here appears perfectly in accordance with the law. It's just an absurd state to be in.
This is such a petty and impotent ruling. If you want to ban them from using culture to make derivative works without proper compensation then do that.
But if you don't want to ban them, telling them to buy one book of each, likely second hand, is complete pettiness that resulted in destructive scanning of millions of books, many of which were already practically available in digital form.
>This is such a petty and impotent ruling. If you want to ban them from using culture to make derivative works without proper compensation then do that.
That's because the judges are supposed to rule on questions of law (ie. "is AI training fair use?"), not whether they think AI's good or not.
how much of your economic output are you comfortable with companies like Anthropic stealing to put you out of work?
at least in Player Piano they paid the workers who made the cassette tapes that made the robots work.
our current LLM overlords demand that they be able to basically steal the sum total of all human knowledge so that they can sell it back to us at a rate they set.
they should have been shunned by society and made penniless when they first announced their goals but we have a bunch of deeply misanthropic people who have money and want to make a world where computer slaves do their bidding.
If you believe information deserves to be free, and if most of your earnings were from information that wasn't given away for free -- well, if you want people to give up their ill gotten gains, maybe you can start by setting an example.
So, mind sending me your bank account information? I'll promise to make good use of it.
Did someone forget to consult with the MPAA and the RIAA on this one? This is a joke of an outcome. $3k per book. How much was it per song for Napster?
The RIAA typically asked for around $2-4 per song to settle without a lawsuit, which would come to a total of a few thousand because they generally only went after people sharing over a thousand songs.
In the couple of few where the party would not agree to a settlement and the RIAA sued, they would pick about 15 of the thousand+ songs to sue over. Statutory damages are a minimum of $750 per infringed work, so the total would now be about 3-5 times what their settlement offer amount had been.
Most parties then got a lawyer, the lawyer told the party that had no chance, and they would then seriously negotiate with the RIAA and get a settlement.
Only a couple would still not settle, went to trial, and did an absolutely terrible job and the judge/jury awarded well above the minimum statutory damages. The RIAA still tried to settle for well below that, but the defendants refused and kept trying to fight and did not have a happy time.
> in practice the RIAA offered defendants the option of establishing a “Clean Slate” by destroying all of their illegally acquired files and paying a settlement of approximately $3 per illegal song.
The only bad thing about OpenAI and Anthropic training on everyone's stuff without their consent is that they didn't give away the model weights afterwards.
The people who espouse copyright abolitionism believe "information wants to be (and should be) free"
So no, for these people including myself, Copyright isn't doing anything good at all. It should be abolished. None of the people in this suit should get a dime. The government should force open weight releases of all foundation models as basically the only regulation that applies to the space at this current time.
That would be my preference, but if people really want to have free as in beer access to information, we can have that conversation. After these companies give people money for the commons that they're strip mining.
This is not even than a slap on the wrist. Publishers who negotiated this really fucked up writers.
According to US federal law, pirating a single copyrighted work and gaining commercial advantage of it (which Anthropic 100% did) represents five years in prison and a $250,000 fine. But it gets worse:
"Penalties for a copyright infringement conviction may increase if the defendant has previous similar convictions, made more than 10 copies of copyrighted works, committed copyright infringement during a period longer than 180 days, or infringed copyrighted material worth more than $2,500."
It's seemingly $3,000 per book, so they could've (and did, partially) just bought the books themselves for way cheaper, and with only a fraction of that money going to the authors
It's valid to not take AI companies' side here but people who think publishers are fighing for the little guy's rights are delusional. Tech companies have been exploiting artists for a few years, publishers/record labels/media companies have been doing it for centuries.
THANK YOU! And the idea that copyright actually helps individuals is such bullshit I can’t even believe anyone believes it! On a site filled with free software advocates.
That is a false statement. Gaining commercial advantage means selling pirated copies which Anthropic absolutely did not do, so none of your following statements are correct either.
I sincerely don't understand what the point of these laws are, when the cost of flagrant violations is no more than a slap on the wrist -- these really meager sums that serve as nothing more than something to point at and say "Look, we did something!"
Cover-your-ass strategy, and nothing more. Who, besides the ones at fault, are ever happy with these mean-nothing fines?
The justice system really needs an overhaul with how it tackles "justice" between the wealthy, the connected, the corporations, and the rest. Though I am unsure what that would look like. Minimum net wealth per category of infraction across the board?
This is a settlement that the authors and Anthropic agreed upon.
They agreed on the amount last year. The judge approved it now.
The lawsuit was for the way the books were acquired. They already ruled that it's not infringement to use the books.
The award was $3,000 per book, which is about 100X higher than it would have cost to buy the books.
It's never going to appease the people who demand companies be sued into collapse, but given that both parties came to an agreement and the damages are 100X higher than what a book costs, it looks reasonable to me.
I dont see the relevance. If Anthropic had bought the book at the store, shredded the spine, scanned the pages and trained on that data instead, there wouldnt have been an issue.
Authors cant simply license away fair use. If it could be dismissed so easily the right wouldn't exist.
Of course it is. If I write a movie review and sell it to a magazine or whatever, it's derived from the movie, and it's fair use, and I don't need to ask the movie owner for permission first, or give them a cut of my sales. Even if I use some reasonable number of screenshots and video clips, as long as the resulting work is "transformative" i.e. actually a new work, a movie review instead of a copy of the movie.
Do you want this to work any other way? I constantly see people in the AI debate working themselves into wildly copyright maximalist positions. I actually don't think that we should give every author veto power over a book review!
>I constantly see people in the AI debate working themselves into wildly copyright maximalist positions
I really dont get this. I know its that conflation fallacy or whatever, but I was under the impression we had sort of gotten over copyright maximalism as a society after Napster etc.
Whats worse is that, meaningful reform in this space has basically been waiting on a multi billion dollar corporation to come along and push it forward. So now that we have an opportunity to expand and globalise fair use, the sudden and quite angry opposition weirds me out to no end.
1. author owns the right to distribute copies of the work
2. this right goes on for faaaaaar too long.
I don't have an issue with 1. You had a good idea, you implemented it, you deserve something for it. Given some people got sued into oblivion with ridiculous dollar value outcomes on a per unit basis - why doesn't this apply here? Sure 1.5 billion is a lot. But the number of infringments is insane and the company is approaching a trillion in valuation. You could make it ten times that number.
I do have an issue with 2. Sure, you had a good idea, you implemented it, you deserve something for it. But after 20 years, you should be able to come up with another idea or just work like the rest of us. Going for 50, 70, 90+ years with the rewards going to estate heirs? Fuck that.
So yeah, I am both against copyright AND surprised at the slap on the wrist for what happened here.
>Given some people got sued into oblivion with ridiculous dollar value outcomes on a per unit basis - why doesn't this apply here?
I mean, it feels to me like one or both of:
1. The class action lawyers werent 100% certain they could win in court.
2. The class action lawyers smelled an easy payday.
They get ~100 million out of this.
I also think that the 1500 bucks going to most of these authors is going to be more than they ever saw in royalties. I read somewhere that 500 - 1500 bucks is roughly what a self pub book makes in its lifetime. Why push the envelope? Anthropic hasnt done anything that deserves to pay for the entire lifetime royalties of most books. Their legal alternative is to cut the spine off and scan the book in. In which case the author and publisher will be splitting 20 bucks instead, assuming Anthropic isnt buying used.
This seems like a donation tbh.
>slap on the wrist for what happened here.
Its not a punishment at all because this is a civil case that has been settled out of court.
IANAL but as an IP creator I have not heard of "derivative products" in the copyright context. There are "derivative works", which are covered by the same copyright as the original. For example, a translation to another language is a derivative work, a novelisation of a movie, a screen adaptation of a book etc. If some author could have proven that any Anthromic model is a derivative work of theirs then they had the copyright on that model and made mad bucks licensing it back to Anthropic.
To sign up for what? The experience of approximately every author on the planet is that they found out that Anthropic did something bad at the same time they were "opted into" the class. The only thing they could do is opt out and litigate on their own against a company with a valuation approaching $1T.
This is a sweet deal for lawyers and for publishers, and nothing else.
Civil justice is primarily about restoring damages, not about punishing wrongdoing (although common law in US it is more punitive than civil law in european countries). Therefore compensations are based on damages, not on profit from wrongdoings.
the irony is that all of this money will go to rent-seeking publishers who won't pass it on to the artists; basically a dispute between the wealthy you're upset with
Default payout is 50/50 author/publisher. If the author and publisher have a contract that states otherwise, then their contract overrides the default.
Source: I’m an author and signed up to be part of the class action, and this was the class action documents said.
> If there is a current publisher(s) (which still possesses an exclusive license), the author(s) will split the $3000 with the publisher. Any co-authors will share the author portion and, if there are multiple publishers (e.g., different publishers have exclusive rights to different formats), they will share the publisher portion. Assume that the co-authors and co-publishers will share the portion equally unless their contracts provide otherwise. The standard default split between publishers and authors of noneducational texts is 50/50, as described below. Authors who are the sole rightsholder in a work—such as self-published authors and authors whose rights have reverted or where the contracts have otherwise terminated—will receive the full award amount.
It is split between the publisher and the author, also publishers will have a large catalog of books they will submit, an author typically will only have a few -- the payout will be going to the lawyers and mostly to publishers.
>It is split between the publisher and the author, also publishers will have a large catalog of books they will submit, an author typically will only have a few -- the payout will be going to the lawyers and mostly to publishers.
This is innumerate. If it's split 50% between authors and publishers, then it won't be "mostly to publishers". Mathematically it will be equal between "authors" and "publishers", and because lawyers are taking their cut, neither would be able to get "most" of it. Yes, the average publisher will get a bigger paycheck, but that's because there's less of them, not because "most going to publishers".
> That means that rightsholders can expect at least $3,000 per title (less costs and fees), which will be shared among the rightsholders for that title (if there is more than one rightsholder)
> if there is more than one rightsholder
Again, a publisher will have a whole catalog of books / titles, a non-negligible portion of that the publisher will own the copyright to (no one to split it with). There's all kinds of books outside of novels, there's media tie-ins, IP franchise books (ie Star Wars), childrens books, textbooks / reference materials, etc etc etc. Yes, with novels the author tends to own the copyright, but you're forgetting all of the other kinds of books out there.
> I sincerely don't understand what the point of these laws are, when the cost of flagrant violations is no more than a slap on the wrist -- these really meager sums that serve as nothing more than something to point at and say "Look, we did something!"
To create a moat around wealth generation. After all, that is the main purpose of all legal systems---to keep the wealthy wealthy and the poor poor. In this case, the settlement is chump change for Anthropic, but ensures that no upstart will be able to compete with them since they will get reamed on copyright charges. It's no different from Google Image search. They can make a product out of republishing others' images. You cannot do it.
To keep users paying for content while companies do whatever they want - and if that's not the reason that's certainly an effect.
> The justice system really needs an overhaul with how it tackles "justice" between the wealthy, the connected, the corporations, and the rest. Though I am unsure what that would look like. Minimum net wealth per category of infraction across the board?
% of annual turnover seems like decent strategy. Caps the amount company can sue mere mortal for copyright infringement while at billion dollar company scale can wipe quite a bit
But main problem is enforcement and lobbying, not the size of the fine
I support anthropics position here, on both learning from and "pirating" books.
The way i see things , the publishers and authors are happy with any policy that makes them more money, and more market control, regardless of what is ethical/just/right.
They would shutdown public libraries , all libraries, if they could.
Aaron Swartz lost his life because he tried to make public knowledge public, and they would be happy to put every information activist to death to protect their monopolies.
IMHO they have no right to stop free access on the internet. The whole copyright system is artificial and monopolistic, and the publishers are complaining yet again, that technology moves information more efficiently than they do, so they want to artificially retard it through goverment action.
The real goverment action that is needed, is to protect private/personal data; not data that is actively traded commercially or publically.
These tech companies are invading personal and private spaces of everyday people, and storing and training with it. Even using it for military targetting and warrantless surveillance.
Anthropic is by no means a good entity, so the way to stick it to them and all tech companies, is to allow their internet scraping, but make it outright criminal to use telemetry or any surveillance techniques they have or will develop.
Also... the ”creators",hollywood,publishers, have no problems scraping themselves, and lift ideas from just about everywhere they can get it.
Almost every hollywood movie is just an assemblage of random memes and topical concerns of everyday ppl, distilled into embelished predictable cheese.
The publishers are the original slop actors. Human Slop.
Aaron swartz lost his life because he committed suicide. Something he had tried multiple times before. If he really only committed suicide because of the legal jeopardy he was in, wouldn't it have made more sense to commit suicide after you're found guilty?
I work as an author. I believe this is total bullshit, from beginning to end - the ruling, the settlement, and the suit itself.
In the UK, we have a thing called the Public Lending Right [1]. This pays authors a fixed sum each time their book is taken out of a library, up to a capped amount.
The cap isn't very high - about $7k - so it is both an OK bit of income for authors who might be making very little money elsewhere, and also doesn't end up all going to authors who are already bestsellers. It's a decent legal system for helping libraries hold niche titles as well as the popular ones. This is, after all, the purpose of a library.
To establish my bias here: My debut novel came out after the period this specific suit concerns. I also uploaded it to LibGen myself.
I strongly believe that books should be available to read, free of charge, to all people. I benefited enormously from libraries and piracy growing up. I think they serve an important educational purpose that does not end when a person leaves school, and I do not think wealth or disposable income is a fair way to decide the breadth of a person's education.
I also have no problem with people making new "language things" using my work. I love sample-based music (like dance music, hip hop, etc) and it'd be hypocritical for me to take issue with anyone doing analogous things using books. Maximising sales is not the end-goal of making art, for me personally. Other artists feel otherwise. They consider training on pirated books stealing. That's OK - it's not for me to tell them what to believe.
The problem for me is that these corporations - undoubtedly still pretraining on pirated material - are, essentially, leeching. By not releasing the model as open-weight, freely available, they are not acting in the same spirit of the system they took advantage of. It's the Spotify model: pirate first, pay a nominal amount that does not meaningfully harm profit later. Now the dust has settled there, we can see the harm it has done to music culture.
A single settlement which does not establish precedent does not solve anything. A tokenistic $3k allows anti-AI authors to wave a cheque in the air and declare a victory. It pays the rent for a month or two. It does nothing for the months after that, when the corporation is still profiting. It does nothing to establish precedent for future artists, who also have to pay rent.
It would be (non-trivial, but) relatively simple to integrate - for example - download figures from Anna's Archive into the PLR. I'd happily dilute my PLR payment appropriately, because I think libraries are important.
You can't stop people pirating digitally replicable things. Digital ownership is not a concept that has held, or will hold.
There are only 23,000 authors in the UK who claim the cash from the PLR. To pay all those authors the national living wage in the UK (£26k) from the PLR, you would need to raise £546 million. That is around 1/34 of Anthropic's reported annual revenue.
I'm of course not arguing Anthropic should be solely responsible. But it's very frustrating that all the pieces of the puzzle for actually paying artists in a sustainable and ongoing way now exist, and one of the major obstacles to this - and the idea of a genuinely free, legal, international library, which creates more authors, writing better books, full-time - are legacy rights holders who remain attached to a completely dysfunctional and outdated concept of ownership.
So - unless part of a sustained and reasonable campaign, which understands the futility of (and damage to the medium and its creators caused by) treating digital ownership in the same way as physical ownership - this suit is close to pointless, and arguably actively harmful in the long term.
As was pointed out, the settlement is for piracy, not training. They had already ruled that Anthropic's use of copyrighted material for training fell within fair use.
As such, if you pirated a book and had to pay $3000 for that one instance, I don't think you'd like it if I said you should have paid $30K or $300K instead. If anything, this is analogous to the ridiculous fines people had to pay when pirating music.
> As such, if you pirated a book and had to pay $3000 for that one instance, I don't think you'd like it if I said you should have paid $30K or $300K instead.
If you pirated a book for personal use the amount of liability wouldn't match a company whose profit could be attributed to pirating the same book. In US copyright law, a copyright infringer could be liable for "any profits of the infringer that are attributable to the infringement" [1] (if the copyright owner elects to recover actual damages and profits instead of statutory damages).
IANAL, but the parent comment quotes "any profits of the infringer that are attributable to the infringement", which I take to mean it's the profit Anthropic stands to make based on its use of the pirated content that's recoverable.
Given the entire global economy is currently bullish on the potential profitability of AI, I dare say they got off incredibly lightly settling for just $3k per book.
None of this matters, this is the judge approving a voluntary settlement reached between the parties last year.
If you think it should be different then you have to make a cogent argument why the public should get to interfere with a settlement the two sides mutually agree on.
Note: I never said it should be different and certainly wasn't arguing for any side. I was merely making an observation that the settlement seemed like a good deal (for both parties) given the potential for Anthropic to be liable for a significantly greater amount depending on how the law would be interpreted if they went to trial.
Because we are mostly discussing a single private person that got caught for maybe 20 songs. I don't want to bring up Aaron but the taste gets saltier the more we see settlements like this.
This is a too big to fail scenario. If these companies fail, so does the US economy. Normal laws for individuals don't apply, so any comparison to that is pointless.
Well, the USA has jurisdiction over USA companies. If the rest of the world's authors can find a way to obtain jurisdiction over the companies in a way that USA courts won't balk at if asked to enforce, then they're welcome to go ahead.
That depends on the answer to a question that hasn't been answered yet.
Is what an AI does similar to a human reading a book, and adding it to their knowledge? Or is it similar to a human plagiarizing a book? If it's the second, for at least some books, no, the damages are not reasonable. They are far too small.
> That depends on the answer to a question that hasn't been answered yet.
It has been answered in a sense, because the courts (so far) have ruled that training is Fair Use. Whether this is similar to a human learning from a book was not quite the question being answered, but AFAICT there is no other relevant doctrine under Copyright law to address it, largely because the question didn't even exist until LLMs came along.
Also, these are not damages, it's a settlement i.e. a negotiated agreement between both parties.
Good question. Can you ask an LLM to repeat the entire contents of a novel, word-for-word, and read that instead of the original book? I haven't tried it, but I would guess it would not be able to do this.
Can you ask it questions about the book and expect it to get them right? Yeah, probably. Same as if I read the book and you asked me questions about it. The LLM would probably answer those questions better than I could, and about every single book in its training data, but still same-same.
See, Judge Alsup should have been the person Biden put on the Supreme Court, that or re-nominate Merrick Garland. Instead, he made a silly promise to sate Black Lives Matter, which even when he took office was fast on its way to ignominy, and now Kagan is stuck being the only competent liberal justice on the court. At least Alsup can continue setting the direction of law as it applies to the tech industry.
What a fucking joke of a country the US is, allowing this kind of behaviour with such a pathetic "punishment". Barely even qualifies as a tap on the wrist, Anthropic should be getting gutted into non-existence for this shit and the execs should be given the Aaron Swartz treatment.
Au contraire! Now the creations of the LLMs stand on legal ground. This was an excellent deal for Anthropic
Perfect - an absolute steal for 1.5B.
Meta also has copyright lawsuits for the open models they released, so open models are not immune.
... unless the line we want to draw is "american orgs pay, others don't", as currently seems to be happening.
That is the case, any regulation increases the cost to enter a market.
But in this case, its irrelevant because the moat of cost to enter is already unfathomable and secondly, they are not adding regulation but fining them for committing a crime.
So yeah, adding that every food compnay needs 3 health inspectors that they pay for would benefit coca cola over you mom and pop bakery. But telling someone they cannot start a Space agency with money laundered from ransom and drug sales payments would not affect much the competition markets
and then sell to all its subscribers, things need to change.
Fixed that for you.
Imagine being able to pay a fraction of your savings to download all Netflix shows and then sell 1 minute chunk of every media to your paid subscribers.
If I put my conspiracy theory hat one and I always get piled on for this theory in other online communities but I think it could be possible. The theory is I think Aaron found some very dark stuff while exploring the MIT private networks, things that he was not supposed to see and could be very damaging to a lot people if they were exposed. The infamous Jeffery Epstein was donating a lot of money to MIT and its Media Labs. I think there is a much deeper story at play that the mainstream narrative is hiding with a “suicide”.
https://storage.courtlistener.com/recap/gov.uscourts.cand.43...
The big deal for publishers and authors is the payout per eligible title is $3k. For a traditional publishing contract involving one author, the amount will be split down the middle.
The other thing which caught my eye is the judge slashed the class counsel's fee by half, from 12.5% ($187.5m) to 6.8% ($101m). The class counsel's unreimbursed litigation expenses were $2.6m.
The three class representatives get just $15k each.
Source: was a licensed real estate agent for a long time.
If this is the 2024 settlement that you are referring to, it did not say anything about the price a Realtor can charge:
https://en.wikipedia.org/wiki/Burnett_v._National_Associatio...
>The cooperative compensation rule has been eliminated as a result of the settlement. Seller's agents are no longer required to offer compensation to buyer's agents when listing a home for sale on a Realtor-owned multiple listing service. In addition, Realtors acting as buyer's agents must enter into contracts with buyers before touring any homes, allowing buyers to negotiate how much they will pay their buyer's agent.
In what sane state does it even get that high?
He's also a longtime hobbyist programmer working in BASIC, much of it in support of his ham radio hobby. Screenshots of his shortwave propagation prediction program here [1].
[1] https://www.theverge.com/2017/10/19/16503076/oracle-vs-googl...
His middle name is Haskell.
It's funny that crimes can be settled in cash. IOW, everything has a price; and the price is always right. Settlement ought to be the euphemism for blood money.
In addition to the settlement, what I'd consider fair is to have these companies pay royalties in perpetuity. Of course, that's not tractable.
No, that's what they got in trouble for - a lack of consent.
If the author consents, it would have been fine. If they bought the books, then it is fine. Digitisation through destruction, like most book scanning systems. As long as the original work is destroyed during the process, and you actually paid for it, then it is fair use.
If it regurgitates, then the author can sue you again. So you are incentivised to make damn sure it doesn't. That's not covered by fair use.
Its only if the original cannot be accessed anymore, and you paid to get the original. Both must be true, for fair use to hold.
> Literary works, including computer programs and databases, protected by access control mechanisms that fail to permit access because of malfunction, damage, or obsoleteness.
DRM being covered under other laws, and being gross, still applies. And still applies to industry giants, too. Which is why most who do this, like Google, actually buy physical copies and scan it destructively, so they don't have to deal with it.
Learning from and building on previous work is civilization. Copyright maximalism is a plague.
Those regulations and principles are for humans.
Either the major LLMs are software tools deployed by ostensibly-profit-seeking companies, and regulations based on the notion that "making humans pay to make use of the things they've learned is profoundly antisocial" don't apply, or the LLM companies have a bigass swarm of unpaid -er- "servants", and labor laws and other human rights regulations do apply.
They are, presumably, human. We can perfectly well say that humans have certain rights without needing to give machines those same rights.
For example, we've more or less all agreed that it's fine for a human to watch a movie and enjoy the memories forever, and be inspired by it forever. But we've also more or less all agreed that that doesn't mean that a human can use a machine to record that movie and keep it forever.
> Learning from and building on previous work is civilization. Copyright maximalism is a plague.
The debate has existed for several generations at this point. You may disagree with the mainstream opinion, but it's disingenuous to frame it as "copyright maximalism".
How does intellectual work get funded in this insane world if yours, pray tell?
If he was not aware, I wonder if he still would have described the process as "exceedingly transformative" had he been aware.
Note also that Sonnet 3.7 had to be jailbroken.
Note also that they got high memorization for a few books that were widely quoted. The books in question can probably also be "retrieved" by putting phrase prefixes into Google, which is probably why Sonnet 3.7 knows them with the precision of a fanboy. Material being widely repeated in the training set is a well-known cause of memorization.
Edit: Apologies, I misread it as "100 pages". My point about copyright still stands, though.
Some of us have a good enough memory.
Police descended upon Kim Dotcom like he was a terrorist or something. They rappelled down helicopters and stormed his home like he was bin Laden.
Then these big techs come along and they make some absurd cost of doing business settlement.
I think it's just a matter of time until everyone learns this. And then it will be the end of the slowly dying liberal democracies.
What was Sean Parker sued for again?
2. There's incredible value in what they stole.
3. IANAL, but I don't believe "but now everyone can write like a terrible version of the writer we fleeced" is a valid legal defense.
That's a very charitable way of saying "someone with deep enough pockets can ignore the law and get away with it."
Obviously exaggerating.. but not by much.
Some of you really don’t deserve good things. You should be blocked from using AI on more than one device without paying an additional subscription plan.
That said, it's been done: https://arxiv.org/abs/2601.02671
> In some cases, jailbroken Claude 3.7 Sonnet outputs entire books near-verbatim (e.g., nv-recall=95.8%).
You don't need to reproduce anything verbatim: a 1/4 resolution copy of a movie is still infringement even though it's only a quarter of the size.
You're thinking civil. They're talking criminal. Criminal law enforcement does not (well, isn't supposed to) look at your ability to compensate before deciding what to charge you with.
[1] https://en.wikipedia.org/wiki/Aaron_Swartz#Arrest_and_prosec...
copyright infringement was enough to get judgements that ruined entire lives when i was in my late teens and early 20s
now you get to be a founder of a trillion dollar business by extremely large copyright infringement
fuck these ghouls fuck LLMs and fuck the waste of money for this shit
- @Nevermark
Most never do.
Maybe publishers should JUST pay authors WELL, and get a book every 2-3 years.
https://authorsguild.org/news/key-takeaways-from-2023-author...
There are a couple problems with this approach.
Firstly, while the median income is 20k, the book business is like films or music; ie not evenly distributed. At the top end are a small number of successful authors. They effectively subsidize the publishing house while the house throws advances at authors hoping for the next big whale.
Many books never earn back their advance. Meaning if the author was paid out of royalties they'd make less, not more.
Making advances bigger would result in fewer advances. The pot of money is finite.
This is all happening in a market where supply is unconstrained (everyone thinks they can write), and demand is very limited.
And before we discuss the value, or lack thereof of having an intermediary at all, it should be noted from your link that the median for published authors is higher than self-published authors. So clearly they seem to be making authors more valuable.
In truth of course, most (published) books aren't terribly valuable. Like music and movies most float to the bottom.
So no, the answer is not "pay authors more".
Thats a nice way of saying publishing houses are paying most authors more than they make from the sales
Well the publisher also pays for the book to be bound, edited, overhead for their staff, cover art. Many books don't sell for the full retail price and are discounted. So net of all of this a 10% profit margin is common, they aren't keeping 90% of the book sales.
Most of their investments fail miserably, but they only need one Google/Stephen King.
I created books3 to help settle the question of whether AI companies should be allowed to train on books. The outcome of "it's okay to pirate books as long as you're only training on them" was a long shot, but it would've let individual hackers train their own AI models (assuming access to sufficient compute, which you can get e.g. via https://sites.research.google/trc/about/).
Now we're in a world where you have to have dozens of millions in capital to do substantial work.
I heard at one point Eleuther was gathering public domain training data. I wonder if they ever built a corpus large enough so that training on books doesn't really matter...
A Gaylord is a type of box that fits on a pallet. There are multiple ways to palletize products, like shrink wrapping or metal banding
(an observation, not agreement)
Let them take on the liability
> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books
This is worse than pirating books to an absurd degree, it's almost a parody - the company that slurps all human knowledge ends up not only metaphorically, but also physically destroying those books, like an information vampire.
Authors don't even receive any financial compensation if the books were bought second hand, either. There's no benefit in doing that. (Not that making one final sale of a hardcover copy would make any difference though)
If Anthropic were at least buying ebooks, this insanity wouldn't need to happen. Unfortunately there is no bulk rates for buying millions of ebooks like you have in the used book market
There's such a thing as fair use and digitizing privately owned printed material is absolutely legal... including for corporations.
There is less publicly available knowledge now on the Internet than there has been 3 years ago.
That would be pirating. So your complaint is that they didn't do more piracy?
To be clear, this isn't a problem with the court process. Everything here appears perfectly in accordance with the law. It's just an absurd state to be in.
the people operating frontier labs are bad people they cannot be trusted in any way
the best solution to them would be to send them to monster island (even though it's really a peninsula)
But if you don't want to ban them, telling them to buy one book of each, likely second hand, is complete pettiness that resulted in destructive scanning of millions of books, many of which were already practically available in digital form.
That's because the judges are supposed to rule on questions of law (ie. "is AI training fair use?"), not whether they think AI's good or not.
at least in Player Piano they paid the workers who made the cassette tapes that made the robots work.
our current LLM overlords demand that they be able to basically steal the sum total of all human knowledge so that they can sell it back to us at a rate they set.
they should have been shunned by society and made penniless when they first announced their goals but we have a bunch of deeply misanthropic people who have money and want to make a world where computer slaves do their bidding.
So, mind sending me your bank account information? I'll promise to make good use of it.
You get a service. The service is using their compute power to run a model and their scientists to build the model.
In the couple of few where the party would not agree to a settlement and the RIAA sued, they would pick about 15 of the thousand+ songs to sue over. Statutory damages are a minimum of $750 per infringed work, so the total would now be about 3-5 times what their settlement offer amount had been.
Most parties then got a lawyer, the lawyer told the party that had no chance, and they would then seriously negotiate with the RIAA and get a settlement.
Only a couple would still not settle, went to trial, and did an absolutely terrible job and the judge/jury awarded well above the minimum statutory damages. The RIAA still tried to settle for well below that, but the defendants refused and kept trying to fight and did not have a happy time.
https://www.history.com/this-day-in-history/september-8/riaa...
> in practice the RIAA offered defendants the option of establishing a “Clean Slate” by destroying all of their illegally acquired files and paying a settlement of approximately $3 per illegal song.
The two notable cases were:
1) https://en.wikipedia.org/wiki/Capitol_Records,_Inc._v._Thoma...
2) https://en.wikipedia.org/wiki/Sony_BMG_Music_Entertainment_v...
The people who espouse copyright abolitionism believe "information wants to be (and should be) free"
So no, for these people including myself, Copyright isn't doing anything good at all. It should be abolished. None of the people in this suit should get a dime. The government should force open weight releases of all foundation models as basically the only regulation that applies to the space at this current time.
A slap on the wrist, that's what it's doing here, isn't it?
According to US federal law, pirating a single copyrighted work and gaining commercial advantage of it (which Anthropic 100% did) represents five years in prison and a $250,000 fine. But it gets worse:
"Penalties for a copyright infringement conviction may increase if the defendant has previous similar convictions, made more than 10 copies of copyrighted works, committed copyright infringement during a period longer than 180 days, or infringed copyrighted material worth more than $2,500."
https://www.justia.com/entertainment-law/piracy-in-the-enter...
It's seemingly $3,000 per book, so they could've (and did, partially) just bought the books themselves for way cheaper, and with only a fraction of that money going to the authors
Absurd.
Cover-your-ass strategy, and nothing more. Who, besides the ones at fault, are ever happy with these mean-nothing fines?
The justice system really needs an overhaul with how it tackles "justice" between the wealthy, the connected, the corporations, and the rest. Though I am unsure what that would look like. Minimum net wealth per category of infraction across the board?
Edit: grammar
They agreed on the amount last year. The judge approved it now.
The lawsuit was for the way the books were acquired. They already ruled that it's not infringement to use the books.
The award was $3,000 per book, which is about 100X higher than it would have cost to buy the books.
It's never going to appease the people who demand companies be sued into collapse, but given that both parties came to an agreement and the damages are 100X higher than what a book costs, it looks reasonable to me.
If you are selling more than 100 books you are clearly losing out
The authors were only owed money for the piracy.
How many of the authors would license their book for endless creation of derivative works for that amount?
Authors cant simply license away fair use. If it could be dismissed so easily the right wouldn't exist.
Do you want this to work any other way? I constantly see people in the AI debate working themselves into wildly copyright maximalist positions. I actually don't think that we should give every author veto power over a book review!
I really dont get this. I know its that conflation fallacy or whatever, but I was under the impression we had sort of gotten over copyright maximalism as a society after Napster etc.
Whats worse is that, meaningful reform in this space has basically been waiting on a multi billion dollar corporation to come along and push it forward. So now that we have an opportunity to expand and globalise fair use, the sudden and quite angry opposition weirds me out to no end.
1. author owns the right to distribute copies of the work
2. this right goes on for faaaaaar too long.
I don't have an issue with 1. You had a good idea, you implemented it, you deserve something for it. Given some people got sued into oblivion with ridiculous dollar value outcomes on a per unit basis - why doesn't this apply here? Sure 1.5 billion is a lot. But the number of infringments is insane and the company is approaching a trillion in valuation. You could make it ten times that number.
I do have an issue with 2. Sure, you had a good idea, you implemented it, you deserve something for it. But after 20 years, you should be able to come up with another idea or just work like the rest of us. Going for 50, 70, 90+ years with the rewards going to estate heirs? Fuck that.
So yeah, I am both against copyright AND surprised at the slap on the wrist for what happened here.
I mean, it feels to me like one or both of:
1. The class action lawyers werent 100% certain they could win in court. 2. The class action lawyers smelled an easy payday.
They get ~100 million out of this.
I also think that the 1500 bucks going to most of these authors is going to be more than they ever saw in royalties. I read somewhere that 500 - 1500 bucks is roughly what a self pub book makes in its lifetime. Why push the envelope? Anthropic hasnt done anything that deserves to pay for the entire lifetime royalties of most books. Their legal alternative is to cut the spine off and scan the book in. In which case the author and publisher will be splitting 20 bucks instead, assuming Anthropic isnt buying used.
This seems like a donation tbh.
>slap on the wrist for what happened here.
Its not a punishment at all because this is a civil case that has been settled out of court.
All US courts so far have ruled yes.
The authors or the publishers?
I have a hard time believing they agreed with the millions of authors they pirated.
If you’re so interested, go read past the headline. Maybe you’ll find that you’re working about what “authors” will agree to.
This is a sweet deal for lawyers and for publishers, and nothing else.
Thats more than it costs to just shred the spine and scan the book in. Which is probably 15 - 20 bucks a piece.
They will be shredding the book not paying the fine.
Source: I’m an author and signed up to be part of the class action, and this was the class action documents said.
It was started by a group of authors, not publishers.
> If there is a current publisher(s) (which still possesses an exclusive license), the author(s) will split the $3000 with the publisher. Any co-authors will share the author portion and, if there are multiple publishers (e.g., different publishers have exclusive rights to different formats), they will share the publisher portion. Assume that the co-authors and co-publishers will share the portion equally unless their contracts provide otherwise. The standard default split between publishers and authors of noneducational texts is 50/50, as described below. Authors who are the sole rightsholder in a work—such as self-published authors and authors whose rights have reverted or where the contracts have otherwise terminated—will receive the full award amount.
It is split between the publisher and the author, also publishers will have a large catalog of books they will submit, an author typically will only have a few -- the payout will be going to the lawyers and mostly to publishers.
This is innumerate. If it's split 50% between authors and publishers, then it won't be "mostly to publishers". Mathematically it will be equal between "authors" and "publishers", and because lawyers are taking their cut, neither would be able to get "most" of it. Yes, the average publisher will get a bigger paycheck, but that's because there's less of them, not because "most going to publishers".
> if there is more than one rightsholder
Again, a publisher will have a whole catalog of books / titles, a non-negligible portion of that the publisher will own the copyright to (no one to split it with). There's all kinds of books outside of novels, there's media tie-ins, IP franchise books (ie Star Wars), childrens books, textbooks / reference materials, etc etc etc. Yes, with novels the author tends to own the copyright, but you're forgetting all of the other kinds of books out there.
To create a moat around wealth generation. After all, that is the main purpose of all legal systems---to keep the wealthy wealthy and the poor poor. In this case, the settlement is chump change for Anthropic, but ensures that no upstart will be able to compete with them since they will get reamed on copyright charges. It's no different from Google Image search. They can make a product out of republishing others' images. You cannot do it.
> The justice system really needs an overhaul with how it tackles "justice" between the wealthy, the connected, the corporations, and the rest. Though I am unsure what that would look like. Minimum net wealth per category of infraction across the board?
% of annual turnover seems like decent strategy. Caps the amount company can sue mere mortal for copyright infringement while at billion dollar company scale can wipe quite a bit
But main problem is enforcement and lobbying, not the size of the fine
The verdict is a joke.
In the UK, we have a thing called the Public Lending Right [1]. This pays authors a fixed sum each time their book is taken out of a library, up to a capped amount.
The cap isn't very high - about $7k - so it is both an OK bit of income for authors who might be making very little money elsewhere, and also doesn't end up all going to authors who are already bestsellers. It's a decent legal system for helping libraries hold niche titles as well as the popular ones. This is, after all, the purpose of a library.
To establish my bias here: My debut novel came out after the period this specific suit concerns. I also uploaded it to LibGen myself.
I strongly believe that books should be available to read, free of charge, to all people. I benefited enormously from libraries and piracy growing up. I think they serve an important educational purpose that does not end when a person leaves school, and I do not think wealth or disposable income is a fair way to decide the breadth of a person's education.
I also have no problem with people making new "language things" using my work. I love sample-based music (like dance music, hip hop, etc) and it'd be hypocritical for me to take issue with anyone doing analogous things using books. Maximising sales is not the end-goal of making art, for me personally. Other artists feel otherwise. They consider training on pirated books stealing. That's OK - it's not for me to tell them what to believe.
The problem for me is that these corporations - undoubtedly still pretraining on pirated material - are, essentially, leeching. By not releasing the model as open-weight, freely available, they are not acting in the same spirit of the system they took advantage of. It's the Spotify model: pirate first, pay a nominal amount that does not meaningfully harm profit later. Now the dust has settled there, we can see the harm it has done to music culture.
A single settlement which does not establish precedent does not solve anything. A tokenistic $3k allows anti-AI authors to wave a cheque in the air and declare a victory. It pays the rent for a month or two. It does nothing for the months after that, when the corporation is still profiting. It does nothing to establish precedent for future artists, who also have to pay rent.
It would be (non-trivial, but) relatively simple to integrate - for example - download figures from Anna's Archive into the PLR. I'd happily dilute my PLR payment appropriately, because I think libraries are important.
You can't stop people pirating digitally replicable things. Digital ownership is not a concept that has held, or will hold.
There are only 23,000 authors in the UK who claim the cash from the PLR. To pay all those authors the national living wage in the UK (£26k) from the PLR, you would need to raise £546 million. That is around 1/34 of Anthropic's reported annual revenue.
I'm of course not arguing Anthropic should be solely responsible. But it's very frustrating that all the pieces of the puzzle for actually paying artists in a sustainable and ongoing way now exist, and one of the major obstacles to this - and the idea of a genuinely free, legal, international library, which creates more authors, writing better books, full-time - are legacy rights holders who remain attached to a completely dysfunctional and outdated concept of ownership.
So - unless part of a sustained and reasonable campaign, which understands the futility of (and damage to the medium and its creators caused by) treating digital ownership in the same way as physical ownership - this suit is close to pointless, and arguably actively harmful in the long term.
[1] https://www.bl.uk/services/plr
As such, if you pirated a book and had to pay $3000 for that one instance, I don't think you'd like it if I said you should have paid $30K or $300K instead. If anything, this is analogous to the ridiculous fines people had to pay when pirating music.
(Not that I'm complaining...)
If you pirated a book for personal use the amount of liability wouldn't match a company whose profit could be attributed to pirating the same book. In US copyright law, a copyright infringer could be liable for "any profits of the infringer that are attributable to the infringement" [1] (if the copyright owner elects to recover actual damages and profits instead of statutory damages).
[1] 17 U.S.C. § 504(b), https://www.law.cornell.edu/uscode/text/17/504
Put another way, their revenues wouldn't drop much if they simply hadn't trained on those 99%.
Given the entire global economy is currently bullish on the potential profitability of AI, I dare say they got off incredibly lightly settling for just $3k per book.
If you think it should be different then you have to make a cogent argument why the public should get to interfere with a settlement the two sides mutually agree on.
Exclude one book from the training dataset.
Did you make a worse model?
We actually know the answer to this, and it is: absolutely not.
The reality is this: your intellectual output is almost always only valuable to any company in existence in aggregate, never in isolation.
If failure means catastrophe for the nation, it shouldn't have been a private, for-profit project in the first place and instead be a public project.
I would be really worried about the US economy then.
Who said that?
Is what an AI does similar to a human reading a book, and adding it to their knowledge? Or is it similar to a human plagiarizing a book? If it's the second, for at least some books, no, the damages are not reasonable. They are far too small.
It has been answered in a sense, because the courts (so far) have ruled that training is Fair Use. Whether this is similar to a human learning from a book was not quite the question being answered, but AFAICT there is no other relevant doctrine under Copyright law to address it, largely because the question didn't even exist until LLMs came along.
Also, these are not damages, it's a settlement i.e. a negotiated agreement between both parties.
Relevant sub-thread here: https://news.ycombinator.com/item?id=48997766
Can you ask it questions about the book and expect it to get them right? Yeah, probably. Same as if I read the book and you asked me questions about it. The LLM would probably answer those questions better than I could, and about every single book in its training data, but still same-same.
I don't think this is plaguarism.