External Link [Archive version]
As David Chisnall says over on Mastodon, this story is terrifying for anyone who is worried about copyright (probably most people).
If Claude, when prompted, occasionally spits out exact copies of code it has in its training data then there’s no way to defend agaisnt copyright suits from the original author.
As David explains:
Companies that have this concern usually do it by ensuring that no one on the team has been exposed to the original and that the code for the original never goes near their systems. But when one of the systems that you use is a language model trained on, among other things, all of the open-source code that it could scrape (and which does not disclose its training set), being able to prove that there was no path from some other codebase to yours is impossible.
It seems like this case may go away quietly, the app developer said he would cease and desist and just link the domain to the original author’s app. However, it will be interesting to see if any high profile instances of this crop up again.
Update: it turns out the original work was generated by Claude too. Which of the AI slop projects gets to claim it was the original? That is probably for a court to decide… Or ignore.