A coalition of news publishers accused OpenAI of allegedly deleting 20 million conversation records and hiding search capabilities in a copyright infringement case that could impact the defense of its language model, ChatGPT.
As the lawsuit against OpenAI and Microsoft continues in court, a coalition of 16 news publishers, including the New York Daily News and the Chicago Tribune, accused OpenAI of allegedly deleting 20 million conversation records crucial for a copyright infringement case that could be determinative for the defense of its language model (ChatGPT). Our team has been monitoring the case since the lawsuit was filed in December 2023, one of the first and highest-profile legal cases against how AI companies obtain their training data. Now, a new chapter is added as the coalition accuses OpenAI of hiding its search capabilities in its training data sets for more than two years. When asked if they could identify copyrighted materials used to train their models, OpenAI responded that they couldn't, while the publishers now claim that response was false. It is essential to note that Microsoft is not a direct target of this particular lawsuit and, although it was included as a defendant in the original case, the focus of this accusation falls on OpenAI, suggesting an intention to directly accuse the company's conversation model.
The publishers argue that the 20 million ChatGPT records deleted could demonstrate that the model reproduced or closely imitated copyrighted news articles. While this may seem like a technical issue, it is essential to analyze its implications for investors or readers. The deletion of these records, along with the claim that OpenAI hid them, could have severe consequences for the company. According to some AI experts, OpenAI's search capabilities could have been crucial to creating its language model, and not revealing them could be considered a grave breach of good faith in the legal process, affecting the company's credibility and reputation in the market.
OpenAI's lawyers have already responded to these accusations, arguing that the deleted records were not relevant to the case and that their team did not break any laws. However, the publishers insist that the ChatGPT records are essential to proving copyright infringement arguments against the company. Therefore, this case continues in court, raising new questions: What consequences would follow if the court declares OpenAI guilty of allegedly deleting essential records? How would this affect the company as a whole? What can be expected in the short and long term for investors and users of its ChatGPT model?
It's essential to remember that this case goes beyond the simple legal dispute. It involves a question of transparency in language model creation and how these companies obtain their training data. Our team will continue to monitor this case and analyze its impact on the company, the AI industry, and investors in the long term.