Stories
250 words
2
Entries

Stories

Legal and technical challenges regarding AI training data and copyright

2 entries over 37 days, from 25 July 2026 to 30 August 2026.

01
25 July

Delhi High Court ruling on OpenAI copyright suit

  • Delhi High Court denied interim injunction to ANI in copyright infringement suit against OpenAI.
  • Court ruled that OpenAI's use of ANI's literary works for training LLMs is covered by the 'fair dealing' exception under Section 52(1)(a) of the Copyright Act.
  • Court noted lack of evidence showing ANI suffered loss of subscribers or news syndication revenue due to OpenAI's activities.
  • ANI had previously offered a license for its content to OpenAI for USD 7.5 million, indicating the claim is quantifiable in monetary terms.
  • Court observed that an interim injunction would cause irreparable prejudice to OpenAI and the public interest.
  • The court highlighted that the development of LLMs depends on vast data availability and requiring licenses from multiple sources would be economically unviable.
  • OpenAI argued that deleting training data would conflict with U.S. legal obligations to preserve such data.
02
30 August

Generative AI training and copyright issues

  • MIT CSAIL research paper, 'Outputs of Generative Diffusion Models are Often Unattributable,' challenges the claim that AI output is causal theft of training data.
  • Ablatable ensembles method allows researchers to remove specific training data to measure the influence of individual samples on AI output.
  • Research indicates that as training datasets grow, the influence of any single image or artist on the final output decays toward zero.
  • Diffusion models (e.g., DALL-E, Stable Diffusion, Midjourney) generate images by removing noise from a canvas, learning geometric and semantic concepts rather than memorizing specific files.
  • Autoregressive models (LLMs like ChatGPT) differ from diffusion models as they predict sequences of tokens, making them more prone to verbatim reproduction of copyrighted text.
  • Scale acts as a technical shield for diffusion models; when trained on billions of images, the statistical contribution of one piece of art becomes insignificant.
  • AI labs are exploring the use of synthetic data generated by current AI models to train future models, further anonymizing the data source.
  • The legal distinction between 'visual similarity' and 'causal theft' remains a central point of contention in copyright litigation.

Questions from this story

Newest first. A story that ran for 37 days is exactly the kind the mains paper asks about as one question.

  1. Discuss the challenges posed by generative AI models to existing Intellectual Property Rights (IPR) frameworks, specifically distinguishing between diffusion-based image generators and autoregressive language models. 150 words · 30 August
  2. The rapid evolution of generative AI necessitates a re-evaluation of copyright laws in the digital age. In light of recent research on the 'unattributability' of AI outputs, analyze the legal and ethical complexities involved in balancing technological innovation with the protection of creative works. 250 words · 30 August
  3. Discuss the challenges posed by Artificial Intelligence development to existing Intellectual Property Rights frameworks in India, with reference to the 'fair dealing' exception under the Copyright Act. 150 words · 25 July
  4. The balance between protecting the rights of content creators and fostering technological innovation is critical for the digital economy. Analyze this statement in the context of Large Language Models (LLMs) and the evolving judicial interpretation of copyright law in India. 250 words · 25 July