Dataset Transparency and Copyright Accountability in AI Training Systems
DOI:
https://doi.org/10.64618/Keywords:
Generative Artificial Intelligence, Copyright Infringement, AI Training Data, Algorithmic Accountability, Fair Use and Fair Dealing, Protection of Creative LabourAbstract
AI companies are reluctant to reveal their training datasets, arguing that it could expose trade secrets, compromise proprietary algorithms, or reveal sensitive personal data. However, without knowing what is being fed to the algorithms, it becomes nearly impossible for the rightsholder or artist to prove that their work has been used without their permission. The case of Hayao Miyazaki is a prime example. Despite his public opposition to AI-generated Ghibli-style art, there was no easy way for him to verify whether his works were used in training models, leaving him powerless to enforce his rights. This paper argues that when artists cannot know whether their works have been used to train AI systems, copyright protection becomes illusory rather than real. Further, the onus of proof for copyright infringement allegations requires that plaintiffs prove unlawful use and sufficient similarity between the work and the copyrighted work being copied. In the context of AI technology, establishing the burden of proof becomes exceedingly difficult, if not impossible, when the underlying data utilized lacks transparency and remains undisclosed. This state of affairs also brings about asymmetry because artists cannot prove their rights without knowledge about the data used by developers of AI. The issue is further complicated by the nature of generative AI models, which do not store or reproduce specific images but learn patterns and styles from large datasets. As a result, proving the case of direct copying is significantly problematic since the outputs are not literal copies but synthetic generations based on learned features. The paper proposes policy reforms and provenance tools (Whitebox AI, watermarking, cryptographic hashing, blockchain tracking) to verify usage, enforce rights, and require dataset disclosure. The absence of these tools and a legal disclosure mandate currently permits unauthorized exploitation of artists' work, thus weakening copyright protection.
Downloads
Published
Issue
Section
Categories
License
Copyright (c) 2026 Journal on Development of Intellectual Property and Research

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
