Blockchain System for Tracking Copyrighted Data in LLM Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing generative AI and Large Language Models (LLMs) lack effective mechanisms to separate and track copyrighted materials, leading to potential copyright infringement and misuse of training data, without proper credit or compensation to content creators.
Innovation Solution
Implement a system that separates copyrighted data from model training, using blockchain technology to track and manage copyrighted materials through non-fungible tokens (NFTs) and smart contracts, ensuring data ownership and control, and providing accurate, reliable information surfacing via chatbots.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If copyrighted data is used to train generative AI models, then the model's knowledge and capabilities are improved, but copyright infringement and misuse of training data occur without proper credit or compensation to content creators
Solution Approach 1:
The patent segments the training data into two distinct categories: public domain data for model training and copyrighted data for reference only. This segmentation allows the model to learn from public data while respecting copyright restrictions on copyrighted materials, thereby reducing copyright infringement risks while maintaining model knowledge through alternative data sources
Solution Approach 2:
The patent introduces an intermediary verification system that checks whether training data is copyrighted before it enters the model training pipeline. This intermediary layer acts as a mediator between data sources and the training process, blocking copyrighted data from training while allowing public domain data to proceed, thus preventing copyright infringement while maintaining training effectiveness
2Reliability
If blockchain technology is implemented to track copyrighted materials, then data ownership and control are protected, but system complexity increases
Solution Approach 1:
The patent extracts the copyright verification and tracking functionality from the main model training system and implements it as a separate blockchain-based verification layer. This extraction allows the core training system to remain relatively simple while adding ownership protection through a dedicated blockchain infrastructure that operates independently but integrates with the training pipeline
3Object-affected harmful factors
If copyrighted data is separated from model training, then copyright compliance is improved, but the model's ability to access accurate information from copyrighted sources is reduced
Solution Approach 1:
The patent makes the verification system universal by enabling it to handle multiple types of data sources (public domain, copyrighted, and intermediate verification) within a single unified framework. This allows the model to access accurate information from copyrighted sources through the verification layer while maintaining copyright compliance, as the same system that blocks copyrighted training data also enables authorized access to verified information
Data Source
AI summary
Tracking data input to a generative artificial intelligence model (generative AI) or a large language model (LLM) involves receiving a plurality of objects comprising the data input to the model, generating a corresponding non-fungible token (NFT) for each object, assigning a corresponding smart contract to each NFT to control interactions with the NFT and its corresponding object, recording the NFT and corresponding smart contract to a block for writing to a blockchain, and writing the block to the blockchain.


