AI Training Data Auditing With Blockchain Validation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing training datasets for AI models are vulnerable to alteration or contamination, which can lead to inaccurate predictions and potential harm, and there is a lack of auditable methods to ensure data integrity and compliance with regulatory requirements.
Innovation Solution
A decentralized peer-to-peer (P2P) computer network using blockchain technology to validate and record data based on pre-defined criteria, including digital image processing and adversarial vulnerability testing, generating immutable data records that ensure data integrity and lineage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If training data is stored in centralized databases, then data access and processing is efficient, but data integrity and security are compromised due to vulnerability to alteration or contamination
Solution Approach 1:
The patent segments the training data into individual records, each with its own cryptographic hash and metadata, stored as separate entries in the blockchain. This segmentation allows each data point to be independently verified and tracked, enhancing data integrity while maintaining manageable system complexity through modular verification processes.
Solution Approach 2:
The patent introduces blockchain technology as an intermediary layer between the training data and the AI model. This intermediary provides immutable verification of data integrity through cryptographic hashing and distributed ledger technology, ensuring data reliability without requiring complex centralized security systems.
2Object-affected harmful factors
If traditional data validation methods are used, then validation process is simple, but adversarial vulnerabilities and data contamination cannot be detected
Solution Approach 1:
The patent implements preliminary validation actions by applying cryptographic hashing and storing verification metadata in the blockchain before the training data is used to train AI models. This preliminary verification ensures that adversarial vulnerabilities and data contamination are detected and prevented before they can affect model training, without adding complexity to the actual training process.
Solution Approach 2:
The patent establishes a feedback mechanism where the blockchain stores verification results and validation metadata that can be retrieved and checked during the AI model training process. This feedback loop provides continuous verification of data integrity, enabling detection of adversarial vulnerabilities while maintaining a relatively simple validation framework through automated cryptographic checks.
3Loss of information
If data lineage tracking is implemented, then data provenance and compliance are improved, but data processing overhead increases
Solution Approach 1:
The patent creates cryptographic copies of data verification information (hashes and metadata) and stores them in the blockchain. These cryptographic copies serve as immutable records of data provenance and lineage, enabling verification without requiring access to or processing of the original training data, thus maintaining data processing speed while improving provenance tracking.
Solution Approach 2:
The patent replaces traditional mechanical or manual data lineage tracking systems with cryptographic hashing and blockchain technology. This substitution eliminates the need for complex manual tracking and verification processes, providing automated, efficient provenance tracking that does not increase data processing overhead.
Data Source
AI summary
A computing node in a P2P computer network obtaining a media item to be verified, determining a hash value for the media item based on a digital cryptographic hash function, retrieving a plurality of data records associated with the media item from a blockchain based at least in part on the hash value, wherein a data record is generated by a validation node in the decentralized P2P computer network, evaluating the plurality of data records to determine whether the media item has satisfied pre-defined validation criteria, and providing information describing the media item based at least in part on the plurality of data records, wherein the information provides an indication as to whether the media item satisfied the pre-defined validation criteria.


