AI Training Data Provenance with TEE Attestation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing AI model training data provenance systems lack integrity and transparency due to user-discretionary ledger information and non-secure compute environments, posing risks to model integrity and compliance with regulatory standards.
Innovation Solution
Implementing a confidential computing environment with a Code Transparency Service (CTS) and Trusted Execution Environment (TEE) to ensure secure, autonomous ledger operations, using cryptographic digests and hardware-backed security to verify and sign training data, firmware, and hardware components, ensuring only attested artifacts are used for model training.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional ledger systems are used for tracking AI training data provenance, then the system is easier to implement and operate, but the integrity and reliability of the tracked information cannot be guaranteed
Solution Approach 1:
The patent introduces a Confidential Computing Platform as an intermediary between the training data provenance tracking system and the underlying hardware infrastructure. This platform includes a Trusted Execution Environment (TEE) that mediates secure computation and a Confidential Ledger that mediates secure data storage and tracking. The intermediary layer provides cryptographic attestation and secure enclaves, ensuring that provenance information is tracked with guaranteed integrity while abstracting the complex security mechanisms from end users.
2Reliability
If information is placed on the ledger at user discretion, then the system is easier to operate and maintain, but the integrity and correctness of the data cannot be guaranteed
Solution Approach 1:
The patent implements a feedback mechanism through cryptographic attestation. The Trusted Execution Environment continuously monitors and attests to the integrity of training data and processing operations, providing feedback to the Confidential Ledger. This attestation process ensures that only verified, un tampered data is recorded on the ledger, guaranteeing information integrity while automating the verification process to maintain ease of operation.
Solution Approach 2:
The system employs self-service mechanisms where the Trusted Execution Environment automatically performs integrity verification and cryptographic signing of provenance information. The secure enclaves autonomously generate attestation reports and update the Confidential Ledger without requiring manual intervention, ensuring data integrity while maintaining operational simplicity.
3Reliability
If traditional compute environments are used for AI model training, then the system is simpler and more accessible, but the security and confidentiality of training data and models are compromised
Solution Approach 1:
The patent implements a nested architecture where Trusted Execution Environments (secure enclaves) are nested within the Confidential Computing Platform, which itself is nested within the broader AI training infrastructure. This nested structure allows secure computation to be embedded within existing systems, providing enhanced security for training data and models while leveraging the existing computational resources and accessibility of traditional AI training environments.
Data Source
AI summary
A method, computer program product, and computing system for generating signed data for training an artificial intelligence (AI) model by processing data stored on a ledger using a signing authority. Signed firmware is generated for training the AI model by processing data stored on the ledger using the signing authority. The AI model is trained with signed data and the signed firmware from the ledger using a data processing unit in response to determining that the signed data and the signed firmware are signed by the signing authority.


