AI Training Data Provenance with TEE Attestation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing AI model training data provenance systems lack integrity and transparency due to user-discretionary ledger information and non-secure compute environments, posing risks to model integrity and compliance with regulatory standards.

Innovation Solution

Implementing a confidential computing environment with a Code Transparency Service (CTS) and Trusted Execution Environment (TEE) to ensure secure, autonomous ledger operations, using cryptographic digests and hardware-backed security to verify and sign training data, firmware, and hardware components, ensuring only attested artifacts are used for model training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional ledger systems are used for tracking AI training data provenance, then the system is easier to implement and operate, but the integrity and reliability of the tracked information cannot be guaranteed

Engineering Contradiction:
Improveintegrity of training data provenanceVSAvoidcomplexity of secure compute environment
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces a Confidential Computing Platform as an intermediary between the training data provenance tracking system and the underlying hardware infrastructure. This platform includes a Trusted Execution Environment (TEE) that mediates secure computation and a Confidential Ledger that mediates secure data storage and tracking. The intermediary layer provides cryptographic attestation and secure enclaves, ensuring that provenance information is tracked with guaranteed integrity while abstracting the complex security mechanisms from end users.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If information is placed on the ledger at user discretion, then the system is easier to operate and maintain, but the integrity and correctness of the data cannot be guaranteed

Engineering Contradiction:
Improveintegrity of ledger informationVSAvoidease of ledger maintenance
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent implements a feedback mechanism through cryptographic attestation. The Trusted Execution Environment continuously monitors and attests to the integrity of training data and processing operations, providing feedback to the Confidential Ledger. This attestation process ensures that only verified, un tampered data is recorded on the ledger, guaranteeing information integrity while automating the verification process to maintain ease of operation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system employs self-service mechanisms where the Trusted Execution Environment automatically performs integrity verification and cryptographic signing of provenance information. The secure enclaves autonomously generate attestation reports and update the Confidential Ledger without requiring manual intervention, ensuring data integrity while maintaining operational simplicity.

Inventive Principle:
Principle #25Self-service

3Reliability

If traditional compute environments are used for AI model training, then the system is simpler and more accessible, but the security and confidentiality of training data and models are compromised

Engineering Contradiction:
Improvesecurity of training data and modelsVSAvoidcomplexity of confidential computing infrastructure
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a nested architecture where Trusted Execution Environments (secure enclaves) are nested within the Confidential Computing Platform, which itself is nested within the broader AI training infrastructure. This nested structure allows secure computation to be embedded within existing systems, providing enhanced security for training data and models while leveraging the existing computational resources and accessibility of traditional AI training environments.

Inventive Principle:
Principle #7Nested doll (Nesting)

Data Source

PatentUS20260064890A1Training Data Provenance System and Method
Publication Date: 2026.03.05 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20260064890A1 patent drawing
  • US20260064890A1 patent drawing
  • US20260064890A1 patent drawing

AI summary

A method, computer program product, and computing system for generating signed data for training an artificial intelligence (AI) model by processing data stored on a ledger using a signing authority. Signed firmware is generated for training the AI model by processing data stored on the ledger using the signing authority. The AI model is trained with signed data and the signed firmware from the ledger using a data processing unit in response to determining that the signed data and the signed firmware are signed by the signing authority.