Blockchain Provenance Tracking for ML Model Inputs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning models lack provenance tracking, which prevents contributors from establishing ownership or controlling the use of outputs, due to unrecorded inputs and their associations with outputs.

Innovation Solution

A data provenance tracking component that stores information related to the creation and use of machine learning models on a blockchain network, broadcasting messages to record inputs and outputs, and associating them for accurate valuation and ownership attribution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If machine learning models operate without provenance tracking, then system simplicity is maintained, but ownership attribution and contributor recognition are lost

Engineering Contradiction:
Improvesystem simplicityVSAvoidprovenance information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent introduces a blockchain as an intermediary system between machine learning models and contributors. The blockchain records provenance information in a decentralized manner, serving as a mediator that tracks ownership and contributions without complicating the core model operation. This resolves the contradiction by adding provenance tracking through a separate, non-intrusive layer.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates digital copies of provenance information on the blockchain, which serve as immutable records of ownership and contributions. These copies allow contributors to establish ownership and control over model outputs without altering the actual model operation. The blockchain stores hashed versions and metadata of provenance data, enabling verification without direct manipulation of the model.

Inventive Principle:
Principle #26Copying

2Reliability

If provenance tracking is implemented using blockchain, then ownership attribution is improved, but system complexity increases

Engineering Contradiction:
Improveownership attributionVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the provenance tracking function from the main machine learning system. The blockchain handles ownership and provenance recording, while the ML model focuses on computation. This segmentation allows reliable ownership attribution without requiring the model itself to become complex. The provenance tracking is isolated in a separate, specialized system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces traditional centralized database systems with a blockchain-based distributed ledger for provenance tracking. This substitution uses cryptographic hashing and consensus mechanisms instead of conventional database management, providing reliable ownership attribution through immutable records while distributing the system complexity across multiple nodes rather than concentrating it in a single complex server.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If inputs and outputs are recorded on blockchain, then provenance accuracy is improved, but data storage requirements increase

Engineering Contradiction:
Improveprovenance accuracyVSAvoiddata storage
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential provenance information (metadata) onto the blockchain rather than storing all input and output data. The blockchain records hashes, timestamps, contributor identifiers, and relationship metadata, while the actual data remains in off-chain storage. This extraction approach maintains provenance accuracy through immutable cryptographic records while dramatically reducing on-chain storage requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent implements a nested storage architecture where provenance metadata is stored on the blockchain and references point to detailed data stored off-chain. The blockchain contains hashed versions and reference information, while the full data resides in external databases or storage systems. This nesting allows accurate provenance tracking through the blockchain layer without requiring all data to be stored on-chain, efficiently managing storage requirements.

Inventive Principle:
Principle #7Nested doll (Nesting)

Data Source

PatentUS20250200334A1Tracking machine learning data provenance via a blockchain
Publication Date: 2025.06.19 COINBASE INC
  • US20250200334A1 patent drawing
  • US20250200334A1 patent drawing
  • US20250200334A1 patent drawing

AI summary

Methods, systems, and devices for data management are described. A middleware component may receive, for generating a machine learning model, one or more user inputs associated with the machine learning model and an indication of a data source for training the machine learning model. After receiving the user inputs and the data source, the middleware component may broadcast one or more first blockchain messages that are configured to store first information associated with the one or more user inputs and the data source on a blockchain network. The middleware component may receive input prompts for the machine learning model and one or more responses generated by the machine learning model, and, after receiving the input prompts, broadcast one or more second blockchain messages that are configured to store second information associated with the one or more input prompts and the one or more responses on the blockchain network.