Machine Learning Training Data Watermarking for Provenance Verification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The quality of machine learning models is compromised when training data is sourced from untrusted or inappropriate sources, leading to inaccuracies and lack of trustworthiness, particularly in mission-critical applications.

Innovation Solution

A method of watermarking training data to ensure legitimacy and trustworthiness by adding noise that can be verified using a one-way function, ensuring the data is from a trusted source and relevant to the application, and removing the noise after verification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If training data is sourced from untrusted or inappropriate sources to improve model training efficiency and data availability, then productivity is improved, but reliability deteriorates due to data quality issues and lack of trustworthiness

Engineering Contradiction:
Improvemodel training efficiencyVSAvoiddata trustworthiness
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies preliminary action by embedding watermarks in training data before the training process begins. The watermarking is performed in advance on the training dataset, allowing verification of data provenance and quality before the model training commences. This ensures that only verified, trustworthy data is used for training, resolving the contradiction between using readily available data and ensuring data reliability.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If watermarking noise is added to training data to verify legitimacy and enhance reliability, then reliability is improved, but device complexity increases due to additional verification mechanisms

Engineering Contradiction:
Improvedata legitimacy verificationVSAvoidverification system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies parameter changes by modifying the training data with embedded watermarks that contain verification information. The watermarking process changes the parameters of the training data by adding subtle noise patterns or metadata that encode provenance information. This allows verification of data legitimacy without requiring complex additional verification systems, as the watermark itself contains the necessary verification parameters.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If training data is verified and watermarked before use to ensure quality and relevance, then reliability is improved, but loss of time increases due to additional verification steps

Engineering Contradiction:
Improvetraining data qualityVSAvoidverification processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by performing watermarking and verification of training data in advance, before the actual model training process. The training data is watermarked and verified once during data preparation, and this verification status can be reused throughout the training process. This approach ensures data quality while minimizing time loss, as the verification is performed once beforehand rather than repeatedly during training.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250278927A1Watermarking Machine Learning Training Data for Verification Thereof
Publication Date: 2025.09.04 DIGICERT INC
  • US20250278927A1 patent drawing
  • US20250278927A1 patent drawing
  • US20250278927A1 patent drawing

AI summary

Systems and methods are disclosed for securing and validating training data used to train machine learning models through watermarking. Initially, watermark entries indistinguishable from genuine data are algorithmically generated from actual training dataset entries. These watermark entries are embedded into the training dataset, forming a watermarked dataset, which is then published for subsequent use. Verification involves applying a predetermined one-way detection function configured to identify watermark entries, confirming dataset authenticity without allowing unauthorized watermark generation. Optionally, the watermarked dataset can also be digitally signed by the dataset provider using a cryptographic certificate, enhancing trust through verified identity. Upon successful watermark detection and optional signature validation, the dataset is approved for machine learning model training, ensuring the integrity and legitimacy of data used in critical ML applications.