AI Model Watermarking With Periodic Signals for Distillation Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing model watermarking solutions are ineffective against ensemble distillation attacks and lack customizability and accuracy in identifying replicated AI models.
Innovation Solution
A method for embedding a periodic watermark signal into AI model outputs during training, which persists through distillation and ensemble averaging, allowing detection through Fourier analysis of model predictions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If model watermarking is implemented to identify replicated models, then model ownership detection capability is improved, but existing solutions are ineffective against ensemble distillation attacks and lack customizability
Solution Approach 1:
The patent embeds periodic watermark signals into the AI model's output probabilities. These periodic signals have specific frequencies that can be detected through Fourier analysis, allowing reliable identification of replicated models even when subjected to ensemble distillation attacks. The periodic nature of the watermark ensures it persists through the distillation process.
Solution Approach 2:
The watermarking system allows customization of watermark parameters including frequency, amplitude, and phase. These parameters can be adjusted to optimize detection capability against different types of model replication attacks, providing both reliability and adaptability to various threat scenarios.
2Measurement precision
If watermark signals are embedded in AI model outputs, then detection accuracy is improved, but the watermark must be robust to ensemble averaging which reduces signal strength
Solution Approach 1:
By using periodic watermarks with specific frequencies, the system can detect watermarks through Fourier analysis even when the signal strength is reduced by ensemble averaging. The periodic structure allows the watermark to stand out from the averaged output through frequency domain analysis.
Solution Approach 2:
The watermark signal is copied across multiple model outputs and persists through the distillation process. Even when multiple watermarked models are averaged, the periodic watermark pattern is replicated and can be detected through spectral analysis, maintaining detection accuracy.
3Reliability
If watermark embedding modifies the AI model, then ownership identification is enabled, but model performance may be degraded
Solution Approach 1:
The watermark is embedded locally in the probability outputs of the AI model rather than modifying the core model architecture or weights. This localized embedding approach enables ownership identification while minimizing impact on the model's primary predictive performance.
Solution Approach 2:
The watermark embedding adjusts output probability parameters with small perturbations that are sufficient for detection but minimal enough to preserve model performance. The amplitude of the watermark signal can be tuned to balance detection reliability with performance preservation.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables accurate identification of replicated models with high customizability, robustness to ensemble distillation, and applicability across various architectures and tasks.
Implementation Method 1
predicting, using the first AI model, a respective set of prediction outputs that each include a probability value, the AI model using a watermark function to insert a periodic watermark signal in the probability values of the prediction outputs
Implementation Method 2
allowing detection through Fourier analysis of model predictions
Data Source
AI summary
Method and system for watermarking prediction outputs generated by a first AI model to enable detection of a target AI model that has been distilled from the prediction outputs. Includes receiving, at the first AI model, a set of input data samples from a requesting device; storing at least a subset of the input data samples to maintain a record of the input data samples; predicting, using the first AI model, a respective set of prediction outputs that each include a probability value, the AI model using a watermark function to insert a periodic watermark signal in the probability values of the prediction outputs; and outputting, from the first AI model, the prediction outputs including the periodic watermark signal.


