Machine Learning Model Watermarking via Fairness Bias
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning model watermarking techniques, particularly backdoor-based methods, are vulnerable to ownership verification attacks, especially when repeated verification attempts are made, as they require secret trigger inputs that can be distinguished and mimicked, leading to potential disclosure of ownership information.
Innovation Solution
Introducing fairness bias into machine learning models by clustering data into groups and modifying labels for specific subsets, using a secret clustering algorithm, which serves as a watermark that is inherently secure and does not require separate verification inputs, allowing for frequent and secure ownership verification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If backdoor-based watermarking is used to embed ownership information, then ownership identification is enabled, but the system becomes vulnerable to verification attacks and repeated queries can disclose ownership information
Solution Approach 1:
The patent converts the harmful concept of backdoor attacks into a beneficial watermarking mechanism. Instead of using secret trigger inputs that can be exploited, the invention embeds ownership information through fairness bias in the model's decision boundaries, transforming a security vulnerability into a robust protection mechanism that withstands repeated verification attempts
2Ease of operation
If secret trigger inputs are used for watermark verification, then ownership can be identified, but the verification process becomes vulnerable to mimicry and disclosure attacks
Solution Approach 1:
The patent extracts the ownership verification mechanism from the vulnerable backdoor approach and embeds it directly into the model's decision-making process. By integrating fairness bias into the training data and model architecture, the verification process no longer requires separate trigger inputs, eliminating the security vulnerability while maintaining ease of verification
Solution Approach 2:
The invention merges the watermark embedding and verification processes into the model's core functionality. The fairness bias that is introduced during training becomes both the watermark embedding mechanism and the verification mechanism, eliminating the need for separate trigger inputs and making the system resistant to mimicry attacks
3Reliability
If fairness bias is introduced through label modification in clustered data subsets, then a secure watermark is embedded, but the data distribution and model fairness are altered
Solution Approach 1:
The patent segments the training data into clusters and applies fairness bias modifications to only specific subsets of these clusters. This selective modification allows the watermark to be embedded in a controlled manner, affecting only the decision boundaries for particular data groups while preserving the overall model functionality and minimizing broad fairness issues
Data Source
AI summary
A machine learning model is watermarked through fairness bias. To do this, an original set of labeled data is obtained and clustered into a plurality of groups using a clustering algorithm. Labels for data in a subset of the groups are modified, inserting fairness bias into the subset. A machine learning model is trained based on the subset of data labeled using the modified labels and the original set of data outside of the subset labeled using the original set of labels. The machine learning model trained as such exhibits the fairness bias when classifying input data belonging to subset of the plurality of groups. A model exhibiting the fairness bias for input data belonging to the subset is a watermark of a machine learning model that was trained using the modified labels for the subset determined based on the subgroup algorithm. The watermark is usable to determine ownership.


