Machine Learning Model Watermarking via Fairness Bias

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning model watermarking techniques, particularly backdoor-based methods, are vulnerable to ownership verification attacks, especially when repeated verification attempts are made, as they require secret trigger inputs that can be distinguished and mimicked, leading to potential disclosure of ownership information.

Innovation Solution

Introducing fairness bias into machine learning models by clustering data into groups and modifying labels for specific subsets, using a secret clustering algorithm, which serves as a watermark that is inherently secure and does not require separate verification inputs, allowing for frequent and secure ownership verification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If backdoor-based watermarking is used to embed ownership information, then ownership identification is enabled, but the system becomes vulnerable to verification attacks and repeated queries can disclose ownership information

Engineering Contradiction:
Improveownership verification securityVSAvoidownership information disclosure
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent converts the harmful concept of backdoor attacks into a beneficial watermarking mechanism. Instead of using secret trigger inputs that can be exploited, the invention embeds ownership information through fairness bias in the model's decision boundaries, transforming a security vulnerability into a robust protection mechanism that withstands repeated verification attempts

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

2Ease of operation

If secret trigger inputs are used for watermark verification, then ownership can be identified, but the verification process becomes vulnerable to mimicry and disclosure attacks

Engineering Contradiction:
Improveownership verificationVSAvoidverification security
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent extracts the ownership verification mechanism from the vulnerable backdoor approach and embeds it directly into the model's decision-making process. By integrating fairness bias into the training data and model architecture, the verification process no longer requires separate trigger inputs, eliminating the security vulnerability while maintaining ease of verification

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The invention merges the watermark embedding and verification processes into the model's core functionality. The fairness bias that is introduced during training becomes both the watermark embedding mechanism and the verification mechanism, eliminating the need for separate trigger inputs and making the system resistant to mimicry attacks

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If fairness bias is introduced through label modification in clustered data subsets, then a secure watermark is embedded, but the data distribution and model fairness are altered

Engineering Contradiction:
Improvewatermark securityVSAvoiddata bias
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent segments the training data into clusters and applies fairness bias modifications to only specific subsets of these clusters. This selective modification allows the watermark to be embedded in a controlled manner, affecting only the decision boundaries for particular data groups while preserving the overall model functionality and minimizing broad fairness issues

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240370741A1Machine learning model watermarking through fairness bias
Publication Date: 2024.11.07 SAP SE
  • US20240370741A1 patent drawing
  • US20240370741A1 patent drawing
  • US20240370741A1 patent drawing

AI summary

A machine learning model is watermarked through fairness bias. To do this, an original set of labeled data is obtained and clustered into a plurality of groups using a clustering algorithm. Labels for data in a subset of the groups are modified, inserting fairness bias into the subset. A machine learning model is trained based on the subset of data labeled using the modified labels and the original set of data outside of the subset labeled using the original set of labels. The machine learning model trained as such exhibits the fairness bias when classifying input data belonging to subset of the plurality of groups. A model exhibiting the fairness bias for input data belonging to the subset is a watermark of a machine learning model that was trained using the modified labels for the subset determined based on the subgroup algorithm. The watermark is usable to determine ownership.