Generative Model Watermarking with White- and Black-Box Protection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing generative models face challenges in protecting ownership and integrity due to unauthorized use, tampering, and the need for robust watermarking solutions that do not affect performance.
Innovation Solution
A method involving the embedding of a white box watermark and a black box watermark into a generative model, with the black box watermark in the probability density function of data abstractions and the white box watermark in the model outputs, to provide double-layer protection against unauthorized use.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If model watermarking is implemented to protect ownership and integrity, then security and authentication are improved, but device complexity and implementation difficulty increase
Solution Approach 1:
The watermarking system is divided into two independent parts: black box watermarking (embedding watermarks in probability density functions without accessing model internals) and white box watermarking (embedding watermarks by accessing model parameters). This segmentation allows each component to be developed and applied independently, reducing overall system complexity while maintaining comprehensive protection
Solution Approach 2:
The patent introduces probability density functions as an intermediary medium between the model and the watermark. Instead of directly embedding watermarks in model parameters (white box) or output only (black box), the system uses probability density functions as a bridge that can be modified to encode watermarks while maintaining model functionality, thus mediating between security requirements and implementation complexity
2Reliability
If watermarks are embedded in model parameters to ensure robustness, then detection reliability is improved, but model performance may be affected
Solution Approach 1:
The patent applies different watermark embedding strategies to different parts of the model: black box watermarks are embedded in probability density functions of data abstractions in respective layers, while white box watermarks are embedded in model parameters. This local differentiation allows watermarks to be placed in locations that minimize impact on overall model performance while maintaining detection reliability
Solution Approach 2:
The system modifies parameters of probability density functions to embed watermarks rather than directly modifying model weights. By changing the parameters of the probability density function (which governs data abstraction distributions) rather than the model parameters themselves, the system achieves watermark embedding with minimal disruption to model performance
3Reliability
If multiple watermarking methods are combined for comprehensive protection, then security coverage is improved, but implementation complexity and computational overhead increase
Solution Approach 1:
The dual watermarks approach segments security protection into two independent systems: black box watermarks for protecting against unauthorized use and white box watermarks for protecting against tampering. Each system can be implemented and verified independently, reducing the computational overhead associated with managing a single complex multi-layer watermarking system
Solution Approach 2:
The patent implements watermarking at specific critical points rather than throughout the entire model uniformly. Black box watermarks are embedded in probability density functions at key layers, and white box watermarks are embedded in specific model parameters. This partial action approach provides comprehensive security coverage while minimizing the total computational overhead compared to uniform watermarking across all model components
Data Source
AI summary
Embodiments of the present disclosure provide a method for determining a generative model. The method includes embedding a white box watermark and a black box watermark into a generative model. The black box watermark is first embedded into a probability density function of data abstractions in respective layers of the generative model. The method further includes embedding, after the embedding of the black box watermark is completed, the white box watermark into respective layers for outputs of the generative model. Model data is generated by the generative model based on predetermined triggering data. The predetermined triggering data includes a predetermined triggering text or a predetermined triggering image. An identity associated with the generative model is determined based on the model data. Advantageously, the illustrative method is capable of providing double-layer protection for a generative model by embedding two complementary and independent watermarks to resist white box and black box attacks.


