Model watermark robustness enhancement method and system based on antagonistic immunity and watermark application

By introducing an adversarial immune mechanism into the model watermarking technology, and using variational autoencoder and adversarial samples for robust training, the existing model watermarking technology is solved, and the problem of insufficient robustness in the face of multiple rounds of processing and careful design attacks is achieved, and stronger watermark extraction effect and resistance are achieved.

CN120107053AActive Publication Date: 2025-06-06XIDIAN UNIV

Patent Information

Application Number
CN202510250766.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-06
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

Existing model watermarking technology is vulnerable to damage in the face of multiple rounds of processing and carefully designed watermark removal attacks, and is difficult to withstand unknown attack methods.

Method used

The model watermark robustness enhancement method based on adversarial immunity is adopted. By using a variational autoencoder as the initial generation of the watermark model, and introducing adversarial samples during the training process for robust training, the extraction effect and resistance of the model watermark are enhanced.

Benefits of technology

It significantly improves the quality and concealment of watermark data, enhances the robustness and generalization capabilities of model watermarks, and can effectively resist various watermark removal attacks, ensuring the integrity and extractability of watermarks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107053A_ABST
    Figure CN120107053A_ABST
Patent Text Reader

Abstract

The invention discloses a model watermark robustness enhancement method based on antagonistic immunity, which mainly solves the problem that the existing model watermark is difficult to resist various watermark removal attacks, and adopts the scheme that a public data set and a convolutional neural network suitable for an image classification task are respectively used as original data and a target model; training a target model by using the original data; taking the generation model as a watermark generation model and training the watermark generation model by using the original data; inputting the original data into the trained generated watermark model for forward propagation to output watermark data; splicing the original data and the watermark data, and inputting the spliced original data and watermark data into the trained target model for forward propagation to obtain a model watermark; tiny disturbance is added to original data to generate an adversarial sample, the adversarial sample is spliced with the original data and then input into the model watermark for robustness training, and the model watermark with enhanced robustness is obtained. The method can effectively enhance the generalization ability of the model watermark under different data distributions, improves the robustness of the model, and can be used in the fields of information security, financial services, medical care and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence security technology, and in particular to a model watermark robustness enhancement method, system and watermark application, which can be used in many fields such as information security industry, financial services, medical care and legal services. Background Art

[0002] With the widespread popularity of artificial intelligence models and the acceleration of commercialization, model watermarking technology has emerged in the field of deep learning and has become an important research direction. Its core is to open up a way for model owners to confirm their ownership.

[0003] At present, model watermarking technology can be roughly divided into two types: black-box watermarking and white-box watermarking, depending on whether the watermark extraction requires the participation of the model. In the black-box model watermarking scheme, it can be divided into data-independent model watermarking, statistical characteristic-based model watermarking, and backdoor model watermarking according to the embedding and verification principles of the watermark. Among them, the backdoor model watermark is an anti-tampering technology that combines backdoor attacks with watermarking technology. It is usually used to protect the intellectual property rights of deep learning models or embed certain information without affecting the performance of the model. The core idea is to add a hidden backdoor to the model and use the backdoor to embed watermark information. This method can not only cleverly hide the watermark information, but also effectively prove the ownership of the intellectual property rights through the behavior of the model.

[0004] The patent document with application number CN202411448920.X discloses a "robust watermarking method for screen-shot images based on attention mechanism and contrastive learning". It first uses an encoder to generate an encoded image containing watermark information; secondly, the encoded image and the carrier image are input into the discriminator, and the discriminator is used to output the expected value; then the encoded image is distorted and simulated; the distorted encoded image is input into the decoder to extract the watermark information hidden in the distorted encoded image; finally, according to the predicted value of the discriminator and the joint loss function, the encoder, discriminator and decoder are trained, and the trained encoder and decoder are used to form a screen-shot watermark model, and the screen-shot image is encoded and decoded using the screen-shot watermark model. This method ensures the invisibility of the encoded image by optimizing the process of watermark image encoding, but its method of generating watermarks based on the attention mechanism can be interfered with and misled by attackers. Attackers can add specific interference information to the image to make the attention mechanism focus on the wrong area, thereby ignoring the important part of the watermark information, making the watermark easy to remove.

[0005] The patent document with application number CN202411011804.1 discloses "a robust generative model watermark processing method and system that is resistant to attack". It first obtains a data set based on the target model to be embedded with the watermark, where the data set is the input and output data pair of the target model to be embedded with the watermark; secondly, the step data set is divided into two parts: a first data set and a second data set, where the first data set is used in the initial training stage of the model watermark, and the second data set is used in the adversarial training stage of the model theft; then in the initial training stage of the model watermark, the first data set is sent to the watermark embedding network for marking, and then the marked image is subjected to simulated attacks and watermark extraction; finally, in the adversarial training stage of the model theft, the second data set is used to train two alternative models, and then the watermark extraction network is fine-tuned to extract the watermark from the output of the alternative model containing the watermark. Although this method can improve the invisibility of the model watermark, there will be a problem of robustness loss in multiple rounds of processing. After multiple rounds of processing, the integrity and extractability of the watermark will be affected. Moreover, this method only simulates and defends against known common attack methods, and lacks sufficient resistance to watermark removal attacks carefully designed by attackers targeting their watermark generation and extraction mechanisms.

[0006] The patent document with application number CN202310525668.7 discloses "a method for constructing a reversible robust watermark embedding and extraction model against image attacks", which mainly includes a first Harr wavelet transformer, a channel coding module, a channel decoding module, a reversible module, an embedding and separation module, a first image attack module and a second image attack module. The channel coding module includes a channel decoder, an upsampler, an expander and a second Harr wavelet transformer connected in sequence; the channel decoding module includes a second separator, a first Harr wavelet inverse transformer, a third separator and a channel decoder connected in sequence; the reversible module includes multiple feature aggregation layers, and after the multiple feature aggregation layers are connected in sequence, they are also connected to a convolutional neural network; the embedding and separation model includes a first separator, a second Harr wavelet inverse transformer and a third Har wavelet transformer; the first image attack module includes a first attack layer, and the second image attack module includes a second attack layer and a third attack layer. This method improves the accuracy of watermark information extraction through the collaborative work of multiple modules, but the embedding strength of the watermark is limited in order to ensure the imperceptibility of the watermark, that is, not to affect the visual quality and usability of the original image. If the embedding strength is too low, the model will face challenges in maintaining the integrity and detectability of the watermark when facing some conventional image processing operations or carefully planned adversarial attacks, which will make the model watermark easy to remove. Summary of the invention

[0007] The purpose of the present invention is to address the deficiencies of the above-mentioned existing model watermarking technology and propose a model watermark robustness enhancement method based on adversarial immunity to improve the extraction effect of the model watermark and ensure that it can effectively resist various watermark removal attacks.

[0008] The technical solution for achieving the purpose of the present invention includes a method and system for enhancing the robustness of a model watermark and the use of the model watermark.

[0009] 1. A model watermark robustness enhancement method based on adversarial immunity, characterized by comprising:

[0010] (1) Select an existing convolutional neural network model suitable for image classification tasks as the initial target model, use a public dataset suitable for image classification tasks as raw data to train the target model, and obtain a trained target model;

[0011] (2) Using the existing variational autoencoder VAE including an encoder and a decoder as the initial generative watermark model, and using the original data to train the initial generative watermark model to obtain a trained generative watermark model;

[0012] (3) Input the original data into the trained watermark generation model to generate watermark data x wm ;

[0013] (4) The watermark data is concatenated with the original data and then input into the trained target model for forward propagation to complete the embedding of the model watermark;

[0014] (5) Add tiny perturbations that are imperceptible to the naked eye to the original data to generate adversarial samples;

[0015] (6) The original data is input into the model watermark after watermark embedding for forward propagation, and then the adversarial sample is concatenated with the original data and input into the model watermark for robustness training to obtain a highly robust model watermark.

[0016] 2. A model watermarking system for enhancing robustness, characterized by comprising:

[0017] The dataset preprocessing module is used to preprocess the public dataset suitable for image classification tasks by horizontal flipping, pixel padding, random cropping, and normalization operations in sequence, and use the preprocessed dataset as the original data;

[0018] The target model training module is used to input the original data into the initial target model for forward propagation to obtain the trained target model;

[0019] The generated watermark model training module is used to input the original data into the initial generated watermark model for forward propagation to obtain a trained generated watermark model;

[0020] The watermark data generation module is used to input the original data into the trained watermark generation model, perform feature conversion and information hiding on the original data through the encoder of the watermark generation model, and then optimize and extract the intermediate data generated in the encoding process through its decoder to obtain the generated watermark data;

[0021] The watermark data embedding module is used to splice the generated watermark data with the original data and input them into the trained target model for forward propagation to obtain the model watermark;

[0022] The model watermark robust training module is used to add tiny perturbations that are imperceptible to the naked eye to the original data to obtain adversarial samples, and then concatenate the original data with the adversarial samples and input them into the model watermark for forward propagation, thereby realizing robust training of the model watermark and obtaining a model watermark with enhanced robustness.

[0023] 3. A method for using a model watermark to enhance robustness, characterized by comprising:

[0024] S1) The user selects a watermark extraction algorithm that matches the specific algorithm used by the developer when embedding the watermark in the model;

[0025] S2) The user loads the model watermark into the set extraction environment and runs the extraction algorithm to obtain the watermark information. During this process, the stability and security of the environment must be ensured to prevent external interference from affecting the accuracy of the extraction results;

[0026] S3) Check whether the key features such as the specific identification of the copyright owner and the copyright statement pre-set in the extracted watermark information can be correctly identified:

[0027] If all key features in the watermark can be correctly identified, it indicates that the watermark is of good integrity, and step S5 is executed;

[0028] Otherwise, it indicates that the watermark is missing or damaged, and an attempt is made to repair the watermark information, and step S4 is executed;

[0029] S4) Repair the watermark information using the algorithm used when the model is embedded in the watermark, and check whether the key features in the repaired watermark can be correctly identified:

[0030] If it is correctly identified, it indicates that the watermark repair is successful, and S5 is executed;

[0031] Otherwise, it indicates that the copyright of the model is suspected of being misappropriated, and further investigation is needed in combination with the dissemination channels of the model;

[0032] S5) Compare the key features of the extracted watermark information with the key features of the original watermark information to determine whether the copyright of the model has been stolen:

[0033] If the two are exactly the same, then the copyright of the model is correct;

[0034] Otherwise, it is determined that the copyright of the model has been plagiarized and corresponding copyright protection measures should be initiated to investigate the infringement liability.

[0035] Compared with the prior art, the present invention has the following advantages:

[0036] Firstly, the present invention uses variational autoencoder VAE as the initial generative watermark model. During the training process of the generative watermark model, the original data is input into the encoder of the generative watermark model, and the perceptual loss between the encoder output data and the original data is used as the key constraint condition for the generative watermark model training. Compared with the existing generative watermark model training method, the present invention significantly improves the quality of the watermark data generated by the generative watermark model, so that the watermark data generated by the training of the present invention is almost visually indistinguishable from the original data, which greatly improves the concealment of the watermark data when it is embedded in the model, effectively avoiding the risk of the model being identified or tampered with because the watermark is easily perceived, and providing more reliable technical support for model copyright protection and watermark application.

[0037] Secondly, the present invention performs robustness training on the trained model watermark, that is, adding tiny perturbations that are imperceptible to the naked eye to the original data to generate adversarial samples, and splicing the adversarial samples with the original data to form adversarial heterogeneous data, which is then input into the model watermark for forward propagation. In this process, a penalty mechanism is set to constrain the model watermark from predicting the original data as watermark data, and at the same time, the robustness training of the model watermark is constrained by the loss function of the adversarial heterogeneous data. Therefore, compared with the existing model watermark, the model watermark generated by the training of the present invention has stronger robustness and generalization ability, and can effectively resist various current watermark removal attacks. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a flow chart of a method for enhancing the robustness of a model watermark based on adversarial immunity according to an embodiment of the present invention;

[0039] Figure 2 It is a block diagram of a model watermark robustness enhancement system based on adversarial immunity according to an embodiment of the present invention;

[0040] Figure 3 This is a diagram of the method of using the model watermark of the present invention. DETAILED DESCRIPTION

[0041] The following is a detailed description of the embodiments of the present invention in conjunction with the accompanying drawings:

[0042] Example 1: Model watermark robustness enhancement method based on adversarial immunity of the present invention

[0043] In view of the current situation that the existing model watermark cannot effectively resist various watermark removal attacks, the present invention enhances the model watermark's resistance to various interferences and attacks through an additional robustness training, enhances the generalization ability of the model watermark under different data distributions, and obtains a highly robust model watermark.

[0044] Reference Figure 1 The implementation steps of this example include the following:

[0045] Step 1: Pre-training dataset.

[0046] 1.1) Select existing public datasets suitable for image classification tasks such as CIFAR10, CIFAR100, MNIST, ImageNet, SVHN, etc. as original data;

[0047] This example uses the CIFAR10 dataset as the original data. The images in this dataset are divided into 10 different categories, each category contains 6,000 images;

[0048] 1.2) The original data is preprocessed by horizontal flipping, pixel filling, random cropping and normalization operations in sequence, and the preprocessed original data is divided into training set and test set in a ratio of 5:1.

[0049] Step 2: Pre-train the target model.

[0050] 2.1) Select existing convolutional neural network models suitable for image classification tasks, such as ResNet18, ResNet50, VGG, AlexNet, etc. as the initial target model;

[0051] This example uses the ResNet18 convolutional neural network as the initial target model, which includes 18 convolutional layers and 1 fully connected layer;

[0052] 2.2) Input the training set data into the initial target model in batches for forward propagation, and use the cross entropy loss to measure the difference between the probability distribution predicted by the model and the label distribution of the original data;

[0053] 2.3) Select the stochastic gradient descent optimizer as the optimizer of the target model, calculate the gradient of the above cross entropy loss with respect to the target model parameters, and then update the parameters of the target model according to the opposite direction of the gradient;

[0054] 2.4) Use the data from the test set to evaluate the target model and calculate the model accuracy on the test set;

[0055] 2.5) Repeat steps 2.2) to 2.4) to train the target model for 100 iterations, select the target models with a model accuracy greater than 85%, sort these target models in descending order according to their accuracy, and save the highest-ranking target model locally.

[0056] Step three: pre-train and generate watermark model.

[0057] 3.1) Select the existing variational autoencoder VAE as the initial watermark generation model. Its structure consists of two parts: encoder and decoder. The encoder is used to map the input data to the latent space, learn the latent feature representation of the input data and calculate the latent variables. The decoder reconstructs the original data according to the latent variables and converts the representation in the latent space back to the space where the input data is located;

[0058] 3.2) Input the original data x into the encoder of the initial watermark generation model, and output the mean μ and logarithmic variance log(σ) of the latent space 2 ), the mean μ represents the central position of the data distribution in the latent space, which reflects the average characteristics of the original data x in the latent space, and the logarithmic variance log(σ 2 ) measures the dispersion of data in the latent space, and the logarithm is taken for the stability and convenience of numerical calculation;

[0059] 3.3) Through the mean μ and logarithmic variance log(σ of the latent space 2 ) determines the distribution of data in the latent space, and its calculation formula is as follows:

[0060]

[0061] Among them, f μ and are two different output branches of the encoder neural network;

[0062] 3.4) Sample a random variable ∈ from the latent space, combining the mean μ and logarithmic variance log(σ) of the encoder output 2 ), calculate the latent variable z:

[0063]

[0064] Among them, ⊙ represents element-by-element multiplication;

[0065] 3.5) Input the sampled latent variable z into the decoder of the initial watermark generation model, and output the reconstructed data with the same dimension as the original data x

[0066]

[0067] Among them, Decoder represents a decoder;

[0068] 3.6) Calculate the original data x and the reconstructed data The reconstruction loss L MSE In this step, you can choose the mean square error loss or the mean absolute error loss as the calculation of the reconstruction loss L MSE The implementations are as follows:

[0069] 3.6.1) Use mean square error loss to calculate reconstruction loss L MSE , traverse the original data x and reconstruct the data For each corresponding pixel in, calculate the original data x and the reconstructed data The average of the squares of the differences between the corresponding elements is taken as

[0070] Reconstruction loss L MSE :

[0071]

[0072] Among them, x i is the value of the i-th pixel in the original data x, is the value of the i-th pixel in the generated image, and N is the total number of pixels in the image;

[0073] 3.6.2) Calculate the reconstruction loss L using the mean absolute error loss MSE , traverse the original data x and reconstruct the data For each corresponding pixel in, calculate the original data x and the reconstructed data The average absolute error of the corresponding elements is taken as the reconstruction loss L MSE :

[0074]

[0075] Among them, x i is the value of the i-th pixel in the original data x, is the value of the i-th pixel in the generated image, and N is the total number of pixels in the image;

[0076] 3.7) Use existing pre-trained convolutional neural networks suitable for image classification tasks, such as VGG, MobileNet, ResNet18 and ResNet50, to measure the original data x and the reconstructed data Differences at the perceptual level;

[0077] This example uses the ResNet50 convolutional neural network as the pre-trained convolutional neural network, which consists of a convolutional layer, a pooling layer, multiple residual block groups, and a fully connected layer;

[0078] 3.8) The original data x and the reconstructed data They are input into the pre-trained ResNet50 convolutional neural network respectively to extract their image feature representation φ at the i-th layer i (x) and Calculate the Euclidean distance between these two feature representations to get the perceptual loss L Per :

[0079]

[0080] in, is the Euclidean distance, which represents the difference between two feature images;

[0081] 3.9) According to the random variable ∈, mean μ, logarithmic variance log(σ) in 3.3) 2 ) and the latent variable z, and calculate the difference between the posterior distribution q(z|x) of the latent variable z and the prior distribution p(z) as the KL divergence loss L KL :

[0082]

[0083] Among them, the posterior distribution q(z|x) satisfies the normal distribution N(μ,σ 2 ), the prior distribution p(z) satisfies the standard normal distribution N(0,1);

[0084] 3.10) The reconstruction loss L calculated above is MSE , Perceptual loss L Per And KL divergence loss L KL Add as the total loss of the generated watermark model The weight of the KL divergence loss is dynamically adjusted in a linear annealing manner, that is, as the training progresses, the annealing factor β(t) gradually increases and eventually reaches 1 in the late stage of training, so that the KL divergence term gradually plays an important role in the late stage of training:

[0085]

[0086] Among them, λ 1 , 2 L MSE and L Per The weight hyperparameters are adjusted according to training needs. t is the current training step, T is the total number of training steps,

[0087] 3.11) Use stochastic gradient descent optimizer to calculate the above total loss Regarding the gradient of the generated watermark model parameters, the model parameters are updated according to the opposite direction of the gradient to obtain the generated watermark model of the current training round;

[0088] 3.12) Repeat steps 3.2) to 3.11) to update the model multiple times until the loss Convergence, and obtain the trained generated watermark model.

[0089] Step 4: Generate watermark data.

[0090] 4.1) Input the original data x into the encoder of the trained watermark generation model, and output the mean μ' and logarithmic variance log(σ') of the latent space 2 ):

[0091]

[0092] Among them, f μ and They are two different output branches of the encoder neural network;

[0093] 4.2) Sample a random variable ∈' from the distribution of the latent space, using the random variable ∈', mean μ' and logarithmic variance log(σ') 2 ) to calculate the latent variable z':

[0094]

[0095] Among them, ⊙ represents element-by-element multiplication;

[0096] 4.3) Input the latent variable z' into the decoder of the trained watermark model and output the watermark data x wm :

[0097] x wm =Decoder(')

[0098] Among them, Decoder represents a decoder;

[0099] Step 5: Embed model watermark.

[0100] 5.1) Combine the original data x and the watermark data x wm Splice along the same dimension at a ratio of 1:1 to obtain the spliced ​​watermark enhancement data

[0101]

[0102] 5.2) Let the label of the original data x be y, and the watermark data x wm The label is y wm , and then the label y of the original data and the label y of the watermark data wm Splice along the same dimension at a ratio of 1:1 to obtain the spliced ​​watermark enhanced data label

[0103]

[0104] 5.3) Watermark Enhanced Data Input into the trained target model for forward propagation, and use the cross entropy loss function to calculate the probability distribution predicted by the model Enhanced data labeling with watermarks The difference between them is taken as the watermark loss L wn :

[0105]

[0106] Among them, L ce represents the cross entropy loss function;

[0107] 5.4) Calculate the watermark loss L using the stochastic gradient descent optimizer wm Regarding the gradient of the model parameters, the trained target model parameters are updated according to the opposite direction of the gradient to complete the watermark embedding training of the current round of target model;

[0108] 5.5) Repeat steps 5.3) to 5.4) to iterate the model multiple times until the watermark loss L wm Convergence, realize the embedding of model watermark.

[0109] Step 6: Generate adversarial samples.

[0110] In this step, adversarial samples can be generated using the optimization-based projected gradient descent method (PGD) or the gradient-based fast gradient sign method (FGSM). The implementation methods are as follows:

[0111] 6.1) Generate adversarial samples using the gradient-based fast gradient sign method (FGSM), which is implemented as follows:

[0112] 6.1.1) Input the original data x into the model watermark to obtain the predicted result f(x), and calculate the loss function L of the original data x based on the difference between the predicted result f(x) and the original data label y x :

[0113] L x =L ce (f(x),y)

[0114] Among them, L ce represents the cross entropy loss function;

[0115] 6.1.2) Calculate the loss function L x The gradient with respect to the original data x

[0116]

[0117] Where c is the number of categories, f i (x) is the i-th component of the model output f(x);

[0118] 6.1.3) According to the gradient of the original data x Using symbolic functions To determine the perturbation direction of the original data x;

[0119] Among them, sign(·) indicates the direction in which the sign function is disturbed. Taking the input data k as an example, its representation is:

[0120]

[0121] If the output result of the sign function is 1, the perturbation direction is positive; if the output result of the sign function is -1, the perturbation direction is reverse; if the output result of the sign function is 0, no perturbation is added to the original data x;

[0122] 6.1.4) Randomly select a value between 0.01 and 0.1 as the perturbation size ω, which is used to determine the degree of difference between the adversarial sample and the original data x;

[0123] 6.1.5) According to the determined disturbance direction And the perturbation size ω calculates the generated adversarial sample x Adv :

[0124]

[0125] 6.2) Generate adversarial samples using the optimization-based projected gradient descent (PGD) method, which is implemented as follows:

[0126] 6.2.1) Initialize a perturbation δ with the same shape as the original data x t =0, t=0,1,2,...,T, T is the maximum step size of iterative update;

[0127] 6.2.2) The disturbance δ t Superimposed on the original data x, we get the current adversarial sample x+δ t , the adversarial sample x+δ t Input into the model watermark to get the prediction result f(x+δ t ), according to the prediction result f(x+δ t ) and the difference between the original data label y to calculate the adversarial sample x+δ t The loss function L δ :

[0128] L δ =Lce (f(x+δ t ),y)

[0129] Among them, L ce represents the cross entropy loss function;

[0130] 6.2.3) Calculate the loss function L δ About the perturbation δ t Gradient

[0131]

[0132] Where c is the number of categories, f i (x+δ t ) is the model output f(x+δ t )’s i-th component;

[0133] 6.2.4) Along the gradient Update the perturbation δ of step t+1 in the direction of t+1 :

[0134]

[0135] Among them, δ t is the perturbation at step t, α is the perturbation step size at each step, and sign(·) indicates the direction in which the sign function obtains the perturbation;

[0136] 6.2.5) Update the perturbation δ t+1 Projected onto L with zero as the center and ∈ as the radius ∞ In the norm ball, we get the current constrained perturbation δ' t+1 :

[0137] δ′ t+1 =clip(δ t+1 ,-∈,∈)

[0138] Among them, the clip function converts δ t+1 Each element of is restricted to the range [-∈,∈];

[0139] 6.2.6) Repeat steps 6.2.2) to 6.2.5) until the maximum number of iterations T is reached, and the final result is

[0140] Constrained perturbation δ' T , and then the disturbance δ' T Add to the original sample x to generate the final adversarial sample x Adv :

[0141] x Adv =x+δ′ T .

[0142] Step 7: Robustness training of model watermark.

[0143] Although the model watermark obtained in the above steps has completely embedded the watermark information, it lacks the ability to learn the distribution of different data features and shows poor robustness under the current adversarial attacks with strong attack capabilities. In order to improve its ability to resist various watermark removal attacks, the model watermark is trained for robustness, and its generalization ability and model performance in different data and environments are enhanced, thereby obtaining a model watermark with enhanced robustness. Its implementation includes the following:

[0144] 7.1) Input the original data x into the model watermark for forward propagation, and calculate the training loss L of the original data x based on the difference between the probability distribution f(x) predicted by the model and the label y of the original data bn :

[0145] L bn =L ce (f(x),y)

[0146] Among them, L ce represents the cross entropy loss function;

[0147] 7.2) During the training process of 7.1), if the original data x is mistakenly identified as watermarked data x by the model watermark wm In the case of , we can strengthen the penalty for such original data x and reduce the model watermark to misclassify it as watermark data x wm The probability of , thereby enhancing the robustness of the model watermark, so the following penalty rules are formulated:

[0148] Set a penalty mask m, which stipulates that when the label y of the original data x is not equal to the watermark data x wm Label y wm , and the predicted label of the original data x in the model Equal to the watermark data x wm Label y wm When , the original data x and its corresponding label y that meet the requirement are marked by the penalty mask m, and the original data and label marked by the penalty mask m are defined as the penalty data x p and penalized data labels y p :

[0149] x p =x[m],y p =y[m]

[0150] in,

[0151] 7.3) Compare the original data x with the adversarial sample x AdvSplice in the same dimension at a ratio of 1:1 to obtain spliced ​​adversarial heterogeneous data

[0152]

[0153] 7.4) Copy the label y of the original data to get y', and concatenate it with the label y of the original data in the same dimension at a ratio of 1:1 to obtain the concatenated adversarial heterogeneous data label

[0154]

[0155] 7.5) The spliced ​​adversarial heterogeneous data Input into the model watermark that completes watermark embedding for forward propagation, according to the probability distribution predicted by the model With heterogeneous data Tags Difference calculation against heterogeneous data The training loss L robusst :

[0156]

[0157] Among them, L ce represents the cross entropy loss function;

[0158] 7.6) According to the probability distribution f(x) predicted by the model based on the penalized data p ) and the penalty data x p Label y p The difference between the two is used to calculate the penalty training loss L penalty :

[0159] L penalty =L ce (f(x p ),y p )

[0160] Among them, L ce represents the cross entropy loss function;

[0161] 7.7) The training loss L of the above original data x is bn , Fighting against heterogeneous data The training loss L robust and the penalty training loss L penalty Add the total loss L as the model watermark retotal ;

[0162] In order to avoid the model watermark falling into the local optimal solution due to the more complex loss function in the early stage of training, and failing to balance its complexity and performance, a linear growth method is used to dynamically adjust Lrobust and L penalty That is, as the number of training rounds increases, the weights of the two losses are gradually increased, so that the model focuses on learning the feature distribution of the original data x in the early stage of training, and focuses more on combating heterogeneous data in the later stage of training. Training and penalty training to improve the overall performance and generalization ability of the model:

[0163] L retotal =L bn +w t L robust +u t L penalty

[0164] Among them, w init For L robust The initial weight, w final For L robust The final weight of

[0165]

[0166] u init For L penalty The initial weight, u final For L penalty The final weight of

[0167]

[0168] t is the current training round, and T is the total training round.

[0169] Example 2: Model watermark robustness system based on adversarial immunity

[0170] Reference Figure 2 This example includes a data set preprocessing module 1, a target model training module 2, a watermark generation model training module 3, a watermark data generation module 4, a watermark data embedding module 5, and a model watermark robust training module 6. Its working principle is as follows:

[0171] The data set preprocessing module 1 is used to preprocess the original data suitable for the image classification task by horizontal flipping, pixel filling, random cropping and normalization operations in sequence, and divide the preprocessed original data into a training set and a test set, and then output the preprocessed original data to the target model training module 2, the watermark model training module 3, the watermark data generation module 4, the watermark data embedding module 5 and the model watermark robust training module 6 respectively;

[0172] The target model training module 2 selects an existing convolutional neural network model suitable for image classification tasks as the initial target model, uses the training set in the original data to perform forward propagation on the initial target model, updates the model parameters, and obtains the trained target model through test set evaluation to output to the watermark data embedding module 5;

[0173] The generative watermark model training module 3 uses a variational autoencoder including an encoder and a decoder as an initial generative watermark model, uses the original data input by the data set preprocessing module 1 to forward propagate the initial generative watermark model, iterates and updates the model parameters multiple times until the loss function converges, and obtains the trained generative watermark model to output to the watermark data generation module 4;

[0174] The watermark data generation module 4 uses the encoder in the trained watermark generation model to perform feature conversion and information hiding on the original data input by the data set preprocessing module 1, and then optimizes and extracts the intermediate data generated in the encoding process through its decoder to obtain the generated watermark data and output it to the watermark data embedding module 5;

[0175] The watermark data embedding module 5 splices the watermark data with the original data input by the data set preprocessing module 1, and then performs forward propagation on the target model input by the target model training module 2, iteratively updates the model parameters for multiple times until the loss function converges, and obtains the model watermark output to the model watermark robust training module 6;

[0176] The model watermark robust training module 6 adds tiny perturbations that are imperceptible to the naked eye to the original data input by the data set preprocessing module 1 to obtain adversarial samples, then forward-propagates the model watermark after splicing the original data with the adversarial samples, and iteratively updates the model parameters multiple times until the loss function converges, thereby realizing robust training of the model watermark and obtaining a model watermark with enhanced robustness.

[0177] Example 3: Method for using model watermark

[0178] In the actual scenario of open source model sharing, in order to prevent the model from being infringed by unauthorized copying, dissemination and use, model watermarking can be used to strengthen copyright protection. Usually, watermark information containing key copyright features, such as the copyright owner's specific logo, copyright statement, etc., is pre-embedded into the model. Once suspected infringement is detected, the user can extract the watermark information in a stable and secure environment based on the watermark embedding method adopted by the developer and select a watermark extraction algorithm that matches it. The user can check whether the key features of the extracted watermark information are complete and consistent with the original watermark. If the verification result shows an abnormality, it can provide the copyright owner with legally effective evidence for rights protection, thereby effectively safeguarding intellectual property rights.

[0179] Reference Figure 3 The implementation steps of this example include the following:

[0180] The first step is to select the watermark extraction method: the user selects a watermark extraction algorithm that matches the specific algorithm used by the developer when embedding the watermark in the model;

[0181] In this step, assuming that the developer uses a discrete cosine transform algorithm to embed the copyright watermark into the parameters of the model, the user needs to select a watermark extraction algorithm based on the discrete cosine transform algorithm to effectively extract the watermark information hidden in the model.

[0182] The second step is to extract the watermark information: the user loads the model watermark into the set extraction environment and runs the extraction algorithm to obtain the watermark information. During this process, the stability and security of the environment must be ensured. You can choose to load the model watermark in a specially built secure server environment that has undergone strict network security protection and system stability testing to prevent external interference from affecting the accuracy of the extraction results.

[0183] The third step is to verify the integrity of the watermark information: verify whether the key features such as the specific identification of the copyright owner and the copyright statement pre-set in the extracted watermark information can be correctly identified:

[0184] If all the key features in the watermark can be correctly identified, it means that the integrity of the watermark is good, and the fifth step is executed;

[0185] Otherwise, it indicates that the watermark is missing or damaged. Try to repair the watermark information and proceed to step 4.

[0186] The fourth step is to repair the watermark information: Use the algorithm used when the model embeds the watermark to repair the watermark information, and then check whether the key features of the repaired watermark can be correctly identified:

[0187] If it is correctly identified, it means that the watermark is repaired successfully, and the fifth step is executed;

[0188] Otherwise, it indicates that the copyright of the model is suspected of being plagiarized, and it is necessary to investigate the model's dissemination channels on major open source platforms, resource websites, etc. to check whether there is any unauthorized dissemination and modification.

[0189] Step 5: Compare watermark information: Compare the key features of the extracted watermark information with the key features of the original watermark information to determine whether the copyright of the model has been stolen:

[0190] If the two are exactly the same, then the copyright of the model is correct;

[0191] Otherwise, it is determined that the copyright of the model has been misappropriated and appropriate copyright protection measures should be initiated to investigate the liability for infringement.

[0192] Application examples:

[0193] Based on the above steps, taking the copyright protection of the game model developed by a game company as an example, the use of the model watermark is explained as follows:

[0194] S1) Select watermark extraction method:

[0195] When developing game models, game companies use a watermark embedding algorithm based on neural network parameter perturbation and a verification method based on parity coding to embed the company's copyright logo, copyright statement and other information into the parameters of the game model in a specific way. When a game company suspects that a similar game on the market has copyright theft, it needs to select a watermark extraction algorithm that matches the algorithm used to embed the watermark.

[0196] S2) Extract watermark information:

[0197] The technical staff of the game company loads the model watermark of the suspected infringing game into a specially built safe and stable extraction environment. This environment has undergone strict network security protection and system stability testing to prevent external interference from affecting the accuracy of the extraction results. Then, the selected watermark extraction algorithm is run to obtain the watermark information from the game model; S3) Verify the integrity of the watermark information:

[0198] The technicians check the extracted watermark information to see whether the specific identification of the copyright owner of the game company, such as the digital code of the company logo, the copyright statement, such as the copyright year, the name of the copyright company, and other key features can be correctly identified;

[0199] S4) Repair watermark information:

[0200] Since the game company originally used a watermark embedding algorithm based on neural network parameter perturbation, the technicians recalculated the position of the missing or damaged watermark information in the neural network parameters according to the principle of the algorithm, and re-embedded the correct watermark information into the corresponding parameter position;

[0201] S5) Check the repaired watermark information:

[0202] The technicians used the parity check code method used by the developer when embedding the watermark to verify the repaired watermark information. All the key features of the watermark information can be correctly identified, indicating that the watermark repair was successful;

[0203] S6) Compare watermark information:

[0204] The technicians compared the key features of the extracted watermark information with the key features of the original watermark information one by one:

[0205] If it is found that the company LOGO digital code in the extracted watermark information is consistent with the company LOGO digital code in the original watermark information, but the company name in its copyright statement has been tampered with, it is determined that the copyright of the game model has been stolen. At this time, the corresponding copyright protection measures are immediately initiated, and an infringement warning letter is sent to the infringing party, requiring it to stop the infringement and bear the corresponding legal responsibility. If the infringing party refuses to cooperate, the game company will protect its copyright rights through legal channels.

[0206] The above descriptions are only a few specific examples of the present invention and do not constitute any limitation to the present invention. It is obvious that for professionals in this field, after understanding the content and principles of the present invention, they may make various modifications and changes in form and details without departing from the principles and structures of the present invention. However, these modifications and changes based on the ideas of the present invention are still within the scope of the claims and protection of the present invention.

[0207] It should be noted that the step numbers in the specification and claims of the present invention are only for a clear description of the implementation scheme of the present invention to facilitate understanding, and the order of the step numbers is not limited.

Claims

1. A model watermark robustness enhancement method based on adversarial immunity, characterized in that: include: (1) Select an existing convolutional neural network model suitable for image classification tasks as the initial target model, use a public dataset suitable for image classification tasks as raw data to train the target model, and obtain a trained target model; (2) Using the existing variational autoencoder VAE including an encoder and a decoder as the initial generative watermark model, and using the original data to train the initial generative watermark model to obtain a trained generative watermark model; (3) Input the original data into the trained watermark generation model to generate watermark data x wm ; (4) The watermark data is concatenated with the original data and then input into the trained target model for forward propagation to complete the embedding of the model watermark; (5) Add tiny perturbations that are imperceptible to the naked eye to the original data to generate adversarial samples; (6) The original data is input into the model watermark after watermark embedding for forward propagation, and then the adversarial sample is concatenated with the original data and input into the model watermark for robustness training to obtain a highly robust model watermark.

2. The method according to claim 1, characterized in that: In step (1), the original data is used to train the initial target model, and its implementation includes the following: 1a) The original data is preprocessed by horizontal flipping, pixel filling, random cropping and normalization operations in sequence, and then the preprocessed original data is divided into a training set and a test set in a ratio of 5:1; 1b) Input the training set data into the initial target model in batches for forward propagation, use the cross entropy loss to measure the difference between the probability distribution predicted by the model and the label distribution of the original data, use the stochastic gradient descent optimizer to calculate the gradient of the cross entropy loss with respect to the model parameters, and then update the model parameters according to the opposite direction of the gradient; 1c) Use the data from the test set to evaluate the model and calculate the model accuracy and loss on the test set; 1d) Repeat 1b) and 1c) to train the model for 100 iterations, select target models with model accuracy greater than 85%, sort these target models in descending order according to their accuracy, and save the highest target model locally.

3. The method according to claim 1, characterized in that: Step (2) uses the original data to train the initial generated watermark model, and its implementation includes the following: 2a) Input the original data x into the encoder of the initial watermark generation model, and output the mean μ and logarithmic variance log(σ) of the latent space 2 ): Among them, f μ and are two different output branches of the encoder neural network; 2b) Sample a random variable ∈ from the distribution of the latent space, using the random variable ∈, mean μ and logarithmic variance log(σ 2 ), calculate the latent variable z: Among them, ⊙ represents element-by-element multiplication; 2c) Input the sampled latent variable z into the decoder of the initial watermark generation model, and output the reconstructed data with the same dimension as the original data x Among them, Decoder represents a decoder; 2d) Set the loss function of the initial watermark generation model Among them, L MSE is the original data x and the reconstructed data The reconstruction loss, L Per is the original data x and the reconstructed data The perceptual loss, L KL is the KL divergence loss, β(t) is the annealing factor of the current training step t; λ1 is L MSE The weight hyperparameter of L Per The weight hyperparameters of 2e) Use the optimizer to calculate the above loss Regarding the gradient of the generated watermark model parameters, the model parameters are updated according to the opposite direction of the gradient to obtain the generated watermark model of the current training round; 2f) Repeat steps 2a) to 2e) to update the model multiple times until the loss Convergence, and obtain the trained generated watermark model.

4. The method according to claim 3, characterized in that: Step 2d) Reconstruction loss L MSE , perceptual loss, L Per , KL divergence loss L KL , respectively expressed as follows: Among them, x i is the i-th pixel of the original data x, To reconstruct data The i-th pixel of , q(z|x) is the posterior distribution of the latent variable z, p(z) is the prior distribution of the latent variable z, φ i (x) is the image feature representation of the original data x in the i-th layer of the pre-trained network ResNet50, To reconstruct data The image feature representation of the i-th layer in the pre-trained network ResNet50, the pre-trained network ResNet50 is used to calculate the perceptual loss L Per Existing convolutional neural networks used when 5. The method according to claim 1, characterized in that: The generation of watermark data in step (3) is implemented as follows: 3a) Input the original data x into the encoder of the trained generative watermark model, and output the mean μ' and logarithmic variance log(σ') of the latent space 2 ): μ'=f μ (x) Among them, f μ and They are two different output branches of the encoder neural network; 3b) Sample a random variable ε' from the distribution in the latent space, using the random variable ∈', mean μ' and logarithmic variance log(σ') 2 ) to calculate the latent variable z': Among them, ⊙ represents element-by-element multiplication; 3c) Input the latent variable z' into the decoder of the trained watermark model and output the watermark data x wm : x wm =Decoder(z') Among them, Decoder represents a decoder.

6. The method according to claim 1, characterized in that: In step (4), the watermark data is concatenated with the original data and then input into the trained target model for forward propagation to complete the embedding of the model watermark. The implementation includes the following: 4a) Combine the original data x and the watermark data x wm Splice along the same dimension at a ratio of 1:1 to obtain the spliced ​​watermark enhancement data 4b) Let the label of the original data x be y, and the watermark data x wm The label is y wm , and then the label y of the original data and the label y of the watermark data wm Splice along the same dimension at a ratio of 1:1 to obtain the spliced ​​watermark enhanced data label 4c) Watermarking Enhanced Data Input into the trained target model for forward propagation, and use the cross entropy loss function to calculate the probability distribution predicted by the model Enhanced data labeling with watermarks The difference between them is taken as the watermark loss L wm : Among them, L ce represents the cross entropy loss function; 4d) Use the optimizer to calculate the watermark loss L wm Regarding the gradient of the model parameters, the trained target model parameters are updated according to the opposite direction of the gradient to complete the watermark embedding training of the current round of target model; 4e) Repeat steps 4c) to 4d) to iterate the model until the watermark loss L wm Convergence, realize the embedding of model watermark.

7. The method according to claim 1, characterized in that: Step (5) adds tiny perturbations that are imperceptible to the naked eye to the original data to generate adversarial samples. The implementation includes the following: 5a) Initialize a perturbation δ with the same shape as the original data x t =0, t=0,1,2,...,T, T is the maximum step size of iterative update; 5b) The disturbance δ t Superimposed on the original data x, we get the current adversarial sample x+δ t , the adversarial sample x+δ t Input into the model watermark to get the prediction result f(x+δ t ), according to the prediction result f(x+δ t ) and the difference between the original data label y to calculate the adversarial sample x+δ t The loss function L δ : L δ =L ce (f(x+δ t ),y) Among them, L ce represents the cross entropy loss function; 5c) Calculate the loss function L δ About the perturbation δ t Gradient Where c is the number of categories, f i (x+δ t ) is the model output f(x+δ t )’s i-th component; 5d) Along the gradient Update the perturbation δ of step t+1 in the direction of t+1 : Among them, δ t is the perturbation at step t, α is the perturbation step size at each step, and sign(·) indicates the direction in which the sign function obtains the perturbation; 5e) Update the perturbation δ t+1 Projected onto L with zero as the center and ∈ as the radius ∞ In the norm ball, we get the current constrained perturbation δ' t+1 : d′ t+1 =clip(δ t+1 ,-∈,∈) Among them, the clip function converts δ t+1 Each element of is restricted to the range [-∈,∈]; 5f) Repeat steps 5b) to 5e) until the maximum number of iterations T is reached, and the final constrained perturbation δ' is obtained. T , and then the disturbance δ' T Add to the original sample x to generate the final adversarial sample x Adv : x Adv =x+δ' T 。 8. The method according to claim 1, characterized in that: Step (6) inputs the original data into the model watermark after watermark embedding for forward propagation, and then concatenates the adversarial sample with the original data and inputs it into the model watermark for robustness training. The implementation includes the following: 6a) Input the original data x into the model watermark after watermark embedding for forward propagation, and calculate the training loss L of the original data x based on the difference between the probability distribution f(x) predicted by the model and the label y of the original data bn : L bn =L ce (f(x),y) Among them, L ce represents the cross entropy loss function, f(x) represents the predicted output of the original data in the model; 6b) Setting penalty rules: Set a penalty mask m, which stipulates that when the label y of the original data x is not equal to the watermark data x wm Label y wm , and the predicted label of the original data x in the model Equal to the watermark data x wm Label y wm When , the original data x and its corresponding label y that meet the requirement are marked by the penalty mask m, and the original data and label marked by the penalty mask m are defined as the penalty data x p and penalized data labels y p : x p =x[m],y p =y[m] in, 6c) Compare the original data x with the adversarial sample x Adv Splice in the same dimension at a ratio of 1:1 to obtain spliced ​​adversarial heterogeneous data 6d) Copy the label y of the original data to get y', and concatenate it with the label y of the original data in the same dimension at a ratio of 1:1 to obtain the concatenated adversarial heterogeneous data label 6e) The spliced ​​adversarial heterogeneous data Input into the model watermark that completes watermark embedding for robustness training 6e1) Let the total loss function of the model watermark be L retotal for: L retotal =L bn +w t L robust +and t L penality in, To combat heterogeneous data The training loss is , which represents the adversarial heterogeneous data of the model watermark prediction The probability distribution of Fighting against heterogeneous data Tags The difference between t Indicates L robust The weight of L penality =L ce (f(x p ),y p ) is the loss of penalty training, which represents the penalty data x predicted by the model watermark p The probability distribution f(x p ) and the penalty data x p Label y p The difference between t Indicates L penalty The weight of 6f) Use the optimizer to calculate the above loss L retotal Regarding the gradient of the model watermark parameters, the trained model watermark parameters are updated according to the opposite direction of the gradient to enhance the robustness of the model watermark of the current round; 6g) Repeat steps 6a) to 6f) to update the model multiple times until the loss L retotal Convergence, completing the robustness training of the model watermark.

9. A model watermarking system with enhanced robustness, characterized in that: include: The dataset preprocessing module is used to preprocess the public dataset suitable for image classification tasks by horizontal flipping, pixel padding, random cropping, and normalization operations in sequence, and use the preprocessed dataset as the original data; The target model training module is used to input the original data into the initial target model for forward propagation to obtain the trained target model; The generated watermark model training module is used to input the original data into the initial generated watermark model for forward propagation to obtain a trained generated watermark model; The watermark data generation module is used to input the original data into the trained watermark generation model, perform feature conversion and information hiding on the original data through the encoder of the watermark generation model, and then optimize and extract the intermediate data generated in the encoding process through its decoder to obtain the generated watermark data; The watermark data embedding module is used to splice the generated watermark data with the original data and input them into the trained target model for forward propagation to obtain the model watermark; The model watermark robust training module is used to add tiny perturbations that are imperceptible to the naked eye to the original data to obtain adversarial samples, and then concatenate the original data with the adversarial samples and input them into the model watermark for forward propagation, thereby realizing robust training of the model watermark and obtaining a model watermark with enhanced robustness.

10. A method for using a model watermark to enhance robustness, characterized in that: include: S1) The user selects a watermark extraction algorithm that matches the specific algorithm used by the developer when embedding the watermark in the model; S2) The user loads the model watermark into the set extraction environment and runs the extraction algorithm to obtain the watermark information. During this process, the stability and security of the environment must be ensured to prevent external interference from affecting the accuracy of the extraction results; S3) Check whether the key features such as the specific identification of the copyright owner and the copyright statement pre-set in the extracted watermark information can be correctly identified: If all key features in the watermark can be correctly identified, it indicates that the watermark is of good integrity, and step S5 is executed; Otherwise, it indicates that the watermark is missing or damaged, and an attempt is made to repair the watermark information, and step S4 is executed; S4) Repair the watermark information using the algorithm used when the model is embedded in the watermark, and check whether the key features in the repaired watermark can be correctly identified: If it is correctly identified, it indicates that the watermark repair is successful, and S5 is executed; Otherwise, it indicates that the copyright of the model is suspected of being misappropriated, and further investigation is needed in combination with the dissemination channels of the model; S5) Compare the key features of the extracted watermark information with the key features of the original watermark information to determine whether the copyright of the model has been stolen: If the two are exactly the same, then the copyright of the model is correct; Otherwise, it is determined that the copyright of the model has been misappropriated and appropriate copyright protection measures should be initiated to investigate the liability for infringement.

Citation Information

Patent Citations

  • Reversible robust watermark embedding and extracting model construction method capable of resisting image attack

    CN116452401A

  • Screen shooting image robust watermarking method based on attention mechanism and contrast learning

    CN118967424A

  • Anti-attack robust generative model watermark processing method and system

    CN119151763A

  • Cloud storage image data ownership verifying method based on multifunction digital watermark

    CN103700059A

  • Method for defending alternative model attack and verifying copyright of DNN model

    CN115828188A

Cited By

  • Safety monitoring method and system based on embedded large model in intelligent screen

    CN120296749A

  • Security monitoring method and system based on large model embedded in smart screen

    CN120296749B

  • Customized generative model copyright protection method based on robust watermark optimization

    CN121544447A

  • Image defogging method and system based on adaptive convolution and texture prior

    CN121998873A