Method for generating backdoor watermark image dataset based on adversarial training network

By adversarial training generator and discriminator networks, it generates indistinguishable fake samples and modifyes the labels, which solves the problem that backdoor watermarks are easily detected and affects model performance in the prior art, and realizes a backdoor watermark image data set with high concealment and high precision.

CN115546003BActive Publication Date: 2025-07-04XIDIAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211242857.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-11
Publication Date
2025-07-04
Estimated Expiration
2042-10-11

AI Technical Summary

Technical Problem

The backdoor watermark generated by the prior art is poor in its secrets and affects the original performance of the model, and cannot meet the needs of concealment and fidelity at the same time.

Method used

The generator and discriminator network are built for adversarial training. The generator generates indistinguishable fake samples. The discriminator distinguishes between true and false samples and modifys the fake sample label to form a backdoor watermark image dataset.

Benefits of technology

The generated backdoor watermark image dataset is highly concealed and is not easily detected by attackers. It does not affect the accuracy of the original task of the model, ensuring the high accuracy of the model on the original task.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115546003B_ABST
    Figure CN115546003B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating a backdoor watermark image dataset based on an adversarial training network. The method is to respectively construct a generator network and a discriminator network, and perform adversarial training on the two networks. The obtained picture samples are determined by the discriminator to be real picture samples with a probability of 50% and fake samples generated by the generator with a probability of 50%. This makes the backdoor watermark image dataset of the present invention have a statistical distribution similar to that of the real picture sample set, is not easily detected by attackers, and has the advantage of strong concealment. At the same time, the present invention modifies the labels of all fake samples generated by the generator network of the backdoor watermark image dataset, does not introduce invalid or incorrect features, does not affect the accuracy of the image classification model in the original task, the decision boundary of the image classification model in the original task remains unchanged, and the image classification model still maintains high accuracy in the original task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing technology, and further relates to a method for generating a backdoor watermark image data set based on an adversarial training network in the field of neural network watermarking. The present invention can be used for copyright protection of image classification models in black box scenarios, and a backdoor watermark image data set is generated in an invisible way. When a model copyright dispute occurs, the model user can declare ownership by verifying the watermark information. Background Art

[0002] As a way to protect model copyright, watermarks are widely used in black box scenarios. At present, the backdoor-based design is to construct a specific backdoor watermarked image dataset, which usually consists of a set of image samples and corresponding specific labels. The mapping between the specific input and its label is regarded as a backdoor and used as a watermark. The trained image classification model is fine-tuned with the backdoor watermarked image dataset so that the model contains watermark information. The image classification model can predict the image samples in the backdoor watermarked image dataset as specific labels; the model user uses the backdoor watermarked image dataset to initiate a prediction query on the suspicious model. If the watermark information is detected, the model user can claim the ownership of the model.

[0003] However, in real scenarios, attackers can detect backdoor samples through a series of means, such as query modification attacks, to evade detection. In addition, the backdoor watermark in the current backdoor watermark technology will inevitably affect the original task of the model, resulting in low classification accuracy of the image classification model containing the backdoor watermark. Therefore, the watermark should be sufficiently hidden and not easily detected by attackers. At the same time, the backdoor watermark should not affect the accuracy of the original model. However, it is difficult for the current neural network model backdoor watermark technology to simultaneously take into account both fidelity and concealment.

[0004] South China Normal University disclosed a neural network watermark embedding method in its patent document "A neural network watermark embedding method, device, electronic device and storage medium" (application number: 202210016799.8 application publication number: CN 114359011 A). This method uses a key acquisition module to obtain a key corresponding to a unique timestamp; then randomly selects part of the image data set in the original training set and scrambles and encrypts the image through the chaotic sequence generated by the key to obtain a trigger set. This invention has a good verification effect on the basis of ensuring that the trigger set is highly invisible to attackers. However, the method still has the disadvantage that, since the method is encrypted on the original data set image, the trigger set obtained by scrambling encryption changes the characteristics of the original image, introduces invalid or erroneous features, distorts the decision boundary of the image classification model on the original task, and reduces the performance of the image classification model on the original task.

[0005] Ryota Namba et al. proposed a backdoor watermarking method with exponential weighting in their published paper "Robust Watermarking of Neural Network with Exponential Weighting" (Proc of the 2019 ACM Asia Conf on Computer and Communications Security). This method randomly selects a certain proportion of training samples from the original training dataset and only changes their labels to obtain a backdoor watermark image dataset. This method improves the invisibility of the backdoor watermark. However, the shortcoming of this method is that it changes the labels of the original pictures, and classifying samples with incorrect labels allows the image classification model to learn bad features, thus changing the decision boundary of the image classification model in the original classification task and resulting in a decline in the original performance, unable to meet the fidelity requirements. Summary of the Invention

[0006] The object of the present invention is to provide a method for generating a backdoor watermark image dataset based on an adversarial training network in view of the deficiencies of the existing technologies, aiming to solve the problems that the backdoor watermark generated by the existing technologies has poor invisibility, the original performance of the model decreases due to the introduction of invalid features, and it cannot meet the fidelity requirements.

[0007] The specific idea for achieving the object of the present invention is as follows: First, a generator network and a discriminator network are respectively constructed. The purpose of the generator network is to make the distribution of the generated fake samples fit the distribution of the real image samples as much as possible, and the purpose of the discriminator network is to distinguish whether the input sample is a real image sample or a fake sample as much as possible; then the two networks are subjected to adversarial training. During the training process, the generator network generates fake samples that look real and are similar to the real image samples to deceive the discriminator network, and the discriminator network distinguishes the fake samples from the real image samples. In this way, the generator network tries to deceive the discriminator network, and the discriminator network tries not to be deceived by the generator network. After the two networks are alternately trained and improve each other, they form a dynamic "game". Finally, the trained generator network can generate picture samples that are "indistinguishable from the real ones". The finally obtained picture samples are judged by the discriminator network to be real picture samples with a probability of 50% and fake samples generated by the generator network with a probability of 50%.

[0008] By modifying the labels of all fake samples generated by the generator network, a backdoor watermark image dataset is formed by combining all the fake samples and their modified labels. Adding new labels not only does not distort the original decision boundary, but also helps the model better learn the features of the picture sample set, overcoming the problem of introducing incorrect mapping relationships and distorting the original decision boundary in existing neural network backdoor watermarking methods, so that the generated backdoor watermark image dataset does not affect the accuracy of the model on the original task.

[0009] The specific steps implemented by the present invention are as follows:

[0010] Step 1, construct a generator network:

[0011] Construct a generator network cascaded by 5 fully connected layers. Set the number of input neurons of the first to fifth fully connected layers to 100, 128, 256, 512, 1024 in sequence, and the number of output neurons to 128, 256, 512, 1024, 784 in sequence; the activation functions of the first to fourth fully connected layers are all implemented using the Relu function, and the activation function of the fifth fully connected layer is implemented using the tanh function;

[0012] Step 2, construct a discriminator network:

[0013] Construct a discriminator network cascaded by 3 fully connected layers. Set the number of input neurons of the first to third fully connected layers to 784, 512, 256 in sequence, and the number of output neurons to 512, 256, 1 in sequence; the activation functions of the first and second fully connected layers are all implemented using the Relu function, and the activation function of the third fully connected layer is implemented using the Sigmoid function;

[0014] Step 3, generate a picture sample set and a noise sample set:

[0015] Step 3.1, form a picture sample set by using half of the N images and their labels in the N images containing C target categories, where C≥2 and N≥2000;

[0016] Step 3.2, randomly generate a noise sample set containing m noises that conform to the Gaussian distribution. The dimension of each noise sample is 100, where the value of m is the same as that of N;

[0017] Step 4, perform adversarial training on the generator network and the discriminator network:

[0018] Step 4.1, input the noise sample set into the generator network, perform non-linear mapping on each noise sample through the generator network, and form a fake sample set by all the mapped noise samples; input the fake sample set into the discriminator network and output the predicted value of each fake sample; input the picture sample set into the discriminator network and output the predicted value of each picture sample;

[0019] Step 4.2, calculate the average loss value of the noise samples output after all noise samples are input into the generator network, calculate the average loss value of the samples output after all picture samples and all fake samples are input into the discriminator network, calculate the gradients of the loss functions of the generator network and the discriminator network respectively, and adopt the gradient descent algorithm to alternately update the parameters of the generator network and the discriminator network until both the average loss value of the noise samples and the average loss value of the samples no longer change, so as to obtain the trained generator network and discriminator network;

[0020] Step 5, generate a backdoor watermark image dataset:

[0021] Modify the label of each fake sample output by the generator network when both the generator network and the discriminator network are trained, and form a backdoor watermark image dataset with all the fake samples and their modified labels.

[0022] Compared with the prior art, the present invention has the following advantages:

[0023] First, by separately constructing a generator network and a discriminator network and performing adversarial training on the two networks, the picture samples obtained by the present invention are determined by the discriminator to be real picture samples with a probability of 50% and fake samples generated by the generator with a probability of 50%; it overcomes the problem in the prior art that the difference between the backdoor watermark image dataset and the real picture sample set is too large and is easily detected by attackers to avoid verification, making the backdoor watermark image dataset of the present invention have a similar statistical distribution to the real picture sample set and not easily detected by attackers, having the advantage of strong concealment. Fine-tuning the trained image classification model with the backdoor watermark image dataset enables the model to contain watermark information. By querying the watermark information in the watermarked model, the model user can claim the ownership of the model.

[0024] Second, the present invention modifies the labels of all fake samples generated by the generator network and changes them to new labels that are all different from the label categories of the original picture samples, overcoming the problem in the prior art that modifying the sample labels in the backdoor watermark image dataset into other labels in the label categories of the original picture samples introduces invalid or incorrect features and distorts the decision boundary of the image classification model in the original task, making the backdoor watermark image dataset of the present invention not affect the accuracy of the image classification model in the original task, and the image classification model still maintains high accuracy in the original task. Description of the Drawings

[0025] Figure 1 is a flowchart of the present invention. Detailed Embodiment

[0026] The following combines the attached Figure 1 and embodiments to further describe the implementation steps of the present invention.

[0027] Step 1, construct the generator network:

[0028] Build a generator network cascaded by 5 fully connected layers, and set the parameters of each layer of the network as follows. The number of input neurons of the first to fifth fully connected layers is set to 100, 128, 256, 512, 1024 in sequence, and the number of output neurons is set to 128, 256, 512, 1024, 784 in sequence. The activation functions of the first to fourth fully connected layers all adopt the Relu function, and the activation function of the fifth fully connected layer adopts the tanh function.

[0029] Step 2, construct the discriminator network:

[0030] Build a discriminator network cascaded by 3 fully connected layers, and set the parameters of each layer of the network as follows. The number of input neurons of the first to third fully connected layers is set to 784, 512, 256 in sequence, and the number of output neurons is set to 512, 256, 1 in sequence. The activation functions of the first and second fully connected layers all adopt the Relu function, and the activation function of the third fully connected layer adopts the Sigmoid function.

[0031] Step 3, generate the picture sample set and the noise sample set:

[0032] Step 3.1, form a picture sample set by using half of the N images containing C target categories and their labels, where C≥2 and N≥200.

[0033] In the embodiment of the present invention, 30,000 images and their labels are selected from 10 categories of the MNIST data set to form a picture sample set. The labels of the MNIST data set are numbers from 0 to 9. The MNIST data set includes 60,000 training image samples and 10,000 test image samples, and each image sample is a grayscale image with a size of 28×28.

[0034] Step 3.2, randomly generate a noise sample set containing m noises that conform to the Gaussian distribution, and the dimension of each noise sample is 100, where the value of m is the same as that of N. In the embodiment of the present invention, m = 30,000.

[0035] Step 4, perform adversarial training on the generator network and the discriminator network:

[0036] Step 4.1: Input the noise sample set into the generator network. Through upsampling with five fully connected layers, map each noise sample with a dimension of 100 into a noise sample with a dimension of 784. Combine all the mapped noise samples to form a fake sample set. Input the fake sample set into the discriminator network. Through downsampling with three fully connected layers, output the prediction value of each fake sample. Input the picture sample set into the discriminator network. Through downsampling with three fully connected layers, output the prediction value of each picture sample.

[0037] Step 4.2: Calculate the average loss value of the noise samples output after all noise samples are input into the generator network. Calculate the average loss value of the samples output after all picture samples and all fake samples are input into the discriminator network. Calculate the gradients of the loss functions of the generator network and the discriminator network respectively. Adopt the gradient descent algorithm to alternately update the parameters of the generator network and the discriminator network until both the average loss value of the noise samples and the average loss value of the samples no longer change, obtaining the trained generator network and discriminator network.

[0038] In the embodiment of the present invention, after 100 times of training, both the average loss value of the noise samples and the average loss value of the samples no longer change. The fake samples output by the generator network have a 50% probability of being judged as real samples by the discriminator network and a 50% probability of being judged as fake samples.

[0039] Step 4.3: Use the following formula to calculate the average loss value of the noise samples output after all noise samples are input into the generator network:

[0040]

[0041] where G loss represents the average loss value of the noise samples output after all noise samples are input into the generator network. i represents the serial number of the samples in the noise sample set, i = 1, 2,..., m, and m represents the total number of samples in the noise sample set. In the embodiment of the present invention, m = 30000. ∑ represents the summation operation, and log represents the logarithm operation with base 2. G(z (i) ) represents the fake sample output after the i-th noise sample z (i) in the noise sample set is input into the generator network. D(G(z (i) )) represents the discrimination probability of the fake sample G(z (i) ) output after being input into the discriminator network.

[0042] Step 4.4: Use the following formula to calculate the average loss value of the samples output after all picture samples and all fake samples are input into the discriminator network:

[0043]

[0044] where Dloss denotes the average loss value of the samples output after all fake samples and all image samples are input into the discriminator network. j represents the sample serial number, j = 1, 2,..., n, where n represents the total number of all fake samples and all image samples. In the embodiment of the present invention, n = 30000, and x j represents the j-th picture sample, represents the j-th fake sample, D(x j ) represents the discrimination probability output after the picture sample x i is input into the discriminator network, represents taking the fake sample as the input to the discriminator network and outputting the discrimination probability.

[0045] Step 5, generate a backdoor watermark image dataset:

[0046] Modify the label of each fake sample output by the generator network when both the generator network and the discriminator network are trained well. Combine all the fake samples and their modified labels to form a backdoor watermark image dataset. In the embodiment of the present invention, each fake sample label is modified to *.

Claims

1. A method for generating a backdoor watermark image dataset based on an adversarial training network, characterized in that, Construct a generator network and a discriminator network respectively, and conduct adversarial training on the generator network and the discriminator network to generate a backdoor watermark image dataset. The steps of this method are as follows: Step 1, construct a generator network: Construct a generator network cascaded by 5 fully connected layers. Set the number of input neurons of the first to fifth fully connected layers to 100, 128, 256, 512, 1024 in sequence, and the number of output neurons to 128, 256, 512, 1024, 784 in sequence; the activation functions of the first to fourth fully connected layers are all implemented using the Relu function, and the activation function of the fifth fully connected layer is implemented using the tanh function; Step 2, construct a discriminator network: Construct a discriminator network cascaded by 3 fully connected layers. Set the number of input neurons of the first to third fully connected layers to 784, 512, 256 in sequence, and the number of output neurons to 512, 256, 1 in sequence; the activation functions of the first and second fully connected layers are all implemented using the Relu function, and the activation function of the third fully connected layer is implemented using the Sigmoid function; Step 3, generate a picture sample set and a noise sample set: Step 3.1, form a picture sample set from half of the N images and their labels in the N images containing C target categories, where C≥2 and N≥2000; Step 3.2, randomly generate a noise sample set containing m noises that conform to the Gaussian distribution. The dimension of each noise sample is 100, where the value of m is the same as that of N; Step 4, conduct adversarial training on the generator network and the discriminator network: Step 4.1, input the noise sample set into the generator network, perform non-linear mapping on each noise sample through the generator network, and form a set of fake samples from all the mapped noise samples; input the set of fake samples into the discriminator network and output the predicted value of each fake sample; input the picture sample set into the discriminator network and output the predicted value of each picture sample; Step 4.2, calculate the average loss value of the noise samples output after all the noise samples are input into the generator network, calculate the average loss value of the samples output after all the picture samples and all the fake samples are input into the discriminator network, calculate the gradients of the loss functions of the generator network and the discriminator network respectively, and use the gradient descent algorithm to alternately update the parameters of the generator network and the discriminator network until both the average loss value of the noise samples and the average loss value of the samples no longer change, and obtain the trained generator network and discriminator network; Step 5, generate a backdoor watermark image dataset: Modify the label of each fake sample output by the generator network when both the generator network and the discriminator network are trained, and form a backdoor watermark image dataset from all the fake samples and their modified labels.

2. The method for generating a backdoor watermark image data set based on an adversarial training network according to claim 1, wherein, The average loss value of the noise samples output after all the noise samples are input into the generator network described in Step 4.2 is obtained by the following formula: Among them, G loss represents the average loss value of the noise samples output after all noise samples are input into the generator network. i represents the serial number of the samples in the noise sample set, i = 1, 2,..., m, m represents the total number of samples in the noise sample set, ∑ represents the summation operation, log represents the logarithm operation with base 2, G(z (i) ) represents the fake sample output after the i-th noise sample z (i) in the noise sample set is input into the generator network. D(G(z (i) )) represents the discrimination probability of the fake sample G(z (i) ) output after the fake sample is input into the discriminator network.

3. The method for generating a backdoor watermark image data set based on an adversarial training network according to claim 2, wherein The average loss value of the samples output after all the picture samples and all the fake samples are input into the discriminator network described in Step 4.2 is obtained by the following formula: Among them, D loss represents the average loss value of the samples output after all fake samples and all image samples are input into the discriminator network. j represents the sample serial number at the corresponding position of all fake samples and all image samples, j = 1, 2,..., n, and n represents the total number of all fake samples and all image samples. x j represents the j-th picture sample, represents the j-th fake sample, D(x j ) represents the discrimination probability output after the picture sample x j is input into the discriminator network, represents the fake sample is input into the discriminator network and the output discrimination probability.

4. The method for generating a backdoor watermark image data set based on an adversarial training network according to claim 1, characterized in that In step 4.2, the parameters of the generator network and the discriminator network are alternately updated using the gradient descent algorithm, and the implementation steps are as follows: Step 1: Use the gradient descent algorithm to update the parameters of the generator network with the loss function value of the generator network; Step 2: Use the gradient descent algorithm to update the parameters of the discriminator network with the loss function value of the discriminator network.

Citation Information

Patent Citations

  • Neural network watermark embedding method and device, electronic equipment and storage medium

    CN114359011A

  • Semi-supervised image classification method based on generative adversarial network

    CN110097103A

  • Anti-printing shot image digital watermarking method based on image noise reduction

    CN111598761A