A data privacy protection method based on adversarial learning
By generating small perturbations through adversarial learning and projective gradient descent, and combining them with a semantic discriminator, the instability problem of face replacement in deep generative adversarial networks is solved. This enables the emergence of controllable semantic features in forged images, protecting facial privacy and improving the stability and transferability of attacks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING UNIV
- Filing Date
- 2022-04-11
- Publication Date
- 2026-05-01
AI Technical Summary
Existing deep generative adversarial network attack methods are unstable and have poor transferability in preventing face replacement and protecting facial privacy, and cannot effectively prevent the synthesis of fake images.
By using an adversarial learning-based approach, projective gradient descent is used to generate adversarial noise with small perturbations. Combined with a semantic discriminator, the attack effect is controlled, so that the generated fake images have identifiable features, such as gender reversal or watermarks, thus disrupting the face-swapping process. The method also shows good transferability across different models.
It effectively disrupts the face-swapping process of fake images without affecting image clarity, and produces controllable semantic features on the fake images to help identify their falsity, protect facial privacy, and improve the stability and transferability of attacks.
Smart Images

Figure CN115936958B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a data privacy protection method based on adversarial learning, specifically a face privacy protection method based on adversarial attack GANs, and relates to the field of computer vision technology, particularly to the field of face image editing and adversarial attack technology. Background Technology
[0002] The emergence of deep generative adversarial networks (GANs) has greatly advanced many fields, including computer vision, natural language processing, and semi-supervised learning. However, it has also made it possible to forge and manipulate images or videos. For example, DeepFake technology, based on autoencoders and GANs, can replace faces in videos with other people, making it indistinguishable to the naked eye. Face editing and face-swapping techniques based on GANs are particularly realistic and convenient. If these face-swapping techniques are used to replace faces in inappropriate backgrounds, it poses a significant threat to the privacy and security of celebrities, politicians, and women. Furthermore, the synthesis of facial images also brings significant security risks to identity verification systems and will lead to increasing social problems. Although some successful DeepFake forgery detection technologies exist, these technologies focus on detecting forged images rather than preventing them from being synthesized. When a forged image is detected, the leakage and misuse of facial information is already a fait accompli. Therefore, protecting facial privacy and preventing the effective synthesis of forged images based on GANs—that is, disrupting face-swapping technology—has become a novel research area.
[0003] One important direction in debunking face-swapping techniques is adversarial attacks within the field of adversarial learning. Traditional adversarial attack methods primarily target deep learning-based classification models, adding small, imperceptible noise perturbations to the original image to cause the classifier to output the wrong category. The process of generating this invisible noise is the adversarial attack process, typically achieved by repeatedly accumulating the adversarial attack loss and backfeeding it to the truncated gradient generation of the original image. The image with the generated adversarial noise added is called an adversarial example, which renders the attacked deep network model ineffective. Furthermore, some attack methods targeting generative adversarial networks have emerged, but the publicly available methods have significant limitations, only adding subtle, meaningless noise to the output image. The attack effect is highly unstable, and the generated adversarial examples show poor transferability when attacking other models. In summary, known attack methods against deep generative adversarial networks are not yet mature enough to prevent face replacement and thus protect facial privacy. Summary of the Invention
[0004] Objective: To address the problems and shortcomings of existing technologies, this invention aims to increase the stability and controllability of attacks on deep generative adversarial networks (GANs), while simultaneously ensuring that images generated by the attacked model exhibit identifiable features. This invention provides a data privacy protection method based on adversarial learning, utilizing adversarial attack GANs to protect facial privacy. It proposes a framework for controllable semantic attacks against GAN models. By concatenating a pre-trained semantic discriminator and using an improved adversarial attack method, adversarial noise with subtle perturbations is generated. Adding this adversarial noise to the face image to be protected causes the face-swapping of forged images generated by the attacked model to fail, resulting in insufficient image realism or localized damage. Furthermore, this invention allows for the manipulation of feature blocks with specific semantic information into the generated forged images, such as gender reversal, aging, or red cross logos, making it easier for reviewers to detect fake images. Moreover, the adversarial samples generated by this invention exhibit a certain degree of transferability when performing adversarial attacks on other similar models.
[0005] By employing adversarial attack techniques, this invention generates minute, invisible protective noise into the facial image to be protected. Adding this noise to the original image does not affect its clarity. The noise-protected facial image disrupts the face-swapping process based on generative adversarial networks (GANs), causing the face-swapping to fail and the forged image to become corrupted and distorted, thus protecting facial privacy. Furthermore, this invention can introduce controllable semantic blocks into the forged image synthesized by the attacked deep GAN. These semantic blocks not only conceal private information but also help identify the forgery. This controllable attack is also more effective and stable than ordinary attacks.
[0006] Technical Solution: This invention presents a data privacy protection method based on adversarial learning. It is an adversarial attack scheme targeting deep generative adversarial networks (GANs)-based face-swapping models, allowing for controlled attack effects. Starting from the semantic representation level of images, this method proposes a way to generate controllable images with semantic features after the attack, building upon traditional adversarial attack methods. Utilizing the semantic decomposability of images, by concatenating the attacked model with a semantic discriminator, semantic labels are modified to simultaneously achieve minor modifications to the original image to be protected and semantic changes to the generated image. The ultimate goal of this scheme is to add micro-perturbations invisible to the human eye to a given face image, causing a significant change in the semantic appearance attributes of the face when the face-swapping model is applied to that face image, leading to face-swapping failure and thus protecting facial privacy. Furthermore, this invention also proposes targeted region attack methods and ensemble attack methods.
[0007] The implementation of the method includes semantic discriminator training and adversarial attacks.
[0008] The face-swapping model training described herein is merely an introduction to a face-swapping algorithm that has been experimentally verified as effective for attack. This algorithm, known as FaceShifter, is also the most convenient and realistic face-swapping model based on deep generative adversarial networks (DGANs). The attack framework described in this invention can attack different DGAN-based face-swapping models in real-world applications. For the desired semantic feature types appearing in the generated image, a corresponding semantic discriminator can be trained in advance. The face-swapping model uses two face images as input, referred to as the source image Xs and the target image Xt. The final output, a synthesized image, contains the identity features of the source image and the facial attribute features and background features of the target image, achieving the effect of transferring and embedding the identity of the source image into the target image while keeping the areas outside the face unchanged.
[0009] The training of the face-swapping model requires a face image dataset centered on the face and normalized in size. To construct this dataset, a face detection model is used to annotate the faces in the original dataset and crop them to the same size. The face detection model can be implemented based on MTCNN (Multi-task Cascaded Convolutional Network). The multi-task cascaded convolutional network model is used to detect the face location and annotate facial feature points in an image containing faces. By cascading P-networks, R-networks, and O-networks, one or more faces are detected and bounded in the image containing faces. Each detected face has a weight representing the salience of the face, and the facial feature points of each face are annotated.
[0010] The adversarial attack method used in this invention is based on projective gradient descent (PGD), a simple and effective white-box attack method considered one of the most effective first-order attack methods for achieving adversarial robustness. This method utilizes only the gradient of the loss function relative to the output result, containing only local first-order information of the model, while attempting to find a perturbation that maximizes the model loss on a specific input, keeping the perturbation size below a specified threshold. The algorithm starts from a random perturbation point near the sample point and iterates multiple times, each iteration descent along the gradient and projecting back into the Lp (p-norm) sphere until convergence. In this invention, projective gradient descent is used for labeled classification models, and its descent direction is such that the output label of the input image changes along the target label direction; that is, its loss is the distance between the image output label and the target label.
[0011] The construction and training steps of the face-swapping model are as follows:
[0012] Step 101: Prepare the encoder model arcface for encoding the identity of the source image;
[0013] Step 102: Prepare a multi-level encoder model (MAE) for encoding facial attribute features of the target image;
[0014] Step 103: Construct a generator model G1 that combines the identity features of the source image with the facial attribute features of the target image to generate multi-level magnified images;
[0015] Step 104: Dataset preprocessing. The face dataset images (approximately 100,000 images) are sequentially labeled with face locations using the multi-task cascaded convolutional network MTCNN model.
[0016] Step 105: If no face was detected in step 104, skip this image;
[0017] Step 106: If a face is detected in step 104, select the face with the highest weight, stretch the bounding box portion of the face to a size of 256×256 and store it in a new dataset;
[0018] Step 107: Use the pre-trained arcface model as the identity encoder for the face-swapping model;
[0019] The ArcFace model described in step 107 is an existing face recognition algorithm, with ArcFace loss added to its core. ArcFace loss is Additive Angular Margin Loss, which normalizes the feature vector and weights and adds an angular margin m to θ. The angular margin has a more direct impact on the angle than the cosine margin. Geometrically, there is a constant linear angle margin. ArcFace directly maximizes the classification boundary in the angle space θ, while CosFace maximizes the classification boundary in the cosine space cos(θ). This algorithm is one of the best in the field of face recognition, with high training efficiency and high performance.
[0020] Step 108: Extract a 216x216 region from the center of the 256x256 source image Xs and downsample it to 112x112. Use this region as the input to the arcface model in step 107 to obtain the identity encoding vector z_id.
[0021] Step 109: Input the target image Xt into the multi-level attribute encoder, which consists of 8 fully connected convolutional layers. Each convolutional layer uses a 4x4 convolutional kernel with dimensions of 32z, 128x128, 64x64x64, 128x32x32, 256x16x16, 512x8x8, 1024x4x4, and 1024x2x2, respectively. Each layer outputs an attribute vector, resulting in a total of 8 attribute vectors z_att.
[0022] Step 110: Input the identity encoding vector z_id and attribute vector z_att obtained in steps 108 and 109 into the multilayer generator G1 to obtain the generated image Y;
[0023] Step 111: Input the generated image Y into the discriminator D to obtain the discrimination result;
[0024] Step 112: Calculate the adversarial loss, identity loss, attribute loss, and reconstruction loss respectively, and sum them up by weight to obtain the generator training loss. Then, backpropagate to optimize the generator network parameters.
[0025] Step 113: Calculate the discriminator loss and optimize the discriminator network parameters through backpropagation;
[0026] Step 114: Repeat steps 108 to 114 until the generator model G1 and the discriminator model D converge;
[0027] Step 115: Training of the face-swapping model is complete.
[0028] In the method described in this invention, the semantic discriminator model is a semantic discriminator model from the image domain editing model based on deep generative adversarial networks. The output of the semantic discriminator model consists of two parts: a judgment of the authenticity of the image and an encoding of the image's domain category. During training, the generator model G2 takes both the real image and the target domain label as input. It regenerates a fake image based on the target domain label from the original image, striving to make the semantic discriminator indistinguishable from the real image within an optimized loss function, while ensuring that the domain classification result matches the given target domain label. The input to the semantic discriminator is either a real image or a fake image generated by the generator model G2. Image features are extracted through six convolutional layers, and the obtained features are then fed into two fully connected layers. The two fully connected layers output the authenticity classification result and the domain classification vector, respectively. The goal of the semantic discriminator is to determine the authenticity of the input image and its domain label. During training, the discriminator is typically updated every N steps, followed by an update of the generator model G2. In the generator model G2, the neighborhood information is first arranged repeatedly according to the image size, and then concatenated with the input image before being input into the convolutional network. The network includes one convolutional layer, two downsampling convolutional layers, six residual layers, two upsampling deconvolutional layers, and one convolutional layer.
[0029] The training steps for the semantic discriminator model are as follows:
[0030] Step 201: Prepare a face dataset containing domain category labels, where the domain category labels include facial appearance semantic attribute features such as age, gender, and hair color.
[0031] Step 202: Train the semantic discriminator. The semantic discriminator distinguishes between real images and fake images generated by the generator model G2. Real images are classified according to their respective domains. The classification loss is calculated and the parameters of the semantic discriminator network are optimized.
[0032] Step 203: Train generator model G2, which uses real images and target domain labels as input to generate fake images;
[0033] Step 204: The generator model G2 uses the domain labels of the fake and real images from step 203 as input to generate a reconstructed image, and calculates the difference between the reconstructed image and the original real image as the reconstruction loss;
[0034] Step 205: Input the fake image generated by the generator model G2 in step 203 into the semantic discriminator to distinguish between real and fake images, calculate the adversarial loss and optimize the parameters of the generator network G2 by combining the reconstruction loss in step 204.
[0035] Step 206: Repeat steps 202 to 205 until the generator model G2 and the semantic discriminator model converge;
[0036] Step 207: The semantic discriminator is trained, and the training ends.
[0037] The controllable semantic adversarial attack method involves concatenating a trained semantic discriminator after the deep generative adversarial network (GAN) model to be attacked. Specifically, the face-swapped image obtained from the face-swapping model is input back into the semantic discriminator model to obtain the semantic features of the face-swapped image. A loss function (e.g., using cross-entropy loss) is calculated between these actual semantic features and the semantic features we want to control (e.g., gender flipping, the appearance of a red cross watermark). This loss is then used to iteratively attack the entire concatenated model, including the face-swapping model and the semantic discriminator model, using projective gradient descent to generate adversarial noise on the original image, thus obtaining the adversarial sample—the image with added protective noise. This method can also achieve implicit editing of GAN models of different depths based on adding noise. The specific steps of the adversarial attack method are as follows:
[0038] Step 301: Let the source image to be protected be Xs, select a face image from the dataset as the target image Xt, and crop the two images into a central 218×218 region;
[0039] Step 302: Xs is processed by the encoder model arcface to obtain the encoded vector vec_s containing the identity features of the source image, and Xt is processed by the multi-task cascaded convolutional network MTCNN to obtain the encoded vector vec_t containing the appearance features of the target image. The two vectors are input to the generator G1 in the generative adversarial network model to be attacked. The generative adversarial network model is a face-swapping model.
[0040] Step 303: The generator outputs the synthesized image Y after face swapping. Y is input to the pre-trained semantic discriminator, which outputs semantic feature labels and realism.
[0041] After the above steps, the original image and background target image are input, and after passing through multiple deep neural network-based models, the semantic feature labels and realism of the synthesized image after face swapping can be output. The encoder model arcface, the generator model G1, and the semantic discriminator form an end-to-end model that can perform gradient backpropagation.
[0042] Step 304: Based on the needs of facial privacy protection, select one or more target semantics (such as gender, aging or youth, whether a red cross appears, etc.) from the semantic categories of the semantic discriminator as facial attributes that need to be controlled and generated;
[0043] Step 305: Iterate the above end-to-end model using the projected gradient descent method. First, set the hyperparameters required for the projected gradient descent method, including the upper bound of the image perturbation eps, the number of iterations n_iter, the perturbation step size per iteration eps_iter, and the loss function loss_fun().
[0044] Step 306: Iteration begins. Create a temporary attack image variable Xtmp, assign the source image Xs to Xtmp, and set the initial perturbation p = 0.
[0045] Step 307: In each iteration, the temporary attack images Xtmp and Xt are input into the end-to-end architecture described above to obtain the output image realism and actual semantic label vector;
[0046] Step 308: Calculate the loss using the loss function based on the actual semantic label vector and the target domain label to be generated in step 304;
[0047] Step 309: Based on the loss value at the last end, the gradient is backpropagated through the above-mentioned interconnected neural networks at each level. The loss is first passed to the semantic discriminator model to generate the gradient. The gradient continues to be backpropagated, passing through the face-swapping generation model and the arcface encoding network in turn, and then back to the source image Xs. At each pixel of Xs, there are some gradients g calculated through backpropagation.
[0048] Step 310: Apply the sign function to the gradient g generated in step 309 on the original image to obtain sign(g), and then multiply sg by the perturbation step size eps_iter to obtain the current perturbation p. tmp ;
[0049] The sign function mentioned in step 310 means that if the input is 0, the output is 0; if the input is greater than 0, the output is 1; if the output is less than 0, the output is -1; if the input is a multi-dimensional vector, the same action is performed on each element of the vector.
[0050] Step 311: Add the current perturbation to the temporary attack image Xtmp to obtain the latest image Xn = Xs + ptmp. After the perturbation upper bound stage, the actual perturbation ptmp = clip(Xn, -eps, eps) - Xtmp is obtained.
[0051] Step 312: Assign the actual disturbance ptmp to the disturbance p;
[0052] Step 313: Repeat steps 306-312 until the set number of iterations is reached;
[0053] Step 314: After a multi-step gradient descent attack, the generated adversarial noise p is added to the original image to obtain an image with added protective noise. This image is then used to replace the center part of the original image to obtain the final attacked image version.
[0054] Step 315: The adversarial attack on the face-swapping model ends.
[0055] The final attack image generated at this point will cause the face-swapping to fail after being processed by the deep generative adversarial network-based face-swapping model mentioned above. Furthermore, the synthesized image will contain some pre-designed semantic features, such as red crosses, gender reversal, and other watermarks, which will help people more easily detect fake images.
[0056] The controllable semantic features described in this invention, in addition to basic facial attributes such as aging, gender, and hair color (red or white), can also be watermark markers such as red crosses, red areas, and black areas. Furthermore, by adding image realism tags to the target tags, non-semantic destructive attacks can be directly achieved.
[0057] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the data privacy protection method based on adversarial learning as described above.
[0058] A computer-readable storage medium storing a computer program that performs the data privacy protection method based on adversarial learning as described above.
[0059] Beneficial Effects: Compared with existing technologies, this invention can achieve controllable semantic attacks on pre-trained generative adversarial models by adding invisible noise to the face image to be protected. This disrupts the process of synthesizing fake images without affecting the quality of the original image. This invention can semantically control the appearance of facial feature changes such as gender reversal, aging, and hair color alterations, or the inclusion of other information such as red crosses, watermarks, and identifiers in the synthesized image. This allows forgeries to be visually identifiable without detection, and the face-swapping fails, thus preventing damage to personal image. Combined with ensemble and targeted attacks, this invention can use a pre-trained semantic discriminator to attack multiple adversarial generative network models of the same type, exhibiting good transferability and stability in applications protecting facial privacy and security. Attached Figure Description
[0060] Figure 1 This is a schematic diagram of face swapping and face swapping after an attack, according to an embodiment of the present invention;
[0061] Figure 2 This is a schematic diagram of the semantic discriminator model architecture according to an embodiment of the present invention;
[0062] Figure 3 This is a flowchart illustrating a controllable semantic attack according to an embodiment of the present invention.
[0063] Figure 4 This is a diagram illustrating the effect of a controllable semantic attack according to an embodiment of the present invention. Detailed Implementation
[0064] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading the present invention, any modifications of the present invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.
[0065] An adversarial learning-based data privacy protection method is an attack scheme targeting deep generative adversarial networks (GANs)-based face-swapping models. It involves adding a tiny perturbation to a given face image to semantically alter the generated face-swapping image when applied to that image, causing the face-swapping to fail. Furthermore, the altered facial appearance attributes can be manually set, making it a controllable attack method. This perturbation is extremely small and imperceptible to the human eye, thus providing stealthy protection of facial privacy. In addition, the controllable alteration of semantic features in the synthesized image makes this protection method more effective, stable, and transferable. The face-swapping process and its effects after the attack are described below. Figure 1 As shown, X s It is the source image that provides identity features and is to be protected, X t It's a background image, Xμ This is the image after adding noise protection. It can be seen that there is a very obvious difference between the face-swapping results before and after adding noise protection; the face-swapping failed.
[0066] Furthermore, the proposed solution is also a novel way to edit deep generative adversarial networks. Compared to previous methods of editing deep generative adversarial networks, which generally operate in the latent space, this invention can directly perform semantic editing on the generated result by adding extremely small noise to the original image.
[0067] In this invention, a face-swapping model based on a deep generative adversarial network (GAN) is first trained to obtain a deep network model capable of swapping faces between two facial images. Next, a semantic discriminator is trained, which needs to possess semantic features that, given an input facial image, output the image's realism and appearance attributes, such as age, gender, and hair color. Finally, adversarial attack techniques are employed, using projective gradient descent to attack the face-swapping model, generating adversarial noise that is added to the image to be protected, resulting in a perturbated face image. This image disables the deep GAN's operation, preventing successful face swapping, and introduces controllable semantic feature blocks into the synthesized fake image, making it easier for people to detect its falsity, thus protecting the given facial image. Because this noise is extremely small and imperceptible to the human eye, it does not affect image quality.
[0068] The face-swapping model here uses a publicly available, convenient, and realistic algorithm—Faceshifter. Note that the controllable semantic attack method of this invention can not only attack this face-swapping algorithm; this invention is also applicable to other face-swapping or face editing models based on deep adversarial networks. The face-swapping model consists of three parts: an identity encoder that extracts identity features from the source image, a multi-level attribute encoder that extracts appearance attribute features from the target image, and a multi-layer generator that combines the above two features to generate a face-swapping image. The identity encoder extracts features representing the face identity of the source image; the multi-level attribute encoder extracts features representing the appearance attributes of the target image, with each layer of the encoder outputting an attribute feature vector; the multi-layer generator increases the resolution of the generated image layer by layer, combining the identity features of the source image and the appearance attribute features of the target image to generate a face-swapping image that possesses features from both. The identity encoder generally uses the publicly available arcface network model; the complete steps for training this model are as follows:
[0069] Step 401: Dataset preprocessing. Since the image sizes in the face dataset are not uniform, a pre-trained MTCNN (Multi-task Cascaded Convolutional Network) model is used to label the face locations contained in each image.
[0070] Step 402: If no face is found in step 401, the process ends. If a face is found, select the face with the highest weight from all the obtained face annotation boxes;
[0071] Step 403: Scale the face bounding box obtained in step 402 to 256×256 to obtain a size-normalized face, and record the position of the facial feature points output by MTCNN.
[0072] Step 404: Train the target attack model (face model). This model refers to the generative adversarial network used to synthesize faces, including a generator G1 that synthesizes a face-swapped forged image using face features from the source image and the background of the target image, and a discriminator D that identifies the authenticity of the image. The internal structure of the target attack model is not unique. Each iteration will be trained using the following steps:
[0073] Step 405: Randomly select source image X from the dataset s With target image X t One of each;
[0074] Step 406: Crop the source image X s The central area is 218×218 in size;
[0075] Step 407: Use bilinear interpolation to downsample the cropped region image obtained in step 406 to a size of 112×112;
[0076] Step 408: Use the arcface pre-trained model to extract the source image identity features z from the downsampled image. id (X s );
[0077] Step 409: X of the target image t Input a multi-level attribute encoder to extract multi-level attribute features z of the target image att (X t );
[0078] Step 410: Combine the source image identity features z obtained in Step 408 and Step 409 id (X s ) and target image multi-level attribute features z att (X t The image Y is generated by inputting it into the multi-layer generator G1.
[0079] Step 411: Input the generated image Y obtained in step 410 into the arcface model to extract the identity features z of the generated image. id (Y);
[0080] Step 412: Calculate the identity loss L idCosine similarity between the identity features of the source image and the identity features of the generated image:
[0081] L id =1-cos(z) id (Y),z id (X s ))
[0082] Step 413: Place X t The generated image Y obtained in step 410 is input into a multi-level attribute encoder to obtain the attribute features z of the corresponding image. att (X t ), z att (Y);
[0083] Step 414: Calculate attribute loss L att The attribute loss is the mean square error between the attribute features of the source image and the attribute features of the generated image, where... This refers to a certain dimension of the encoding vector:
[0084]
[0085] Step 415: If the source image and the target image represent the same person, then according to...
[0086] The reconstruction loss L is calculated by comparing the generated image with the source image. rec Otherwise, the reconstruction loss is 0:
[0087]
[0088] Step 416: Calculate the adversarial loss L based on the output of discriminator D. adv The adversarial loss uses the hinge loss function f. hinge The smaller the adversarial loss, the more realistic the image Y is:
[0089] L adv =f hinge (D(Y),1)
[0090] The hinge loss function f described in step 416 hinge This refers to:
[0091] f hinge (y,y′)=max(0,1-y*y′)
[0092] Step 417: Calculate the weighted sum of the losses obtained in steps 414-416 to obtain the total loss function of generator G1:
[0093] L G =L adv +10L att +5L id +10Lrec
[0094] Step 418: Use the Adam optimizer to optimize the generator G based on the loss, where the optimizer parameter beta is 0 and 0.999 respectively;
[0095] Step 419: Use the real image Y respectively real The generated image Y of generator G fake As input to the discriminator D, its optimization objective is to enable the discriminator to correctly identify forged images as fake images and real images as real images. hinge Let L represent the hinge loss function described in step 416, and the discriminant loss L of D. D The calculation is as follows:
[0096]
[0097] Step 420: Optimize the discriminator using the Adam optimizer based on the discriminative loss of D, where the optimizer parameter beta is 0 and 0.999 respectively;
[0098] Step 421: Repeat steps 405 to 420 until the model converges;
[0099] Step 422: Process complete.
[0100] After training the generative model in the face-swapping model using the above method, it is necessary to train a semantic discriminator model, the neural network structure of which is as follows: Figure 3 As shown, the main function of this model is to extract highly discriminative, identifiable, and describable features from images, serving as the basis for subsequent controllable attacks on image-generated features. The semantic discriminator model used in this invention takes a face image as input and outputs the image's realism and a 0 / 1 encoding of a given facial appearance attribute feature, such as age, gender, and hair color, indicating whether the face possesses that attribute feature. Furthermore, after specific training, this semantic discriminator model can also determine the presence of other semantic information such as red crosses, red eyes, and specific markers for post-attack identification. The semantic discriminator model training process is as follows:
[0101] Step 501: Prepare a face dataset in which face images are labeled with facial semantic tags, such as age, gender, hair color, nose size, lip thickness, whether glasses are present, and whether red crosses are present. Prepare a generator similar to StarGAN G2 and the semantic discriminator described above.
[0102] Step 502: At each step of the training process, a batch (e.g., 32 images) of images and their corresponding category labels are obtained from the dataset, and a set of appearance attribute vectors c are randomly generated in the set image classification domain space.
[0103] Step 503: Use generator G2 to generate a set of fake images G(x,c) based on the image x in step 502 and the random appearance attribute c;
[0104] Step 504: Input the real image and the fake image generated in step 503 into the semantic discriminator. The semantic discriminator distinguishes between the real image x and the fake image G(x,c) generated by the generator to obtain D. src (x), D src (G(x,c)) , classifying the domain to which the real image belongs to obtain D cls (x);
[0105] Step 505: Interpolate the real image and the fake image, input the interpolated image into the semantic discriminator, and calculate the gradient of the discrimination result with respect to the interpolated image to obtain the gradient penalty GP;
[0106] Step 506: Determine the result D based on the actual image. src (x), D cls (x) Forged image detection result D src The adversarial loss is calculated using (G(x,c)) and the gradient penalty GP obtained in step 505. Classification loss Total loss L D And optimize the semantic discriminator network parameters:
[0107]
[0108]
[0109]
[0110] Where x is the input image, c is the input domain label, c' is the domain label of the input image, H is the binary cross-entropy function, E denotes the expectation, and λ gp =10 is the gradient penalty coefficient, λ cls =1 is the classification penalty coefficient.
[0111] Step 507: Repeat steps 502 to 506; every 5 rounds, perform the following steps to train generator G2;
[0112] Step 508: Using the real image x from step 502 and the random appearance attribute c, use generator G2 to generate a fake image G(x,c). Use the generated fake image and the real image label c' to generate a reconstructed image G(G(x,c),c') again through generator G2.
[0113] Step 509: Calculate the reconstruction loss using the L1 norm based on the difference between the reconstructed image from Step 508 and the real image x.
[0114]
[0115] Step 510: Input the forged image G(x,c) into the semantic discriminator, and calculate the adversarial loss based on the discrimination result:
[0116]
[0117] Step 511: Use a semantic discriminator to generate the domain category label D for the real image x. cls The classification loss is calculated by comparing (x) with the random appearance attribute label c used to generate the fake image:
[0118]
[0119] Step 512: Apply the above reconstruction loss Combating losses Classification loss The total loss L of generator G2 is obtained by weighted summation. G , where λ cls =1, λ rec =10 is the weighting parameter, used to optimize the parameters of the generator G2 network:
[0120]
[0121] Step 513: Repeat steps 502 to 513 until the semantic discriminator converges;
[0122] Step 514: Process complete.
[0123] Controllable semantic attacks mainly consist of three parts: generating a face-swapped forged image Y; using a semantic discriminator to output the semantic category (class) of the face-swapped forged image; and using the projective gradient descent method to perform the attack. The target semantic category (class_target) that needs to be controlled to appear on the forged image is selected, the loss is calculated compared to the semantic category (class) of the actual face-swapped forged image, and adversarial noise (μ) is generated using the projective gradient descent attack on the source image X to be protected. s Adding adversarial noise results in a protected image, also known as an adversary example.
[0124] Controlled semantic attacks first require constructing the model to be attacked using the projective gradient descent method. The previously trained face-swapping model and semantic discriminator are concatenated as an end-to-end model. This involves using two face images as input, referred to as the source image and the target image. First, the source image is passed through the identity encoder in the face-swapping model to obtain identity features. Then, the target image is passed through a multi-level attribute encoder to obtain multi-level attribute features. These identity features and multi-level attribute features are input into a generator to produce the face-swapped image. Finally, the face-swapped image is passed through a semantic discriminator to obtain its semantic appearance features, which serve as the output of the overall model being attacked. In each iteration, a small amount of noise is added to the original image. After several iterations, the forged image generated by the generative network will selectively display the semantic information of the modified parts of the label.
[0125] Figure 3 This is a flowchart of a controllable semantic attack on a face-swapping model, explained in detail below:
[0126] Step 601: Prepare the generator model G1 and the discriminator model D that can output target semantic labels from the previously trained face-swapping model to be attacked.
[0127] Step 602: Denote the source image to be protected as X. s A face image is randomly selected from the dataset as the target image X. t The two images are cropped into a central 218×218 area;
[0128] Step 603: Connect the rows of each model, X s The encoded vector vec_s,X is obtained through the encoder model arcface. t The encoded vector vec_t is obtained by passing it through the multi-task cascaded convolutional network MTCNN;
[0129] Step 604: Input the two encoded vectors vec_s and vec_t into the generator G1 in the face-swapping model to be attacked;
[0130] Step 605: Generator G1 outputs the synthesized image Y after face swapping. Y is input to the pre-trained semantic discriminator, which outputs semantic feature labels and realism.
[0131] After the above steps, the original image of the identity to be protected and the background target image are input. After passing through multiple deep neural network-based models, the semantic feature labels and realism of the synthesized image after face swapping can be output, forming an end-to-end channel architecture that can perform gradient backpropagation.
[0132] Step 606: Based on the need for facial privacy protection, select one or more target labels c_t composed of target semantics from the semantic categories of the semantic discriminator, such as gender, aging and youth, hair color, nose size, etc., as the facial attributes that need to be controlled for generation;
[0133] In this example, the original semantic labels of the output are blonde, female, small nose, and not wearing glasses. During the attack, these labels are reversed or replaced to obtain the target semantic labels black hair, male, big nose, and wearing glasses.
[0134] Step 607: Iterate the above end-to-end model using the projective gradient descent method. First, set the hyperparameters required for the projective gradient descent method, including the upper bound of the image perturbation eps = 1, the number of iterations n_iter = 100, and the perturbation step size per iteration eps_iter = 7 * 10e. -6 loss function fun () represents the cross-entropy loss function, i.e.
[0135]
[0136] Step 608: Iteration begins, create temporary attack image variable X tmp , source image X s Assigned to X tmp The initial disturbance p = 0;
[0137] Step 609: In each iteration, the temporary attack image X... tmp and X t Input is given to the above end-to-end architecture;
[0138] Step 610: Forward propagation of the neural network, outputting the image realism tr and semantic label c from the final semantic discriminator;
[0139] Step 611: Calculate the loss using the loss function based on the actual semantic label vector c and the target domain label c_t selected in step 606, obtaining loss = loss fun (c,c t );
[0140] In step 611, a fully controlled semantic attack is performed. In addition, using -tr directly as the loss can also directly destroy the realism of the generated image. -tr and the loss in step 611 can also be used in combination.
[0141] Step 612: Based on the final loss value, the gradient is backpropagated through the above-mentioned interconnected neural networks. The loss is first passed to the semantic discriminator model to generate the gradient, and the gradient continues to be backpropagated, passing through the face-swapping generation model and the arcface encoding network before being passed to the source image X. s Above, in Xs Each pixel has a gradient g calculated through backpropagation;
[0142] Step 613: Apply the sign function to the gradient of the original image obtained in step 612 to obtain sign(g), and then multiply it by the perturbation step size of each step to obtain the current perturbation ptmp = eps_iter * sign(g).
[0143] The sign function mentioned in step 613 means that if the input is 0, the output is 0; if the input is greater than 0, the output is 1; if the output is less than 0, the output is -1; if the input is a multi-dimensional vector, the same action is performed on each element of the vector.
[0144] Step 614: Transfer the temporary attack image X tmp The latest image X is obtained by adding the current perturbation. n =X s +p tmp After the upper bound of the perturbation stage, the actual perturbation p is obtained. tmp =clip(X) n ,-eps,eps)-X tmp .
[0145] Step 615: Convert the actual disturbance p tmp Assign the value to the disturbance p;
[0146] Step 616: Repeat steps 609-615 until the set number of iterations is reached;
[0147] Step 617: After a multi-step gradient descent attack, the generated adversarial noise p is added to the original image to obtain an image with added protective noise. This image is then used to replace the center part of the original image to obtain the final attacked image version.
[0148] Step 618: The model attack ends.
[0149] Figure 4This image shows the effect of a controlled semantic attack on the Faceshifter face-swapping algorithm. The first column shows the source and background images, the second column shows the image after a normal face swap, and the first two rows of the following columns show the source image with added protective noise and the synthesized fake image after face swapping. It's clear that the source image with added protective noise is almost identical to the original image; the human eye cannot distinguish the difference, thus not affecting image quality. However, when the protected source image is then face-swapped by the attacked deep adversarial generative model, the generated fake image exhibits some pre-controlled semantic features, such as aging, a large nose, glasses, dark skin, and gender reversal. This causes the face swap to fail, and these generated semantic features are more easily recognized by AI algorithms. Furthermore, designing these semantic features as red crosses or attacking specific areas can also help the human eye more easily identify the fake image.
[0150] Obviously, those skilled in the art should understand that the steps of the face privacy protection method based on adversarial attack GANs in the above embodiments of the present invention can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using device-executable program code, thereby storing them in a storage device for execution by a computing device. Furthermore, in some cases, the steps shown or described can be performed in a different order than presented herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the embodiments of the present invention are not limited to any particular hardware and software combination.
Claims
1. A data privacy protection method based on adversarial learning, characterized in that, This paper proposes an adversarial attack method that controls the attack effect by using a deep generative adversarial network-based face-swapping model as the attacked model. Starting from the semantic representation level of face images, it achieves the generation of controllable images with semantic features after the attack. Utilizing the semantic decomposability of images, by concatenating the attacked model with a semantic discriminator, the semantic labels are modified to simultaneously achieve minor modifications to the original image to be protected and semantic changes to the generated image. Perturbations are added to a given face image so that when the face-swapping model is applied to the face image, the generated face-swapping image undergoes significant changes in the semantic appearance attributes, leading to face-swapping failure. The method includes semantic discriminator training and adversarial attacks. For the semantic feature types that are desired to appear in the generated images, the corresponding semantic discriminator is trained in advance. The face-swapping model uses two face images as input, referred to as source image X. s With target image X t The final output of the face-swapping result will contain the identity features of the source image and the facial attribute features and background features of the target image, achieving the effect of transferring and embedding the identity of the source image into the target image while keeping the area outside the face of the target image unchanged. The adversarial attack is based on the projective gradient descent method. For a labeled classification model, the descent direction is to make the output label of the input image change along the direction of the target label. That is, the loss based on the projective gradient descent method is the distance between the image output label and the target label. The method for generating controllable images with semantic features after an attack consists of three parts: generating a face-swapped forgery image Y; Use a semantic discriminator to output the semantic category of the face-swapping forged image; use projective gradient descent to launch the attack. Select the target semantic category (class_target) that needs to be controlled to appear on the fake image, calculate the loss between the semantic category (class) and the actual face-swapped fake image, and use a projection gradient attack to generate adversarial noise (μ) on the source image X to be protected. s Adding adversarial noise results in a protected image, also known as an adversarial example. To achieve a controllable semantic attack, which generates controllable images with semantic features after the attack, the first step is to construct the model required for the projective gradient descent method. The previously trained face-swapping model and semantic discriminator are concatenated as an end-to-end model, using two face images as inputs, referred to as the source image and the target image, respectively. First, the source image is passed through the identity encoder in the face-swapping model to obtain identity features. Then, the target image is passed through the multi-level attribute encoder to obtain multi-level attribute features. The obtained identity features and multi-level attribute features are input into the generator to obtain the face-swapped image. The face-swapped image is then passed through the semantic discriminator to obtain the semantic appearance features of the face-swapped image, which is used as the output of the overall model being attacked. In each iteration, a small amount of noise is added to the original image. After several iterations, the forged image generated by the generative network will selectively show the semantic information of the modified parts in the label. The adversarial attack method involves concatenating a trained semantic discriminator after the deep generative adversarial network model to be attacked. Specifically, the face-swapped image obtained from the face-swapping model is input back into the semantic discriminator model to obtain the semantic features of the face-swapped image. A loss function is calculated using the actual semantic features and the semantic features to be controlled. The calculated loss is then used to perform an iterative attack on the entire concatenated model, including the face-swapping model and the semantic discriminator model, using the projective gradient descent method to generate adversarial noise on the original image, thereby obtaining the adversarial sample, i.e., the image after adding protective noise.
2. The data privacy protection method based on adversarial learning according to claim 1, characterized in that, The construction and training steps of the face-swapping model are as follows: Step 101: Prepare the encoder model arcface for encoding the identity of the source image; Step 102: Prepare a multi-level encoder model (MAE) for encoding facial attribute features of the target image; Step 103: Construct a generator model G1 that combines the identity features of the source image with the facial attribute features of the target image to generate multi-level magnified images; Step 104: Dataset preprocessing, the face dataset images are sequentially annotated with the face locations using the multi-task cascaded convolutional network MTCNN model; Step 105: If no face was detected in step 104, skip this image; Step 106: If a face is detected in step 104, select the face with the highest weight, stretch the bounding portion of the face to the set size, and store it in a new dataset; Step 107: Use the encoder model arcface as the identity encoder for the face-swapping model; Step 108: Transfer the source image X from the new dataset s The central region is extracted and downsampled to a set size, and used as the input to the encoder model arcface in step 107 to obtain the identity encoding vector z. id ; Step 109: X of the target image t The input is a multi-level attribute encoder, which yields multiple attribute vectors z. att ; Step 110: Combine the identity encoding vector z obtained in Step 108 and Step 109 id With attribute vector z att The image Y is generated by inputting it into the multi-layer generator model G1. Step 111: Input the generated image Y into the discriminator D to obtain the discrimination result; Step 112: Calculate the adversarial loss, identity loss, attribute loss, and reconstruction loss respectively, and sum them up by weight to obtain the generator training loss. Then, backpropagate to optimize the parameters of the generator model G1 network. Step 113: Calculate the discriminator D loss and backpropagate to optimize the discriminator D network parameters; Step 114: Repeat steps 108 to 114 until the generator model G1 and the discriminator model D converge; Step 115: Training of the face-swapping model is complete.
3. The data privacy protection method based on adversarial learning according to claim 1, characterized in that, The semantic discriminator model used is the discriminator model in the image domain editing model based on deep generative adversarial networks. The output of the semantic discriminator model is divided into two parts: the judgment of the authenticity of the image and the domain category encoding of the image. During the training of the semantic discriminator model, the generator model G2 takes the original image and the target domain label of the image as input. It regenerates the fake image based on the target domain label according to the original image, and in the optimized loss function, it makes the semantic discriminator unable to distinguish between the fake image and the real image, while the domain classification result of the image conforms to the given target domain label. The input to the semantic discriminator is either the original image or a forged image generated by the generator model G2. The image features are extracted through six convolutional layers, and the obtained features are fed into two fully connected layers. The two fully connected layers output the true / false classification result and the domain classification vector, respectively. The goal of the semantic discriminator is to determine the authenticity and domain label of the input image. During training, the semantic discriminator is updated every N steps, and the generator model G2 is updated once. The generator model G2 first arranges the domain information according to the image size, concatenates it with the input image, and then inputs it into the convolutional network. The network includes one convolutional layer, two downsampling convolutional layers, six residual layers, two upsampling deconvolutional layers, and one convolutional layer.
4. The data privacy protection method based on adversarial learning according to claim 1, characterized in that, The training steps for the semantic discriminator model are as follows: Step 201: Prepare a face dataset containing domain category labels, where the domain category labels are semantic attribute features of face appearance; Prepare the image editing generator model G2 and the semantic discriminator; Step 202: Train the semantic discriminator. The semantic discriminator distinguishes between real images and fake images generated by the generator model G2. Real images are classified according to their respective domains. The classification loss is calculated and the parameters of the semantic discriminator network are optimized. Step 203: Train generator model G2, which uses real images and target domain labels as input to generate fake images; Step 204: Use the domain labels of the fake image and the real image generated by the generator model G2 in step 203 as input to generate a reconstructed image, and calculate the difference between the reconstructed image and the original real image as the reconstruction loss; Step 205: Input the fake image generated by the generator model G2 in step 203 into the semantic discriminator to distinguish between real and fake images, calculate the adversarial loss, and optimize the network parameters of the generator model G2 by combining the reconstruction loss in step 204. Step 206: Repeat steps 202 to 205 until the generator model G2 and the semantic discriminator model converge; Step 207: The semantic discriminator is trained, and the training ends.
5. The data privacy protection method based on adversarial learning according to claim 1, characterized in that, Countermeasures against attacks include the following steps: Step 301: Denote the source image to be protected as X. s Select a face image from the dataset as the target image X. t Cropping the center regions of both images separately, and then cropping the source image to X. s and target image X t For step 302; Step 302: X s The encoder model arcface generates an encoded vector vec_s containing the identity features of the source image. t The encoding vector vec_t containing the appearance features of the target image is obtained through the multi-task cascaded convolutional network MTCNN. The two vectors are then input into the generator model G1 in the generative adversarial network model to be attacked. The generative adversarial network model is a face-swapping model. Step 303: The generator outputs the synthesized image Y after face swapping. Y is input to the trained semantic discriminator, which outputs semantic feature labels and realism. After the above steps, the original image and background target image are input. After passing through multiple deep neural network-based models, the semantic feature labels and realism of the synthesized image after face swapping can be output. The encoder model arcface, the generator model G1, and the semantic discriminator form an end-to-end model that can perform gradient backpropagation. Step 304: Based on the needs of face privacy protection, select one or more target semantics from the semantic categories of the semantic discriminator as the face attributes to be controlled and generated, that is, the target domain labels to be controlled and generated; Step 305: Iterate the above end-to-end model using the projected gradient descent method. First, set the hyperparameters required for the projected gradient descent method, including the upper bound of the image perturbation eps, the number of iterations n_iter, the perturbation step size per iteration eps_iter, and the loss function loss_fun(). Step 306: Iteration begins, create temporary attack image variable X tmp , source image X s Assigned to X tmp The initial perturbation is p=0; Step 307: In each iteration, the temporary attack image X... tmp and X t The input is fed into the above end-to-end model to obtain the image realism and actual semantic label vector. Step 308: Calculate the loss using the loss function based on the actual semantic label vector and the target domain label to be generated in step 304; Step 309: Based on the loss value, gradient information is backpropagated through the neural networks strung together in the end-to-end model described above. The loss value is first passed to the semantic discriminator model to generate the gradient, and the gradient continues to be backpropagated, passing through the generator model G1 and the encoder model arcface in the face-swapping model before being passed to the source image X. s Above, in X s Each pixel has a gradient g calculated through backpropagation; Step 310: Apply the sign function to the gradient g generated in step 309 on the original image to obtain sign(g), and then multiply sg by the perturbation step size eps_iter to obtain the current perturbation p. tmp ; The sign function mentioned in step 310 means that if the input is 0, the output is 0; if the input is greater than 0, the output is 1; if the output is less than 0, the output is -1. If the input is a multidimensional vector, the same action is performed on each element of the vector; Step 311: Transfer the temporary attack image X tmp The latest image X is obtained by adding the current perturbation. n =X s +p tmp After the upper bound of the perturbation stage, the actual perturbation p is obtained. tmp =clip(X n , -eps,eps)-X tmp ; Step 312: Convert the actual disturbance p tmp Assign the value to the disturbance p; Step 313: Repeat steps 306-312 until the set number of iterations is reached; Step 314: After a multi-step gradient descent attack, the generated adversarial noise p is added to the original image to obtain an image with added protective noise. This image is then used to replace the center part of the original image to obtain the final attacked image version. Step 315: The adversarial attack on the face-swapping model ends.
6. The data privacy protection method based on adversarial learning according to claim 5, characterized in that, The target semantics are facial attributes or watermarks; the facial attributes include: aging or youth, gender, and hair color; the watermarks include red crosses, red areas, and black areas.
7. The data privacy protection method based on adversarial learning according to claim 5, characterized in that, Adding image authenticity tags to target tags allows for direct, non-semantic destructive attacks.
8. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the data privacy protection method based on adversarial learning as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program that performs the data privacy protection method based on adversarial learning as described in any one of claims 1-7.