Multi-module collaboration-based antagonistic facial privacy protection method

By employing a multi-module collaborative adversarial generative framework, the problems of facial geometric distortion and low visual fidelity in facial recognition systems are solved, generating high-quality adversarial examples and achieving synergistic optimization of facial privacy protection and attack effectiveness.

CN121502800APending Publication Date: 2026-02-10CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511613277.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing generative adversarial networks suffer from facial geometry distortion and low visual fidelity in facial recognition systems, making it difficult to generate high-quality and aggressive adversarial examples that are easily detected.

Method used

A multi-module collaborative adversarial generation framework is adopted, including a generator, a discriminator, a regularization module, a transferability enhancement module, and a facial key point structure preservation module. Through multi-task joint optimization of target training, the generator acquires and fuses source images and reference makeup images to generate adversarial examples, the discriminator distinguishes between real images and adversarial examples, the regularization module cleans up adversarial noise, and the facial key point structure preservation module quantifies geometric structure bias.

Benefits of technology

Generating natural and realistic protected images can effectively mislead facial recognition systems, maintain the consistency of facial geometry, achieve high attack efficiency and visual quality, and provide reliable privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121502800A_ABST
    Figure CN121502800A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of face recognition and information protection, in particular to a multi-module cooperation-based antagonistic face privacy protection method, which comprises the following steps of: training an antagonistic generation framework to obtain an antagonistic generation model fused with key point constraints; a protection image capable of effectively misleading a face recognition system is generated by adopting an antagonism generation model, so that face privacy protection of the user is realized; the antagonism generation framework comprises a generator, a discriminator, a regularization module, a mobility enhancement module and a face key point structure keeping module, and a multi-task joint loss function is adopted for training. According to the method, collaborative optimization is carried out on the adversarial sample visual quality, the geometric structure fidelity and the black box attack efficiency, a protection image which is natural and vivid and can effectively mislead a face recognition system can be generated, and the problems of face geometric distortion, adversarial noise and content keeping conflicts in a traditional adversarial face privacy protection method can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of facial recognition and information protection technology, specifically to an adversarial facial privacy protection method based on multi-module collaboration. Background Technology

[0002] With the rapid development of deep learning technology, facial recognition systems are widely used in scenarios such as identity authentication and security monitoring. However, these systems also bring serious risks of personal privacy leaks. Malicious parties can scrape facial images from public platforms and use facial recognition technology to track identities and mine social relationships, seriously infringing on user privacy.

[0003] In existing technologies, methods based on Generative Adversarial Networks (GANs) interfere with facial recognition models by generating more natural-looking facial images, thereby protecting user privacy. However, existing GANs generally suffer from drawbacks when performing complex tasks such as makeup transfer: they easily lead to distortion of facial geometry and misalignment of key facial features. The introduction of adversarial noise often conflicts with content preservation mechanisms (such as cycle consistency), resulting in unnatural shifts in the relative positions of the eyes, nose, and mouth in the generated images, leading to low visual fidelity and easy detection by the human eye. This highlights the problem of poor visual quality of adversarial samples, making them easily detectable by humans or detectors.

[0004] Therefore, there is an urgent need for a facial privacy protection method based on adversarial image generation that possesses the core characteristics of high quality and high attack power, while maintaining the consistency of facial geometry. Summary of the Invention

[0005] In view of this, this application discloses an adversarial facial privacy protection method based on multi-module collaboration to solve the above problems; including: training an adversarial generation framework to obtain an adversarial generation model that integrates key point constraints; using the adversarial generation model to generate a protected image that can effectively mislead the facial recognition system, thereby achieving facial privacy protection for users;

[0006] The adversarial generation framework includes: a generator, a discriminator, a regularization module, a transferability enhancement module, and a facial key point structure preservation module;

[0007] The generator is used to acquire and fuse the source image and the reference makeup image, embed adversarial perturbations, and generate adversarial examples;

[0008] The discriminator, as the evaluator of the resistance generation framework, is used to distinguish between real images and adversarial examples;

[0009] The regularization module is used to remove or weaken adversarial perturbations in adversarial examples while preserving the content of the source image and the reference makeup image.

[0010] The transferability enhancement module guides the generator to learn adversarial features with generalization and transferability; the transferability enhancement module consists of different face recognition models; the different face recognition models need to be pre-trained and have fixed parameters;

[0011] The facial key point structure preservation module is used to quantify the deviation of adversarial examples from the source image in terms of geometric structure, and provides a supervision signal for structure preservation during training. It needs to be pre-trained and has fixed parameters.

[0012] The beneficial effects of this application include:

[0013] An adversarial facial privacy protection framework based on multi-module collaboration is proposed. Through the design of generator, discriminator, regularization module, transferability enhancement module and facial key point structure preservation module, the visual quality of adversarial examples, geometric structure fidelity and black box attack effectiveness are synergistically optimized.

[0014] Training is performed based on a multi-task joint optimization objective, including regularized cycle consistency loss and key point structure loss. This method can generate natural and realistic protected images that can effectively mislead facial recognition systems, and can solve the problems of facial geometric distortion, adversarial noise and content preservation conflict in traditional adversarial facial privacy protection methods.

[0015] The adversarial facial privacy protection method designed in this application provides a reliable means of privacy protection for users to securely share personal photos on public platforms such as social media. It achieves an excellent balance between privacy protection and visual usability, and provides a new design idea for those skilled in the art. Attached Figure Description

[0016] Figure 1 A schematic diagram illustrating the overall concept of the adversarial facial privacy protection method based on multi-module collaboration designed for this application;

[0017] Figure 2 This is a schematic diagram of the adversarial generation framework in an embodiment of this application;

[0018] Figure 3 This is a schematic diagram of the generator structure in an embodiment of this application;

[0019] Figure 4 This is a schematic diagram of the regularization module in an embodiment of this application;

[0020] Figure 5 This is a schematic diagram of the facial key point structure preservation module in the embodiments of this application;

[0021] Figure 6 This is a comparative test result of the adversarial generative model in the embodiments of this application and the adversarial attack methods in the prior art. Detailed Implementation

[0022] To make the objectives, technical solutions, features, and advantages of this application clearer and to enable those skilled in the art to better understand the technical solutions of this application, the following detailed description of this application is provided in conjunction with the accompanying drawings and embodiments.

[0023] Example 1:

[0024] This embodiment includes an adversarial facial privacy protection method based on multi-module collaboration, comprising: training an adversarial generative framework to obtain an adversarial generative model that integrates keypoint constraints; and using the adversarial generative model to generate a protected image that can effectively mislead the facial recognition system, thereby achieving user facial privacy protection. Figure 1 The diagram shows the overall concept of the adversarial facial privacy protection method based on multi-module collaboration designed in this application.

[0025] The adversarial generation framework includes: a generator, a discriminator, a regularization module, a transferability enhancement module, and a facial key point structure preservation module.

[0026] The generator is used to acquire and fuse source images and reference makeup images, embed adversarial perturbations, and generate adversarial examples. The discriminator, as the evaluator of the adversarial generation framework, is used to distinguish between real images and adversarial examples. The regularization module is used to remove or weaken adversarial perturbations in adversarial examples while preserving the content of source images and reference makeup images. The transferability enhancement module guides the generator to learn adversarial features with generalization and transferability. The transferability enhancement module consists of different face recognition models. The different face recognition models need to be pre-trained and have fixed parameters. The facial key point structure preservation module is used to quantify the deviation of adversarial examples from source images in terms of geometric structure, providing supervision signals for structure preservation during training. It also needs to be pre-trained and have fixed parameters.

[0027] The generator G designed in this application, namely the makeup generation network, has the following structure: Figure 3 As shown, a dual-branch encoder-decoder architecture is adopted, consisting of two core sub-networks: the makeup parameter extraction network PNet and the makeup application network TNet. In this embodiment, the input image size obtained by the generator G is 3×256×256.

[0028] The makeup parameter extraction network PNet is used to decouple and extract the style parameters of the makeup from the reference makeup image y. Specifically, PNet is composed of a PNet encoder, a PNet residual block, and a PNet decoder connected in sequence.

[0029] Specifically, the PNet encoder consists of interconnected 7×7 convolutional layers and two downsampling modules. The downsampling module is composed of a 4×4 convolution with a stride of 2, an instance normalization module, and a ReLU activation function module connected together, used to encode the image into deep features. The PNet residual block consists of three interconnected residual blocks used to perform deep feature transformation on the feature map; each residual block consists of two 3×3 convolutional layers connected together. The PNet decoder consists of two parallel 1×1 convolutional layers used to decode the final spatially variable style parameters γ and β from the deep features.

[0030] The TNet application network is used to fuse the identity information of the source image x with the style parameters γ and β extracted by PNet to obtain the adversarial reconstructed image, i.e., the adversarial example. Specifically, TNet consists of a TNet encoder, a content residual block group, a bottleneck fusion layer, a TNet residual block, a TNet decoder, and a TNet output layer connected in sequence.

[0031] Specifically, the TNet encoder adopts the same design as the PNet encoder, consisting of a 7×7 convolutional layer and two downsampling modules, mainly used to extract the multi-scale content features fcontent of x; the content residual block group consists of three connected content residual blocks; the bottleneck fusion layer adopts a spatial adaptive normalization (SPADE) mechanism, guided by a non-local attention module, to effectively inject style parameters γ and β into the content features; the TNet residual block consists of three connected residual blocks for information integration; the TNet decoder consists of two upsampling modules, each consisting of a 4×4 transposed convolution with a stride of 2, an instance normalization module, and a ReLU activation function module, used to gradually restore deep features to the original resolution; the TNet output layer consists of a 7×7 convolutional layer and a Tanh activation function, used to output the final adversarial image y. fake .

[0032] In this embodiment, the discriminator D adopts the PatchGAN architecture to judge the authenticity of local regions of the input image, thereby guiding the generator G to generate more realistic local details.

[0033] Specifically, the discriminator consists of an input layer, an intermediate layer, and an output layer. The input layer is a 4×4 convolutional layer used to map the input image of the discriminator to the initial feature space. The intermediate layer consists of five sequentially connected downsampling modules, each of which is composed of a 4×4 convolution with a stride of 2, an instance normalization layer, and a LeakyReLU activation function layer connected in sequence. The intermediate layer is used to halve the size of the feature map layer by layer while doubling the number of feature channels layer by layer to extract discriminative features at different scales. The output layer is a convolutional layer used to map the deep feature map into a single-channel discriminant score map, where each element in the discriminant score map represents the probability that a local region of the input image is true.

[0034] The regularization module H, i.e., the anti-noise decanting network, has the following structure: Figure 4 As shown, it is used to achieve effective separation of adversarial disturbances and image content, and its input image size is the same as that of the generator.

[0035] Specifically, the regularization module consists of a shallow feature extraction layer, an RRDB module, and an image reconstruction layer. The shallow feature extraction layer, composed of a 3×3 convolutional layer, captures the initial features of the image. The RRDB module, consisting of 23 stacked residual dense blocks (RRDBs), forms the backbone network of the regularization module. Each RRDB block comprises a densely connected convolutional layer and a residual connection layer. This structure maximizes information flow and effectively fuses multi-scale features, thereby separating high-frequency adversarial noise. The image reconstruction layer, composed of a 3×3 convolutional layer, outputs the final cleaned image y reconstructed from the clean features. clean .

[0036] The transferability enhancement module M, also known as the multi-model feature extractor, is used to simulate black-box attacks and guide G to learn more generalizable and transferable adversarial features.

[0037] The transferability enhancement module M consists of different face recognition (FR) models, such as IR-152, IRSE-50, and FaceNet. These FR models need to be pre-trained with fixed parameters. FR models pre-trained on large face datasets can extract deep features with strong identity discrimination capabilities.

[0038] Furthermore, the M module designed in this application does not require training itself. Its function is to provide diverse, black-box perspectives of identity features, providing a basis for calculating the loss against adversarial attacks.

[0039] Facial Key Point Structure Preservation Module (FLSP), structure as follows Figure 5 As shown, the FLSP module is designed in this application to address the problem of geometric distortion and is used to objectively score facial geometry.

[0040] Furthermore, FLSP is a pre-trained facial keypoint detector with fixed parameters. The FLSP module processes the data as follows: acquiring the input image, i.e., the source image or adversarial example; performing forward propagation on the input image and outputting the coordinate matrix P of the N facial keypoints of the image.

[0041] P is used to quantify geometric bias and provide a supervisory signal. It uses coordinate values ​​to accurately measure the difference in facial geometry between the adversarial sample and the source image, and backpropagates this bias as a loss function. This forces the generator to maintain facial geometry consistent with the source image during training, suppressing distortion and facial feature misalignment. P can be constructed simply by loading and freezing the pre-trained weights.

[0042] This embodiment uses a detector trained on the 300W-LP dataset that can detect 68 facial keypoints. Its parameters are fixed to ensure an objective and neutral geometric evaluation standard, preventing it from catering to the erroneous output of the generator G during training in an attempt to reduce loss.

[0043] Furthermore, a multi-task joint loss function is used when training the adversarial generative framework. The multi-task joint loss function includes at least: regularized cycle consistency loss and keypoint structure loss.

[0044] This application specifically designs a multi-task joint loss function, employing an alternating optimization strategy to train the discriminator D, generator G, and regularization module H separately. For ease of description, intermediate variables generated as data flows through the adversarial generative framework are defined, such as... Figure 2 As shown: Generator G receives source image x and reference dummy image y, and outputs preliminary adversarial sample y. fake ; Regularization module H for y fake Perform purification processing and output purified image y clean In a circularly consistent path, G receives y. clean Combine x with the reconstructed image to output the reconstructed adversarial image x. rec The regularization module H again applies the regularization to x. rec After purification, the final reconstructed image x is output. cycle x cycle Used for subsequent calculation of the loss function.

[0045] The multi-task joint loss function designed in this application includes the following key components:

[0046] Discriminator Loss Function: The goal of discriminator D is to distinguish between the real, made-up image y and the fake, initial adversarial image y. fake Its loss function is:

[0047]

[0048] in, This represents the discriminator loss function. This represents the expected value sampled from the true image distribution Y. This represents the logarithm of the probability that the discriminator classifies a real image as "real". This represents the expected value sampled from the source image distribution X and the reference image distribution Y. This represents the logarithm of the probability that the discriminator will classify a forged image as "fake".

[0049] Generator Loss Function: The adversarial loss objective of the generator G and the regularization module H is to deceive the discriminator D into classifying their respective outputs as "true". This loss also applies to the output y of G. fake and the output y of H clean The calculation formula is as follows:

[0050]

[0051] in, Represents the generator loss function. This indicates that the discriminator will forge images. The logarithm of the probability of being judged as "true". This represents the expected value sampled from the X and Y distributions. This indicates that the discriminator will clean the image. The logarithm of the probability of being judged as "true".

[0052] Further adversarial discriminant loss is obtained:

[0053]

[0054] in, This indicates a loss due to the opposing judgment.

[0055] Regularization cycle consistency loss: Used to ensure that the generator and regularization module work together to ensure that the entire generation and cleanup process preserves content. This loss simultaneously constrains G and H, forcing the generator and regularization module to work together to accurately reconstruct the original content. The calculation formula is:

[0056]

[0057] in, This represents the loss of regularization cycle consistency. Indicates a cyclically reconstructed image. This represents the source image.

[0058] Identity Preservation Loss: This loss aims to constrain the generator to not change the identity features of an image when unnecessary. Ideally, the output should be completely identical to the input x when the input source image x and the reference image y are the same image. Overall content similarity is guaranteed by calculating the pixel-level L1 norm distances between G(x,x) and H(G(x,x)) and x, respectively.

[0059]

[0060] in, Let G(x,x) represent the identity preservation loss, G(x,x) represent the generator output when the source image x is used as both content and style input, and H(G(x,x)) represent the cleaned output of G(x,x).

[0061] Keypoint Structure Loss: The coordinate matrix P of the facial keypoints output by the FLSP module is used to force the output y of the generator G. fake To preserve the facial geometry of the source image x, two sets of keypoint coordinates are first extracted using FLSP with fixed parameters, and then the L1 distance is calculated point by point and averaged.

[0062]

[0063] in, Indicates the structural loss at key points. This represents the facial keypoint structure preservation module, where N represents the total number of keypoints in the coordinate matrix P. and They represent the first The coordinates of key points on the source and generated images.

[0064] Counter-attack loss: used to enhance the portability of black-box attacks, for y fake and y clean Apply a diversity transformation to the input (such as random scaling, Gaussian noise), and then input it into multiple FR models M in the M module. k The formula for calculating the loss is:

[0065]

[0066] in, Indicates losses incurred in combating attacks. This represents calculating the expected value of a sample pair of x, y. Let represent the k-th FR model in the portability enhancement module. Indicates the index of the FR model.

[0067] Makeup style loss: used to ensure the generated image y fakeTo accurately replicate the makeup style of the reference image y, this embodiment employs a histogram loss based on facial semantic regions. This loss is calculated independently on pre-segmented facial regions (such as lips, skin, and eyes). By matching the color histogram distribution of the reference makeup image with adversarial examples in the corresponding regions, the generator is forced to learn and replicate the color and lighting style of the reference makeup image. The calculation formula is as follows:

[0068]

[0069] in, This indicates a loss of makeup and styling style. This indicates a pre-segmented facial region. This indicates semantic region division of an image. Masking operation, This represents the color histogram of the calculated image region.

[0070] Furthermore, the generator G and the regularization module H are trained end-to-end under a unified multi-task joint loss function. The total loss of G is as follows:

[0071]

[0072] in, This indicates the multi-task joint loss function ultimately used in training the adversarial generative framework. , , , , , These represent the adversarial discrimination loss, regularized cycle consistency loss, adversarial attack loss, makeup style loss, identity preservation loss, and keypoint structure loss, respectively. , , , , , represents the weights of the loss functions for adversarial discrimination loss, regularized cycle consistency loss, adversarial attack loss, makeup style loss, identity preservation loss, and keypoint structure loss, respectively.

[0073] Example 2:

[0074] 1. Preprocess the makeup transfer dataset to obtain an adversarial facial privacy protection dataset.

[0075] This embodiment uses the publicly available Makeup Transfer (MT) dataset for training, which contains a large number of facial images with and without makeup. Testing and evaluation use the independent, high-quality face dataset CelebA-HQ to ensure objectivity. All images undergo uniform preprocessing before being input into the network, including face alignment and cropping using a facial landmark detector, and scaling all images to a uniform 256×256 pixel resolution. The adversarial face privacy-preserving dataset includes both training and testing sets.

[0076] 2. Construct an adversarial generative framework.

[0077] In this embodiment, the adversarial generation framework comprises five modules: a generator (G), a regularization module (H), a discriminator (D), a transferability enhancement module (M), and a facial keypoint structure preservation module (FLSP). G employs... Figure 3 The dual-branch encoder-decoder structure shown uses H. Figure 4 The RRDB architecture is shown. The FLSP module uses a detector pre-trained on the 300W-LP dataset that can detect 68 keypoints, and its weights are fixed during training. The M module is an integration of pre-trained FR models such as IR-152, IRSE-50, and FaceNet. The overall framework of the model in this invention is as follows: Figure 1 As shown.

[0078] 3. Train the adversarial generative framework to obtain an adversarial generative model that incorporates keypoint constraints.

[0079] This embodiment was implemented using the Ubuntu 20.04 operating system, with a hardware platform configuration of an Intel Xeon Silver4210R CPU and an NVIDIA GeForce RTX 3090 GPU (24GB VRAM). The software development environment was based on Python 3.8 and PyTorch 1.11.0.

[0080] The training parameters in this embodiment are set as follows: a total of 100 training epochs, a batch size of 4, an optimizer named Adam, and a learning rate of 0.0002. The weight coefficients of each component loss in the joint loss function (Equation 4) are set as follows: , , , , , .

[0081] The specific training process includes: loading the training set from the adversarial face privacy-preserving dataset and iteratively training using an alternating optimization strategy. Each iteration includes:

[0082] Step 1: Fix the parameters of the generator G and the regularization module H, and train the discriminator D using real images and images generated by G.

[0083] Step 2: Fix the parameters of the discriminator D, and train G and H using a multi-task joint loss function.

[0084] Step 3: Adjust the weights of each module in the adversarial generative framework through backpropagation and gradient descent to end this iteration.

[0085] After each training cycle, the model's performance metrics are evaluated on the validation set. After training, the best-performing model weights are saved and deployed to the adversarial generative framework to obtain an adversarial generative model that incorporates keypoint constraints.

[0086] 4. Test the adversarial generative model.

[0087] The test set from the adversarial facial privacy protection dataset is loaded to test the adversarial generative model with fused keypoint constraints. To accurately evaluate the performance of the proposed method, several detection performance metrics are employed, including Attack Success Rate (ASR), Fréchet Inception Distance (FID), Structural Similarity (SSIM), and Landmark Mean Distance (LMD), which is the focus of this application. The formulas for calculating these detection performance metrics are given below:

[0088] The formula for attack success rate is:

[0089]

[0090] in, This represents the number of images that were successfully misidentified as the target identity by the black-box model. This represents the total number of test images.

[0091] The formula for FID is:

[0092]

[0093] in,( , ) and( , ) represent the mean and covariance matrices of the real image features and the generated image features, respectively, and Tr represents the trace of the matrix.

[0094] The formula for structural similarity is:

[0095]

[0096] in, The mean, and Represents variance. This represents the covariance.

[0097] The formula for the average distance between key points is:

[0098]

[0099] Where N is the total number of facial key points. and They represent the first The coordinates of key points on the source and generated images.

[0100] Furthermore, to test the advancement of the design method in this application, the adversarial generative model integrating key point constraints in this embodiment is compared with current state-of-the-art adversarial attack methods in a black-box attack scenario. The black-box attack model chosen is MobileFace, and the test set consists of 1000 pairs of faces with different identities randomly selected from CelebA-HQ to evaluate attack performance. The comparative test results of the adversarial generative model and existing adversarial attack methods are as follows: Figure 6 As shown, from Figure 6 As can be seen, the attack confidence of the method designed in this application is significantly higher than that of existing state-of-the-art methods. Furthermore, this invention demonstrates significant advantages in both SSIM and LMD metrics, proving that it can better maintain the stability of facial geometry and generate images of higher visual quality while maintaining high attack power.

[0101] Finally, it should be noted that the above description only depicts some embodiments of this application. For those skilled in the art, various changes, modifications, substitutions, and variations can be conceived of these embodiments without departing from the principles and spirit of this application. The scope of protection of this application is defined by the appended claims and their equivalents, and all the above-mentioned behaviors should be covered within the scope of protection of this application.

Claims

1. An adversarial facial privacy protection method based on multi-module collaboration, characterized in that, include: Train the adversarial generative framework to obtain an adversarial generative model that incorporates keypoint constraints; An adversarial generative model is used to generate protective images that can effectively mislead facial recognition systems, thereby protecting users' facial privacy. The adversarial generation framework includes: a generator, a discriminator, a regularization module, a transferability enhancement module, and a facial key point structure preservation module; The generator is used to acquire and fuse the source image and the reference makeup image, embed adversarial perturbations, and generate adversarial examples; The discriminator, as the evaluator of the resistance generation framework, is used to distinguish between real images and adversarial examples; The regularization module is used to remove or weaken adversarial perturbations in adversarial examples while preserving the content of the source image and the reference makeup image. The transferability enhancement module guides the generator to learn adversarial features with generalization and transferability; the transferability enhancement module consists of different face recognition models; the different face recognition models need to be pre-trained and have fixed parameters; The facial key point structure preservation module is used to quantify the deviation of adversarial examples from the source image in terms of geometric structure, and provides a supervision signal for structure preservation during training. It needs to be pre-trained and has fixed parameters.

2. The adversarial facial privacy protection method based on multi-module collaboration according to claim 1, characterized in that, The generator adopts a dual-branch encoder-decoder architecture, consisting of two core sub-networks: a makeup parameter extraction network and a makeup application network. The makeup parameter extraction network is used to decouple and extract the style parameters of the makeup from the reference makeup image; the makeup application network is used to fuse the identity information and style parameters of the source image to obtain the adversarial reconstruction image, i.e., the adversarial sample.

3. The adversarial facial privacy protection method based on multi-module collaboration according to claim 1, characterized in that, The regularization module is composed of a shallow feature extraction layer, an RRDB module, and an image reconstruction layer connected together. The shallow feature extraction layer consists of a 3×3 convolutional layer; the RRDB module consists of 23 stacked residual dense blocks; the residual dense blocks consist of densely connected convolutional layers and residual connected layers; the image reconstruction layer consists of a 3×3 convolutional layer.

4. The adversarial facial privacy protection method based on multi-module collaboration according to claim 1, characterized in that, The portability enhancement module itself does not require training.

5. The adversarial facial privacy protection method based on multi-module collaboration according to claim 1, characterized in that, The facial key point structure preservation module processes the data as follows: acquiring the input image, i.e., the source image or adversarial sample; performing forward propagation on the input image and outputting the coordinate matrix P of N facial key points of the image.

6. The adversarial facial privacy protection method based on multi-module collaboration according to claim 1, characterized in that, The training of the adversarial generative framework includes: loading a training set and iteratively training using an alternating optimization strategy; each iteration includes: Step 1: Fix the parameters of the generator and regularization module, and train the discriminator using real images and adversarial examples; Step 2: Fix the parameters of the discriminator and train the generator and regularization module using a multi-task joint loss function; Step 3: Adjust the weights of each module in the adversarial generative framework using backpropagation and gradient descent to end this iteration; After each training cycle, the model's performance metrics are evaluated on the validation set. After training, the best-performing model weights are saved and deployed to the adversarial generative framework to obtain an adversarial generative model that incorporates keypoint constraints.

7. The adversarial facial privacy protection method based on multi-module collaboration according to claim 6, characterized in that, The multi-task joint loss function includes at least regularized cycle consistency loss and key point structure loss. The regularization cycle consistency loss is used to ensure that the generator and the regularization module work together. The keypoint structure loss, through the coordinate matrix of the facial keypoints output by the facial keypoint structure preservation module, forces the generator output to maintain the facial geometry of the source image.

8. The adversarial facial privacy protection method based on multi-module collaboration according to claim 7, characterized in that, The multi-task joint loss function is calculated using the following formula: ; ; in, This indicates the multi-task joint loss function ultimately used in training the adversarial generative framework. , , , , , These represent the adversarial discrimination loss, regularized cycle consistency loss, adversarial attack loss, makeup style loss, identity preservation loss, and keypoint structure loss, respectively. , , , , , denoted as the weights of the loss functions for adversarial discrimination loss, regularized cycle consistency loss, adversarial attack loss, makeup style loss, identity preservation loss, and keypoint structure loss, respectively. This represents the discriminator loss function. This represents the generator loss function.

9. The adversarial facial privacy protection method based on multi-module collaboration according to claim 8, characterized in that, The makeup style loss adopts histogram loss based on facial semantic regions, which is calculated independently on pre-segmented facial regions. By matching the color histogram distribution of adversarial samples and reference makeup images in the corresponding regions, the generator is forced to learn and replicate the color and lighting style of the reference makeup images.