A face prototype reconstruction method and system fusing an anti-attack defense mechanism

By constructing a prototype reconstruction generative adversarial network and an adversarial attack defense layer, the problem of insufficient adversarial attack suppression capability in existing face prototype reconstruction methods is solved, achieving improved adaptability to various facial variations and enhanced security of face recognition.

CN120726704BActive Publication Date: 2025-11-25JIANGXI POLICE COLLEGE +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511207127.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-11-25
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Existing methods for reconstructing facial prototypes are ineffective at suppressing adversarial attacks, resulting in low system security and robustness, and an inability to flexibly handle combinations of various facial variations.

Method used

A prototype reconstruction generative adversarial network is constructed. It is trained by a multi-objective joint loss function of the generator and discriminator, combined with an image denoising network and a matrix estimation layer, to build an adversarial attack defense layer. This enables end-to-end training to generate standard face prototypes without facial changes and suppress adversarial attacks.

Benefits of technology

It significantly improves the adaptability of the face prototype reconstruction model to various facial variations, enhances the accuracy, robustness, and security of face recognition, and can effectively suppress interference from adversarial attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726704B_ABST
    Figure CN120726704B_ABST
Patent Text Reader

Abstract

The application discloses a face prototype reconstruction method and system fusing an anti-attack defense mechanism, relates to the technical field of computer vision, and comprises the following steps: firstly, a prototype reconstruction generative adversarial network containing a generator and a discriminator is constructed, and the two are trained through preset loss functions respectively; subsequently, an anti-attack defense layer is built, including an image denoising network layer and a matrix estimation layer, the former performs preliminary denoising on an input image, and the latter further optimizes to obtain a target denoised image; a defense loss function is built by calculating the difference between the target denoised image and the original image, and the defense layer is trained in an end-to-end mode; finally, the trained generator, discriminator and defense layer are used to process an input face image, and a face prototype reconstruction result is generated. The application can effectively inhibit an adversarial attack, so that the safety and robustness of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision technology, specifically relating to a method and system for reconstructing human face prototypes that integrates adversarial attack defense mechanisms. Background Technology

[0002] With the development of computer vision, facial recognition technology has been widely used in various scenarios such as security monitoring, social media, and criminal investigation. Currently, most facial recognition systems focus on accurately identifying individuals from standard facial images that are free from occlusion, facial expressions, poses, and other facial variations and are well-lit. However, in reality, especially in criminal investigation and security monitoring scenarios, facial images captured by cameras are often contaminated by facial variations such as expressions, poses, occlusion, and lighting conditions. These interfering factors severely hinder facial recognition models from effectively extracting facial identity features, leading to a significant decline in recognition performance. A facial prototype refers to a standard facial image that is free from occlusion, facial expressions, poses, and other facial variations and is well-lit. Therefore, reconstructing a facial prototype from contaminated facial images to extract effective facial identity features is crucial for facial recognition tasks. Furthermore, when carefully designed adversarial perturbations are embedded in the input image, the reconstruction effect of the facial prototype is usually significantly negatively affected, leading to identification errors and potential security risks. Therefore, improving the robustness of facial prototype reconstruction systems against adversarial attacks is a critical challenge that urgently needs to be addressed.

[0003] Existing face prototype reconstruction methods mainly fall into two categories: The first category introduces samples from the query set as auxiliary information into the enrolment database and estimates the face prototypes of contaminated samples in the enrolment database using methods such as Gaussian Mixture Model (GMM) and semi-supervised low-rank representation. Although this type of method can achieve effective prototype reconstruction, it requires joint prototype estimation with an unknown query set, which is difficult to meet the real-time requirements of real-world face recognition systems. The second category utilizes Generative Adversarial Networks (GANs) to learn the complex mapping relationship between contaminated samples and standard prototypes on the generic set, and then reconstructs face prototypes from contaminated samples in the enrolment database. Thanks to the powerful mapping capabilities of generative adversarial networks (GANs), these methods achieve good reconstruction results for single-type facial variations such as lighting, occlusion camouflage, expression, or pose. However, they require prior knowledge of the types of facial variations contained in the input image for accurate modeling and cannot flexibly handle combinations of multiple facial variations, thus limiting their practicality and adaptability. Furthermore, existing face prototyping methods do not consider robustness against adversarial attacks, making their reconstruction results susceptible to the negative impact of adversarial perturbations.

[0004] Therefore, there is an urgent need to provide a solution to improve the above problems. Summary of the Invention

[0005] The purpose of this invention is to provide a face prototype reconstruction method and system that integrates adversarial attack defense mechanisms, which can solve the problem that existing technologies are unable to effectively suppress adversarial attacks, resulting in low system security and robustness.

[0006] To solve the above-mentioned technical problems, the present invention is implemented as follows:

[0007] In a first aspect, embodiments of the present invention provide a method for reconstructing a human face prototype by incorporating adversarial attack defense mechanisms, the method comprising:

[0008] Construct a prototype to reconstruct a generative adversarial network, which includes a generator and a discriminator;

[0009] The generator and discriminator are trained based on preset generator loss functions and preset discriminator loss functions, respectively, to obtain trained generator and trained discriminator.

[0010] An adversarial attack defense layer is constructed, which includes an image denoising network layer and a matrix estimation layer. The image denoising network layer is used to denoise the noise in the input image to obtain an initial denoised image, and the matrix estimation layer is used to denoise the initial denoised image again to obtain the target denoised image.

[0011] Construct an adversarial attack defense loss function based on the difference between the denoised target image and the original image;

[0012] Based on the adversarial attack defense loss function, the adversarial attack defense layer is trained end-to-end to obtain a well-trained adversarial attack defense layer.

[0013] The pre-trained generator, discriminator, and adversarial attack defense layer are used to process the preset input face image to generate a face prototype reconstruction result.

[0014] As an optional implementation of the first aspect of this application, the generator includes an encoder and a decoder. The encoder is used to acquire facial identity features of the input face image. The encoder consists of multiple consecutive convolutional layers and an average pooling layer. Each convolutional layer is followed by a batch normalization layer and an exponential linear unit. The decoder is used to concatenate and decode the facial identity features with a random noise vector to generate a face prototype. The decoder consists of multiple consecutive deconvolutional layers and a fully connected layer. Each deconvolutional layer is followed by a batch normalization layer and an exponential linear unit layer. The fully connected layer is used to expand the dimension of the encoder output features to obtain a dimensionally expanded feature map. The deconvolutional layer is used to upsample the dimensionally expanded feature map to obtain a reconstructed face image.

[0015] As an optional implementation of the first aspect of this application, the discriminator includes a first sub-discriminator, a second sub-discriminator, and a third sub-discriminator. The first sub-discriminator is responsible for predicting the identity of the input image, the second sub-discriminator is responsible for determining whether the input image contains facial changes, and the third sub-discriminator is responsible for evaluating the input image based on the similarity score between the input image and the real prototype to generate a judgment result.

[0016] As an optional implementation of the first aspect of this application, the discriminator further includes multiple consecutive convolutional layers, an average pooling layer, and a fully connected layer. The convolutional layers are used to extract features from the input image to obtain convolutional features. The average pooling layer is used to compress the convolutional features to obtain compressed features. The fully connected layer is used to perform dimensionality transformation on the compressed features to obtain a multidimensional feature vector. The dimension of the multidimensional feature vector is... ,in, The first dimension represents the number of identities contained in the training set, and the second dimension represents the face identity label predicted by the first sub-discriminator. The last two dimensions are used to represent the results of the second sub-discriminator in judging whether the input image contains facial changes and the results of the third sub-discriminator in judging whether the input image is a real prototype.

[0017] As an optional implementation of the first aspect of this application, the specific process of training the generator based on a preset generator loss function to obtain a trained generator includes:

[0018] Define the generator identity preservation loss function, where the identity preservation loss function is... The mathematical expression is:

[0019]

[0020] in, This represents the output of the generator. Indicates the input image. This represents a noise vector randomly sampled from a uniform distribution. This represents the prototype generated by the generator, i.e. , This represents the output of the sub-discriminator responsible for identity prediction. Indicates the input image identity tags Sub-discriminator Predicting the generation of prototypes Identity label The probability, Represents logarithmic operations. Indicates to Maximum likelihood estimation;

[0021] Define the generator facial change loss function, where the generator facial change loss function is... The mathematical expression is:

[0022] ;

[0023] in, This represents the output of the generator. Indicates the input image. This represents a noise vector randomly sampled from a uniform distribution. This represents the prototype generated by the generator, i.e. , This represents the output of the sub-discriminator responsible for determining whether the input image contains facial changes. Indicates the input image Does it include tags indicating facial changes? Sub-discriminator Predicting the generation of prototypes Facial changes tagged as The probability, Represents logarithmic operations. Indicates to Maximum likelihood estimation;

[0024] Define a generator adversarial generation loss function, where the generator adversarial generation loss function is... The mathematical expression is:

[0025] ;

[0026] in, This represents the output of the generator. Indicates the input image. This represents a noise vector randomly sampled from a uniform distribution. This represents the prototype generated by the generator, i.e. , This represents the output of the sub-discriminator responsible for determining whether the input image is the real prototype or the generated prototype. Sub-discriminator A prototype will be generated. The probability of identifying it as the real prototype. Represents logarithmic operations. Indicates to Maximum likelihood estimation;

[0027] Define the generator reconstruction loss function, where the generator reconstruction loss function is... The mathematical expression is:

[0028] ;

[0029] in, This represents the output of the generator. This represents the input real-world prototype image. This represents a noise vector randomly sampled from a uniform distribution. This indicates that the generator produces a fake prototype based on the real prototype. This represents the calculation of the square of the Frobenius norm, where the Frobenius norm is the square root of the sum of the squares of all elements in the input vector. Indicates to Maximum likelihood estimation;

[0030] Define the generator-aware loss function. The mathematical expression is:

[0031] ;

[0032] in, This represents the output of the generator. Indicates the input image. This represents a noise vector randomly sampled from a uniform distribution. This represents the prototype generated by the generator, i.e. , This represents the feature extractor, responsible for extracting features from the real prototype. and generating prototypes Extract features from them, This represents the calculation of the square of the Frobenius norm, where the Frobenius norm is the square root of the sum of the squares of all elements in the input vector. Indicates to Maximum likelihood estimation;

[0033] A multi-objective joint loss function is established based on the generator identity preservation loss function, generator facial change loss function, generator adversarial generation loss function, generator reconstruction loss function, and generator perception loss function;

[0034] The generator is trained using a multi-objective joint loss function, where the mathematical expression for the multi-objective joint loss function is:

[0035]

[0036] in, This represents the adversarial generation loss function of the generator. , , and These are all weight parameters used to balance the loss function. Represents the joint loss function for multiple objectives. This indicates that the generator is optimized during training. The network parameters are used to maximize The corresponding loss function value.

[0037] As an optional implementation of the first aspect of this application, the specific process of training the discriminator based on a preset discriminator loss function to obtain a trained discriminator includes:

[0038] Define the identity prediction loss function for the discriminator, where the mathematical expression for the identity prediction loss function is:

[0039] ;

[0040] in, Indicates the input image. This represents the output of the sub-discriminator responsible for identity prediction. Indicates the input image identity tags Sub-discriminator Predict the input image Identity label The probability, Represents logarithmic operations. Indicates to The maximum likelihood estimate, This represents the identity prediction loss function of the discriminator;

[0041] Define the facial change loss function for the discriminator, where the mathematical expression for the facial change loss function is:

[0042] ;

[0043] in, Indicates the input image. This represents the output of the sub-discriminator responsible for determining whether an image contains facial changes. Indicates the input image Does it include a label indicating facial changes? Sub-discriminator Determine the input image Does the result include facial changes? The probability, Represents logarithmic operations. Indicates to The maximum likelihood estimate, This represents the facial change loss function of the discriminator;

[0044] Define the adversarial loss function for the discriminator, where the mathematical expression for the adversarial loss function is:

[0045] ;

[0046] in, Indicates the input image. This represents the output of the generator. This represents a noise vector randomly sampled from a uniform distribution. This represents the prototype generated by the generator. Represents the real prototype. This represents the output of the sub-discriminator responsible for determining whether an image is the real prototype or the generated prototype. Sub-discriminator A prototype will be generated. The probability of identifying it as the real prototype. Sub-discriminator real prototype The probability of judging it as true. Represents logarithmic operations. Indicates to The maximum likelihood estimate, Indicates to The maximum likelihood estimate, This represents the adversarial loss function of the discriminator;

[0047] A multi-objective joint loss function is established based on the discriminant's identity prediction loss function, facial change loss function, and adversarial loss function, with the following expression:

[0048] ;

[0049] in, This represents the adversarial loss function of the discriminator. This represents the identity prediction loss function of the discriminator. This represents the facial change loss function of the discriminator. and These are all weight parameters used to balance the loss function. This represents the multi-objective joint loss function for the discriminator. This indicates that the discriminator is optimized during training. The network parameters are used to maximize The corresponding loss function value;

[0050] After training the pre-defined discriminator based on the multi-objective joint loss function, a trained discriminator is generated, so that the first sub-discriminator can accurately predict the identity label of the input image, the second sub-discriminator can accurately predict the facial change label of the input image, and the third sub-discriminator can accurately determine whether the input image is a real prototype or a fake prototype generated by the generator.

[0051] As an optional implementation of the first aspect of this application, the image denoising network layer includes a feature enhancement layer, an attention layer, and a residual reconstruction layer.

[0052] The feature enhancement layer comprises multiple convolutional layers and skip connection layers. Each of the first three convolutional layers is followed by a batch normalization layer and a rectified linear unit. Skip connections concatenate the input image of the network with the output of the fourth convolutional layer to obtain concatenated features. These concatenated features are then processed using the Tanh activation function to obtain the output of the feature enhancement layer, expressed as:

[0053] ;

[0054] in, Indicates the input image. This indicates the first three convolutional layers in the feature enhancement layer, each followed by a batch normalization (BN) layer and a rectified linear unit (ReLU). Indicates a convolutional layer; Indicates features and characteristics splicing; This represents the Tanh activation function, which aims to more effectively propagate training gradients and make the training process more stable. This represents the output of the feature enhancement layer FEM;

[0055] The attention layer is used to filter noise in the output of the feature enhancement layer to obtain enhanced noise, and its expression is:

[0056] ;

[0057] in, This represents the output features of the feature enhancement layer (FEM). This represents the output feature of the fourth convolutional layer in the Feature Enhancement Model (FEM). Indicates 1 1 convolutional layer, This indicates an element-wise multiplication operation. This represents the output of the attention layer AM, i.e., the predicted noise;

[0058] The residual reconstruction layer is used to connect the input image of the network with the noise predicted by the attention layer AM through skip connections, thereby achieving element-wise subtraction between the input image and the predicted noise to reconstruct the denoised image. The expression is:

[0059] ;

[0060] in, This represents the input image of the denoising network. This represents the noise predicted by the attention layer AM. This indicates an element-by-element subtraction operation. This represents the output of the residual reconstruction layer, i.e., the denoised image output by the image denoising network.

[0061] As an optional implementation of the first aspect of this application, the matrix estimation layer is used to perform denoising processing on the initial denoised image again to obtain the target denoised image. The specific process includes:

[0062] The input noisy image is normalized to obtain a normalized image matrix, where the mathematical expression for the normalized image matrix is:

[0063] ;

[0064] in, This represents the input noisy image. This represents the normalized image matrix to be decomposed. Through the normalization operation, the numerical range of the image matrix is ​​changed from... It became ;

[0065] The normalized image matrix is ​​subjected to singular value decomposition and singular value thresholding sequentially to obtain the thresholded singular values. The mathematical expression for singular value decomposition is as follows:

[0066] ;

[0067] ;

[0068] in, This represents the normalized image matrix to be decomposed. Describes a left singular vector matrix. Describes a right singular vector matrix. This represents the transpose of a right singular vector matrix. This represents a diagonal matrix, whose diagonal elements... These are called singular values, and they satisfy... ;

[0069] The mathematical expression for singular value thresholding is:

[0070]

[0071] in, This indicates the number of singular values ​​before thresholding. This indicates the proportion of singular values ​​retained. This represents the number of singular values ​​retained after thresholding, i.e., the number of singular values ​​after thresholding. ;

[0072] The normalized image matrix is ​​reconstructed based on the thresholded singular values ​​to generate the target denoised image, as expressed by:

[0073] ;

[0074] in, Let represent a diagonal matrix composed of the retained singular values. Describes a left singular vector matrix. Describes a right singular vector matrix. This represents the transpose of a right singular vector matrix. This represents the output of the matrix estimation, i.e., the denoised image.

[0075] As an optional implementation of the first aspect of this application, the mathematical expression of the adversarial attack defense loss function is:

[0076] ;

[0077] in, This represents the training set samples without added noise. This represents the training set samples after noise has been added. This represents the total number of samples in the training set. This indicates the output of the defense layer against attacks. Let represent the square of the L2 norm, where the L2 norm is the square root of the sum of the squares of the elements of the input vector. This represents the loss function for defense against attacks.

[0078] Secondly, embodiments of the present invention provide a face prototype reconstruction system that integrates adversarial attack defense mechanisms, the system comprising:

[0079] The first building module is used to construct a prototype reconstructed generative adversarial network, which includes a generator and a discriminator.

[0080] The first training module is used to train the generator and the discriminator based on the preset generator loss function and the preset discriminator loss function, respectively, so as to obtain the trained generator and the trained discriminator.

[0081] The second building module is used to build an adversarial attack defense layer. The adversarial attack defense layer includes an image denoising network layer and a matrix estimation layer. The image denoising network layer is used to denoise the noise in the input image to obtain an initial denoised image. The matrix estimation layer is used to denoise the initial denoised image again to obtain the target denoised image.

[0082] The defense loss function construction module is used to construct an adversarial attack defense loss function based on the difference between the target denoised image and the original image;

[0083] The second training module is used to perform end-to-end training of the adversarial attack defense layer based on the adversarial attack defense loss function to obtain the trained adversarial attack defense layer.

[0084] The face prototype reconstruction module is used to process a preset input face image through a trained generator, a trained discriminator, and a trained adversarial attack defense layer to generate a face prototype reconstruction result.

[0085] Compared with existing technologies, the beneficial effects of the face prototype reconstruction method that integrates adversarial attack defense mechanisms proposed in this invention are as follows:

[0086] First, the face prototype reconstruction network of this invention consists of a generator with an encoder-decoder structure and a multi-task discriminator. It trains the generator and discriminator to engage in adversarial game through an original multi-objective joint loss function, enabling the network to collaboratively learn face prototypes and identity representations within a unified framework. This generates standardized face prototypes that retain identity features, avoiding the limitations of traditional methods that can only model a single type of facial change and require prior knowledge of the specific type of facial change. This significantly improves the adaptability of the prototype reconstruction model to various combinations of facial changes, making the model more practical and applicable to a wider range of scenarios.

[0087] Secondly, the prototype reconstruction generative adversarial network constructed in this invention collaboratively learns face prototypes and identity representations within a unified framework, directly generating standardized face prototypes that do not contain facial changes such as pose, expression, occlusion, or disguise and are adequately lit. These high-quality face prototypes can be used in practical scenarios such as security identification and criminal investigation. At the same time, the identity prototype features learned by the network can be directly used for face recognition tasks, enabling accurate identification of contaminated face images and helping to solve the negative impact of facial change interference on face recognition performance.

[0088] Third, this invention innovatively integrates an adversarial attack defense mechanism into the prototype reconstruction generative adversarial network. It employs an original design of a defense module combining an image denoising network and matrix estimation. The image denoising network predicts noise in the image through attention mechanisms and residual learning to achieve denoising, while matrix estimation further eliminates noise and restores the global structure of the image through a general singular value thresholding method. During the training phase, this invention uses perturbation-enhanced training samples and an original adversarial attack defense loss to train the defense module end-to-end. This effectively suppresses the interference of adversarial perturbations on the prototype reconstruction network, reduces the threat of adversarial attacks to subsequent face recognition tasks, significantly improves the system's security and robustness, and overcomes the deficiency of existing prototype reconstruction methods that do not consider adversarial attack defense.

[0089] In summary, this invention achieves prototype reconstruction and identity prototype feature learning of face images contaminated by various facial variations through a unified framework, and effectively suppresses the interference of adversarial attacks on the prototype reconstruction network by integrating adversarial defense mechanisms. It significantly improves the accuracy, robustness, and security of face prototype reconstruction and face recognition, and has good practical value and application prospects. Attached Figure Description

[0090] Figure 1 This is a flowchart illustrating a face prototype reconstruction method that integrates adversarial attack defense mechanisms, as provided in the first embodiment of the present invention.

[0091] Figure 2 This diagram illustrates the prototype reconstruction generative adversarial network structure provided in the first embodiment of the present invention.

[0092] Figure 3 This is a schematic diagram illustrating the principle of the training generator provided in the first embodiment of the present invention;

[0093] Figure 4 This is a schematic diagram illustrating the principle of the training discriminator provided in the first embodiment of the present invention;

[0094] Figure 5 This is a schematic diagram of the image denoising network structure provided in the first embodiment of the present invention;

[0095] Figure 6 This is a schematic diagram illustrating the principle of the matrix estimation layer provided in the first embodiment of the present invention;

[0096] Figure 7 This is a schematic diagram illustrating the principle of the training adversarial attack defense layer provided in the first embodiment of the present invention;

[0097] Figure 8 This image shows the face prototype reconstruction result of the face prototype reconstruction network provided by the first embodiment of the present invention.

[0098] Figure 9 This is a diagram of a face prototype reconstruction system provided by the second embodiment of the present invention, which integrates an adversarial attack defense mechanism. Detailed Implementation

[0099] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0100] The terms "first," "second," etc., used in the specification and claims of this invention are used to distinguish similar objects and are not used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0101] The following description, in conjunction with the accompanying drawings, details a face prototype reconstruction method and system that integrates adversarial attack defense mechanisms, provided by the present invention, through specific embodiments and application scenarios.

[0102] Example 1

[0103] It is worth noting that this invention proposes a face prototype reconstruction method that integrates adversarial attack defense mechanisms, mainly consisting of four parts: The first part constructs an original prototype reconstruction generative adversarial network, which consists of a generator with an encoder-decoder structure and a multi-task discriminator. The discriminator aims to predict the identity of the input face image, detect whether it contains facial variations, and determine the authenticity of the prototype. The generator aims to generate fake prototypes that retain identity information to deceive the discriminator. The second part uses an original multi-objective joint loss function to collaboratively and alternately train the generator and discriminator in the prototype reconstruction generative adversarial network. The generator and discriminator are collaboratively optimized through an adversarial game process, enabling the generator to learn facial representations while simultaneously... The first part preserves identity information, enabling the trained network to reconstruct a standard human face prototype without occlusion, facial expressions, poses, or other facial variations, and with sufficient lighting, for face recognition tasks. The second part constructs an original adversarial attack defense layer and integrates it with the prototype reconstruction generative adversarial network constructed in the first part. This defense layer preprocesses adversarial attack samples by combining image denoising networks and matrix estimation methods, thereby suppressing the negative impact of adversarial attacks on prototype reconstruction results and improving the system's security and robustness. The third part uses an original adversarial attack defense loss to train the adversarial attack defense layer end-to-end based on perturbation-enhanced training samples, enabling it to effectively suppress the interference of adversarial perturbations on the prototype reconstruction network.

[0104] Specifically, please refer to Figure 1 This is a flowchart of a face prototype reconstruction method that integrates adversarial attack defense mechanisms provided by the present invention, the method including steps S1 to S6.

[0105] S1: Construct a prototype to reconstruct a generative adversarial network, which includes a generator and a discriminator.

[0106] For details, please refer to Figure 2 The prototype reconstruction generative adversarial network constructed in this invention is the prototype reconstruction generative adversarial network in DF-PRGAN, which integrates adversarial attack defense mechanisms. This prototype reconstruction generative adversarial network consists of a generator... and a discriminator constitute.

[0107] Please see Figure 3 The generator of the present invention The network consists of an encoder and a decoder. The encoder is responsible for extracting facial identity features from the input face image. It consists of nine consecutive convolutional layers and one average pooling layer. Each convolutional layer is followed by a batch normalization (BN) layer and an exponential linear unit (ELU). The consecutive convolutional layers are designed to encode features from shallow to deep layers in the input image. The average pooling layer is responsible for downsampling the feature map after convolution to reduce the size of the feature map and reduce the computational cost, while retaining the main features and making the features more robust. The batch normalization layer is designed to perform standardization processing on the batch data input to the network layer, making the training gradient more stable. The exponential linear unit is a piecewise non-linear activation function that can accelerate the convergence of network parameters. The process of the encoder extracting facial identity features from the input face image is shown in Equation (1).

[0108] (1)

[0109] In formula (1), This indicates that the input is a face image. Indicates encoder, This represents the identity features extracted by the encoder from the input image, which are... A dimensional vector, where Represents identity feature vector Dimensions.

[0110] The decoder is responsible for concatenating the identity features extracted by the encoder with a noise vector randomly sampled from a uniform distribution, and then decoding the concatenated vector to generate a face prototype. The decoder consists of a fully connected layer and nine consecutive deconvolutional layers. Each deconvolutional layer is followed by a batch normalization layer and an exponential linear unit. The fully connected layer is responsible for expanding the dimension of the encoder output features to prepare for image reconstruction by the subsequent deconvolutional layers. The consecutive deconvolutional layers are responsible for gradually restoring the spatial dimension of the feature map through upsampling until the face image is reconstructed. The process of the decoder concatenating features and decoding the concatenated vector to generate a face prototype is shown in formula (2):

[0111] (2)

[0112] In formula (2), This represents the generated human face prototype. Indicates decoder, Represents the noise vector. Represents the noise vector Dimensions Representing an interval Uniform distribution on Represents the noise vector From the interval Random sampling from a uniform distribution Dimensional vector.

[0113] Please see Figure 4 Discriminator Includes , and Three sub-discriminators, among which the sub-discriminators Sub-discriminator is responsible for predicting the identity of the input image. The sub-discriminator is responsible for determining whether the input image contains facial variations. Responsible for determining whether the input image is a real human face prototype or a generator. The generated fake prototype. Specifically, the sub-discriminator. Evaluation is based on the similarity score between the input image and the real prototype. A higher similarity score indicates a greater similarity between the input image and the real prototype. The higher the probability that the input image is a real human face prototype, the better. From the perspective of network structure, the discriminator... The structure consists of nine consecutive convolutional layers, one average pooling layer, and one fully connected layer. The convolutional and average pooling layers are responsible for feature extraction and compression of the input image, while the fully connected layer transforms the features processed by the convolutional and average pooling layers, outputting the final feature. dimensional feature vectors, where This represents the number of identities contained in the training set. (This is the output of the fully connected layer.) In the dimensional feature vector, the first The vector dimension is used to represent the sub-discriminator. The predicted face identity label, with the latter two vector dimensions used to represent the sub-discriminator. The result of determining whether the input image contains facial variations is represented by the last vector dimension, which indicates the sub-discriminator. The result of determining whether the input image is the real prototype.

[0114] Step S2: Train the generator and discriminator respectively based on the preset generator loss function and the preset discriminator loss function to obtain the trained generator and the trained discriminator.

[0115] Specifically, in step S2, the present invention trains the generator and discriminator in the prototype reconstruction generative adversarial network constructed in step S1 using two different original multi-objective joint loss functions. Given a training set of face images, the number of identities contained therein is represented as... The label for each image is represented as ,in Indicates identity label, To indicate whether an image contains facial variations, the standard facial images in the training set (i.e., prototype faces without facial variations) are represented as follows: .

[0116] For the generator, this invention sets the following five training objectives:

[0117] 1) Train the generator to make the sub-discriminator For generating prototypes Identity prediction results and input image The identity labels are the same. Correspondingly, this invention designs an identity preservation loss function for the generator, the calculation method of which can be expressed by formula (3):

[0118] (3)

[0119] In formula (3), This represents the output of the generator. Indicates the input image. This represents a noise vector randomly sampled from a uniform distribution. This represents the prototype generated by the generator, i.e. , This represents the output of the sub-discriminator responsible for identity prediction. Indicates the input image identity tags Sub-discriminator Predicting the generation of prototypes Identity label The probability, Represents logarithmic operations. Indicates to The maximum likelihood estimate, This represents the identity preservation loss function of the generator.

[0120] 2) Train the generator to enable the sub-discriminator In generating prototypes No facial changes were detected. Correspondingly, this invention designs a facial change loss function for the generator, the calculation method of which can be expressed by formula (4):

[0121] (4)

[0122] In formula (4), This represents the output of the generator. Indicates the input image. This represents a noise vector randomly sampled from a uniform distribution. This represents the prototype generated by the generator, i.e. , This represents the output of the sub-discriminator responsible for determining whether the input image contains facial changes. Indicates the input image Does it include tags indicating facial changes? Sub-discriminator Predicting the generation of prototypes Facial changes tagged as The probability, Represents logarithmic operations. Indicates to The maximum likelihood estimate, This represents the facial variation loss function of the generator.

[0123] 3) Train the generator to detect cheaters. This will enable the generation of a prototype. It is determined to be a true prototype. Correspondingly, this invention designs an adversarial generation loss function for the generator, the calculation method of which can be expressed by formula (5):

[0124] (5)

[0125] In formula (5), This represents the output of the generator. Indicates the input image. This represents a noise vector randomly sampled from a uniform distribution. This represents the prototype generated by the generator, i.e. , This represents the output of the sub-discriminator responsible for determining whether the input image is the real prototype or the generated prototype. Sub-discriminator A prototype will be generated. The probability of identifying it as the real prototype. Represents logarithmic operations. Indicates to The maximum likelihood estimate, This represents the adversarial generation loss function of the generator.

[0126] 4) Train the generator so that it can process each input image that does not contain facial variations. The corresponding generated prototype should be as consistent as possible with the input. Correspondingly, this invention designs a reconstruction loss function for the generator, the calculation method of which can be expressed by formula (6):

[0127] (6)

[0128] In formula (6), This represents the output of the generator. This represents the input real-world prototype image. This represents a noise vector randomly sampled from a uniform distribution. This indicates that the generator produces a fake prototype based on the real prototype. This represents the calculation of the square of the Frobenius norm, where the Frobenius norm is the square root of the sum of the squares of all elements in the input vector. Indicates to The maximum likelihood estimate, This represents the reconstruction loss function of the generator.

[0129] 5) Train the generator to produce prototypes. With the real prototype The features should be as similar as possible in the feature space. Correspondingly, this invention designs a perceptual loss function for the generator, the calculation method of which can be expressed by formula (7):

[0130] (7)

[0131] In formula (7), This represents the output of the generator. Indicates the input image. This represents a noise vector randomly sampled from a uniform distribution. This represents the prototype generated by the generator, i.e. , This represents the feature extractor, responsible for extracting features from the real prototype. and generating prototypes To extract features, this invention uses a ResNet-50 network pre-trained on the MS1M face dataset for face recognition as the feature extractor. , This represents the calculation of the square of the Frobenius norm, where the Frobenius norm is the square root of the sum of the squares of all elements in the input vector. Indicates to The maximum likelihood estimate, This represents the perceptual loss function of the generator.

[0132] Based on the above five loss functions for generators, this invention designs an original multi-objective joint loss function for generators to train prototypes to reconstruct generators in generative adversarial networks to achieve the above five objectives. The calculation method of this multi-objective joint loss function for generators can be expressed by formula (8):

[0133] (8)

[0134] In formula (8), This represents the adversarial generation loss function of the generator. This represents the identity preservation loss function of the generator. This represents the facial variation loss function of the generator. This represents the reconstruction loss function of the generator. This represents the perceptual loss function of the generator. , , and These are all weight parameters used to balance the loss function. This represents the multi-objective joint loss function for the generator. This indicates that the generator is optimized during training. The network parameters are used to maximize The corresponding loss function value.

[0135] Furthermore, regarding the discriminator This invention sets the following three training objectives:

[0136] 1) Training the first sub-discriminator This enables it to accurately predict the input image. Identity tags Correspondingly, this invention designs an identity prediction loss function for the discriminator, the calculation method of which can be expressed by formula (9):

[0137] (9)

[0138] In formula (9), Indicates the input image. This represents the output of the sub-discriminator responsible for identity prediction. Indicates the input image identity tags Sub-discriminator Predict the input image Identity label The probability, Represents logarithmic operations. Indicates to The maximum likelihood estimate, This represents the identity prediction loss function of the discriminator.

[0139] 2) Training the second sub-discriminator This enables it to accurately judge the input image. Whether it contains facial variations, i.e., whether it can accurately predict the input image. Facial change tags Correspondingly, this invention designs a facial change loss function for the discriminator, the calculation method of which can be expressed by formula (10):

[0140] (10)

[0141] In formula (10), Indicates the input image. This represents the output of the sub-discriminator responsible for determining whether an image contains facial changes. Indicates the input image Does it include a label indicating facial changes? Sub-discriminator Determine the input image Does the result include facial changes? The probability, Represents logarithmic operations. Indicates to The maximum likelihood estimate, This represents the facial change loss function of the discriminator.

[0142] 3) Training the third sub-discriminator This enables it to accurately determine whether the input image is real or generated, that is, to accurately determine whether the input image is a real prototype or a fake prototype generated by the generator. Correspondingly, this invention designs an adversarial loss function for the discriminator, the calculation method of which can be expressed by formula (11):

[0143] (11)

[0144] In formula (11), Indicates the input image. This represents the output of the generator. This represents a noise vector randomly sampled from a uniform distribution. This represents the prototype generated by the generator. Represents the real prototype. This represents the output of the sub-discriminator responsible for determining whether an image is the real prototype or the generated prototype. Sub-discriminator A prototype will be generated. The probability of identifying it as the real prototype. Sub-discriminator real prototype The probability of judging it as true. Represents logarithmic operations. Indicates to The maximum likelihood estimate, Indicates to The maximum likelihood estimate, This represents the adversarial loss function of the discriminator.

[0145] Based on the above three loss functions for the discriminator, this invention designs an original multi-objective joint loss function for the discriminator to train the prototype to reconstruct the discriminator in the generative adversarial network and achieve the above three objectives. The calculation method of the multi-objective joint loss function can be expressed by formula (12):

[0146] (12)

[0147] In formula (12), This represents the adversarial loss function of the discriminator. This represents the identity prediction loss function of the discriminator. This represents the facial change loss function of the discriminator. and These are all weight parameters used to balance the loss function. This represents the multi-objective joint loss function for the discriminator. This indicates that the discriminator is optimized during training. The network parameters are used to maximize The corresponding loss function value.

[0148] During training, the generator and discriminator Adversarial training is performed by alternately optimizing the loss functions in equations (8) and (12). As the discriminator... The generator enhances capabilities in identity recognition, facial change detection, and prototyping. They also work to generate realistic facial prototypes that retain identity information and do not contain facial alterations, in order to deceive the discriminator. With the sub-discriminator With the gradual improvement of identity recognition capabilities, Will guide generator encoder in They learn more distinctive facial features; at the same time... and Mutual cooperation enables the generator Generate a human face prototype that does not contain facial changes, thereby decoupling from facial changes.

[0149] S3: Build an adversarial attack defense layer, which includes an image denoising network layer and a matrix estimation layer. The image denoising network layer is used to denoise the noise in the input image to obtain an initial denoised image. The matrix estimation layer is used to denoise the initial denoised image again to obtain the target denoised image.

[0150] For details, please refer to Figure 7This diagram illustrates the principle of training the adversarial attack defense layer. During step S3, when constructing the attack defense layer, this invention builds the adversarial attack defense layer within the Prototype Reconstruction Generative Adversarial Network (DF-PRGAN), which integrates adversarial attack defense mechanisms. By combining an image denoising network and matrix estimation, adversarial attack samples are preprocessed, thereby suppressing the negative impact of adversarial attacks on the prototype reconstruction effect and improving the system's security and robustness. This adversarial attack defense layer mainly consists of two parts: an image denoising network and a matrix estimation layer.

[0151] Please refer to Figure 5 This represents a schematic diagram of an image denoising network structure. The image denoising network denoises the input image by predicting the noise contained in the input image and combining it with residual learning. It mainly consists of three parts: a feature enhancement layer, an attention layer, and a residual reconstruction layer.

[0152] The Feature Enhancement Layer (FEM) mainly consists of four convolutional layers and skip connections. The first three convolutional layers are followed by batch normalization layers and rectified linear units (RLUs). The skip connections concatenate the input image of the network with the output of the fourth convolutional layer. The concatenated features are then processed by the Tanh activation function to obtain the output of the Feature Enhancement Layer. The feature processing procedure of the Feature Enhancement Layer (FEM) is shown in formula (13):

[0153] (13)

[0154] In formula (13), Indicates the input image. This indicates the first three convolutional layers in the feature enhancement layer, each followed by a batch normalization (BN) layer and a rectified linear unit (ReLU). Indicates a convolutional layer; Indicates feature splicing processing; This represents the Tanh activation function, which aims to more effectively propagate training gradients and make the training process more stable. This represents the output of the feature enhancement layer FEM.

[0155] Attention layer AM mainly consists of a 1 It consists of 1 convolutional layer. In the attention layer, it first consists of 1... The convolutional layer compresses the input features to obtain the attention weights corresponding to the input features. Then, the attention weights are multiplied element-wise with the input features to enhance the noise-related parts of the features and suppress the noise-independent parts. The feature processing of the attention layer AM is shown in formula (14):

[0156] (14)

[0157] In formula (14), This represents the output features of the feature enhancement layer (FEM). This represents the output feature of the fourth convolutional layer in the Feature Enhancement Model (FEM). Indicates 1 1 convolutional layer, This indicates an element-wise multiplication operation. This represents the output of the attention layer AM, i.e., the predicted noise.

[0158] The Residual Reconstruction Layer (RRM) is responsible for connecting the input image of the network with the noise predicted by the Attention Layer (AM) through skip connections, thereby achieving element-wise subtraction between the input image and the predicted noise to reconstruct the denoised image. The feature processing of the Residual Reconstruction Layer (RRM) is shown in Equation (15):

[0159] (15)

[0160] In formula (15), This represents the input image of the denoising network. This represents the noise predicted by the attention layer AM. This indicates an element-by-element subtraction operation. This represents the output of the residual reconstruction layer, i.e., the denoised image output by the image denoising network.

[0161] Please see Figure 6 This diagram illustrates the principle of the matrix estimation layer. The matrix estimation layer of this invention aims to further denoise the image output by the denoising network and restore the global structural characteristics of the image. The matrix estimation layer in this invention is mainly implemented through a general singular value thresholding method. Its core is to identify and retain significant signal components by singular value decomposition, while eliminating noise-related singular values, thereby achieving image denoising. The core steps of the matrix estimation layer of this invention consist of four steps: image normalization, singular value decomposition, singular value thresholding, and matrix reconstruction. Specifically, these include:

[0162] The first step is image normalization, which normalizes the input noisy image. This process can be represented by formula (16):

[0163] (16)

[0164] In formula (16), This represents the input noisy image. This represents the normalized image matrix to be decomposed. Through normalization, the numerical range of the image matrix is ​​expanded from... It became .

[0165] The second step is singular value decomposition, which involves performing singular value decomposition (SVD) on the normalized matrix. This process can be represented by formula (17):

[0166] (17)

[0167] In formula (17), This represents the normalized image matrix to be decomposed. Describes a left singular vector matrix. Describes a right singular vector matrix. This represents the transpose of a right singular vector matrix. Represents the transpose of a matrix. This represents a diagonal matrix, whose diagonal elements... These are called singular values, and they satisfy... .

[0168] The third step is singular value thresholding, which retains only a specific number of singular values, thereby preserving significant signal components and eliminating noise-related singular values. This process can be expressed by formula (18):

[0169] (18)

[0170] In formula (18), This indicates the number of singular values ​​before thresholding. This indicates the proportion of singular values ​​retained. This represents the number of singular values ​​retained after thresholding, i.e., the number of singular values ​​after thresholding. .

[0171] The fourth step is matrix reconstruction, which uses the thresholded singular values ​​to reconstruct the image matrix, resulting in a denoised image. This process can be represented by formula (19):

[0172] (19)

[0173] In formula (19), This represents the output of matrix estimation, i.e., the denoised image. Let represent a diagonal matrix composed of the retained singular values. Describes a left singular vector matrix. Describes a right singular vector matrix. This represents the transpose of a matrix.

[0174] S4: Construct an adversarial attack defense loss function based on the difference between the target denoised image and the original image.

[0175] Specifically, during step S4, this invention trains the adversarial attack defense layer from step S3 by adding Gaussian noise to the face images in the training set to obtain noisy images. The adversarial attack defense layer constructed in step S3 is then used to denoise the noisy images, and the adversarial attack defense loss function is calculated based on the difference between the original image and the denoised image. This adversarial attack defense loss function... The calculation method can be expressed by formula (20):

[0176] (20)

[0177] In formula (20), This represents the training set samples without added noise. This represents the training set samples after noise has been added. This represents the total number of samples in the training set. This indicates the output of the defense layer against attacks. Let represent the square of the L2 norm, where the L2 norm is the square root of the sum of the squares of the elements of the input vector.

[0178] In summary, during step S4, this invention enhances the original training data by perturbation and uses an original adversarial attack defense loss function to train the adversarial attack defense layer constructed in step S3 end-to-end. This enables the adversarial attack defense layer to learn to reconstruct clean original face images from perturbed input samples, thereby possessing the ability to defend against adversarial perturbations and effectively reducing the negative impact of adversarial attacks on face prototype reconstruction.

[0179] S5: Based on the adversarial attack defense loss function, perform end-to-end training on the adversarial attack defense layer to obtain a well-trained adversarial attack defense layer.

[0180] S6: The pre-trained generator, discriminator, and adversarial attack defense layer are used to process the preset input face image to generate the face prototype reconstruction result.

[0181] Specifically, the input face image is sequentially passed through a trained generator, a trained discriminator, and a trained adversarial attack defense layer to generate a face prototype reconstruction result.

[0182] To verify the feasibility of the proposed face prototype reconstruction method that integrates adversarial attack defense mechanism, simulation experiments were also conducted, including face prototype reconstruction experiments and face recognition and reconstruction prototype quality evaluation experiments.

[0183] In conducting the face prototype reconstruction experiment, this invention selected the public LAMP-HQ dataset and the public Multi-PIE dataset as experimental datasets. Specifically, for the LAMP-HQ dataset, this invention selected facial images of 180 individuals under visible light conditions, including variations in expression, pose, lighting, occlusion, and background. Of these, 140 individuals were used for training, and 40 were used for testing. For the Multi-PIE dataset, this invention selected facial images of 152 individuals, including variations in expression, pose, and lighting. Of these, 120 individuals were used for training, and 32 were used for testing.

[0184] The face prototype reconstruction network designed in this invention incorporates an adversarial attack defense mechanism. During training, the Adam optimizer is used, with 100 training epochs and a training learning rate of [value missing]. The identity prototype feature dimension extracted by the generator in the network (i.e., in formula (1)). The dimension of the noise vector spliced ​​in the network (i.e., the dimension in formula (2)) is set to 320; The weight parameters of the joint loss function of the generator in formula (8) are set to 50. , , and The values ​​are set to 5.0, 0.5, 0.1, and 0.2 respectively; for the weight parameters of the joint loss function of the discriminator in formula (12) and They were set to 5.0 and 0.5 respectively.

[0185] To test the effectiveness of this invention in reconstructing human face prototypes, this invention selected face images containing one or more facial variations for prototype reconstruction experiments. To test the defensive effect of this invention against adversarial attacks, this invention used three currently popular adversarial attack methods—FGSM, PGD, and TRM-UAP—to add adversarial perturbations to the test samples, with perturbation strengths set to three levels: 0.10, 0.15, and 0.20.

[0186] The face prototype reconstruction result of the face prototype reconstruction network with the fusion adversarial attack defense mechanism designed in this invention is as follows: Figure 8 As shown, Figure 8 The image, from left to right, shows the real prototype, the contaminated sample, the reconstructed prototype of this invention, and the reconstructed prototype of this invention without incorporating adversarial attack defense mechanisms. Figure 8As can be seen, the present invention can effectively generate standard face prototypes that do not contain facial changes and retain identity information. However, without the adversarial attack defense mechanism, the face prototypes reconstructed by the present invention show a decrease in both identity preservation and image quality. This demonstrates the effectiveness of the adversarial attack defense mechanism in the present invention. This defense mechanism helps to mitigate the negative impact of adversarial attacks on prototype reconstruction and improves the robustness of the system.

[0187] Furthermore, in the process of face recognition and prototype reconstruction quality evaluation experiments, this invention employs the Modified Laplacian Variance (MLV) method to quantitatively evaluate the image sharpness of the reconstructed face prototype; a higher MLV value indicates a sharper image. Additionally, this invention utilizes the identity prototype features learned by the prototype reconstruction network in this invention for face recognition, using the Rank-10 recognition rate to quantitatively evaluate face recognition performance. Tables 1 and 2 show the experimental results of this invention on the LAMP-HQ dataset and the Multi-PIE dataset, respectively.

[0188] Table 1. Quality evaluation results of face recognition and reconstruction prototypes on the LAMP-HQ dataset.

[0189]

[0190] Table 2. Quality evaluation results of face recognition and reconstruction prototypes on the Multi-PIE dataset.

[0191]

[0192] As can be seen from Tables 1 and 2, the present invention can achieve a high Rank-10 face recognition rate without adding adversarial perturbations. This indicates that the prototype reconstruction network of the present invention can directly use the identity prototype features learned by it to perform face recognition tasks on contaminated images while reconstructing the face prototype. As the intensity of adversarial attacks increases, both face recognition rate and image clarity show a decreasing trend. It is noteworthy that, compared to the case without adversarial attack defense mechanisms, the face recognition rate and image clarity of this invention are both higher, indicating that the performance degradation caused by adversarial perturbations is smaller. For example, under an FGSM attack with a perturbation intensity of 0.15 (Table 1), the Rank-10 recognition rate of this invention on the LAMP-HQ dataset is improved by 10.27% and the MLV is improved by 2.84 compared to the case without adversarial attack defense mechanisms. Under a PGD attack with a perturbation intensity of 0.20 (Table 2), the Rank-10 recognition rate of this invention on the Multi-PIE dataset is improved by 13.23% and the MLV is improved by 4.39 compared to the case without adversarial attack defense mechanisms. This demonstrates that the adversarial attack defense mechanism integrated in this invention can effectively reduce the negative impact of adversarial attacks on prototype reconstruction quality and face recognition performance, improving the system's security and robustness.

[0193] Example 2

[0194] Please see Figure 9 This invention provides a face prototype reconstruction system that integrates adversarial attack defense mechanisms. The system includes:

[0195] The first building module 100 is used to build a prototype reconstruction generative adversarial network, which includes a generator and a discriminator.

[0196] The first training module 200 is used to train the generator and the discriminator based on the preset generator loss function and the preset discriminator loss function respectively, so as to obtain the trained generator and the trained discriminator.

[0197] The second building module 300 is used to build an adversarial attack defense layer. The adversarial attack defense layer includes an image denoising network layer and a matrix estimation layer. The image denoising network layer is used to denoise the noise in the input image to obtain an initial denoised image. The matrix estimation layer is used to denoise the initial denoised image again to obtain a target denoised image.

[0198] Defense loss function construction module 400 is used to construct an adversarial attack defense loss function based on the difference between the target denoised image and the original image;

[0199] The second training module 500 is used to perform end-to-end training of the adversarial attack defense layer based on the adversarial attack defense loss function to obtain the trained adversarial attack defense layer.

[0200] The face prototype reconstruction module 600 is used to process a preset input face image through a trained generator, a trained discriminator, and a trained adversarial attack defense layer to generate a face prototype reconstruction result.

[0201] The beneficial effects of the face prototype reconstruction system that integrates adversarial attack defense mechanisms provided by this invention are as follows:

[0202] First, this invention uses a first construction module 100 to build a prototype reconstruction generative adversarial network (GAN). It trains the generator and discriminator to engage in adversarial game through an original multi-objective joint loss function, enabling the network to collaboratively learn facial prototypes and identity representations within a unified framework. This generates standardized facial prototypes that retain identity features, avoiding the limitations of traditional methods that can only model single facial variations and require prior knowledge of specific facial variation types. This significantly improves the prototype reconstruction model's adaptability to various combinations of facial variations, making the model more practical and applicable to a wider range of scenarios. Furthermore, the prototype reconstruction GAN, through collaborative learning of facial prototypes and identity representations within a unified framework, directly generates standardized facial prototypes that do not contain facial variations such as pose, expression, or occlusion, and are adequately lit. These high-quality facial prototypes can be used in practical scenarios such as security identification and criminal investigation. Simultaneously, the identity prototype features learned by the network can be directly used for facial recognition tasks, achieving accurate identification of contaminated facial images and helping to address the negative impact of facial variation interference on facial recognition performance.

[0203] Secondly, this invention constructs an adversarial attack defense layer through the second building module 300. The image denoising network predicts noise in the image through an attention mechanism and residual learning to achieve denoising. Matrix estimation further eliminates noise in the image and restores the global structure of the image through a general singular value thresholding method. During the training phase, this invention uses perturbation-enhanced training samples and an original adversarial attack defense loss to train the defense module end-to-end. This effectively suppresses the interference of adversarial perturbations on the prototype reconstruction network, reduces the threat of adversarial attacks to subsequent face recognition tasks, significantly improves the system's security and robustness, and solves the deficiency of existing prototype reconstruction methods that do not consider adversarial attack defense.

[0204] The face prototype reconstruction system integrating adversarial attack defense mechanisms in this embodiment of the invention can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can refer to mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can refer to servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This embodiment of the invention does not impose specific limitations.

[0205] The face prototype reconstruction system that integrates adversarial attack defense mechanisms in this embodiment of the invention can represent a device with an operating system. This operating system can represent Android, iOS, or other possible operating systems; this embodiment of the invention does not specifically limit the definition.

[0206] The face prototype reconstruction system provided in this embodiment of the invention can achieve... Figures 1 to 8 The various processes of a face prototype reconstruction method that integrates adversarial attack defense mechanisms are described in the method embodiments below. To avoid repetition, they will not be repeated here.

[0207] Optionally, embodiments of the present invention also provide an electronic device, including a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the various processes of the above-described embodiment of a face prototype reconstruction method with a fusion adversarial attack defense mechanism and achieve the same technical effect. To avoid repetition, further details are omitted here.

[0208] This invention also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described embodiment of a face prototype reconstruction method with a fusion adversarial attack defense mechanism, and achieve the same technical effect. To avoid repetition, this will not be elaborated further here.

[0209] The processor refers to the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0210] It should be noted that, in this invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatus in the embodiments of this invention is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0211] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0212] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of the present invention.

Claims

1. A method for reconstructing a human face prototype by integrating adversarial attack defense mechanisms, characterized in that, include: Construct a prototype reconstruction generative adversarial network, wherein the adversarial network includes a generator and a discriminator; The generator and the discriminator are trained based on preset generator loss functions and preset discriminator loss functions, respectively, to obtain trained generator and trained discriminator. An adversarial attack defense layer is constructed, comprising an image denoising network layer and a matrix estimation layer. The image denoising network layer denoises the noise in the input image to obtain an initial denoised image, and the matrix estimation layer further denoises the initial denoised image to obtain a target denoised image. The image denoising network layer includes a feature enhancement layer, an attention layer, and a residual reconstruction layer. The feature enhancement layer comprises multiple convolutional layers and skip connection layers. Each of the first three convolutional layers is followed by a batch normalization layer and a rectified linear unit. The skip connection concatenates the input image of the network with the output of the fourth convolutional layer to obtain a concatenated feature. This concatenated feature is then processed using the Tanh activation function to obtain the output of the feature enhancement layer, expressed as: ; in, Indicates the input image. This indicates the first three convolutional layers in the feature enhancement layer, each followed by a batch normalization (BN) layer and a rectified linear unit (ReLU). Indicates a convolutional layer; Indicates features and characteristics splicing; This represents the Tanh activation function, which aims to more effectively propagate training gradients and make the training process more stable. This represents the output of the feature enhancement layer FEM; The attention layer is used to filter noise in the output of the feature enhancement layer to obtain enhanced noise, as expressed in the following expression: ; in, This represents the output features of the feature enhancement layer (FEM). This represents the output feature of the fourth convolutional layer in the Feature Enhancement Model (FEM). Indicates 1 1 convolutional layer, This indicates an element-wise multiplication operation. This represents the output of the attention layer AM, i.e., the predicted noise; The residual reconstruction layer is used to connect the input image of the network with the noise predicted by the attention layer AM through skip connections, thereby achieving element-wise subtraction between the input image and the predicted noise to reconstruct the denoised image. The expression is as follows: ; in, This represents the input image of the denoising network. This represents the noise predicted by the attention layer AM. This indicates an element-wise subtraction operation. This represents the output of the residual reconstruction layer, i.e., the denoised image output by the image denoising network; An adversarial attack defense loss function is constructed based on the difference between the denoised target image and the original image; Based on the adversarial attack defense loss function, the adversarial attack defense layer is trained end-to-end to obtain a trained adversarial attack defense layer. The trained generator, the trained discriminator, and the trained adversarial attack defense layer are used to process the preset input face image to generate a face prototype reconstruction result.

2. The face prototype reconstruction method according to claim 1, characterized in that, The generator includes an encoder and a decoder. The encoder is used to obtain facial identity features of the input face image. The encoder consists of multiple consecutive convolutional layers and an average pooling layer. Each convolutional layer is followed by a batch normalization layer and an exponential linear unit. The decoder is used to concatenate and decode the facial identity features with random noise vectors to generate a face prototype. The decoder consists of multiple consecutive deconvolutional layers and a fully connected layer. Each deconvolutional layer is followed by a batch normalization layer and an exponential linear unit layer. The fully connected layer is used to expand the dimension of the encoder output features to obtain a dimension-expanded feature map. The deconvolutional layer is used to upsample the dimension-expanded feature map to obtain a reconstructed face image.

3. The face prototype reconstruction method according to claim 1, characterized in that, The discriminator includes a first sub-discriminator, a second sub-discriminator, and a third sub-discriminator. The first sub-discriminator is responsible for predicting the identity of the input image. The second sub-discriminator is used to determine whether the input image contains facial changes. The third sub-discriminator is used to evaluate the input image based on the similarity score between the input image and the real prototype to generate a judgment result.

4. The face prototype reconstruction method integrating adversarial attack defense mechanisms according to claim 3, characterized in that, The discriminator further includes multiple consecutive convolutional layers, an average pooling layer, and a fully connected layer. The convolutional layers extract features from the input image to obtain convolutional features. The average pooling layer compresses the convolutional features to obtain compressed features. The fully connected layer performs dimensionality transformation on the compressed features to obtain a multidimensional feature vector. The dimension of the multidimensional feature vector is... ,in, The first dimension represents the number of identities contained in the training set, and the second dimension represents the face identity label predicted by the first sub-discriminator. The last two dimensions are used to represent the results of the second sub-discriminator in judging whether the input image contains facial changes and the results of the third sub-discriminator in judging whether the input image is a real prototype.

5. The face prototype reconstruction method according to claim 1, characterized in that, The specific process of training the generator based on a preset generator loss function to obtain a trained generator includes: Define a generator identity preservation loss function, wherein the identity preservation loss function... The mathematical expression is: ; in, This represents the output of the generator. Indicates the input image. This represents a noise vector randomly sampled from a uniform distribution. This represents the prototype generated by the generator, i.e. , This represents the output of the sub-discriminator responsible for identity prediction. Indicates the input image identity tags Sub-discriminator Predicting the generation of prototypes Identity label The probability, Represents logarithmic operations. Indicates to Maximum likelihood estimation; Define a facial change loss function for the generator, wherein the facial change loss function for the generator... The mathematical expression is: ; in, This represents the output of the generator. Indicates the input image. This represents a noise vector randomly sampled from a uniform distribution. This represents the prototype generated by the generator, i.e. , This represents the output of the sub-discriminator responsible for determining whether the input image contains facial changes. Indicates the input image Does it include tags indicating facial changes? Sub-discriminator Predicting the generation of prototypes Facial changes tagged as The probability, Represents logarithmic operations. Indicates to Maximum likelihood estimation; Define a generator adversarial generation loss function, wherein the generator adversarial generation loss function The mathematical expression is: ; in, This represents the output of the generator. Indicates the input image. This represents a noise vector randomly sampled from a uniform distribution. This represents the prototype generated by the generator, i.e. , This represents the output of the sub-discriminator responsible for determining whether the input image is the real prototype or the generated prototype. Sub-discriminator A prototype will be generated. The probability of identifying it as the real prototype. Represents logarithmic operations. Indicates to Maximum likelihood estimation; Define a generator reconstruction loss function, wherein the generator reconstruction loss function The mathematical expression is: ; in, This represents the output of the generator. This represents the input real-world prototype image. This represents a noise vector randomly sampled from a uniform distribution. This indicates that the generator produces a fake prototype based on the real prototype. This represents the calculation of the square of the Frobenius norm, where the Frobenius norm is the square root of the sum of the squares of all elements in the input vector. Indicates to Maximum likelihood estimation; Define a generator-aware loss function, wherein the generator-aware loss function The mathematical expression is: ; in, This represents the output of the generator. Indicates the input image. This represents a noise vector randomly sampled from a uniform distribution. This represents the prototype generated by the generator, i.e. , This represents the feature extractor, responsible for extracting features from the real prototype. and generating prototypes Extract features from them, This represents the calculation of the square of the Frobenius norm, where the Frobenius norm is the square root of the sum of the squares of all elements in the input vector. Indicates to Maximum likelihood estimation; A multi-objective joint loss function is established based on the generator identity preservation loss function, the generator facial change loss function, the generator adversarial generation loss function, the generator reconstruction loss function, and the generator perception loss function; The generator is trained based on the multi-objective joint loss function to obtain a trained generator, wherein the mathematical expression of the multi-objective joint loss function is: in, This represents the adversarial generation loss function of the generator. , , and These are all weight parameters used to balance the loss function. Represents the joint loss function for multiple objectives. This indicates that the generator is optimized during training. The network parameters are used to maximize The corresponding loss function value.

6. The face prototype reconstruction method according to claim 1, characterized in that, The specific process of training the discriminator based on a preset discriminator loss function to obtain a trained discriminator includes: Define the identity prediction loss function of the discriminator, wherein the mathematical expression of the identity prediction loss function of the discriminator is: ; in, Indicates the input image. This represents the output of the sub-discriminator responsible for identity prediction. Indicates the input image identity tags Sub-discriminator Predict the input image Identity label The probability, Represents logarithmic operations. Indicates to The maximum likelihood estimate, This represents the identity prediction loss function of the discriminator; Define a facial change loss function for the discriminator, wherein the mathematical expression of the facial change loss function is: ; in, Indicates the input image. This represents the output of the sub-discriminator responsible for determining whether an image contains facial changes. Indicates the input image Does it include a label indicating facial changes? Sub-discriminator Determine the input image Does the result include facial changes? The probability, Represents logarithmic operations. Indicates to The maximum likelihood estimate, This represents the facial change loss function of the discriminator; Define the adversarial loss function of the discriminator, wherein the mathematical expression of the adversarial loss function of the discriminator is: ; in, Indicates the input image. This represents the output of the generator. This represents a noise vector randomly sampled from a uniform distribution. This represents the prototype generated by the generator. Represents the real prototype. This represents the output of the sub-discriminator responsible for determining whether an image is the real prototype or the generated prototype. Sub-discriminator A prototype will be generated. The probability of identifying it as the real prototype. Sub-discriminator real prototype The probability of judging it as true. Represents logarithmic operations. Indicates to The maximum likelihood estimate, Indicates to The maximum likelihood estimate, This represents the adversarial loss function of the discriminator; A multi-objective joint loss function is established based on the identity prediction loss function, facial change loss function, and adversarial loss function of the discriminator, with the following expression: ; in, This represents the adversarial loss function of the discriminator. This represents the identity prediction loss function of the discriminator. This represents the facial change loss function of the discriminator. and These are all weight parameters used to balance the loss function. This represents the multi-objective joint loss function for the discriminator. This indicates that the discriminator is optimized during training. The network parameters are used to maximize The corresponding loss function value; After training the preset discriminator based on the multi-objective joint loss function, a trained discriminator is generated, so that the first sub-discriminator can accurately predict the identity label of the input image, the second sub-discriminator can accurately predict the facial change label of the input image, and the third sub-discriminator can accurately determine whether the input image is a real prototype or a fake prototype generated by the generator.

7. The face prototype reconstruction method according to claim 1, characterized in that, The matrix estimation layer is used to further denoise the initial denoised image to obtain the target denoised image. The specific process includes: The input noisy image is normalized to obtain a normalized image matrix, wherein the mathematical expression of the normalized image matrix is: ; in, This represents the input noisy image. This represents the normalized image matrix to be decomposed. Through the normalization operation, the numerical range of the image matrix is ​​changed from... It became ; The normalized image matrix is ​​subjected to singular value decomposition and singular value thresholding sequentially to obtain the thresholded singular values. The mathematical expression for the singular value decomposition is: ; ; in, This represents the normalized image matrix to be decomposed. Describes a left singular vector matrix. Describes a right singular vector matrix. This represents the transpose of a right singular vector matrix. This represents a diagonal matrix, whose diagonal elements... These are called singular values, and they satisfy... ; The mathematical expression for the singular value thresholding screening is: in, This indicates the number of singular values ​​before thresholding. This indicates the proportion of singular values ​​retained. This represents the number of singular values ​​retained after thresholding, i.e., the number of singular values ​​after thresholding. ; The normalized image matrix is ​​reconstructed based on the thresholded singular values ​​to generate the target denoised image, as expressed by: ; in, Let represent a diagonal matrix composed of the retained singular values. Describes a left singular vector matrix. Describes a right singular vector matrix. This represents the transpose of a right singular vector matrix. This represents the output of the matrix estimation, i.e., the denoised image.

8. The face prototype reconstruction method according to claim 1, characterized in that, The mathematical expression for the adversarial attack defense loss function is: ; in, This represents the training set samples without added noise. This represents the training set samples after noise has been added. This represents the total number of samples in the training set. This indicates the output of the defense layer against attacks. Let represent the square of the L2 norm, where the L2 norm is the square root of the sum of the squares of the elements of the input vector. This represents the loss function for defense against attacks.

9. A face prototype reconstruction system integrating adversarial attack defense mechanisms, used to implement the method as described in any one of claims 1-8, characterized in that, include: The first construction module is used to construct a prototype reconstruction generative adversarial network, wherein the adversarial network includes a generator and a discriminator; The first training module is used to train the generator and the discriminator based on preset generator loss functions and preset discriminator loss functions respectively, so as to obtain the trained generator and the trained discriminator. The second construction module is used to build an adversarial attack defense layer. The adversarial attack defense layer includes an image denoising network layer and a matrix estimation layer. The image denoising network layer is used to denoise the noise in the input image to obtain an initial denoised image. The matrix estimation layer is used to denoise the initial denoised image again to obtain a target denoised image. The defense loss function construction module is used to construct an adversarial attack defense loss function based on the difference between the target denoised image and the original image; The second training module is used to perform end-to-end training on the adversarial attack defense layer based on the adversarial attack defense loss function to obtain a trained adversarial attack defense layer. The face prototype reconstruction module is used to process a preset input face image through the trained generator, the trained discriminator, and the trained adversarial attack defense layer to generate a face prototype reconstruction result.

Citation Information

Patent Citations

  • Defense method for high-hidden poisoning attack based on generative adversarial network and application

    CN110598400A

  • Face recognition anti-attack defense method based on GAN

    CN119006974A