A Heterogeneous Face Recognition Method and System Based on Identity-Attribute Decoupled Synthesis

By adopting the identity-attribute decoupling method in heterogeneous face recognition technology, identity and attribute features are extracted, and training data is generated through image synthesis, the problems of insufficient data and modal differences in heterogeneous face recognition are solved, and the recognition accuracy is significantly improved.

CN114764939BActive Publication Date: 2025-05-27INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210322834.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-29
Publication Date
2025-05-27
Estimated Expiration
2042-03-29

AI Technical Summary

Technical Problem

The existing heterogeneous face recognition technology faces problems such as insufficient heterogeneous data, huge modal differences and complex facial attributes, resulting in insufficient recognition accuracy and generalization.

Method used

A heterogeneous face recognition method based on identity-attribute decoupling is proposed. Identity features are extracted through the pre-trained face recognition model, and attribute features are extracted using a modal-oriented face attribute encoder. Identity-attribute decoupling algorithm is used to decouple identity and attribute features. At the same time, heterogeneous face images are generated through an image synthesis algorithm, which is used to train heterogeneous face recognition models.

Benefits of technology

By generating a large number of heterogeneous face images with rich facial attributes, reducing the cross-modal distance of features within the class, effectively solving the problems of insufficient heterogeneous data, huge modal differences and complex facial attributes, significantly improving the accuracy of heterogeneous face recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114764939B_ABST
    Figure CN114764939B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of biometric recognition, and discloses a heterogeneous face recognition method and system based on identity-attribute decoupling synthesis. First, the present invention designs a face identity-attribute decoupling algorithm to decouple face identity features and attribute features, and constructs a large number of different combinations of identity features and attribute features; further, a face image synthesis algorithm is designed to generate heterogeneous face images using different combinations of identity features and attribute features for training a heterogeneous face recognition model. The present invention solves the challenges faced by heterogeneous face recognition, such as insufficient image data, huge cross-domain differences, and complex facial attributes, and has excellent heterogeneous face recognition performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of biometric recognition, and particularly relates to a heterogeneous face recognition method and system based on identity-attribute decoupling synthesis. Background Art

[0002] In recent years, face recognition has been widely applied to scenarios such as identity authentication, mobile payment, and security monitoring, providing great convenience for fields such as e-government, intelligent transportation, e-commerce, and public security. However, in actual scenarios, face images are usually collected using different types of imaging devices according to different needs. For example, near-infrared images of mobile phones, thermal infrared images of security gates, and sketch portraits of criminal investigations. This results in a huge modal difference between the collected face images and the face images in the identity database, and the identity cannot be directly recognized. Therefore, an effective heterogeneous face recognition method (HFR) is required. Among them, heterogeneous images refer to images of different modalities, such as visible light, near-infrared, and thermal infrared modalities. The goal of heterogeneous face recognition is to match the identities of faces of different modalities, which is of great significance for enhancing the usability and accuracy of face recognition.

[0003] Currently, heterogeneous face recognition faces three major challenges: insufficient heterogeneous data, huge modal differences, and complex facial attributes. Different from visible light faces, collecting non-visible light faces requires special imaging devices. Therefore, the number of non-visible light face images is far less than that of visible light face images. Insufficient heterogeneous data easily leads to overfitting of face recognition algorithms, affecting the accuracy and generalization of face recognition. The modal differences between visible light and non-visible light faces are huge, increasing the difficulty of heterogeneous face recognition. The attributes of the human face are complex and variable, such as attributes like angle, expression, decoration, and ambient light, further increasing the difficulty of heterogeneous face recognition.

[0004] In response to this, research scholars have continuously improved heterogeneous face recognition methods. Existing methods are mainly divided into three categories: modal invariant feature methods, subspace learning methods, and image synthesis methods.

[0005] (1) The main idea of the modal invariant feature method is to design or learn a face feature extraction algorithm to obtain identity discriminant features independent of modality from heterogeneous face images, thereby reducing the modal differences of heterogeneous faces. In the early stage, this type of method mainly relied on manually designed features, with low accuracy. Deep convolutional neural networks have greatly promoted the learning of image feature representation, which is conducive to extracting modality-invariant identity features. The modal invariant feature method solves the challenge of huge modal differences, but still faces the two challenges of insufficient heterogeneous data and complex facial attributes.

[0006] (2) The main idea of the subspace learning method is to project heterogeneous face images onto the same subspace and identify identities based on the distances of face features within the same subspace. The feature dimensionality reduction methods in machine learning are the early tools for subspace projection. The feature decoupling algorithms in deep learning effectively eliminate the differences in facial attribute features, thereby learning a feature subspace that is independent of modalities and facial attributes, further improving the performance of heterogeneous face recognition. However, this type of method cannot alleviate the impact brought by insufficient heterogeneous data.

[0007] (3) Image synthesis methods are divided into two categories: one is conditional synthesis, which converts images in different data domains into the same data domain and matches face identities within the same domain to solve the modality differences of heterogeneous faces; the other is unconditional synthesis, which synthesizes a large number of heterogeneous face images to solve the problem of the current lack of heterogeneous face data and avoid the phenomenon of overfitting in training, thereby improving the accuracy of face recognition. Existing heterogeneous face recognition methods based on image synthesis methods, although achieving good recognition performance on publicly available heterogeneous face video databases, have not yet solved the impact of complex facial attributes on heterogeneous face recognition. Summary of the Invention

[0008] Aiming at the deficiencies of existing methods, the present invention proposes a heterogeneous face recognition method and system based on identity-attribute decoupling synthesis, with leading recognition accuracy. The present invention is specifically implemented according to the following technical solutions:

[0009] The first aspect of the present invention provides a heterogeneous face recognition method based on identity-attribute decoupling synthesis, including the following steps:

[0010] Step S100, randomly sample a pair of heterogeneous face images I N and I V from a heterogeneous face database, and this pair of face images belongs to the same person. First, use a pre-trained face recognition model to extract the identity feature z V of the face image from the face image I id . Among them, the identity refers to the person label to which the face belongs. Second, propose a face attribute encoder oriented to modalities, and extract face attribute features N and V from the heterogeneous face images I and respectively. Third, propose an identity-attribute decoupling algorithm, fix the face identity feature z id , and constrain the face attribute features to be orthogonal to the face identity feature, so as to decouple the features of identity and attributes. The face attribute features specifically include information such as facial angle, expression, skin color, and illumination.

[0011] Step S200, randomly sample a pair of heterogeneous face images X N and X V, using the modality-oriented face attribute encoder in step S100, from X N and X V Extract attribute features and Combine the identity feature z in step S100 id with the attribute features in step S200 and to obtain a feature combination and Then propose a face image synthesis algorithm, and the face image generator synthesizes heterogeneous face images from the feature combinations and respectively and Store the synthesized heterogeneous face images as the training data of the heterogeneous face recognition model.

[0012] Step S300, loop steps S100 and S200 until all images in the heterogeneous face database are sampled, then proceed to step S400.

[0013] Step S400, combine the original images in the heterogeneous face database and the heterogeneous face images generated in step S300 to train the heterogeneous face recognition model. For the original images in the heterogeneous face database, use the cross-entropy function as the classification loss function; for the heterogeneous face images generated in step S300, propose a cross-modal contrast loss function. The values of the classification loss function and the cross-modal contrast loss function are added together as the training loss of the heterogeneous face recognition model to update the model parameters. The heterogeneous face recognition model includes a convolutional neural network and a fully connected layer, where the convolutional neural network extracts identity features from the face image to be recognized, and the fully connected layer classifies according to the extracted identity features and outputs the label number of the identity.

[0014] Step S500, the model trained in step S400 can be used to identify the identity of heterogeneous face images. Input the face image to be recognized into the heterogeneous face recognition model, and the convolutional neural network inside the model encodes the image into an identity feature vector v; similarly, input all the face images in the identity database into the heterogeneous face recognition model in turn and encode them into feature vectors v i , i ∈ [1, N], where N is the number of people in the identity database. Calculate the similarity between the feature vector v and v i in turn. If the similarity between the feature vector v and v j , j ∈ [1, N] is the highest, then it is determined that the face to be recognized belongs to the j-th person.

[0015] The second aspect of the present invention proposes a heterogeneous face recognition system based on identity-attribute decoupling synthesis, including a preprocessing module, a feature decoupling module, an image synthesis module, an image storage medium, and a training module.

[0016] The preprocessing module is configured for data reading and preprocessing of face images. Among them, data reading refers to reading the original data from a heterogeneous face database; data preprocessing refers to performing face detection, localization, alignment, cropping, and size normalization on the original face data to obtain face images with consistent sizes.

[0017] The feature decoupling module is configured to decouple the identity features and attribute features of a face. It includes a pre-trained face recognition model and two modality-oriented face attribute encoders. The pre-trained face recognition model is used to extract the identity features of face images. The modality-oriented face attribute encoders extract face attribute features independent of identity from two face images of different modalities respectively, so as to separate the face images into identity features and attribute features.

[0018] The image synthesis module is configured to generate heterogeneous face images. It includes a generator that receives the feature combination of face identity and attributes and synthesizes heterogeneous face images.

[0019] The image storage medium is used to store the heterogeneous face database and the heterogeneous face images generated by the image synthesis module, and provide training sample data for the training module.

[0020] The training module is configured to train the model parameters of a neural network. This module updates the neural network model parameters of the feature decoupling module, the image synthesis module, and the heterogeneous face recognition model through the gradient backpropagation algorithm.

[0021] The beneficial effects of the technical solution of the present invention are as follows:

[0022] The present invention effectively solves the problems of insufficient heterogeneous data, huge modality differences, and complex facial attributes by synthesizing a large number of heterogeneous face images with rich facial attributes and reducing the cross-modal distance of intra-class features, and improves the accuracy of heterogeneous face recognition. Description of the Drawings

[0023] Figure 1 It is a schematic structural diagram of a heterogeneous face recognition model based on identity-attribute decoupling synthesis in one or more embodiments of the present invention.

[0024] Figure 2 It is a schematic diagram of the ROC curve of the test results of the heterogeneous face recognition model in the CASIA NIR-VIS 2.0 database in multiple embodiments of the present invention.

[0025] Figure 3 It is a schematic diagram of the ROC curve of the test results of the heterogeneous face recognition model in the BUAA-VisNir database in multiple embodiments of the present invention.

[0026] Figure 4 Schematic diagram of the ROC curve of the test results of the heterogeneous face recognition model in multiple embodiments of the present invention in the Oulu-CASIA NIR-VIS database.

[0027] Figure 5 Schematic diagram of the ROC curve of the test results of the heterogeneous face recognition model in multiple embodiments of the present invention in the LAMP-HQ database. Detailed implementation manners

[0028] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0029] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the present invention otherwise clearly indicates, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0030] As introduced in the background art, in view of the deficiencies in the prior art, the present invention provides a heterogeneous face recognition method based on identity-attribute decoupled synthesis for the temporal relationship features between different regions of the human face, and the accuracy of heterogeneous face recognition is high.

[0031] For the convenience of description, Table 1 introduces the symbols used in the present invention.

[0032] Table 1

[0033]

[0034]

[0035] Embodiment 1

[0036] A typical implementation manner of the present invention, as Figure 1 shown, Embodiment 1 discloses a heterogeneous face recognition method based on identity-attribute decoupled synthesis, including the following steps:

[0037] Step S100, extract the identity features of the face image, and then use the identity-attribute decoupling algorithm to learn the attribute features of the face image and extract the face facial attribute information independent of the identity. The specific process of this step includes:

[0038] Step S111: Use the pre-trained face recognition model LightCNN (refer to Wu X, He R, Sun Z, et al. A Light CNN for Deep Face Representation with Noisy Labels[J]. IEEE Transactions on Information Forensics and Security, 2018, 13(11): 2884-2896) as the face identity feature encoder E id , and extract identity features from heterogeneous face images.

[0039] Step S112: Propose two modality-oriented face attribute encoders E N and E V . E N performs feature extraction on the face image I N of modality one to obtain face attribute features irrelevant to identity E V performs feature extraction on the face image I V of modality two to obtain face attribute features irrelevant to identity The neural network structure of the face attribute encoder is shown in Table 2. The convolutional layer specifically includes one 2D convolutional operation, one instance normalization operation, and one leaky rectified linear unit (Leaky ReLU) activation function. After the face image is input into the face attribute encoder, it is first processed by the convolutional layer, and the feature map is obtained through convolutional operations in sequence. Data regularization is performed through the normalization operation, and non-linear transformation of the feature map is performed by the activation function. Similarly, the feature map is processed by subsequent convolutional layers. Finally, the fully connected layer transforms the feature map into two feature vectors μ and σ, representing the mean and standard deviation of the face attribute distribution respectively. Since the prior distribution of face attributes follows a multivariate Gaussian distribution I is the identity matrix. Use the reparameterization trick (Kingma D P, Welling M. Auto-Encoding Variational Bayes[C]. International Conference on Learning Representations, 2014) to transform μ and σ into attribute feature vectors: ∈ is Gaussian noise.

[0040] Table 2

[0041]

[0042]

[0043] Step S200: Select a large number of face images from the heterogeneous face database, extract a large number of identity features and attribute features using the identity-attribute decoupling algorithm, and randomly combine the identity and attribute features. Then, use the face image synthesis algorithm to generate heterogeneous face images according to the feature combinations of identity and attribute. The specific process of this step includes:

[0044] Step S211: Select two pairs of heterogeneous face images from the heterogeneous face database. The first pair of heterogeneous face images is I N and I V , and the second pair of heterogeneous face images is X N and X V . Use E id to extract the identity feature z V from I id , and use E N to extract the attribute features N and N from I and respectively. Use E V to extract the attribute features V and V from I and

[0045] Step S212: Combine the identity feature z id and the attribute feature . Use the generator G (refer to Goodfellow I, Pouget-Abadie J, Mirza M, et al. Generative Adversarial Networks[J]. Communications of the ACM, 2020, 63(11):139-144) to generate the face image from Combine the identity feature z id and the attribute feature . Use the generator G to generate the face image from The neural network structure of the generator is shown in Table 3. Among them, the transposed convolutional layer specifically includes 1 2D transposed convolutional operation, 1 Adaptive Instance Normalization (refer to Huang X, Belongie S. Arbitrary StyleTransfer in Real-time with Adaptive Instance Normalization[C]. IEEEInternational Conference on Computer Vision. 2017: 1501-1510) operation, and 1 activation function with Leaky ReLU; Tanh is the hyperbolic tangent activation function. The specific processing flow of the generator for generating face images is as follows, where the symbol z X represents the attribute feature or To avoid redundancy. The face identity feature z id and the attribute feature z X are input into the generator. First, they are mapped to the feature vector x by the fully connected layer; then processed by the transposed convolutional layer. After the transposed convolutional operation, a feature map is obtained. After the adaptive instance normalization operation, data normalization is performed on the feature map, and the regularized feature map is non-linearly transformed by the activation function. Further, the feature map is successively processed by the subsequent transposed convolutional layer and convolutional layer. The processing process is the same as above and will not be elaborated. Finally, the hyperbolic tangent activation function transforms the feature map into a face image, and the generated face image is obtained.

[0046] Table 3

[0047]

[0048]

[0049] Step S213, loop steps S211-212 until all heterogeneous face images are selected.

[0050] Step S300, combine the original images in the heterogeneous face database and the heterogeneous face images generated in step S300 to train the heterogeneous face recognition model. For the original images in the heterogeneous face database, the cross-entropy function is used as the classification loss function; for the heterogeneous face images generated in step S300, a cross-modal contrast loss function is proposed. The values of the classification loss function and the cross-modal contrast loss function are added together as the training loss of the heterogeneous face recognition model to update the model parameters. The specific training steps and the details of the cross-modal contrast loss function are introduced in Embodiment 2.

[0051] To verify the effectiveness of the method proposed in the present invention, experiments on heterogeneous face recognition were carried out using the CASIA NIR-VIS 2.0, BUAA-VisNir, Oulu-CASIA NIR-VIS, and LAMP-HQ databases.

[0052] The experiments used the Rank-1 Accuracy, VR@FAR = 1%, VR@FAR = 0.1%, and VR@FAR = 0.01% metrics to test the performance of the method. The experimental results of the model on the CASIA NIR-VIS 2.0, BUAA-VisNir, Oulu-CASIA NIR-VIS, and LAMP-HQ databases are shown in Tables 4-7. The schematic diagram of the ROC curve for the experimental results of the model on the CASIA NIR-VIS 2.0, BUAA-VisNir, Oulu-CASIA NIR-VIS, and LAMP-HQ databases is as Figures 2 - 5 shown. FSIAD is the English abbreviation of the method of the present invention.

[0053] Table 4 Ten-fold cross-validation results of the model on CASIA NIR-VIS 2.0

[0054] Method Rank-1 Accuracy(%) VR@FAR=0.1%(%) VR@FAR=0.01%(%) TRIVET 95.7±0.5 91.0±1.3 74.5±0.7 IDR 97.3±0.4 95.7±0.7 - ADFL 98.2±0.3 97.2±0.5 - W-CNN 98.7±0.3 98.4±0.4 94.3±0.4 PACH 98.9±0.2 98.3±0.2 - DVR 99.7±0.1 99.6±0.3 98.6±0.3 DVG 99.8±0.1 99.8±0.1 98.8±0.2 LightCNN 96.7±0.2 94.8±0.4 88.5±0.2 FSIAD 99.9±0.0 99.9±0.0 99.2±0.1

[0055] Table 5 Experimental results of the model on BUAA-VisNir

[0056] Method Rank-1 Accuracy(%) VR@FAR=1%(%) VR@FAR=0.1%(%) TRIVET 93.9 93.0 80.9 IDR 94.3 93.4 84.7 ADFL 95.2 95.3 88.0 W-CNN 97.4 96.0 91.9 PACH 98.6 98.0 93.5 DVR 99.2 98.5 96.9 DVG 99.3 98.5 97.3 LightCNN 96.5 95.4 86.7 FSIAD 99.8 99.7 99.1

[0057] Table 6 Experimental results of the model on Oulu-CASIA NIR-VIS

[0058] Method Rank-1 Accuracy(%) VR@FAR=1%(%) VR@FAR=0.1%(%) TRIVET 92.2 67.9 33.6 IDR 94.3 73.4 46.2 ADFL 95.5 83.0 60.7 W-CNN 98.0 81.5 54.6 PACH 100.0 97.9 88.2 DVR 100.0 97.2 84.9 DVG 100.0 97.5 90.6 LightCNN 96.7 92.4 65.1 FSIAD 100.0 98.1 92.0

[0059] Table 7 Experimental results of the model on LAMP-HQ

[0060] Method Rank-1 Acc(%) VR@FAR=1%(%) VR@FAR=0.1%(%) VR@FAR=0.01%(%) LightCNN 96.2 96.1 95.3 69.3 ADFL 95.8 91.5 71.0 - PACH 96.9 93.9 78.7 - DVG 98.3 98.8 96.0 88.2 FSIAD 98.8 99.1 97.9 93.0

[0061] The experimental results of the model show that the method proposed in the present invention achieves the highest face recognition accuracy, and the VR@FAR = 0.01% on the LAMP-HQ dataset is 4.8% higher than the existing best method.

[0062] Example 2

[0063] Example 2 discloses a heterogeneous face recognition system based on identity-attribute decoupled synthesis, including a preprocessing module, a feature decoupling module, an image synthesis module, an image storage medium, and a training module.

[0064] The preprocessing module completes the data reading and preprocessing of face images. Among them, data reading refers to reading the original data from a heterogeneous face database; data preprocessing refers to performing face detection, localization, alignment, cropping, and size normalization on the original face data to obtain face images with consistent sizes.

[0065] The feature decoupling module works according to step S100. It decouples the identity features and attribute features of the face. It includes a pre-trained face recognition model and two modality-oriented face attribute encoders. The pre-trained face recognition model is used to extract the identity features of the face image. The two modality-oriented face attribute encoders extract attribute features independent of identity from face images in two different domains respectively.

[0066] The face attribute encoder is implemented using a Variational Auto-Encoder (VAE, refer to Kingma D P, Welling M. Auto-Encoding Variational Bayes[C]. International Conference on Learning Representations, 2014). and The prior distribution and obey a multivariate Gaussian distribution It is necessary to use the KL divergence loss function to train the face attribute encoder to make and The posterior distribution and approximate the prior distribution, so as to learn the face attribute features. The KL divergence loss function is expressed as: Furthermore, in order to make the face attribute features orthogonal to the face identity features, a decoupling loss function is proposed to train the face attribute encoder, which is expressed as: When decreases to 0, z id is orthogonal to

[0067] Combining the KL divergence loss function and the decoupling loss function, a feature decoupling loss function is proposed to train the face attribute encoder, where λ dis is a hyperparameter and is set to 2.

[0068] The image synthesis module generates heterogeneous face images according to step S200. Image generation includes face reconstruction and face generation. Face reconstruction combines I N and I V ​Identity features z id Combined with attribute features to obtain a feature combination and Input generator G to generate a reconstructed face image and Use a reconstruction loss function to reduce the difference between the reconstructed face and the original face.

[0069]

[0070] Face generation takes I N and I V 's identity features z id Combined with X N and X V 's attribute features and to obtain a feature combination and Input generator G to generate a face image and The generated face image should satisfy two properties: (1) The identity is consistent with z id ; (2) The attributes are similar to X N and X V . Therefore, an identity-preserving loss function is proposed to train the generator G to reduce the distance between and 's identity features and z id .

[0071]

[0072] For the similarity of attributes, constraints are imposed from both the image and feature aspects to train the generator G.

[0073] First, combine the L 1 loss (L 1 (x 1 , x 2 ) = ‖x 1 - x 2 ‖ 1) and the multi-scale structural similarity (MSSSIM, refer to Wang Z, Bovik A C, Sheikh H R, et al. Image Quality Assessment: From Error Visibility to Structural Similarity[J]. IEEE Transactions on Image Processing, 2004, 13(4): 600 - 612), an image similarity loss function is proposed

[0074]

[0075] Among them, The smaller it is, the more similar x 1 is to x 2 ;

[0076] MSSSIM(x 1 , x 2 ) is the multi-scale structural similarity function; α is a hyperparameter that adjusts the ratio of the L 1 loss to the multi-scale structural similarity function, and is set to 0.84. The specific calculation process of MSSSIM(x 1 , x 2 ) is as follows:

[0077] Calculate the means μ 1 and μ 2 of images x 1 and x 2 respectively, calculate the standard deviations σ 1 and σ 2 of images x 1 and x 2 , and calculate the covariance σ 1 and σ 2 of images x 12 ;

[0078] Estimate the brightness of images x 1 and x 2 : C 1 is a constant, which is 6.5025;

[0079] Estimate the contrast of images x 1 and x 2 : C 2 is a constant, which is 58.5225;

[0080] Estimate the structural similarity of images x 1 and x 2 : C 3 is a constant, which is 29.26125;

[0081] Using a low-pass filter and a downsampler, perform low-pass filtering and 1 / 2 downsampling on the images x 1 and x 2 iteratively for M times, and calculate the contrast c(x 1 , x 2 ) and the structural similarity s(x 1 , x 2 ) after each iteration. The contrast and structural similarity after the j-th iteration are denoted as c j (x 1 , x 2 ) and s j (x 1 , x 2 ). Until the image is downsampled to a size of 1*1, calculate the luminance similarity l M (x 1 , x 2 ), where M is the number of iterations.

[0082] where α M , β j and γ j are hyperparameters, set as α M = β j = γ j = 1.

[0083] Second, propose an attribute feature loss function to reduce the distance between attribute features.

[0084]

[0085] where and are the attribute features of and respectively.

[0086] To further improve the quality of the generated images, a discriminator D is introduced to train the generator G in an adversarial manner. The goal of the discriminator D is to correctly distinguish between real images and generated images, while the goal of the generator G is to generate high-quality images so that D cannot correctly distinguish their authenticity, thereby optimizing the generator G. Propose an adversarial loss function where D(I) is the probability that the discriminator judges the image I as a generated image. The image I is sampled from the real image data, that is, the heterogeneous face image I N and I V , denotes the real image data; is the probability that the discriminator judges the image for the generated image. The image is sampled from the generated image data, that is, the generated heterogeneous face image and represents the generated image data.

[0087] The comprehensive reconstruction loss function The identity preservation loss function The image similarity loss function and the adversarial loss function propose an image synthesis loss function for training the image synthesis module.

[0088] The image storage medium is used to store the heterogeneous face database and the heterogeneous face images generated by the image synthesis module, and provide training sample data for the training module.

[0089] The training module trains the heterogeneous face recognition model F according to step S400. For the original heterogeneous face images, the cross-entropy loss function is used as the classification loss to train the heterogeneous face recognition model, where n is the total number of heterogeneous face images; is the i-th pair of heterogeneous face images, i ∈ [1, n], including 2 face images of different modalities and y i is the identity label of the face image . For the heterogeneous face images generated in step S300 and the identity features are respectively extracted using the heterogeneous face recognition model and Then the cross-modal intra-class loss function is used to reduce the intra-class identity feature distance of images of different modalities,

[0090] The comprehensive cross-entropy loss function and the cross-modal intra-class loss function propose a heterogeneous face recognition loss function for training the heterogeneous face recognition model F. where γ is a hyperparameter, set to 0.001.

[0091] Example 3

[0092] To verify the effectiveness of the method proposed by the present invention, experiments are carried out in this embodiment. The present invention and the loss function proposed by the present invention are cancelled to test the heterogeneous face recognition performance of the model. Figure 5 As shown in Table 7, cancelling the present invention or the loss function proposed by the present invention both results in a decline in the heterogeneous face recognition performance, verifying that the present invention helps to improve the accuracy of heterogeneous face recognition.

[0093] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A heterogeneous face recognition method based on identity-attribute decoupled synthesis, the steps of which include: 1) Randomly extract a sample Y from a heterogeneous face database, where each sample in the heterogeneous face database includes two heterogeneous face images belonging to the same person; this sample Y includes a pair of heterogeneous face images I N and I V . Use a face identity feature encoder to extract the identity feature z V from I id , and use a face attribute encoder to extract face attribute features from I N respectively Extract face attribute features from I V respectively Then fix the face identity feature z id , constrain the face attribute features to be orthogonal to the face identity feature, and extract face facial attribute information independent of identity; 2) Randomly sample a sample X from the heterogeneous face database, which includes a pair of heterogeneous face images X N and X V ; Extract the attribute features from X N Extract the attribute features from X Extract the attribute features from X V Extract the attribute features from X Combine the identity feature z id with the attribute features respectively to obtain the feature combinations and Then synthesize the heterogeneous face image from the feature combination Synthesize the heterogeneous face image Synthesize the heterogeneous face image from the feature combination Synthesize the heterogeneous face image For the synthesized heterogeneous face image As training data for a heterogeneous face recognition model; 3) Repeat steps 1) to 2) to obtain a training data set; 4) Combine the original images of the heterogeneous face database and the training data set to train the heterogeneous face recognition model; wherein, the heterogeneous face recognition model includes a convolutional neural network and a fully connected layer, the convolutional neural network is used to extract identity features from the input face image and input them into the fully connected layer, and the fully connected layer is used to classify the face image according to the identity features and output the identity label corresponding to the face image; when the input image is a sample Y in the heterogeneous face database, the cross-entropy loss function is used as the classification loss function to calculate the loss value; when the input image is the synthesized image corresponding to the sample Y in the training data set, the cross-modal contrast loss function is used to calculate the loss value; add the loss value calculated by the classification loss function and the loss value calculated by the cross-modal contrast loss function as the training loss of the heterogeneous face recognition model, and update the parameters of the heterogeneous face recognition model; 5) Input the face image to be recognized into the trained heterogeneous face recognition model, and the convolutional neural network encodes the face image to be recognized into an identity feature vector v; Then calculate the similarity between the feature vector v and the feature vectors corresponding to each face image in the identity database. If the similarity between the feature vector v and the feature vector v j corresponding to the j-th person in the identity database is the highest, it is determined that the face image to be recognized belongs to the j-th person.

2. The method according to claim 1, wherein, In step 1), the identity-attribute decoupling algorithm is used to extract face facial attribute information independent of identity. The steps include: selecting two modality-oriented face attribute encoders, that is, the face attribute encoder for the first modality is E N and the face attribute encoder for the second modality is E V ; then using E N to extract features from the face image I N belonging to modality one, obtaining face attribute features independent of identity Using E V to extract features from the face image I V belonging to modality two, obtaining face attribute features independent of identity 3. The method according to claim 1, wherein, The steps of generating a heterogeneous face image include: First, use a fully connected layer to map the face identity feature z id and the attribute feature z X to a feature vector x, where z X is or Then, use a transposed convolutional layer to perform a transposed convolutional operation on the feature vector x to obtain a feature map, perform data regularization on the feature map through an adaptive instance normalization operation, and then use an activation function to perform a non-linear transformation on the regularized feature map to obtain a synthesized heterogeneous face image.

4. The method according to claim 1 or 2 or 3, wherein, The face attribute encoder selects a variational autoencoder; the loss function used for training the face attribute encoder is where is the KL divergence loss function, the decoupling loss function When drops to 0, z id is orthogonal to , where λ dis is a hyperparameter.

5. The method according to claim 1 or 2 or 3, wherein, Use the generator G to generate synthetic heterogeneous face images The loss function used to train the generator G is Among them, the reconstruction loss function During the training process, the generator G generates a reconstructed face image according to the input feature combination Generate a reconstructed face image and Identity preservation loss function is the identity feature extracted by the face identity feature encoder from the input ; is the identity feature extracted by the face identity feature encoder from the input ; The image similarity loss function is the multi-scale structural similarity function, and α is a hyperparameter; The attribute feature loss function and are respectively and 's attribute features; The adversarial loss function where D(I) is the probability that the discriminator judges the image I as a generated image, and the image I is sampled from the real image data, represents the real image data; is the probability that the discriminator determines the image to be the generated image. The image is sampled from the generated image data, which represents the generated image data.

6. The method according to claim 1 or 2 or 3, wherein, The cross-modal contrastive loss function is the synthetic heterogeneous face image The identity features output when inputting into the heterogeneous face recognition model is the synthetic heterogeneous face image Extract the identity features from the identity features output when inputting the synthetic heterogeneous face image into the heterogeneous face recognition model.

7. The method according to claim 1 or 2 or 3, wherein, The face attribute features include facial angle, expression, skin color, and illumination; the identity feature z id is the person label to which the face belongs.

8. A heterogeneous face recognition system based on identity-attribute decoupled synthesis, wherein, includes a feature decoupling module and an image synthesis module; wherein, The feature decoupling module is configured to decouple the identity features and attribute features of the face, and it includes a pre-trained face recognition model and two modality-oriented face attribute encoders; the pre-trained face recognition model is used to extract the identity features of the face image, and the two modality-oriented face attribute encoders respectively extract face attribute features independent of the identity from two different modality face images, so as to separate the face image into identity features and attribute features; The image synthesis module is configured to generate heterogeneous face images, and it includes a generator that receives the feature combination of face identity and attributes and synthesizes heterogeneous face images.

9. A server, wherein, includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the steps in any one of claims 1 to 7.

10. A computer-readable storage medium, on which a computer program is stored, wherein, the computer program realizes the steps of the method according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Patent Citations

  • Face recognition method based on combining with face attribute information

    CN107766850A

  • Image processing method, image processing model training method and equipment

    CN111553267A