A Cross-Domain Face Generation Method Based on Adversarial Network and Correlation Analysis

Through the cross-domain face generation method of adversarial network and correlation analysis, the problem of image quality and identity information preservation in cross-domain face generation is solved, high-quality visible face images are generated and face matching accuracy is improved.

CN116311448BActive Publication Date: 2025-07-18CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310247353.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-15
Publication Date
2025-07-18
Estimated Expiration
2043-03-15

AI Technical Summary

Technical Problem

The existing cross-domain face generation methods have low quality in generating visible face images and are difficult to maintain the identity information of the original domain face, especially in the influence of insufficient data volume and noise.

Method used

Using an adversarial network and correlation analysis method, through feature extraction, reconstruction, generation and discrimination networks, combined with typical correlation analysis and adversarial training, the mapping and reconstruction of visible and non-visible face features is achieved, high-quality visible face images are generated and identity information is maintained.

Benefits of technology

The generated visible-light face images are of higher quality, can reflect the original domain face characteristics more realistically, and the accuracy rate is significantly improved in the face matching task.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116311448B_ABST
    Figure CN116311448B_ABST
Patent Text Reader

Abstract

The present invention relates to a cross-domain face generation method based on adversarial networks and correlation analysis, belonging to the field of computers. The method includes the following steps: preprocess paired cross-domain pictures, obtain the location of the face through face recognition of visible light pictures and cut to obtain paired visible light and non-visible light face pictures; input the visible light face and the non-visible light face into the model, and the feature extraction module of the model extracts visible light face features and non-visible light face features from the visible light face picture and the non-visible light face picture respectively; analyze and calculate the correlation of the face features to obtain a reconstructed face picture, obtain a face picture in the target domain, and the discrimination module makes a discrimination. The visible light face images generated by the present invention have higher quality, can generate more realistic images, and can better preserve the original domain face identity information than other methods, and the accuracy rate in the face matching task is significantly better than other methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computers and relates to a cross-domain face generation method based on adversarial networks and correlation analysis. Background Art

[0002] In real life, limited by the environment and conditions, it is often impossible to obtain the required information, so a different perspective or method is needed. For example, due to poor lighting conditions, an ordinary camera cannot normally capture the human subject and cannot obtain the target information. However, since the human subject can generate heat, a thermal imaging camera can be used to capture and obtain information; when the police handle a case, since there is no camera at the crime scene and the face image of the suspect cannot be obtained, the portrait of the suspect can be drawn based on the description of the witness about the appearance characteristics of the suspect. Although some of the required information can be obtained under harsh conditions, since this information does not belong to the visible light domain that conforms to what the human eye sees, there will be a large difference from what the human eye actually sees. For example, the face in thermal imaging lacks texture details and only has a general outline. In this case, it is difficult and challenging to recognize a face. To solve this problem, people adopt an idea: convert the portrait in the non-visible light domain to the visible light domain, and then use the face matching method in the visible light domain to identify the information to which the face belongs. The problem of converting from the non-visible light domain to the visible light domain is called the image conversion problem. For example, infer the general appearance of a real face based on a thermal imaging face; help the police generate a visible light face from a sketch portrait and then perform matching in the face database to determine the suspect information.

[0003] Image conversion involves the mapping from the original domain to the target domain. During the conversion process, the model removes the attributes of the original domain in the image and reassigns the attributes of the target domain. Although image conversion has achieved remarkable results in many fields, it still needs to be improved in the field of cross-domain face conversion. First, different from other applications, the details of the face are relatively rich and contain more feature information, so the requirements for mapping are higher. A small generation difference may lead to an unsatisfactory generation result. Second, the generated face also needs to ensure the face features of the original domain, otherwise, due to the large similarity between the generated face and the face in the original domain, the matching effect in face matching will be poor. Finally, limited by the equipment and shooting methods, the amount of cross-domain paired face data is small. The lack of data volume may lead to inaccurate parameters obtained after model training, resulting in overfitting, thus deteriorating the accuracy of the model, and the model is easily affected by noise, which may also affect the stability of the model. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a cross-domain face generation method based on adversarial networks and correlation analysis.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A cross-domain face generation method based on adversarial network and correlation analysis, the method comprising the following steps:

[0007] S1: Preprocess the paired cross-domain images, obtain the position of the face through the visible light image of face recognition and cut to obtain paired visible light and non-visible light face images;

[0008] S2: Input the visible light face and the non-visible light face into the model, and the feature extraction module of the model extracts visible light face features and non-visible light face features from the visible light face image and the non-visible light face image respectively;

[0009] S3: Calculate the correlation between the visible light face features and the non-visible light face features in step S2 by canonical correlation analysis;

[0010] S4: The reconstruction module in the model maps the face features obtained in step S2 back to the original domain to obtain a reconstructed face image;

[0011] S5: The generation module in the model maps the face features obtained in step S2 to the target domain to obtain a face image in the target domain;

[0012] S6: Input the generated face image obtained in S4 and the original domain face image in S1 into the discriminant network of the model, and the discriminant module discriminates between the two.

[0013] Optionally, in S2, the obtained visible light face features and non-visible light face features are respectively:

[0014] fea x = E x (x)

[0015] fea y = E y (y)

[0016] where x is the input visible light face image, y is the input non-visible light image, E x is the visible light feature extraction network, E y is the non-visible light feature extraction network, fea x is the extracted visible light face feature, fea y is the extracted non-visible light face feature.

[0017] Optionally, in S3, the correlation loss obtained according to canonical correlation analysis is:

[0018] L cca (E x ,Ey ) = -corr(E x (x), E y (y)) = -||T|| tr = -tr(T'T) 1 / 2

[0019] Where, let H1 = E x (x), H2 = E y (y), is the de - centralization matrix, corresponding to Similarly, define and corresponding to where r1, r2 > 0 are regularization constants; The total correlation of the first k components of H1 and H2 is the sum of the first k singular values of matrix T, k is an optional parameter,

[0020] Optionally, in S4, mapping the face features back to the original domain to obtain the reconstructed face image, then the visible - light reconstruction loss L rec_x and the non - visible - light reconstruction loss L rec_y are respectively:

[0021]

[0022] where D rec_x and D rec_y are the visible - light reconstruction network and the non - visible - light reconstruction network respectively, ||·||1 is the L1 - norm.

[0023] Optionally, in S5, mapping the face features to the target domain to obtain the face picture in the target domain, then the visible - light generation perception loss L per_y and the visible - light generation perception loss L per_x are respectively:

[0024]

[0025] where D x is the visible - light generation network, D y is the non - visible - light generation network, φ(·) is the operation of extracting features of each dimension by the pre - trained model, extracting the features of the real picture and the generated picture respectively, and then using the L1 norm to calculate the difference size of the extracted features.

[0026] Optionally, the pre - trained model includes VGG - 19 and ResNet - 50.

[0027] Optionally, in S6, inputting the generated face picture and the original face picture into the discriminant network for discrimination, the feature extraction networks E x 、E y and the generation network Gx , G y respectively perform adversarial training by competing with the discriminant networks Dis x and Dis y . The visible light adversarial loss L LSGAN_x and the non-visible light adversarial loss L LSGAN_y are respectively as follows:

[0028]

[0029] The beneficial effects of the present invention are as follows: The cross-domain face generation method based on generative adversarial and canonical correlation analysis proposed by the present invention, compared with the existing cross-domain face generation methods, generates visible light face images with higher quality, and can generate more realistic images and can better maintain the original domain face identity information than other methods. The accuracy rate in the face matching task is significantly better than other methods.

[0030] Other advantages, objectives and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, wherein:

[0032] Figure 1 is the flow chart of the present invention;

[0033] Figure 2 is the schematic diagram of the CRC-pix2pix model;

[0034] Figure 3 are the results of generating visible light face pictures by each method on the TFW dataset;

[0035] Figure 4 are the results of generating visible light face pictures by each method on the BUAA VisNir dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0037] Among them, the attached drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the attached drawings will be omitted, enlarged or reduced, which does not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the attached drawings may be omitted.

[0038] In the attached drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the attached drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the attached drawings are only for illustrative purposes and should not be construed as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0039] Please refer to Figure 1 , which is a cross-domain face generation method based on generative adversarial and canonical correlation analysis, including the following steps:

[0040] Step S1: Preprocess the paired cross-domain pictures, obtain the location of the face through the visible light picture of face recognition and cut to obtain paired visible light and non-visible light face pictures.

[0041] Step S2: Input the visible light face and the non-visible light face into the model. The feature extraction module of the model extracts visible light face features and non-visible light face features from the visible light face picture and the non-visible light face picture respectively.

[0042] Step S3: Use canonical correlation analysis to calculate the correlation between the visible light face features and the non-visible light face features in Step S2.

[0043] Step S4: The reconstruction module in the model maps the face features obtained in Step S2 back to the original domain to obtain the reconstructed face pictures.

[0044] Step S5: The generation module in the model maps the face features obtained in Step S2 to the target domain to obtain a face image in the target domain.

[0045] Step S6: Input the generated face image obtained in Step S4 and the original domain face image in S1 into the discriminant network of the model, and the discriminant module discriminates between the two.

[0046] Further, in the said Step S2, the obtained visible light face features and non-visible light face features are respectively:

[0047] fea x = E x (x)

[0048] fea y = E y (y)

[0049] where x is the input visible light face image, y is the input non-visible light image, E x is the visible light feature extraction network, E y is the non-visible light feature extraction network, fea x is the extracted visible light face feature, fea y is the extracted non-visible light face feature.

[0050] Further, in the said Step S3, the correlation loss obtained according to canonical correlation analysis is:

[0051] L cca (E x , E y ) = -corr(E x (x), E y (y)) = -||T|| tr = -tr(T'T) 1 / 2

[0052] where let H1 = E x (x), H2 = E y (y), is the de-centralized matrix, corresponding to Similarly, define and corresponding to where r1, r2 > 0 are regularization constants; the total correlation of the first k components of H1 and H2 is the sum of the first k singular values of the matrix T, k is an optional parameter,

[0053] Further, in the said Step S4, when mapping the face features back to the original domain to obtain a reconstructed face image, the visible light reconstruction loss L rec_xand the non-visible light reconstruction loss L rec_y are respectively:

[0054]

[0055] where D rec_x and D rec_y are the visible light reconstruction network and the non-visible light reconstruction network respectively, and ||·||1 is the L1-norm.

[0056] Furthermore, in the step S5, mapping the face features to the target domain to obtain the face picture in the target domain, then the visible light generation perception loss L per_y and the visible light generation perception loss L per_x are respectively:

[0057]

[0058] where D x is the visible light generation network, D y is the non-visible light generation network, φ(·) is the operation of extracting features of each dimension by the pre-trained model, extracting the features of the real picture and the generated picture respectively, and then using the L1 norm to calculate the difference size of the extracted features.

[0059] Furthermore, in the step S6, inputting the generated face picture and the original face picture into the discriminant network for discrimination, the feature extraction networks E x 、E y and the generation networks G x 、G y respectively perform adversarial training by competing with the discriminant networks Dis x and Dis y The visible light adversarial loss L LSGAN_x and the non-visible light adversarial loss L LSGAN_y are respectively:

[0060]

[0061] Embodiment 1

[0062] The overall flowchart of the cross-domain face generation method based on generative adversarial and canonical correlation analysis of the present invention may specifically include:

[0063] Step S1: Preprocess the paired cross-domain pictures, obtain the position of the face by recognizing the visible light picture of the face and cut to obtain the paired visible light and non-visible light face pictures.

[0064] This embodiment uses the Thermal Faces in the Wild (TFW) and BUAA VisNir datasets. In addition, face matching tests are conducted after generating visible light images in TFW. The model uses an input image size of 256×256, adopts MobileFaceNet as the feature extraction network, and both the generation network and the reconstruction network are composed of 6 residual network modules. Train 150 epochs for each generation network and discriminant network, with a batch size of 1. Among them, a learning rate of 0.0002 is used in the first 100 epochs, and the learning rate linearly decays from 0.0002 to 0 in the last 50 epochs. The Adam optimizer with a momentum term of β1 = 0.9 is used to optimize the network.

[0065] Step S2: Input the visible light face and non-visible light face into the model. The feature extraction module of the model extracts visible light face features and non-visible light face features from the visible light face image and non-visible light face image respectively.

[0066] Specifically, the obtained visible light face features and non-visible light face features are respectively:

[0067] fea x = E x (x)

[0068] fea y = E y (y)

[0069] where x is the input visible light face image, y is the input non-visible light image, E x is the visible light feature extraction network, E y is the non-visible light feature extraction network, fea x is the extracted visible light face feature, fea y is the extracted non-visible light face feature.

[0070] Step S3: Calculate the correlation between the visible light face feature and non-visible light face feature in Step S2 using canonical correlation analysis.

[0071] Specifically, the correlation loss obtained according to canonical correlation analysis is:

[0072] L cca (E x , E y ) = -corr(E x (x), E y (y)) = -||T|| tr = -tr(T'T) 1 / 2

[0073] Among them, let H1 = E x (x), H2 = E y (y), is a decentralized matrix, corresponding to Similarly, define and corresponding to where r1, r2 > 0 are regularization constants; the total correlation of the first k components of H1 and H2 is the sum of the first k singular values of matrix T, and k is an optional parameter.

[0074] Step S4: The reconstruction module in the model maps the face features obtained in step S2 back to the original domain to obtain a reconstructed face image.

[0075] Specifically, the visible light reconstruction loss L rec_x and the non-visible light reconstruction loss L rec_y are respectively:

[0076]

[0077] where D rec_x and D rec_y are the visible light reconstruction network and the non-visible light reconstruction network respectively, and ||·||1 is the L1-norm.

[0078] Step S5: The generation module maps the face features obtained in step S2 to the target domain to obtain a face image in the target domain.

[0079] Specifically, the visible light generation perception loss L per_y and the visible light generation perception loss L per_x are respectively:

[0080]

[0081] where D x is the visible light generation network, D y is the non-visible light generation network, φ(·) is the operation of extracting features of each dimension by a pre-trained model (such as VGG-19, ResNet-50), extracting features from the real image and the generated image respectively, and then using the L1 norm to calculate the difference size of the extracted features.

[0082] At present, most methods still use a loss function based on pixel points for the generated image. However, for faces, the similarity between two faces is evaluated by comparing the perception such as contours and textures. Therefore, the perception loss explores the differences between the high-dimensional representations of images extracted from a pre-trained classifier.

[0083] Step S6: Input the generated face image obtained in Step S4 and the original domain face image in S1 into the discriminative network of the model, and the discrimination module discriminates between the two.

[0084] Specifically, the feature extraction networks E x 、E y and the generation network G x 、G y respectively perform adversarial training by competing with the discriminative network Dis x and Dis y The visible light adversarial loss L LSGAN_x and the non-visible light adversarial loss L LSGAN_y are respectively:

[0085]

[0086]

[0087] Figure 2 is the schematic diagram of the CRC-pix2pix model; Figure 3 are the results of generating visible light face images by each method on the TFW dataset; Figure 4 are the results of generating visible light face images by each method on the BUAAVisNir dataset.

[0088] Figures 2 to 4 The person photos involved in [] are from the publicly available dataset Thermal Faces in the Wild (TFW), and are example samples formed after being processed according to the procedures of this patent. The cited literature in the GB / T 7714 version of the paper of this publicly available dataset is: Kuzdeuov A, Aubakirova D, Koishigarina D, et al. TFW: Annotated thermal faces in the wild dataset [J]. IEEE Transactions on Information Forensics and Security, 2022, 17: 2084 - 2094. And HUANG D, SUN J, WANG Y H. The BUAAVisNir face database instructions [R]. Laboratory of Intelligent Recognition and Image Processing, School of Computer Science and Engineering, Beihang University, Beijing, China, 2012.

[0089] For the Cross-view Reconstruction Correlated pix2pix (CRC-pix2pix) model obtained by adversarial training, different networks are used on the same training dataset to achieve the image-to-image conversion task, and the generated results are evaluated by indicators. The models involved in the comparison are pix2pix, CycleGAN, CDGAN, FFE-GAN, dual-directional GAN and PAN. Structural Similarity Index (SSIM), Peak Signal to Noise Ratio (PSNR), Mean Square Error (MSE) and Learned Perceptual Image Patch Similarity (LPIPS) are used as the quality evaluation indicators of the generated target domain image. These evaluation indicators are very common in the image-to-image conversion problem and are used to judge the similarity between the generated image and the real image. In addition, in order to test the preservation of identity information in the generation process, in the test set of the TFW dataset, 30 individuals' visible frontal faces are used as the matching face library, and face matching is performed on the visible light images generated by thermal imaging images. First, a visible light frontal face image is extracted from each of the 30 individuals, a total of 30 faces as an image pool, and then the VGG-Face model is used to extract the features of the images in the image pool. Then, for the generated image, the VGG-Face model is also used to extract features, and then the Euclidean distance between its features and the features of all the library images is calculated. Finally, the n images with the shortest distance are used as the matching results. If the n images contain the correct image, it is judged to be a correct match, and n is set to 1, 3, 5. In addition, the feature distance (feature distance) between the frontal face image of each individual and the generated photo of the same individual is calculated separately.

[0090] As can be seen from Table 1 and Table 2, whether on the TFW dataset or the BUAAVisNir dataset, this method is generally better than other cross-domain generation methods. This method is better than other methods in terms of image generation quality and similarity with the original image. In the face matching test in the TFW test set, the face matching rate of this method is still higher than that of other methods. From the feature distance metric, it can be concluded that this method can shorten the distance between the generated image and the target image at the feature level.

[0091] Table 1 TFW dataset test results

[0092]

[0093] Table 2 Test Results of the BUAA VisNir Dataset

[0094]

[0095] The above experiments and related result analysis verify the effectiveness of the cross-domain face generation method proposed in the present invention.

[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A cross-domain face generation method based on adversarial network and correlation analysis, characterized in that: The method includes the following steps: S1: Preprocess the paired cross-domain images, obtain the location of the face through face recognition of visible light images, and cut to obtain paired visible light and non-visible light face images; S2: Input the visible light face and non-visible light face into the model. The feature extraction module of the model extracts visible light face features and non-visible light face features from the visible light face image and the non-visible light face image respectively; S3: Calculate the correlation between the visible light face features and the non-visible light face features in step S2 using canonical correlation analysis, and use the calculated correlation loss in the training process of the model to optimize the performance of the feature extraction module; S4: The reconstruction module in the model maps the face features obtained in step S2 back to the original domain to obtain reconstructed face images; S5: The generation module in the model maps the face features obtained in step S2 to the target domain to obtain face images in the target domain, and inputs the generated face images in the target domain and the real target domain images into the pre-trained model to extract features, and calculates the perceptual loss by calculating the feature difference through the L1 norm; S6: Input the generated face images obtained in S4 and the original domain face images in S1 into the discriminant network of the model, and the discriminant module discriminates between the two.

2. The cross-domain face generation method based on adversarial network and correlation analysis according to claim 1, wherein: In S2, the obtained visible light face features and non-visible light face features are respectively: fea x = E x (x) fea y = E y (y) Among them, x is the input visible-light face image, y is the input non-visible-light image, E x is the visible-light feature extraction network, E y is the non-visible-light feature extraction network, fea x is the extracted visible-light face feature, fea y is the extracted non-visible-light face feature.

3. The cross-domain face generation method based on adversarial network and correlation analysis according to claim 2, characterized in that: In S3, the correlation loss obtained according to canonical correlation analysis is: L cca (E x ,E y )=-corr(E x (x),E y (y))=-||T|| tr =-tr(T'T) 12 where, let H1 = E x (x), H2 = E y (y), is a decentralized matrix, corresponding to Similarly, define and corresponding to where r1, r2 > 0 are regularization constants; the total correlation of the first k components of H1 and H2 is the sum of the first k singular values of matrix T, and k is an optional parameter, 4. The cross-domain face generation method based on adversarial network and correlation analysis according to claim 3, characterized in that: In S4, the face features are mapped back to the original domain to obtain a reconstructed face image, and the visible light reconstruction loss L rec_x and the non-visible light reconstruction loss L rec_y are respectively: Among them, D rec_x and D rec_y are the visible light reconstruction network and the non-visible light reconstruction network respectively, and ||·||1 is the L1-norm.

5. A cross-domain face generation method based on adversarial network and correlation analysis according to claim 4, characterized in that: In S5, the face features are mapped to the target domain to obtain a face image in the target domain, and the visible light generation perceptual loss L per_y and the visible light generation perceptual loss L per_x are respectively as follows: Among them, D x is a visible light generation network, D y is a non-visible light generation network, φ(·) is an operation to extract features of each dimension by a pre-trained model. Features are extracted from the real image and the generated image respectively, and then the L1 norm is used to calculate the difference in the extracted features.

6. The cross-domain face generation method based on adversarial network and correlation analysis according to claim 5, characterized in that: The pre-trained model includes VGG-19 and ResNet-50.

7. A cross-domain face generation method based on adversarial network and correlation analysis according to claim 5, characterized in that: In S6, the generated face image and the original face image are input into the discriminant network for discrimination, and the feature extraction networks E x , E y and the generation network G x , G y are respectively trained adversarially by competing with the discriminant networks Dis x and Dis y , and the visible light adversarial loss L LSGAN_x and the non-visible light adversarial loss L LSGAN_y are respectively: