A cross-domain face generation method based on multi-stage attention correlation learning

Through the multi-stage attention correlation learning method, combined with attention correlation analysis and cross-connection fusion module, the problem of feature differences and insufficient connections in cross-domain face generation is solved, and high-quality visible face images are generated, improving image details and feature alignment effects.

CN116798102BActive Publication Date: 2025-08-08XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310937737.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-27
Publication Date
2025-08-08
Estimated Expiration
2043-07-27

AI Technical Summary

Technical Problem

The existing cross-domain face generation methods have insufficient learning of cross-domain face feature differences and connections, resulting in low image quality, especially in the conversion of thermal imaging face images to visible light faces.

Method used

Using the multi-stage attention correlation learning method, the one-way cross-domain mapping relationship between thermal imaging and visible light is learned through the first adversarial network model, and the cross-domain feature connection between visible faces and thermal imaging images is learned through the second adversarial network model. The attention correlation analysis module and the cross-connection fusion module are used to reduce the cross-domain feature differences and realize cross-domain feature alignment.

Benefits of technology

The quality of cross-domain face generated images is improved, the connection between facial features and facial contours is enhanced, the differences in cross-domain features are reduced, and the generated images are more realistic and rich in details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116798102B_ABST
    Figure CN116798102B_ABST
Patent Text Reader

Abstract

The present invention provides a cross-domain face generation method based on multi-stage attention correlation learning, including: inputting a thermal imaging face image into a first adversarial network model; generating a visible light face image based on a constructed first attention correlation analysis module and a first channel-spatial attention cross-connection fusion module; combining the visible light face image and the thermal imaging face image into a cross-domain image pair; inputting the cross-domain image pair into a second adversarial network model; performing preliminary cross-domain facial feature fusion on the cross-domain image pair; and alternatingly training the first and second adversarial network models; generating a cross-domain visible light composite face image based on a constructed second attention correlation analysis module and a second channel-spatial attention cross-connection fusion module. This method solves the problem of low quality composite face images caused by existing adversarial network models ignoring the differences and connections between cross-domain facial features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image synthesis, and in particular to a cross-domain face generation method based on multi-stage attention correlation learning. Background Art

[0002] Cross-domain face generation involves applying facial features to images through artificial intelligence and machine learning techniques to create a synthetic image. Cross-domain face generation is a hot research topic in generative modeling today. Cross-domain face generation involves many domain transformations, such as converting thermal images to visible light, converting near-infrared images to visible light, and converting portraits to real photos. It is a challenging and realistic research topic. In recent years, both traditional and deep learning methods have made significant progress in cross-domain face generation. Traditional methods typically rely on statistical models or feature representations to bridge the gap between different domains. These methods generally fall into three categories: subspace learning methods, sparse representation methods, and Bayesian inference methods.

[0003] In recent years, face generation methods based on deep learning models have become a research hotspot for cross-domain face generation. Deep learning models include adversarial networks (GANs), which consist of a generator and a discriminator. The generator is responsible for generating synthetic images, while the discriminator determines whether the generated images are realistic. Through continuous iterative training, the generator and discriminator compete with each other, ultimately producing realistic synthetic images.

[0004] Existing GAN models are based on single-stage networks. Cross-domain face generation plays an important role in computer vision applications across various fields, including security, entertainment, and healthcare. However, as a challenging task, existing cross-domain face generation methods lack sufficient in-depth understanding of facial features, facial contours, and color and lighting conditions, resulting in certain limitations. This is particularly true for thermal face images, where each pixel represents temperature information. Thermal face images cannot be aligned pixel-wise with visible light faces, resulting in significant differences in data distribution. Therefore, thermal face images provide relatively little information. Therefore, in practical applications, cross-domain face generation generally fails to achieve ideal results. Summary of the Invention

[0005] The purpose of the embodiments of the present invention is to provide a cross-domain face generation method based on multi-stage attention correlation learning. First, a first adversarial network model uses adversarial game learning to learn the unidirectional cross-domain mapping relationship from thermal imaging images to visible light facial images. Then, a second adversarial network model learns the cross-domain feature connections between visible light facial images and thermal imaging images to form cross-domain image pairs, thereby reducing the cross-domain feature differences between cross-domain image pairs and achieving cross-domain feature alignment. This method solves the problem of low quality face synthesis images caused by existing adversarial network models ignoring the differences and connections between cross-domain facial features.

[0006] The specific technical solutions are as follows:

[0007] A cross-domain face generation method based on multi-stage attention correlation learning, including:

[0008] Inputting the thermal imaging face image into the first adversarial network model, the first adversarial network model generates a visible light face image based on the constructed first attention correlation analysis module and the first channel-spatial attention cross-connection fusion module;

[0009] Combining the visible light facial image and the thermal imaging facial image into a cross-domain image pair;

[0010] Inputting the cross-domain image pair into a second adversarial network model, and the second adversarial network model performs preliminary cross-domain facial feature fusion on the cross-domain image pair;

[0011] The alternating training of the first adversarial network model and the second adversarial network model is started. Based on the constructed second attention correlation analysis module and the second channel-spatial attention cross-connection fusion module, the second adversarial network model generates a cross-domain visible light face synthetic image.

[0012] In one embodiment of the present invention, before the thermal imaging face image is input into the first adversarial network model, a thermal imaging image is first acquired and preprocessed into the thermal imaging face image;

[0013] The preprocessing of thermal imaging images includes: adjusting image size, dividing training set and test set, and image data enhancement.

[0014] In one embodiment of the present invention, the preprocessing process of the thermal imaging image includes:

[0015] Adjust the image size of thermal imaging images;

[0016] Normalize and enhance thermal imaging images;

[0017] The preprocessing of thermal imaging images is completed to obtain the training set and test set.

[0018] In one embodiment of the present invention, the first attention correlation analysis module and the first channel-spatial attention cross-connection fusion module are constructed, including:

[0019] Build the attention score scoring of the channel attention module and the spatial attention module corresponding to the first attention correlation analysis module, set up the ACL attention correlation loss module to calculate and improve the correlation between the two feature score maps through backpropagation;

[0020] Building a first channel-spatial attention cross-connection fusion module, in which the two feature score maps obtained by the channel attention module and the spatial attention module are cross-connected to extract attention information;

[0021] The construction of the first attention correlation analysis module and the first channel-spatial attention cross-connection fusion module is completed, and the core module of the generator in the first adversarial network model is obtained.

[0022] In one embodiment of the present invention, the first adversarial network model generates a visible light face image, including:

[0023] The thermal imaging face image X is input into the first adversarial network model, and the generator of the first adversarial network model is used as G. The formula for generating a visible light face image is as follows:

[0024]

[0025] In formula (1), G(X) represents the visible light face image generated by the first adversarial network model; X is the thermal imaging face image, It means that the visible light face image has n different resolutions. When n=3, The visible light face image is a 128*128 resolution image. The visible light face image is a 64*64 resolution image. The visible light face image with a resolution of 32*32;

[0026] The discriminator D of the first adversarial network model has n independent discriminators D i , i=1…n;

[0027] The adversarial loss formula between the generator G and the discriminator D of the first adversarial network model is as follows:

[0028]

[0029] In formula (2), min G max DV(D,G) means maximizing the loss from the perspective of the discriminator and minimizing the loss from the perspective of the generator, so that the discriminator and the generator can achieve confrontation while sharing the loss; where E represents expectation, represents the expectation of logD(x) when all x are real data; represents the expectation of log(1-D(G(z)) when all data are generated data; by sharing the loss function of formula (2), the generator and discriminator of the first adversarial network model can compete with each other during the training process.

[0030] In one embodiment of the present invention, forming a cross-domain image pair from the visible light facial image and the thermal imaging facial image includes:

[0031] A cross-domain fusion block CDF is constructed to perform preliminary cross-domain information fusion. The generated target domain image and the input source domain image are taken as a cross-domain input pair and combined in the cross-domain fusion block CDF as auxiliary image information and main image information respectively.

[0032] In one embodiment of the present invention, starting alternating training of the first adversarial network model and the second adversarial network model includes:

[0033] Before the 100th round, train the first adversarial network model;

[0034] Starting from the 100th round, the first adversarial network model and the second adversarial network model are trained alternately;

[0035] The first adversarial network model and the second adversarial network model perform parameter learning through error back propagation.

[0036] In one embodiment of the present invention, the second attention correlation analysis module and the second channel-spatial attention cross-connection fusion module are constructed, including:

[0037] Constructing a second attention correlation analysis module for the second adversarial network model. This module is designed by combining the attention mechanism and correlation analysis. It includes two Self-Attention Blocks with the same structure. It increases the correlation between cross-domain image features by calculating ACL and backpropagating.

[0038] Constructing a second channel-spatial attention cross-connection fusion module of the second adversarial network model; the second channel-spatial attention cross-connection fusion module performs attention scoring through the attention mechanism of the channel attention dimension and the spatial attention dimension;

[0039] The second channel-spatial attention cross-connection fusion module calculates the correlation between different attention score maps through attention correlation loss.

[0040] In one embodiment of the present invention, the first adversarial network model generates n target domain face images with the same scale readings as the second adversarial network model, using the formula:

[0041]

[0042] In formula (3), G'(X) in G'(X) represents the entire second adversarial network model, and Y represents the input cross-domain image pair; It shows that the cross-domain image has n different resolutions. When n=3, The cross-domain image is of 128*128 resolution, The cross-domain image is a 64*64 resolution image. The cross-domain image has a resolution of 32*32.

[0043] In one embodiment of the present invention, the second adversarial network model generates a cross-domain visible light face synthetic image, including:

[0044] Inputting the target domain features obtained by the target domain encoder of the second adversarial network model and the cross-domain visible light face synthesis image into the discriminator of the second adversarial network model for discrimination;

[0045] The output of the discriminator is divided into two branches, and the discriminator sends the preliminary features obtained by downsampling the image to the two branches respectively;

[0046] In the first branch, the preliminary features of the cross-domain visible light face synthesis image are further downsampled to determine the authenticity of the image. In the second branch, the target domain features and the preliminary features are spliced in the channel dimension.

[0047] The discriminator judges the authenticity and domain distribution matching of the image through downsampling.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] 1. The present invention first uses a first adversarial network model to learn the unidirectional cross-domain mapping relationship from thermal imaging images to visible light facial images through adversarial game, and then uses a second adversarial network model to learn the cross-domain feature connection of visible light facial images and thermal imaging images to form a cross-domain image pair, thereby reducing the cross-domain feature differences of the cross-domain image pairs, thereby achieving the purpose of cross-domain feature alignment, and solving the problem that the existing adversarial network model ignores the differences and connections between cross-domain facial features, resulting in low quality of facial synthesis images.

[0050] 2. The first adversarial network model and the second adversarial network model of the present invention both combine the attention mechanism and the correlation analysis idea to construct a first attention correlation analysis module, a first channel-spatial attention cross-connection fusion module, a second attention correlation analysis module and a second channel-spatial attention cross-connection fusion module to remove background noise interference from facial images and improve the connection between facial features and the correlation between cross-domain features.

[0051] 3. The first attention correlation analysis module of the present invention performs attention scoring through attention mechanisms of two different dimensions, channel attention and spatial attention, and calculates the correlation between different attention score maps through attention correlation loss, thereby improving the connection between the facial features and facial contours of the thermal imaging input face image; the second attention correlation analysis module and the second channel-spatial attention cross-connection fusion module of the present invention perform attention scoring on cross-domain face images through two identical self-attention modules, and increase the cross-domain feature connection through attention correlation loss, thereby reducing the difference in cross-domain features. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 A flow chart of a cross-domain face generation method based on multi-stage attention correlation learning provided by the present invention;

[0053] Figure 2 Structural diagrams of the first adversarial network model AC-GAN and the second adversarial network model MAC-GAN provided by the present invention;

[0054] Figure 3 This is a structural diagram of the channel-spatial attention cross-connection fusion module CSACF provided by the present invention;

[0055] Figure 4 This is the CDF structure diagram of the cross-domain fusion block provided by the present invention;

[0056] Figure 5 Flowchart of the discriminator 2 provided by the present invention;

[0057] Figure 4 In the image, Ia is auxiliary image information and Im is main image information. DETAILED DESCRIPTION

[0058] The present invention is described in detail below with reference to the various embodiments shown in the accompanying drawings, but it should be noted that these embodiments are not limitations of the present invention, and any equivalent transformations or substitutions in functions, methods, or structures made by ordinary technicians in this field based on these embodiments are all within the scope of protection of the present invention.

[0059] Example 1:

[0060] This paper aims to solve the problem that the existing generative adversarial network model only learns one-way cross-domain mapping in a single-stage learning process, but does not fully learn the differences and connections between cross-domain features, and does not learn much cross-domain facial feature information. Figure 1 , an embodiment of the present invention provides a cross-domain face generation method based on multi-stage attention correlation learning, comprising:

[0061] Step S110: Input the thermal imaging face image into the first adversarial network model. Based on the constructed first attention correlation analysis module and the first channel-spatial attention cross-connection fusion module, the first adversarial network model generates a visible light face image.

[0062] The first attention correlation analysis module and the first channel-spatial attention cross-connection fusion module constructed in step S110 above include:

[0063] Step S111: Building the attention score scoring of the channel attention module and the spatial attention module in the first attention correlation analysis module, setting the ACL attention correlation loss module to calculate and improve the correlation between the two feature score maps through back propagation;

[0064] Step S112: Building a first channel-spatial attention cross-connection fusion module, in which the two feature score maps obtained by the channel attention module and the spatial attention module are cross-connected to extract attention information;

[0065] Step S113, end the construction of the first attention correlation analysis module and the first channel-spatial attention cross-connection fusion module, and obtain the core module of the generator in the first adversarial network model.

[0066] In practical application, step S110 constructs the attention correlation analysis module HAFCA and the channel-spatial attention cross connection fusion module CSACF, including the following steps:

[0067] Step (1) Build the channel attention module CAM and the spatial attention module SAM to score the attention scores of two different dimensions, set the ACL attention correlation loss module to calculate and improve the correlation between the two feature score maps through backpropagation. In CSACF, CAM and SAM are cross-connected to further improve the acquired attention information.

[0068] If the input is X, set and Represents the two sets of extracted single-domain features. Where B is the batch size, C, H, and W are the number of channels, the height and width of the feature map, respectively. X1 and X2 are input into the spatial attention module (SAM) and channel attention module (CAM) under the hybrid attention feature correlation analysis module (HAFCA), respectively, to simulate the semantic dependencies in the spatial and channel dimensions. In SAM, X1 obtains a spatial attention map with global context semantic information and selectively aggregated context. This spatial attention map represents the face distribution attention map after removing background interference. The formula is as follows:

[0069]

[0070] In formula (4), ɑ is initialized to 0 and gradually learns to assign more weights. P, Q, and R are feature maps obtained by three different convolutional layers, where i, j are the position indices of each feature map; N = H·W represents the number of pixels. is the spatial attention map.

[0071] In CAM, X2 explicitly models the interdependencies between channels and improves the feature representation of specific semantics through a series of operations. The formula is as follows:

[0072]

[0073] where β gradually learns weights from zero. represents the channel attention map.

[0074] In order to update the weight parameters of CAM, SAM and encoder, the attention correlation loss between the two attention maps is calculated and back-propagated. Let C 11 , C 22 is the covariance of U and V, then the cross covariance of U, V is C 12 .set up Update U,V by calculating the relevant loss as follows:

[0075]

[0076] Afterwards, the encoder, CAM, and SAM parameters can be updated by minimizing the attention CCA loss, i.e., maximizing the correlation between U and V.

[0077] Step (2) builds a channel-spatial attention cross-connection fusion module CSACF, cross-connects the feature maps U and V obtained in (1) to extract attention information, and the formula is as follows:

[0078]

[0079] in Represents element-wise multiplication; It represents the intelligent summation of elements; E2 is the final fusion output; C is the channel attention module; S is the spatial attention module; hafca has a sam and a cam sub-module, and the attention score maps output by these two sub-modules are u and v.

[0080] Step (3) completes the construction of the attention correlation analysis module HAFCA and the channel-spatial attention cross-connection fusion module CSACF, and obtains the core module of the generator in the first stage.

[0081] In step S110, the first adversarial network model generates a visible light face image, including:

[0082] Step S114: The thermal imaging face image X is input into the first adversarial network model, and the generator of the first adversarial network model is used as G. The formula for generating a visible light face image is as follows:

[0083]

[0084] In formula (1), G(X) represents the visible light face image generated by the first adversarial network model; X is the thermal imaging face image, It means that there are n different resolutions of visible light face images. When n=3, It is a visible light face image with a resolution of 128*128. It is a visible light face image with a resolution of 64*64. A visible light face image with a resolution of 32*32;

[0085] Step S115: The discriminator D of the first adversarial network model has n independent discriminators D i , i=1…n;

[0086] The adversarial loss formula between the generator G and the discriminator D of the first adversarial network model is as follows:

[0087]

[0088] In formula (2), min G max D V(D,G) means maximizing the loss from the perspective of the discriminator and minimizing the loss from the perspective of the generator, so that the discriminator and the generator can achieve confrontation while sharing the loss; where E represents expectation, represents the expectation of logD(x) when all x are real data; represents the expectation of log(1-D(G(z)) when all data are generated data; by sharing the loss function of formula (2), the generator and discriminator of the first adversarial network model can compete with each other during the training process.

[0089] Step S120: Combine the visible light face image and the thermal imaging face image into a cross-domain image pair.

[0090] In the above step S120, the visible light face image and the thermal imaging face image are combined into a cross-domain image pair, including: constructing a cross-domain fusion block CDF to perform preliminary cross-domain information fusion, and combining the generated target domain image and the input source domain image as a cross-domain input pair, as auxiliary image information and main image information respectively, in the cross-domain fusion block CDF.

[0091] Step S130: input the cross-domain image pair into the second adversarial network model, and the second adversarial network model performs preliminary cross-domain facial feature fusion on the cross-domain image.

[0092] In the embodiment of the present invention, the first adversarial network model generates n target domain face images with the same scale readings as the second adversarial network model, and the formula is:

[0093]

[0094] In formula (3), G'(X) represents the entire second adversarial network model, and Y represents the input cross-domain image pair; It shows that there are n different resolutions of cross-domain images. When n=3, For a cross-domain image with a resolution of 128*128, For cross-domain images with a resolution of 64*64, It is a cross-domain image with a resolution of 32*32.

[0095] Step S140: Start alternating training of the first adversarial network model and the second adversarial network model. Based on the constructed second attention correlation analysis module and the second channel-spatial attention cross-connection fusion module, the second adversarial network model generates a cross-domain visible light face synthetic image.

[0096] Initiating alternating training of the first adversarial network model and the second adversarial network model in step S140 includes:

[0097] Step S141: Before the 100th round, train the first adversarial network model;

[0098] Step S142: Starting from the 100th round, the first adversarial network model and the second adversarial network model are trained alternately;

[0099] Step S143: The first adversarial network model and the second adversarial network model perform parameter learning through error back propagation.

[0100] The second attention correlation analysis module and the second channel-spatial attention cross-connection fusion module constructed in step S140 include:

[0101] Step S144: Construct a second attention correlation analysis module of the second adversarial network model. The second attention correlation analysis module is designed by combining the attention mechanism and correlation analysis ideas. The second attention correlation analysis module includes two Self-AttentionBlock attention mechanisms with the same structure. The correlation between cross-domain image features is increased by calculating ACL and backpropagating.

[0102] Step S145: construct a second channel-spatial attention cross-connection fusion module of the second adversarial network model; the second channel-spatial attention cross-connection fusion module performs attention scoring through the attention mechanism of the channel attention dimension and the spatial attention dimension;

[0103] Step S146: The second channel-spatial attention cross-connection fusion module calculates the correlation between different attention score maps through attention correlation loss.

[0104] In step S140, the second adversarial network model generates a cross-domain visible light face composite image, including:

[0105] Step S147: input the target domain features obtained by the target domain encoder of the second adversarial network model and the cross-domain visible light face synthetic image into the discriminator of the second adversarial network model for discrimination;

[0106] Step S148: The output of the discriminator is divided into two branches, and the discriminator sends the preliminary features obtained by downsampling the image to the two branches respectively;

[0107] Step S149: In the first branch, further downsample the preliminary features of the cross-domain visible light face synthesis image to determine the authenticity of the image. In the second branch, splice the target domain features and the preliminary features in the channel dimension.

[0108] Step S1410: The discriminator determines the authenticity and domain distribution matching of the image through downsampling.

[0109] It should be noted that, before the thermal imaging face image is input into the first adversarial network model in step S110 of the embodiment of the present invention, the thermal imaging image is first acquired and preprocessed into a thermal imaging face image; wherein the preprocessing of the thermal imaging image includes: adjusting the image size, dividing the training set and the test set, and image data enhancement.

[0110] The preprocessing process of the above thermal imaging images includes:

[0111] Adjust the image size of thermal imaging images;

[0112] Normalize and enhance thermal imaging images;

[0113] The preprocessing of thermal imaging images is completed to obtain the training set and test set.

[0114] The embodiment of the present invention generates higher quality target domain face images by designing a two-stage GAN network. The first stage is used to learn a cross-domain unidirectional mapping relationship, and the second stage is used to further learn the connections and differences between cross-domain features based on the learned cross-domain mapping relationship, and optimize the learned cross-domain mapping relationship. It is proposed to combine the attention mechanism and the correlation analysis idea to design a hybrid attention correlation analysis module to remove the background noise of the face image and focus the model's attention on the facial features and facial contours. More specifically, (1) Since the existing GAN model is based on a single-stage network, the generator and discriminator networks are trained, and only the unidirectional mapping relationship of cross-domain faces is learned, ignoring the differences and connections between cross-domain facial features, the quality of the generated target domain face images is not high. The embodiment of the present invention designs a two-stage GAN model. The first stage learns the unidirectional cross-domain mapping relationship from thermal imaging to visible light face through generative adversarial game. The second stage forms a cross-domain input pair by combining the generated visible light face and the input thermal imaging face, learns the cross-domain feature connection, reduces the cross-domain feature difference, and thus achieves the purpose of cross-domain feature alignment. (2) Since the main modules of the existing generators do not fully learn the correlation between cross-domain features, the impact of background noise in facial images on cross-domain face generation is ignored. The embodiment of the present invention combines the attention mechanism and the correlation analysis idea to design a hybrid attention correlation analysis module HAFCA. The first stage of HAFCA uses two different dimensions of attention mechanisms, channel attention and spatial attention, to score attention. The correlation between different attention score maps is calculated through attention correlation loss, thereby improving the connection between the facial features and facial contours of the thermal imaging input face image. The second stage of HAFCA uses two identical self-attention modules to score the cross-domain face image and increases the cross-domain feature connection through attention correlation loss, thereby reducing the difference in cross-domain features. (3) Since the discriminator is only used to judge the generated face and the real face, it lacks certain consideration of whether the face conforms to the target domain feature distribution. At the same time, there are few publicly available cross-domain face datasets from thermal imaging to visible light faces. In an embodiment of the present invention, an output stream is added to the second-stage discriminator to further judge whether the data distribution of the image conforms to the target domain feature distribution, thereby indirectly constraining the generator, thereby obtaining a visible light facial image whose facial features are more consistent with the target domain distribution.

[0115] Through the above-mentioned solution of the embodiment of the present invention, a higher-quality target domain face image can be generated, and the facial features and facial contours can be further refined. Experiments have shown that the present invention well ensures the details of the face image and achieves efficient cross-domain face generation.

[0116] Example 2:

[0117] Based on the solution disclosed in Example 1, this embodiment combines Figure 2 、 Figure 3 、 Figure 4 and Figure 5 , Figure 2 Structural diagram of the first adversarial network model AC-GAN and the second adversarial network model MAC-GAN provided in an embodiment of the present invention; Figure 3 A structural diagram of the channel-spatial attention cross-connection fusion module CSACF provided in an embodiment of the present invention; Figure 4 The CDF structure diagram of the cross-domain fusion block provided by the embodiment of the present invention is as follows: Figure 4 In the image, Ia is the auxiliary image information, and Im is the main image information; Figure 5 Flowchart of the discriminator 2 provided in an embodiment of the present invention. A cross-domain face generation method based on multi-stage attention correlation learning is disclosed in an embodiment of the present invention. By learning the unidirectional cross-domain mapping relationship from thermal imaging to visible light in the first stage, the connections and differences between cross-domain features are learned in the second stage. The attention mechanism and the correlation analysis idea are combined to remove background noise interference from facial images, improve the connection between facial features and the correlation between cross-domain features. The generation capability of the generator is enhanced through feature domain identification by the discriminator. The following steps are included:

[0118] Step 100: Start cross-domain face generation based on a multi-stage attention correlation analysis generative adversarial network.

[0119] Step 200: Construct an attention correlation analysis module and a channel-spatial attention cross-connection fusion module to preprocess the thermal imaging face image to be input, such as adjusting the image size to divide it into training set and test set, and performing data enhancement such as adding Gaussian noise, flipping left and right, and adjusting contrast.

[0120] More specifically, step 200 includes the following steps:

[0121] Step 210: Construct an attention correlation analysis module HAFCA and a channel-spatial attention cross-connection fusion module CSACF. It should be noted that step 210 includes:

[0122] Step 211: Build the channel attention module (CAM) and the spatial attention module (SAM) to score attention scores in two different dimensions. Set up the ACL attention correlation loss module to calculate and improve the correlation between the two feature score maps through backpropagation. In CSACF, cross-connect the CAM and SAM to further improve the acquired attention information.

[0123] If the input is X, set and Represents the two sets of extracted single-domain features. Where B is the batch size, C, H, and W are the number of channels, the height and width of the feature map, respectively. X1 and X2 are input into the spatial attention module (SAM) and channel attention module (CAM) under the hybrid attention feature correlation analysis module (HAFCA), respectively, to model the semantic dependencies in the spatial and channel dimensions. In SAM, X1 obtains a spatial attention map with global context semantic information and selectively aggregated context. This spatial attention map represents the face distribution attention map after removing background interference. The formula is as follows:

[0124]

[0125] α is initialized to 0 and gradually learns to assign more weights. P, Q, and R are feature maps obtained from three different convolutional layers. i,j are the position indexes of each feature map. N=

[0126] HW represents the number of pixels, is the spatial attention map.

[0127] In CAM, X2 explicitly models the interdependencies between channels and improves the feature representation of specific semantics through a series of operations. The formula is as follows:

[0128]

[0129] where β gradually learns weights from zero. represents the channel attention map.

[0130] To update the weight parameters of CAM, SAM and encoder, we compute and backpropagate the attention correlation loss between the two attention maps. Let C 11 , C 22 is the covariance of U and V, then the cross covariance of U, V is C 12 .set up We update U,V by calculating the relevant loss as follows:

[0131]

[0132] Afterwards, the encoder, CAM, and SAM parameters can be updated by minimizing the attention CCA loss, i.e., maximizing the correlation between U and V.

[0133] Step 212: Build a channel-spatial attention cross-connection fusion module CSACF to cross-connect the feature maps U and V obtained in step 211 to extract attention information. The formula is as follows:

[0134]

[0135] in represents element-wise multiplication, represents the element-wise summation, E2 is the final fusion output, C is the channel attention module, and S is the spatial attention module.

[0136] Step 213: End the construction of the attention correlation analysis module HAFCA and the channel-spatial attention cross-connection fusion module CSACF to obtain the core module of the generator in the first stage.

[0137] Step 220: Preprocess the input image data.

[0138] Step 230: normalize the image by channel and perform data enhancement, such as left-right flipping of the portrait, adding random Gaussian noise, and adjusting contrast and brightness.

[0139] Step 240: End the preprocessing of the input image to obtain the training set and the test set.

[0140] Step 300: Train a single-stage (first stage) AC-GAN for 100 rounds, so that the visible light face image generated in the first stage is as close as possible to the target domain face, thereby providing more target domain feature information for the second stage.

[0141] More specifically, step 300 includes:

[0142] For the input thermal imaging face image X, the generator of AC-GAN is used as G, and the formula for generating the target domain face image is as follows:

[0143]

[0144] Indicates that the generated target domain faces have n different resolutions, and n=3 is set, where For a 128*128 face image, is 64*64, is 32*32, X is the input image. The structure of the AC-GAN discriminator has no dropout. It needs to distinguish the generated images of n different scales. Therefore, we design n independent discriminators D i , i = 1…n. Finally, the size of the bottleneck feature map obtained by the discriminator downsampling is consistent. The adversarial loss of stage 1 is as follows:

[0145]

[0146] Step 400: The visible light face generated in the first stage and the thermal imaging input face are combined into a cross-domain face input pair, which is input into the second stage and subjected to preliminary cross-domain feature fusion. Then, alternating training of the first and second stages is started.

[0147] More specifically, step 400 includes:

[0148] Step 410: construct a cross-domain fusion block CDF to perform preliminary cross-domain information fusion, and use the generated target domain image and the input source domain image as a cross-domain input pair, and combine them in the CDF as auxiliary image information and main image information respectively.

[0149] Step 420: Starting from the 100th round, the first and second phases of the MAC-GAN model are trained alternately, and the entire network model learns parameters through error back propagation.

[0150] Step 500: Construct the attention correlation analysis module and the channel-spatial attention cross-connection fusion module of the second stage, retain the target domain features obtained by the visible light encoder for use by the discriminator, and finally obtain the visible light face image generated in the second stage.

[0151] More specifically, step 500 includes:

[0152] Step 510: Construct the second-stage attention correlation analysis module, in which the attention modules are changed from SAM and CAM to two Self-AttentionBlocks with the same structure. By calculating ACL and backpropagating, the correlation between cross-domain image features is increased and the cross-domain feature differences are reduced, thereby achieving the purpose of cross-domain feature alignment.

[0153] Step 520: MAC-GAN generates n target domain face images with the same scale as AC-GAN.

[0154]

[0155] Finally, after multi-stage training, MAC-GAN obtains more detailed and realistic target domain faces than AC-GAN.

[0156] Step 530: The target domain features obtained by the target domain encoder are input into the discriminator along with the image for discrimination. Here, the discriminator's output is divided into two branches. The discriminator sends the preliminary features obtained by downsampling the image to each branch. In the first branch, the preliminary features are further downsampled to determine the image's authenticity. In the second branch, the domain features and preliminary features are concatenated along the channel dimension. Finally, patches are obtained through downsampling to determine the image's authenticity and domain distribution match.

[0157] Step 540: Train the entire MAC-GAN.

[0158] Step 600: Input the test thermal imaging face image data into the trained network model MAC-GAN, and end the cross-domain face generation based on the multi-stage attention correlation analysis generative adversarial network.

[0159] More specifically, the test thermal imaging face image data is input into the network model trained in step 540, and finally a high-quality target domain face image is obtained.

[0160] This embodiment of the present invention incorporates multi-stage learning into the GAN model. The first stage is used to learn a cross-domain unidirectional mapping relationship. The second stage combines the generated image with the input image to form a cross-domain input pair to analyze cross-domain feature connections, reduce cross-domain feature differences, and achieve cross-domain feature alignment. This multi-stage learning can produce higher-quality and more detailed target domain facial images. This embodiment of the present invention combines the attention mechanism with correlation analysis to construct an attention correlation analysis module. The first stage enhances the connection between facial features in thermal imaging by improving the correlation between attention score feature maps of different dimensions. The second stage reduces cross-domain feature differences by improving the correlation between cross-domain feature score maps, achieving feature alignment. To enhance the generative capabilities of the generator, this embodiment of the present invention uses a zero-sum game between the generator and the discriminator. Therefore, enhancing the discriminator's capabilities indirectly helps enhance the generator's generation capabilities. By identifying the distribution of facial features in the image, the generator is improved in generating facial images that are more consistent with the target domain distribution, thereby improving the network's generation capabilities.

[0161] The embodiment of the present invention uses the idea of multi-stage learning to optimize the cross-domain unidirectional mapping relationship, and on this basis, increases the learning of the connection between cross-domain features, improves the network generation ability, optimizes the facial contour details of the facial features, and reduces ID differences. In order to increase the network's learning of the details of the facial features and the connection between the facial features, thereby reducing problems such as misalignment, blurring and distortion in the generated facial features, the embodiment of the present invention combines the attention mechanism and the idea of correlation analysis, removes the influence of background noise information in the facial image, and focuses on the connection between the facial features and the correlation learning of cross-domain features, thereby improving the learning ability of the network. The embodiment of the present invention has the ability to generalize across data. Due to the maximization of cross-domain feature correlation and effective feature representation, the embodiment of the present invention has good generation effects on cross-datasets. At the same time, the model also performs well in the extended experiment, namely the near-infrared to visible light face task, which reflects the robustness of the model. The embodiment of the present invention designs the two-stage discriminator as a triplet discriminator, while improving the generation capability of the generator, avoiding simple style transfer such as coloring the input face image, and focusing on learning the target domain data distribution and facial features as well as cross-domain connections, thereby improving the interpretability of the network model.

[0162] The embodiment of the present invention can generate a target domain face image with more obvious facial features, clearer facial contours, and smaller ID differences.

[0163] The series of detailed descriptions listed above are only specific descriptions of feasible implementation methods of the present invention. They are not intended to limit the scope of protection of the present invention. Any equivalent implementation methods or changes that do not deviate from the technical spirit of the present invention should be included in the scope of protection of the present invention.

[0164] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

[0165] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

Claims

1. A cross-domain face generation method based on multi-stage attention correlation learning, characterized in that: include: Inputting the thermal imaging face image into the first adversarial network model, the first adversarial network model generates a visible light face image based on the constructed first attention correlation analysis module and the first channel-spatial attention cross-connection fusion module; Combining the visible light facial image and the thermal imaging facial image into a cross-domain image pair; Inputting the cross-domain image pair into a second adversarial network model, and the second adversarial network model performs preliminary cross-domain facial feature fusion on the cross-domain image pair; Alternating training of the first adversarial network model and the second adversarial network model is started. Based on the constructed second attention correlation analysis module and the second channel-spatial attention cross-connection fusion module, the second adversarial network model generates a cross-domain visible light face synthetic image. The constructed first attention correlation analysis module and first channel-spatial attention cross-connection fusion module include: Build the attention score scoring of the channel attention module and the spatial attention module corresponding to the first attention correlation analysis module, set up the ACL attention correlation loss module to calculate and improve the correlation between the two feature score maps through backpropagation; Building a first channel-spatial attention cross-connection fusion module, in which the two feature score maps obtained by the channel attention module and the spatial attention module are cross-connected to extract attention information; Complete the construction of the first attention correlation analysis module and the first channel-spatial attention cross-connection fusion module to obtain the core module of the generator in the first adversarial network model; The constructed second attention correlation analysis module and second channel-spatial attention cross-connection fusion module include: Constructing a second attention correlation analysis module for the second adversarial network model. This module is designed by combining the attention mechanism and correlation analysis. It includes two Self-Attention Blocks with the same structure. It increases the correlation between features of cross-domain image pairs by calculating the ACL and backpropagating it. Constructing a second channel-spatial attention cross-connection fusion module of the second adversarial network model; the second channel-spatial attention cross-connection fusion module performs attention scoring through the attention mechanism of the channel attention dimension and the spatial attention dimension; The second channel-spatial attention cross-connection fusion module calculates the correlation between different attention score maps through attention correlation loss.

2. The cross-domain face generation method based on multi-stage attention correlation learning according to claim 1 is characterized in that Before the thermal imaging face image is input into the first adversarial network model, a thermal imaging image is first acquired and preprocessed into the thermal imaging face image; The preprocessing of thermal imaging images includes: adjusting image size, dividing training set and test set, and image data enhancement.

3. The cross-domain face generation method based on multi-stage attention correlation learning according to claim 2 is characterized in that: The preprocessing process of thermal imaging images includes: Adjust the image size of thermal imaging images; Normalize and enhance thermal imaging images; The preprocessing of thermal imaging images is completed to obtain the training set and test set.

4. The cross-domain face generation method based on multi-stage attention correlation learning according to claim 1 is characterized in that The first adversarial network model generates a visible light face image, including: The thermal imaging face image X is input into the first adversarial network model, and the generator of the first adversarial network model is used as G. The formula for generating a visible light face image is as follows: Formula (1) In formula (1), G(X) represents the visible light face image generated by the first adversarial network model; X is the thermal imaging face image, It means that the visible light face image has n different resolutions. When n=3, The visible light face image is a 128*128 resolution image. The visible light face image is a 64*64 resolution image. The visible light face image with a resolution of 32*32; The discriminator D of the first adversarial network model has n independent discriminators , i=1…n; The adversarial loss formula between the generator G and the discriminator D of the first adversarial network model is as follows: Formula (2) In formula (2), It means maximizing the loss from the perspective of the discriminator and minimizing the loss from the perspective of the generator, so that the discriminator and the generator can achieve confrontation under the condition of shared loss; where E represents expectation, Indicates that all x are real data expectations; Indicates that all data is generated data expectation; z is the input noise of the generator G; G(z) represents the forged image synthesized by the generator G based on the input noise z; by sharing the loss function of formula (2), the generator and discriminator of the first adversarial network model can compete with each other during the training process.

5. The cross-domain face generation method based on multi-stage attention correlation learning according to claim 1 is characterized in that Starting alternating training of the first adversarial network model and the second adversarial network model includes: Before the 100th round, train the first adversarial network model; Starting from the 100th round, the first adversarial network model and the second adversarial network model are trained alternately; The first adversarial network model and the second adversarial network model perform parameter learning through error back propagation.

6. The cross-domain face generation method based on multi-stage attention correlation learning according to claim 1, characterized in that: The first adversarial network model generates n target domain face images with the same scale reading as the second adversarial network model, and the formula is: Formula (3) In formula (3), G'(Y) represents the entire second adversarial network model, and Y represents the input cross-domain image pair; Indicates that the cross-domain image pair has n different resolutions. When n=3, The cross-domain image pair is of 128*128 resolution, The cross-domain image pair is a 64*64 resolution image. The cross-domain image pair has a resolution of 32*32.

7. The cross-domain face generation method based on multi-stage attention correlation learning according to claim 1 is characterized in that The second adversarial network model generates a cross-domain visible light face synthetic image, including: Inputting the target domain features obtained by the target domain encoder of the second adversarial network model and the cross-domain visible light face synthesis image into the discriminator of the second adversarial network model for discrimination; The output of the discriminator is divided into two branches, and the discriminator sends the preliminary features obtained by downsampling the image to the two branches respectively; In the first branch, the preliminary features of the cross-domain visible light face synthesis image are further downsampled to determine the authenticity of the image. In the second branch, the target domain features and the preliminary features are spliced in the channel dimension. The discriminator judges the authenticity and domain distribution matching of the image through downsampling.

Citation Information

Patent Citations

  • Generative adversarial network-based grayscale picture colorizing method

    CN108711138A

  • Infrared and visible light image fusion method based on coupling generative adversarial network

    CN112488970A