Generative Adversarial Network and Generation Method for Face Super-Resolution Based on Pose Constraints

By using a face super-resolution generative adversarial network based on pose constraints, combined with a side-face attention guidance and attention fusion transformation module, high-resolution face images can be generated efficiently under pose changes. This solves the problem of poor restoration effect of existing methods under pose changes and improves image quality and recognition accuracy.

CN119359544BActive Publication Date: 2025-12-02SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411673676.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-21
Publication Date
2025-12-02
Estimated Expiration
2044-11-21

AI Technical Summary

Technical Problem

Existing face super-resolution methods based on generative adversarial networks struggle to achieve satisfactory restoration results when faced with low-resolution face images exhibiting significant pose variations, resulting in a substantial decrease in model performance.

Method used

A generative adversarial network (GAN) for face super-resolution based on pose constraints is proposed, which includes a profile attention guidance module, an attention fusion transformation module, a guidance transformation sub-network, and a discriminant network. The GAN performs pose correction and image reconstruction to generate high-resolution face images.

Benefits of technology

Under low resolution and different pose conditions, it can effectively restore high-quality frontal face images, improve the realism and robustness of the images, and adapt to changes in facial expressions and lighting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119359544B_ABST
    Figure CN119359544B_ABST
Patent Text Reader

Abstract

This invention discloses a generative adversarial network (GAN) and generation method for face super-resolution based on pose constraints. It includes at least a profile attention guidance module, an attention fusion transformation module, a guidance transformation sub-network, an upsampling sub-network, and a discriminant network. A unified GAN is established to jointly perform face super-resolution and pose correction tasks, restoring a realistic high-resolution face image while simultaneously correcting the face pose. This invention can simultaneously achieve face super-resolution reconstruction and pose correction, restoring a realistic high-resolution frontal face image from a low-resolution profile face image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of face restoration, and mainly relates to a face super-resolution generative adversarial network and generation method based on pose condition constraints. Background Technology

[0002] In real-world scenarios, facial recognition technology plays an increasingly important role in various fields, including traffic management, law enforcement, and public surveillance. On one hand, cameras installed at traffic intersections not only identify license plate numbers but also collect facial image data of drivers for comparison, facilitating traffic authorities in verifying the identities of traffic violators. On the other hand, the nation's emphasis on urban security has led to the widespread deployment of surveillance cameras throughout cities. These cameras can not only capture various illegal and criminal activities but also quickly track criminal suspects using advanced facial recognition technology, providing strong law enforcement assistance to public security departments. However, facial recognition technology often relies on high-quality facial image input. Facial images obtained in uncontrolled scenarios such as traffic intersections typically have low resolution, directly leading to a decrease in the accuracy and reliability of facial recognition algorithms. To address this issue, facial super-resolution technology has emerged. Facial super-resolution involves reconstructing high-resolution facial images with rich details and high clarity by refining and enhancing low-resolution facial images. Improving image quality through super-resolution technology can effectively enhance the accuracy and reliability of facial recognition.

[0003] In recent years, with the rise of deep learning, neural network-based face super-resolution methods have made significant progress. Among them, Generative Adversarial Networks (GANs) can generate high-quality face images without explicit assumptions about the parameter distribution of face images. Therefore, more and more researchers have proposed face super-resolution methods based on GANs. GAN-based face super-resolution methods typically consist of two stages: adversarial training and generation. The adversarial training stage involves training the generator network and the discriminator network against each other. The generator network tries to deceive the discriminator network as much as possible, while the discriminator network tries to distinguish between real and fake images reconstructed by the generator network, ultimately making the data distribution generated by the generator network approximate the real distribution. The generation stage uses the trained generator network to reconstruct high-resolution face images from low-resolution face images.

[0004] Because face super-resolution reconstruction targets are highly deterministic, requiring the generation of high-resolution faces with specific identities, generative networks are often conditionally constrained. Existing generative adversarial networks (GANs) have only achieved good results in face super-resolution tasks with frontal face poses. However, in practical applications, due to factors such as shooting angle and face motion, captured face images often have different poses, containing very limited facial feature information. Most existing GAN-based face super-resolution methods assume a fixed face pose, ignoring the impact of pose variations on the reconstruction results. Therefore, when faced with low-resolution face images with significant pose variations, these methods often fail to achieve satisfactory reconstruction results, and the model's performance will significantly degrade. Summary of the Invention

[0005] This invention addresses the limitations and poor performance of existing face super-resolution generation technologies by proposing a generative adversarial network (GAN) and generation method based on pose constraints. The method includes at least a profile attention guidance module, an attention fusion transformation module, a guidance transformation sub-network, an upsampling sub-network, and a discriminant network. A unified GAN is established to jointly perform face super-resolution and pose correction tasks, restoring a realistic high-resolution face image while simultaneously correcting the face pose. This invention can simultaneously achieve face super-resolution reconstruction and pose correction, restoring a realistic high-resolution frontal face image from a low-resolution profile face image.

[0006] To achieve the above objectives, the technical solution adopted by this invention is: a face super-resolution generative adversarial network based on pose-constrained features, comprising at least a side-face attention guidance module, an attention fusion transformation module, a guidance transformation sub-network, an upsampling sub-network, and a discriminant network.

[0007] The side profile attention guidance module generates a Gaussian-distributed feature heatmap based on prior facial feature points of the input low-resolution side profile image, which is used to guide the generative adversarial network to focus on key facial organs.

[0008] The attention fusion transformation module consists of channel attention, spatial transformation network and spatial attention sequence, and is used to embed into the generative adversarial network to assist the generative adversarial network in gradually refining the processing of face images, thereby improving the realism and structural similarity of the generated images.

[0009] The guided transform subnetwork adopts a U-Net structure, consisting of a downsampling encoder and an upsampling decoder, and is used to process low-resolution side face images and their facial feature heatmaps, outputting a frontal face feature map; the downsampling encoder and the upsampling decoder are connected to each other through connection layers.

[0010] The upsampling sub-network consists of a deconvolutional layer and an attention fusion transformation module. Finally, a convolutional layer is used to transform the feature map into an RGB face image, which is used to upsample the frontal face feature map, enlarge the image and enhance key information such as subtle facial expressions, textures or contours.

[0011] The discriminative network is used to input generated face images and real face images, and to train the generative adversarial network (GAN) against it, thereby supervising the GAN's generation performance.

[0012] As an improvement of the present invention, the channel attention uses global average pooling to obtain the weight map of each channel through Sigmoid, and multiplies it with the original feature to obtain the feature map after channel attention, thereby realizing the adaptive calibration of features in the channel dimension.

[0013] The spatial transformation network compensates for in-plane rotation, translation, and scaling of the face by predicting and applying affine transformation parameters.

[0014] The spatial attention is achieved by using global max pooling and global average pooling, concatenating and convolving the data according to the channels, and then obtaining the spatial attention weight matrix through the Sigmoid activation function. The weight matrix is ​​then multiplied by the original feature map to obtain the spatial attention feature map.

[0015] As another improvement of the present invention, in the downsampling encoder of the guided transform subnetwork, pooling layers, convolutional layers, and attention fusion transform modules are stacked in sequence, and in the upsampling decoder, connection layers, deconvolutional layers, and attention fusion transform modules are stacked in sequence.

[0016] To achieve the above objectives, the present invention also adopts the following technical solution: a face super-resolution generation method based on pose condition constraints, comprising the following steps:

[0017] S1. Collect and preprocess data: Collect and preprocess face images under different pose conditions. The face images are divided into frontal face images and side face images. Construct training data pairs of low-resolution side face images and high-resolution frontal face images.

[0018] S2. Generate facial feature heatmap: The low-resolution side face image constructed in step S1 is sent to the side face attention guidance module for processing. A Gaussian distributed feature heatmap is generated based on the prior facial feature points and then connected to the original image in the channel direction before being input into the generative adversarial network.

[0019] S3. Correcting facial pose: The facial feature heatmap generated in step S2 is concatenated with the corresponding original image in the channel dimension and used as input. This input is then fed into the guided transformation subnetwork in the generative adversarial network. Combined with the attention fusion transformation module and the skip connections used for multi-scale feature fusion, the pose-corrected facial feature map is obtained.

[0020] S4. Face image reconstruction: The face feature map output by the guided transformation subnetwork in step S3 is fed into the upsampling subnetwork in the generative adversarial network to refine and reconstruct the face feature map after pose correction.

[0021] S5. Model Training: The refined and reconstructed face image obtained in step S4 and the corresponding real face image are fed into the discriminant network to train the generative adversarial network. The loss function is used to supervise the model parameters of the generative adversarial network and improve the network performance.

[0022] S6. Output the result: Input the low-resolution side face image constructed in step S1 into the adversarial network trained in step S5 to reconstruct a high-resolution frontal face image.

[0023] As an improvement of the present invention, the preprocessing in step S1 includes at least face extraction and alignment of the collected face images, and downsampling of the side face images to obtain low-resolution side face images.

[0024] As another improvement of the present invention, the prior facial feature points in step S2 include at least five pairs, with positions corresponding to the center of the left eye, the center of the right eye, the tip of the nose, the left corner of the mouth, and the right corner of the mouth, respectively. A two-dimensional Gaussian distribution is used to fill the values ​​of the pixel points corresponding to the coordinates of each pair of feature points and their nearby pixel points, with the standard deviation set to 0.8.

[0025] As another improvement of the present invention, in the upsampling network of step S4, the upsampling rate of the deconvolution layer is 2, the kernel size is 5×5, and the stride is 2. A convolutional layer is used to transform the feature map into a 128×3×3 RGB face image, with a kernel size of 5×5 and a stride of 2.

[0026] As a further improvement of the present invention, step S5 specifically includes the following steps:

[0027] S51: Use the generated face image as input to the discriminator network, perform forward propagation to obtain the discrimination result, and perform the same operation using the label image. The two discrimination results are used to construct the adversarial loss.

[0028] S52: Use the generated face image as input to the VGG19 network, perform forward propagation to obtain feature maps, and then perform the same operation on the label image to construct the identity preservation loss from the two feature maps.

[0029] S53: Use the l2 loss function to calculate the pixel loss between the generated image and the label image;

[0030] S54: Construct the total loss function from the three losses in steps S51, S52, and S53, and improve the parameters of the generative adversarial network model by minimizing the loss.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] (1) This invention proposes an innovative generative adversarial network that restores a real high-resolution face image while performing face pose correction. Specifically, it proposes a generative adversarial network (GT-GAN) based on pose condition constraints. In the side face attention guidance module, a Gaussian distributed feature heatmap is generated based on prior facial feature points and connected to the original image in the channel dimension before being sent to the network to guide the network to focus on key facial organs of the side face, which is more efficient and accurate.

[0033] (2) This invention embeds the attention fusion transformation module into a generative adversarial network, utilizes the attention mechanism and spatial transformation network, and combines sequential feature enhancement strategies to enhance facial reconstruction details and improve the realism of network-generated images and network performance.

[0034] (3) The generative adversarial network of the present invention can reconstruct high-quality frontal face images well under low resolution and different face pose conditions, and has good robustness to facial expressions and slight lighting changes, thus improving the quality of image reconstruction. Attached Figure Description

[0035] Figure 1 This is an example image of a low-resolution side-view face in Embodiment 2 of the present invention;

[0036] Figure 2 This is an example diagram of the profile attention guidance module of the present invention;

[0037] Figure 3 This is a schematic diagram of the structure of the generative adversarial network of this invention;

[0038] Figure 4 This is a schematic diagram of the attention fusion transformation module of the present invention;

[0039] Figure 5 is a schematic diagram of the face restoration result in Example 2, wherein,

[0040] Figure (a) is a schematic diagram of the input low-resolution side profile face image;

[0041] Figure (b) is a schematic diagram of the result after the method of the present invention;

[0042] Figure (c) is the label image;

[0043] Figure 6 This is a flowchart of the steps of the generation method of the present invention. Detailed Implementation

[0044] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention.

[0045] Example 1

[0046] A face super-resolution generative adversarial network based on pose constraints includes at least a side face attention guidance module, an attention fusion transformation module, a guidance transformation sub-network, an upsampling sub-network, and a discriminant network.

[0047] The profile attention guidance module first locates the coordinates of the eyes, nose, and mouth in the profile image based on prior facial feature points, and fills them with a Gaussian distribution, where the pixel with the mean is the corresponding feature point of the organ, and the standard deviation is set to 0.8. After obtaining the feature heatmap, it is connected to the original image in the channel direction and sent together into the generator network.

[0048] The attention fusion transformation module consists of channel attention, a spatial transformation network, and a spatial attention sequence. Channel attention uses global max pooling and global average pooling, and then uses a sigmoid function to obtain the weight map for each channel. Multiplying this weight map by the original feature map yields the feature map after channel attention, achieving adaptive calibration of features along the channel dimension. The spatial transformation network compensates for in-plane rotation, translation, and scaling of the face by predicting and applying affine transformation parameters. Spatial attention uses global max pooling and global average pooling, and performs concatenation and convolution according to the channels. Then, a sigmoid activation function is used to obtain the spatial attention weight matrix, which is then multiplied by the original feature map to obtain the spatial attention feature map. The guided transformation sub-network adopts a U-Net structure, consisting of a downsampling encoder and an upsampling decoder. The downsampling encoder stacks pooling layers, convolutional layers, and the attention fusion transformation module in sequence, while the upsampling decoder stacks connection layers, deconvolutional layers, and the attention fusion transformation module in sequence. Each layer of the downsampling encoder and upsampling decoder is connected by a connection layer.

[0049] The upsampling subnetwork consists of deconvolutional layers and attention fusion transformation modules, and finally uses a convolutional layer to transform the feature map into an RGB face image.

[0050] The generated face image is used as input to the discriminative network for forward propagation to obtain the discrimination result. The same process is then performed using the labeled image, and the two discrimination results are used to construct the discriminator loss, which in turn yields the adversarial loss. Similarly, the generated face image is used as input to the VGG19 network for forward propagation to obtain feature maps. The same process is then performed using the labeled image, and the two feature maps are used to construct the identity preservation loss. The L2 loss function is used to calculate the pixel loss between the generated and labeled images. These three losses are then combined into a total loss function. By minimizing this total loss, the generator network learns model parameters that are more similar to those under high-quality input conditions, even under low-quality input conditions.

[0051] Finally, the trained generative network is used to process the low-resolution side-view face image, ultimately reconstructing a high-resolution frontal face image. This invention can simultaneously achieve face super-resolution reconstruction and pose correction, restoring a realistic high-resolution frontal face image from a low-resolution side-view face image.

[0052] Example 2

[0053] The original input image in this embodiment is a degraded image (16×16 pixels) of a side-view face image (128×128 pixels) after being downsampled by 8 times. Figure 1 As shown, it can be seen that face images are severely degraded at low resolution and in different poses, making it difficult to associate real faces with existing images, especially the texture and features of local details. This will have a great impact on downstream recognition tasks. In addition, the lack of side face feature information makes the restoration task extremely challenging.

[0054] This embodiment uses a similar approach. Figure 1 The image distribution shown was used to train 250 subjects on a total of 38,987 images, with the first 200 used for training and the rest for testing. Face super-resolution generation methods based on pose constraints, such as... Figure 6 As shown, the specific steps include:

[0055] Step S1: Collect face images under different pose conditions, divide them into frontal face images and side face images, and construct training data pairs of low-resolution side face images and high-resolution frontal face images.

[0056] We collect face images under different poses, extract and align faces, and scale them to 128×128 pixels. Based on the pose, we divide them into frontal face images and side face images. Then, we perform 8x downsampling on the side face to obtain a low-resolution side face, and use the corresponding high-resolution frontal face image as the label image.

[0057] Step S2: The low-resolution side-view face image is fed into the side-view attention guidance module for processing. A Gaussian-distributed feature heatmap is generated based on prior facial feature points and concatenated with the original image along the channel direction before being input into the generation network. Figure 2 As shown.

[0058] First, read the coordinates of the prior profile facial feature points. Each profile image corresponds to a set of feature point coordinates. Each set of feature point coordinates has five pairs, and their positions correspond to the center of the left eye, the center of the right eye, the tip of the nose, the left corner of the mouth, and the right corner of the mouth, respectively.

[0059] For each pair of coordinates, the values ​​of the corresponding pixel and its neighboring pixels are filled using a two-dimensional Gaussian distribution, where the pixel with the mean value represents the location of the corresponding feature point of the organ, and the standard deviation is set to 0.8. After obtaining the feature heatmap, it is concatenated with the original image along the channel direction and then fed into the generator network.

[0060] Step S3: Process the image using the guided transform subnetwork in the generative network, combined with the attention fusion transform module, while using skip connections for multi-scale feature fusion, such as... Figure 3 and Figure 4 As shown, Figure 3 This is a schematic diagram of the structure of the Generative Adversarial Network in this embodiment, where PAG (Profile Attention Guidance) is the profile attention guidance module and AFTM (Attention Fusion Transformation Module) is the attention fusion transformation module. Figure 4 This is a schematic diagram of the attention fusion transformation module.

[0061] First, the proposed attention fusion transformation module consists of Channel Attention (CA), Spatial Transformation Network (STN), and Spatial Attention (SA) in that order. CA uses global max pooling and global average pooling, followed by channel compression via 1×1 convolution to obtain two feature maps. These two maps are then summed and passed through a sigmoid function to obtain the weight map for each channel. Multiplying this weight map by the original feature map yields the feature map after channel attention, achieving adaptive calibration of features along the channel dimension. STN consists of a localization network and a mesh generator. The localization network is a small convolutional neural network that outputs a set of parameters based on the input image. These parameters define... A spatial transformation can be translation, rotation, scaling, or more complex affine or nonlinear transformations. The grid generator receives the transformation parameters predicted by the localization network and generates a coordinate grid that represents the new position of each pixel in the input image mapped to the output image. The STN compensates for the in-plane rotation, translation, and scaling of the face by predicting and applying the affine transformation parameters. SA uses global max pooling and global average pooling, and concatenates and convolves according to channels. The convolution kernel size is 3×3. Then, the spatial attention weight matrix is ​​obtained by passing the sigmoid activation function. The weight matrix is ​​then multiplied by the original feature map to obtain the spatial attention feature map.

[0062] Next, the original image is concatenated with the feature heatmap and fed into the guided transform subnetwork, which adopts a U-Net structure consisting of a downsampling encoder and an upsampling decoder. The downsampling encoder sequentially stacks pooling layers, convolutional layers, and an attention fusion transform module, with pooling layer kernels of 2×2 and convolutional layer kernels of 3×3, ultimately yielding a 512×2×2 feature map. This map is then unfolded and reconstructed into a 512×2×2 feature map through a fully connected layer, and then input into the upsampling decoder. The upsampling decoder sequentially stacks connection layers, deconvolutional layers, and an attention fusion transform module, with the deconvolutional layers having an upsampling rate of 2 and a convolutional kernel size of 3×3, ultimately yielding a 512×32×32 feature map. Each layer of the downsampling encoder and upsampling decoder is connected via a connection layer.

[0063] Step S4: The face feature map output by the guided transformation sub-network is fed into the upsampling sub-network in the generator network to gradually refine and reconstruct the face feature map after pose correction. The main body of the upsampling sub-network consists of a deconvolution layer and an attention fusion transformation module. The upsampling rate of the deconvolution layer is 2, the kernel size is 5×5, and the stride is 2. Finally, a convolution layer is used to transform the feature map into a 128×3×3 RGB face image with a kernel size of 5×5 and a stride of 2.

[0064] Step S5: Feed the reconstructed face and its corresponding real face into the discriminative network to perform adversarial training against the generative network, and use a loss function to supervise the generative performance of the generative network. The specific steps are as follows:

[0065] (1) Using the generated face image as input to the discriminator network, perform forward propagation to obtain the discrimination result, and then perform the same operation using the label image. Construct the discriminator loss L from the two discrimination results. D The formula is shown below:

[0066]

[0067] In the formula, Indicates the generated image With label image h i The joint distribution between them, D and d, represent the discriminant network and its parameters. During training, the minimum value of L is minimized. D And update the parameters of the discriminator.

[0068] The goal of the generative network is to deceive the discriminative network as much as possible, so that the generated image passes the discriminative network's detection; therefore, it has an adversarial loss L. adv It can be represented as:

[0069]

[0070] During training, minimize L D With L adv The discrimination network and the generator network are updated alternately.

[0071] (2) Using the generated face image as input to the VGG19 network, perform forward propagation to obtain the feature map output by the ReLU32 layer. Then, perform the same operation using the labeled image to construct the identity preservation loss L from the two feature maps. id The formula is shown below:

[0072]

[0073] In the formula, Φ(·) represents the feature map extracted by the ReLU32 layer in VGG19.

[0074] (3) Calculate the pixel loss L between the generated image and the label image using the l2 loss function. mse The formula is shown below:

[0075]

[0076] (4) Construct the above three types of losses into a generator network loss function L G By minimizing L GThis allows the generative network to learn model parameters that are more similar to those under high-quality input conditions, even with low-quality input. The formula is shown below:

[0077] L G =L mse +αL id +βL adv

[0078] In the formula, α and β are constants, both set to 0.01.

[0079] Step S6: Use the trained generative network to process the low-resolution side face image and finally reconstruct a high-resolution frontal face image.

[0080] Figure 5 shows a schematic diagram of the restoration results of some face images after the method of the present invention. Figure 5(a) is the input low-resolution side face image. It can be seen from the figure that the original low-resolution side face image is blurry and has very low resolution. Figure 5(b) is the result output image after the method of the present invention. Figure 5(c) is the label image. It can be seen from the figure that after the method of the present invention, a high-resolution frontal face image can be restored from the low-resolution side face image, and the effect is obvious.

[0081] In summary, the method of this invention first collects face images in different poses and divides them into two categories: frontal face images and side face images. Side face images are downsampled to degenerate into low-resolution images, thus constructing training data pairs of low-resolution side face images and high-resolution frontal face images. The low-resolution side face images are then fed into a side face attention guidance module for processing. A Gaussian-distributed feature heatmap is generated based on prior facial feature points and concatenated with the original image along the channel direction before being input into a generative adversarial network (GAN) to guide the network to focus on key facial organs of the side face. Finally, a guided transformation subnetwork within the GAN is used. The invention employs a network and an attention fusion transformation module to process images, using skip connections for multi-scale feature fusion to perform pose correction of the face in the feature space. The face feature map output from the guided transformation sub-network is fed into the upsampling sub-network of the generative adversarial network (GAN) to progressively refine and reconstruct the pose-corrected face feature map. The reconstructed face and its corresponding real face are then fed into a discriminant network for adversarial training against the GAN, with a loss function used to supervise the GAN's generative performance. The trained generative network is then used to process low-resolution side-view face images, ultimately reconstructing a high-resolution frontal face image. This invention can simultaneously achieve face super-resolution reconstruction and pose correction, restoring a realistic high-resolution frontal face image from a low-resolution side-view face image.

[0082] It should be noted that the above content merely illustrates the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. For those skilled in the art, various improvements and modifications can be made without departing from the principle of the present invention, and all such improvements and modifications fall within the scope of protection of the claims of the present invention.

Claims

1. A face super-resolution generative adversarial network based on pose-constrained conditions, characterized in that, It includes at least a profile attention guidance module, an attention fusion and transformation module, a guidance and transformation subnetwork, an upsampling subnetwork, and a discriminant network. The side profile attention guidance module generates a Gaussian-distributed feature heatmap based on prior facial feature points of the input low-resolution side profile image, which is used to guide the generative adversarial network to focus on key facial organs. The attention fusion transformation module consists of channel attention, a spatial transformation network, and a spatial attention sequence, embedded in a generative adversarial network (GAN) to assist the GAN in processing face images; wherein... The channel attention: global average pooling is used to obtain the weight map of each channel through Sigmoid, and multiply it with the original feature to obtain the feature map after channel attention, so as to achieve adaptive calibration of features in the channel dimension; The spatial transformation network compensates for in-plane rotation, translation, and scaling of the face by predicting and applying affine transformation parameters. The spatial attention is achieved by using global max pooling and global average pooling, concatenating and convolving according to channels, and then obtaining the spatial attention weight matrix through the Sigmoid activation function. The weight matrix is ​​multiplied by the original feature map to obtain the spatial attention feature map. The guided transform subnetwork adopts a U-Net structure, consisting of a downsampling encoder and an upsampling decoder, and is used to process low-resolution side face images and their facial feature heatmaps, outputting a frontal face feature map; the downsampling encoder and the upsampling decoder are connected to each other through connection layers. The upsampling sub-network consists of a deconvolutional layer and an attention fusion transformation module. Finally, a convolutional layer is used to transform the feature map into an RGB face image, which is used to upsample the frontal face feature map, enlarge the image and enhance key facial information. The discriminative network is used to input generated face images and real face images, and to train the generative adversarial network (GAN) against it, thereby supervising the GAN's generation performance.

2. The face super-resolution generative adversarial network based on pose condition constraints as described in claim 1, characterized in that: In the downsampling encoder of the guided transform subnetwork, pooling layers, convolutional layers, and attention fusion transform modules are stacked in sequence, and in the upsampling decoder, connection layers, deconvolutional layers, and attention fusion transform modules are stacked in sequence.

3. A face super-resolution generation method based on pose-conditional constraints using adversarial networks as described in claim 1, characterized in that, Includes the following steps: S1. Collect and preprocess data: Collect and preprocess face images under different pose conditions. The face images are divided into frontal face images and side face images. Construct training data pairs of low-resolution side face images and high-resolution frontal face images. S2. Generate facial feature heatmap: The low-resolution side face image constructed in step S1 is sent to the side face attention guidance module for processing, and a Gaussian distributed feature heatmap is generated based on the prior facial feature points. S3. Correcting facial pose: The facial feature heatmap generated in step S2 is concatenated with the corresponding original image in the channel dimension and used as input. This input is then fed into the guided transformation subnetwork in the generative adversarial network. Combined with the attention fusion transformation module and the skip connections used for multi-scale feature fusion, the pose-corrected facial feature map is obtained. S4. Face image reconstruction: The face feature map output by the guided transformation subnetwork in step S3 is fed into the upsampling subnetwork in the generative adversarial network to refine and reconstruct the face feature map after pose correction. S5. Model Training: The refined and reconstructed face image obtained in step S4 and the corresponding real face image are fed into the discriminant network to train the generative adversarial network model, and the loss function is used to supervise the model parameters of the generative adversarial network. S6. Output the result: Input the low-resolution side face image constructed in step S1 into the generative adversarial network trained in step S5 to reconstruct a high-resolution frontal face image.

4. The face super-resolution generation method based on pose condition constraints as described in claim 3, characterized in that: The preprocessing in step S1 includes at least facial extraction and alignment of the collected face images, and downsampling of the side face images to obtain low-resolution side face images.

5. The face super-resolution generation method based on pose condition constraints as described in claim 3, characterized in that: The prior facial feature points in step S2 include at least five pairs, with positions corresponding to the center of the left eye, the center of the right eye, the tip of the nose, the left corner of the mouth, and the right corner of the mouth, respectively. A two-dimensional Gaussian distribution is used to fill the values ​​of the pixels corresponding to the coordinates of each pair of feature points and their nearby pixels, with the standard deviation set to 0.

8.

6. The face super-resolution generation method based on pose condition constraints as described in claim 3, characterized in that: In the upsampling network of step S4, the upsampling rate of the deconvolution layer is 2, the kernel size is 5×5, and the stride is 2. A convolutional layer is used to transform the feature map into a 128×3×3 RGB face image with a kernel size of 5×5 and a stride of 2.

7. The face super-resolution generation method based on pose condition constraints as described in claim 3, characterized in that: Step S5 specifically includes the following steps: S51: Use the generated face image as the input of the discriminator network, perform forward propagation to obtain the discrimination result, and perform the same operation using the label image. The two discrimination results are used to construct the adversarial loss. S52: Use the generated face image as input to the VGG19 network, perform forward propagation to obtain feature maps, and then perform the same operation on the label image to construct the identity preservation loss from the two feature maps. S53: Use The loss function calculates the pixel loss between the generated image and the label image; S54: Construct the total loss function from the three losses in steps S51, S52, and S53, and improve the parameters of the generative adversarial network model by minimizing the loss.

Citation Information

Patent Citations

  • Face super-resolution reconstruction method

    CN113379597A

  • Video crowd counting method based on combination of attention and spatial transformation network

    CN116385964A