Dual-branch face image super-resolution reconstruction network based on prior information

By designing a dual-branch structure and cross-attention mechanism in the super-resolution model of face images, the prior information is integrated with local and global features, and the problem of insufficient utilization of prior information in the existing technology is solved, which significantly improves the effect of super-resolution reconstruction of face images.

CN120070176APending Publication Date: 2025-05-30NORTHEASTERN UNIV AT QINHUANGDAO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510046664.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When using prior information, the existing facial image super-resolution model only fuses the prior information with local features, resulting in insufficient utilization of prior information and reducing the effect of super-segment reconstruction.

Method used

A dual-branch face image super-defense reconstruction network based on prior information is designed, and the prior information is fused with low-resolution features through the cross-attention mechanism, and the features of local branches and global branches are collaboratively fused in the dual-branch structure, and finally high-resolution recovery results are generated through the convolution layer.

Benefits of technology

By making full use of the fusion of prior information and the features of different depth layers, the model's understanding of semantic features is improved, and the effect of super-resolution reconstruction of face images is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070176A_ABST
    Figure CN120070176A_ABST
Patent Text Reader

Abstract

The invention provides a dual-branch face image super-resolution reconstruction network based on prior information, and relates to the technical field of computer vision and image processing. The prior information-based double-branch face image super-resolution reconstruction network comprises the following contents: 1) firstly, a low-resolution image ILR passes through a convolutional layer and a PFB module, and the PFB module comprises two PAFB modules; (2) the PAFB module is responsible for fusing prior information (a human face analysis graph is adopted) and a low-resolution original feature F0 to obtain a feature F1, and the fusion mode is a cross-attention mode; according to the method, the prior information is fused with the local features and the global features respectively, and the local features and the global features fused with the prior information are fused again, so that the prior information is fully utilized; features of different depth layers are fused in a fusion mode that the features are added firstly and then pass through a CRC sequence, so that the understanding ability of the model to semantic features is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and image processing, and specifically to a dual-branch face image super-resolution reconstruction network based on prior information. Background Art

[0002] In non-restricted scenarios, since the monitoring camera is too far away from the target, the captured facial images have too low a resolution, affecting some downstream tasks such as facial attribute analysis and facial recognition. Face image super-resolution reconstruction refers to converting a low-resolution blurred face image into a high-resolution clear face image through super-resolution reconstruction technology. The super-resolution model obtains the restoration result through inference. As a specific type of image, a face image has unique features and contains rich prior knowledge, such as a series of features including facial contours, nose shapes, whether wearing glasses, and whether having a beard.

[0003] Currently, methods for face super-resolution are mainly divided into two categories: unsupervised (without using prior knowledge) and supervised (using prior knowledge). Currently, the prior knowledge mainly includes facial parsing maps, face heat maps, facial key points, etc., which are the prior knowledge used by most researchers. The method using prior knowledge mainly fuses the prior information with the intermediate features of the network through an attention mechanism. And the intermediate features of the network mainly solve the problem of gradient disappearance through residual connections.

[0004] Although using prior information can improve the restoration effect of the face super-resolution model, previous methods only fuse the prior information with local features (because the network's modeling ability is not strong and can only extract local features), which will result in insufficient utilization of the prior information and thus reduce the effect of super-resolution reconstruction. Because the modeling ability of local features is limited, and when restoring a face image, it is also important to be able to extract global features, which can maintain the accuracy of the facial contour.

[0005] Although using skip connections can indeed solve the problems of gradient explosion or gradient disappearance, in an image super-resolution network, there are a large number of convolutional layers, which will change the abstraction degree of semantic features. Shallow features are relatively simple and the semantic features they contain are more basic; as the number of layers increases, the semantic features become very abstract and difficult to understand, but this abstract semantic feature becomes more precise and easier to be understood by the model, so the cooperation between shallow features and deep features is very important. However, this is not reflected in previous work. Summary of the Invention

[0006] (1) Technical Problems to be Solved

[0007] Aiming at the deficiencies of the prior art, the present invention provides a dual-branch face image super-resolution reconstruction network based on prior information, solving the defects and deficiencies existing in the prior art.

[0008] (2) Technical solution

[0009] To achieve the above objectives, the present invention is implemented through the following technical solutions: A dual-branch face image super-resolution reconstruction network based on prior information, including the following:

[0010] 1) First, the low-resolution image I LR will pass through a convolutional layer and a PFB module, and the PFB module contains two PAFB modules;

[0011] 2) The PAFB module is responsible for fusing the prior information (we use a face parsing map) and the low-resolution original feature F 0 to obtain the feature F 1 , and the fusion method is through the cross-attention method;

[0012] 3) Next, the feature F 1 will pass through the core structure of this model: the dual-branch structure; the dual-branch structure includes a local branch and a global branch;

[0013] 4) Each branch contains three downsampling stages, a "neck" stage, and three upsampling stages. Finally, the final feature obtained by the local branch passes through a convolutional layer to obtain the final high-resolution restoration result I SR .

[0014] Preferably, the global branch is one of the branches in the dual-branch model. Its structure is similar to that of the local branch. Replacing the FRB module in it with an SPB module and adding a module for the cooperation of global features and local features constitutes the local branch, and the output of the global branch is different from that of the local branch. The output of the local branch is the finally adopted high-resolution restoration result.

[0015] Preferably, the global branch starts with a convolutional layer and a PFB module. The convolutional layer extracts the original low-resolution feature, and the PFB module preliminarily fuses the original low-resolution feature with the prior information. The two processes are described by the formula:

[0016] F fre = f Conv (I LR )(1)

[0017]

[0018] where f Conv is the convolutional layer, F fre is the extracted original feature, P is the face parsing map. In this model, the face parsing map uses a 0-1 matrix, that is, the facial skin is 0 and the facial components are 1; fPFB is the PFB module, which is the result of the initial fusion of the original features and the prior information.

[0019] Preferably, the downsampling stage includes a downsampling module, two FRB modules and a PFB module. The purpose of the downsampling module is to reduce the resolution. The FRB module includes a Fourier transform operation, an inverse Fourier transform operation and two sets of convolutional layers. Each set of convolutional layers contains two convolutional layers, which are responsible for processing frequency information and phase information. The Fourier transform extracts the amplitude information and phase information in the features, enlarges the receptive field in this way, and extracts global features. The inverse Fourier transform transforms the phase information and amplitude information into features. The formula description of the downsampling stage is as follows:

[0020]

[0021] where f downsample represents the downsampling module, and f FRB represents the FRB module, represents the output feature obtained in the first downsampling stage. Similarly, in the following text represents the output feature obtained in the third downsampling stage.

[0022] Preferably, the neck stage includes a DSIB module, two FRB modules and a PFB module. The DSIB module fuses the input feature of the DSIB module with the features of different resolutions before, so as to cooperate different semantic features. The process is described by the formula:

[0023]

[0024] where Concat represents concatenating the results obtained from the three FEB modules, and f FEB represents the FEB module, represents the output feature of the second FRB module in the third downsampling stage, represents the output feature of the Concat module in the DSIB module;

[0025] After the concatenation, since the number of channels becomes three times the original, a 1×1 convolutional layer is used to adjust the number of channels to 3. After adjusting the number of channels, the channel weights are adjusted through the Channel Attention layer. The process is expressed as:

[0026]

[0027] where CA represents the Channel Attention layer, represents the output feature of the DSIB module in the neck stage.

[0028] Preferably, in the FEB module, the pairwise feature fusion method is addition. Since the resolutions of the two features are different, bicubic interpolation is used to adjust the resolution. After adjusting the resolution, the two features are added together, and then after passing through the CRC sequence, the output feature of the final FEB is obtained through the channel attention layer and the spatial attention layer. This process is described by the formula:

[0029]

[0030] where SpaA represents the spatial attention layer, f CRC represents the CRC sequence, represents the output feature of the FEB.

[0031] (III) Beneficial Effects

[0032] The present invention provides a dual-branch face image super-resolution reconstruction network based on prior information, which has the following beneficial effects:

[0033] By separately fusing the prior information with the local features and the global features, and then fusing the local features and the global features that have been fused with the prior information again, the prior information is fully utilized; the features of different depth layers are fused together. The fusion method is to add first and then pass through the CRC sequence, thereby improving the model's ability to understand semantic features; when using the GAN model for fine-tuning, a new loss function is used for constraint, and this loss function combines the style loss, the pixel-level loss function, and the frequency-level + spatial-level adversarial loss. Description of the Drawings

[0034] Figure 1 is the PDBNet model diagram of the present invention;

[0035] Figure 2 is the schematic diagram of the global branch of the present invention;

[0036] Figure 3 is the schematic diagram of the DSIB module of the present invention;

[0037] Figure 4 is the schematic diagram of the FEB module of the present invention;

[0038] Figure 5 is the schematic diagram of the comparative experiment results of the present invention;

[0039] Figure 6 is the schematic diagram of the visualization results of ablation experiment 2 of the present invention;

[0040] Figure 7 is the schematic diagram of the visualization results of ablation experiment 3 of the present invention;

[0041] Figure 8Schematic diagram of the visualization result of ablation experiment 4 of the present invention. Detailed implementation manners

[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0043] Embodiment:

[0044] As Figure 1-8 shown, the embodiment of the present invention provides a dual-branch face image super-resolution reconstruction network based on prior information. The three core innovations of the model are: 1. A new prior knowledge fusion framework (HPFF) is proposed to fully fuse prior knowledge; 2. The DSIB module is designed, and its core function is to fuse semantic features of different depth layers to improve the model's understanding ability of semantic features; 3. A new loss function is designed, which combines adversarial loss and style loss to improve the authenticity and fine-grainedness of facial textures.

[0045] As Figure 1 shown, the global branch is one of the branches in the dual-branch model. Its structure is similar to that of the local branch. Replacing the FRB module in it with the SPB module and adding a module for the cooperation of global features and local features constitutes the local branch. Moreover, the outputs of the global branch and the local branch are different. The output of the local branch is the finally adopted high-resolution restoration result. Here, the global branch is taken as an example.

[0046] First, the global branch starts with a convolutional layer and a PFB module. The convolutional layer extracts the original low-resolution features, and the PFB module preliminarily fuses the original low-resolution features with prior information (facial parsing map). The two processes are described by the formulas as follows:

[0047] F fre = f Conv (I LR )(1)

[0048]

[0049] where f Conv is the convolutional layer, F fre is the extracted original feature, P is the facial parsing map (prior information). In this model, the facial parsing map is a 0-1 matrix, that is, the facial skin is 0 (blank) and the facial components are 1 (black color blocks); f PFB is the PFB module. It is the result of the initial fusion of the original features and the prior information.

[0050] Next, it goes through three downsampling stages, one "neck" stage, and three upsampling stages.

[0051] The downsampling stage contains one downsampling module, two FRB modules, and one PFB module. The purpose of the downsampling module is to reduce the resolution. The FRB module contains one Fourier transform operation, one inverse Fourier transform operation, and two sets of convolutional layers. Each set of convolutional layers contains two convolutional layers, which are responsible for processing frequency information and phase information. The Fourier transform extracts the amplitude information and phase information in the features, enlarging the receptive field and extracting global features in this way. The inverse Fourier transform transforms the phase information and amplitude information into features. The downsampling stage is described by the formula:

[0052]

[0053] where, f downsample represents the downsampling module, f FRB represents the FRB module, represents the output feature obtained from the first downsampling stage. Similarly, in the following text represents the output feature obtained from the third downsampling stage.

[0054] The neck stage contains one DSIB module, two FRB modules, and one PFB module. The DSIB module is as Figure 3 、 4 shown. The DSIB module fuses the input feature of the DSIB module with the features of different resolutions before, thus collaborating different semantic features. This process is described by the formula:

[0055]

[0056] where, Concat represents concatenating the results obtained from the three FEB modules, f FEB represents the FEB module, represents the output feature of the second FRB module in the third downsampling stage, represents the output feature of the Concat module in the DSIB module.

[0057] After the concatenation, since the number of channels becomes three times the original, a 1×1 convolutional layer is used to adjust the number of channels to 3. After adjusting the number of channels, the weights of the channels are adjusted through the ChannelAttention layer. This process is expressed as:

[0058]

[0059] Among them, CA represents the channel attention layer, represents the output feature of the DSIB module in the neck stage.

[0060] In the FEB module, the way of pairwise feature fusion is by addition. Since the resolutions of the two features are different, in the present invention, the resolution is adjusted by bicubic interpolation (i.e., Figure 4 the resize module in it). After adjusting the resolution, the two features are added, and then after passing through the CRC sequence (convolutional layer, ReLU activation function, convolutional layer), the output feature of the final FEB is obtained through the channel attention layer and the spatial attention layer. This process is described by the formula:

[0061]

[0062] Among them, SpaA represents the spatial attention layer, f CRC represents the CRC sequence, represents the output feature of the FEB.

[0063] After passing through the neck stage, through three upsampling stages, the principle is similar to the downsampling stage. The upsampling stage includes an upsampling module (to increase the resolution), two FRB modules, one PFB module, and one DSIB module. After three upsampling stages, an image can be obtained through a convolutional layer, which is the high-resolution recovery result in the frequency domain, but this is not the final recovery result adopted.

[0064] As Figure 1 shown, the structure of the local branch is similar to that of the global branch. Just replace the FRB with SPB. The SPB module contains 8 resblocks, aiming to extract local features (because the receptive field of the convolutional layer is limited), and the purpose of the SFMLM module is to fuse the local features captured by the local branch and the global features captured by the global branch. The fusion method is also cross-attention. The structure of the SFMLM module can illustrate that this model is a dual-branch structure.

[0065] Next, the implementation principle of the HPFF framework is described. The original intention of proposing the HPFF framework is to make full use of prior knowledge. First, from the above description, it can be seen that the PFB module exists and appears multiple times in the dual-branch structure. Appearing multiple times is to fuse the prior knowledge with the intermediate features of the network multiple times to make full use of it, and existing in both branches is also to improve the utilization rate of prior knowledge. As mentioned before, the local branch extracts local features, and the global branch extracts global features. The prior knowledge is fused with the local features and global features multiple times respectively, and the local features and global features fused with prior information are fused again, so as to make full use of prior information.

[0066] The above is all the principles of the PSNR model. The loss function of the PSNR model is a pixel-level loss function, as shown in formulas (7)(8)(9). The pixel-level loss function includes a spatial loss function and a frequency loss function, both of which are L1 loss functions. In the spatial loss function, I HR represents the real image (GT image), represents the high-resolution restored image in the spatial domain, represents the high-resolution restored image in the frequency domain, represents the amplitude information of the high-resolution restored image in the frequency domain, represents the phase information of the high-resolution restored image in the frequency domain.

[0067]

[0068] A disadvantage of the high-resolution restored image generated using the pixel-level loss function is that the image is too smooth. Therefore, in the present invention, the GAN model is used for fine-tuning, and the proposed new loss function is used for constraint. This loss function combines the pixel-level loss function, the frequency-level adversarial loss function, and the style loss function.

[0069]

[0070] Among them, and are the adversarial loss functions in the spatial domain and the frequency domain respectively. SD and FD are the spatial discriminator and the frequency discriminator, and the architecture of the discriminator is the LightCNN-9 model. Φ is the VGG model, and L Style is the style loss function, and are the Gram matrices of I SR and I HR respectively. C, H, and W represent the number of channels, the height of the image, and the width of the image respectively.

[0071] The experimental scheme of this model is described below. Following the work of DIC, in this invention, OpenFace is used to detect 68 key points of the face. Based on the estimated key points, a square region is cropped in each image to remove the redundant background and its size is adjusted to 128×128 pixels without any pre-alignment as the ground truth (GT) image. Next, these GT images are downsampled to 32×32 and 16×16, and they are resized to the same size as the LR images (128×128) by bicubic interpolation, corresponding to the ×4 FSR task and the ×8 face super-resolution task respectively. In the training phase, the training dataset includes 80,000 images selected from the CelebA dataset, while the validation dataset includes 100 images from the same source. Finally, 1,000 images from the CelebA dataset and 50 images from the Helen dataset are used for testing. These three datasets are mutually exclusive. In this invention, four widely used reference metrics are adopted to quantitatively prove the effectiveness of super-resolution recovery: PSNR, SSIM, LPIPS, and FID. The PSNR and SSIM values are positively correlated with image similarity. On the contrary, the LPIPS and FID values are negatively correlated with image similarity.

[0072] In this invention, after retraining ParsingNet with 80,000 images, the face parsing map estimated by ParsingNet is input into the PFB module. The PFB and DSIB modules in PDBNet exist in both the frequency and spatial branches simultaneously. The last convolutional layer of both branches will finally generate an image. The output image of the spatial branch is selected as the final high-resolution recovery result. For the selection of the loss function weights, extensive experiments have proven that the optimal weight combination is γ1 = 0.01, γ2 = 0.0005, γ3 = 0.001, γ4 = 0.1, γ5 = 0.01. The Adam optimizer is selected to train both the PSNR-based model (PDBNet) and the GAN-based model (PDBGAN) simultaneously, where β1 = 0.9 and β2 = 0.99. The learning rate is set to 0.0001 and the batch size is 8. The PSNR model is trained for 500K iterations, while the GAN model is trained for 100K iterations. All experiments are conducted using the PyTorch framework on an NVIDIA GeForce RTX 4090.

[0073] The visualization experiment results are given below. Among them, (a) represents the low-resolution image, and (b)-(k) represent the SRCNN [9], RCAN

[10] , SAN

[11] , DIC [6], SPARNet

[12] , SISN

[13] , FishFSRNet [3], SFMNet [4], PDBNet, and PDBGAN models respectively, and (l) represents the groundtruth (GT) image.

[0074] The following are the quantization experiment results. Among them, ×4 and ×8 represent the scales respectively. ×4 means changing the low-resolution image with a resolution of 32×32 to a high-resolution image of 128×128, and ×8 means changing the low-resolution image with a resolution of 16×16 to a high-resolution image of 128×128. Bicubic is the baseline, the benchmark method. Bicubic is the bicubic interpolation method. CelebA and Helen are the test sets. PDBNet and PDBGAN represent our PSNR-based model and GAN-based model respectively. Blue represents the second-best performance, and red represents the best performance. As shown in Table 1, PDBNet can obtain the best PSNR and SSIM metrics. This is because the optimization objective of the PSNR-based model is to obtain higher PSNR and SSIM values, so the restored image is relatively smooth (as Figure 5 shown), while the optimization objective of PDBGAN is to obtain more realistic textures, so PDBGAN can obtain the lowest LPIPS value.

[0075]

[0076] Table 1 Quantization Experiment Results

[0077] The following ablation experiments are used to illustrate the role of the innovation points proposed by the present invention.

[0078] 1. Regarding the use of prior knowledge and the DSIB module;

[0079] To analyze the effectiveness of face prior knowledge and the DSIB module, a series of experiments were conducted at two scales, ×4 and ×8. First, in the present invention, the PFB and DSIB modules were removed from the frequency and spatial branches, resulting in Model 1. Second, the DSIB module was separately added back to the two branches, resulting in Model 2. Third, face prior and the PFB module were added to Model 1, called Model 3. Finally, face prior, the PFB module, and the DSIB module were combined into Model 1 to generate PDBNet. The effectiveness of PFB and DSIB is shown in Table 2. For the ×8 scale on Helen, as shown in Table 2, adding DSIB and PFB can increase PSNR and SSIM by 0.18 dB and 0.0025, 0.19 dB and 0.0061 respectively. When PFB and DSIB are used in combination, PDBNet achieves the highest quantitative results, with PSNR and SSIM increasing by 0.3 dB and 0.0080 respectively. The reasons are analyzed as follows: Face prior is the key to face super-resolution reconstruction, providing a reasonable reference for the inference stage of the FSR model. And the combined multi-scale semantic features can improve the feature representation ability.

[0080]

[0081] Table 2 Ablation Experiment 1 Quantitative Results

[0082] 2. Regarding the effectiveness of HPFF;

[0083] To illustrate the effectiveness of HPFF, in the present invention, experiments were conducted on the ×8 FSR task. To compare the HPFF proposed in the present invention with the strategy of only fusing face prior with local features, the present invention modified the model by removing the PFB module from the frequency branch and using it as the latter strategy. This modified model was named Model 4. The comparison of the restoration results between Model 4 and PDBNet is shown in Table 3. The present invention measures the FSR performance from three metrics: PSNR, SSIM, and LPIPS. In Table III, we note that PDBNet can increase PSNR and SSIM by 0.11 dB and 0.0033 respectively, while reducing LPIPS by 0.0060. Figure 6 Shows the effects of restored images under different face prior fusion strategies. The present invention can note that adopting the strategy of separately incorporating face prior into local and global features can obtain better visual effects (for example, in the second row, the eyes and teeth parts in the image restored by PDBNet are clearer).

[0084]

[0085] Table 3 Ablation Experiment 2 Quantitative Experiment Results

[0086] 3. Regarding the effectiveness of using a better generator;

[0087] In this section, to illustrate the effectiveness of generator improvement on adversarial training, the present invention uses Model 1 as Generator 1, the discriminator is based on LightCNN9, and the loss function is modified by removing the style loss in Formula 10. The trained model is named SFMGAN. Subsequently, PDBNet is used as Generator 2, and the discriminator and loss function are the same as those of SFMGAN. The obtained trained model is named Model 5. The effectiveness of the generator is shown in Table 4, where G represents the generator. Using a performance-enhanced generator can enhance the metrics of the two test datasets on the ×8 task, indicating better restored image quality. Among them, the LPIPS and FID values of Model 5 are 0.0038 and 4.39 lower than those of SFMGAN, respectively. Figure 8 The visual effect is shown, and it can be noted that improving the generator can enhance the authenticity of facial textures.

[0088]

[0089] Table 4 Quantification Results of Ablation Experiment 3

[0090] Among them, the lower the LPIPS and FID, the better the performance.

[0091] 4. Regarding the role of the style loss;

[0092] In the present invention, the style loss function is added to the loss function of Model 5 to generate PDBGAN. Experiments on the ×8 FSR task are carried out on two test datasets (CelebA and Helen). The effectiveness of the style loss is shown in Table 5. It can be noted that on Helen, the LPIPS of PDBGAN is 0.0015 lower than that of Model 6, and the FID is 3.17 lower than that of Model 6, indicating the effectiveness of incorporating the style loss. From the visual results, Figure 8 it is shown that the hair texture in the face images restored by PDBGAN appears clearer.

[0093]

[0094] Table 5 Quantification Results of Ablation Experiment 4

[0095] It should be noted that, in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a reference structure" does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0096] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A dual-branch face image super-resolution reconstruction network based on prior information, characterized by: Includes the following: 1) First, a low-resolution image I LR It passes through a convolutional layer and a PFB module, which contains two PAFB modules; 2) The PAFB module is responsible for fusing the prior information with the low-resolution original feature F0 to obtain the feature F1. The fusion method is through the cross-attention method; 3) Next, feature F1 will pass through the core structure of this model: the dual-branch structure; the dual-branch structure includes local branches and global branches; 4) Each branch includes three downsampling stages, one "neck" stage and three upsampling stages. Finally, the final features obtained by the local branch pass through a convolution layer to obtain the final high-resolution restoration result I SR .

2. The dual-branch face image super-resolution reconstruction network based on prior information according to claim 1, characterized in that: The global branch is one of the branches in the dual-branch model, and its structure is similar to that of the local branch. The FRB module is replaced by the SPB module, and a module that collaborates with global features and local features is added to form a local branch. The output of the global branch is different from that of the local branch, and the output of the local branch is the final high-resolution restoration result.

3. The dual-branch face image super-resolution reconstruction network based on prior information according to claim 2, characterized in that: The global branch starts with a convolutional layer and a PFB module. The convolutional layer extracts the original low-resolution features, and the PFB module preliminarily fuses the original low-resolution features with the prior information. The two processes are described by the formula: F fre =f Conv (I LR )(1) Among them, f Conv is the convolutional layer, F fre is the original feature extracted, P is the facial analysis map. In this model, the facial analysis map uses a 0-1 matrix, that is, facial skin is 0 and facial components are 1; f PFB For the PFB module, It is the result of preliminary fusion of original features and prior information.

4. The dual-branch face image super-resolution reconstruction network based on prior information according to claim 3, characterized in that: The downsampling stage includes a downsampling module, two FRB modules and a PFB module. The purpose of the downsampling module is to reduce the resolution. The FRB module includes a Fourier transform operation, an inverse Fourier transform operation and two groups of convolutional layers. Each group of convolutional layers includes two convolutional layers, which are responsible for processing frequency information and phase information. The Fourier transform extracts the amplitude information and phase information in the feature, thereby increasing the receptive field and extracting the global feature. The inverse Fourier transform transforms the phase information and amplitude information into features. The formula of the downsampling stage is described as: Among them, f downsample represents the downsampling module, f FRB represents the FRB module, represents the output features obtained in the first downsampling stage. Similarly, Represents the output features obtained in the third downsampling stage.

5. The dual-branch face image super-resolution reconstruction network based on prior information according to claim 4, characterized in that: The neck stage includes a DSIB module, two FRB modules and a PFB module. The DSIB module fuses the input features of the DSIB module with the previous features of different resolutions, thereby coordinating different semantic features. The process is described by the formula: Concat represents the concatenation of the results obtained from the three FEB modules. FEB Represents the FEB module, represents the output features of the second FRB module in the third downsampling stage, Represents the output features of the Concat module in the DSIB module; After the splicing is completed, since the number of channels becomes three times the original, a 1×1 convolution layer is used to adjust the number of channels to 3. After adjusting the number of channels, the channel weights are adjusted through the channel attention layer (ChannelAttention). The process is expressed as: Among them, CA represents the channel attention layer, Represents the output features of the DSIB module in the neck stage.

6. The dual-branch face image super-resolution reconstruction network based on prior information according to claim 5, characterized in that: In the FEB module, the way to fuse two features is by addition. Because the two features have different resolutions, the resolution is adjusted by bicubic interpolation. After the resolution is adjusted, the two features are added, and then after the CRC sequence, the final FEB output features are obtained through the channel attention layer and the spatial attention layer. The process is described by the formula: Among them, SpaA represents the spatial attention layer, f CRC Represents the CRC sequence, Represents the output characteristics of FEB.