A Lightweight Face Super-Resolution Reconstruction Method Based on Hybrid Residual Attention

By combining local convolution and lightweight self-attention module, the problem of global and local feature fusion in face image reconstruction is solved, and high-quality and low-complexity face super-resolution reconstruction is achieved.

CN118941449BActive Publication Date: 2025-07-04BEIJING ELECTRONICS SCI & TECH INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410984595.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-22
Publication Date
2025-07-04
Estimated Expiration
2044-07-22

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate the global and local feature information of face images, and at the same time, the computational complexity is high, resulting in poor reconstruction quality and excessive resource consumption.

Method used

The hybrid residual attention module is adopted, combined with the local convolution module and the lightweight self-attention module, feature fusion is performed through the residual structure, and the reverse order fusion module is used to improve the multi-layer feature fusion capability to generate high-resolution face images.

Benefits of technology

Improves the quality of face image reconstruction, reduces computing complexity, and is suitable for resource-constrained devices and real-time scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118941449B_ABST
    Figure CN118941449B_ABST
Patent Text Reader

Abstract

This application proposes a lightweight face super-resolution reconstruction method based on hybrid residual attention. The method includes: capturing local detail information and global dependency relationships of the shallow feature map through a local convolution module and a lightweight self-attention module respectively to form local features and global features; integrating the local convolution module and the lightweight self-attention module to construct a hybrid residual attention module, and performing feature fusion on the local features and global features through a residual structure; extracting initial features of the low-resolution face image, concatenating multiple hybrid residual attention modules in series, and fusing multi-layer features through a reverse fusion module to obtain the target deep comprehensive features; inputting the target deep comprehensive features into an image upsampling module to generate a high-resolution face reconstruction image. This application improves the reconstruction quality of face images by combining global and local features, effectively reducing the computational complexity, and is suitable for applications on resource-constrained devices and real-time scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image super - resolution, and particularly to a lightweight face super - resolution reconstruction method based on hybrid residual attention. Background Art

[0002] Image super - resolution reconstruction technology has wide applications in the field of image processing, such as satellite image enhancement, medical image analysis, and video surveillance. In the field of security, especially in face recognition and surveillance systems, the super - resolution reconstruction of face images is of great significance. By reconstructing low - resolution face images into high - resolution images, the clarity and details of the images can be significantly improved, thereby enhancing the accuracy of face recognition and the effectiveness of surveillance systems.

[0003] Traditional face super - resolution reconstruction methods mainly rely on convolutional neural networks (CNNs) to extract and reconstruct image features. Convolutional neural networks gradually extract multi - level features of images by stacking convolutional layers, pooling layers, and activation functions. However, due to the characteristics of local convolution operations, the effective receptive field of convolutional neural networks is limited, and there are certain limitations in capturing global dependencies and long - distance information in images, which easily leads to blurred reconstructed images and loss of details.

[0004] In recent years, Transformer models based on attention mechanisms have achieved remarkable results in the fields of natural language processing and computer vision. The self - attention mechanism can effectively capture global dependencies by calculating the similarity between each position in the input feature map. However, the standard Transformer model has a high computational complexity. Especially when processing high - resolution images, it requires a large amount of computational resources and memory space, and is not suitable for applications in resource - constrained devices and real - time scenarios.

[0005] To reduce the computational complexity, researchers have proposed various lightweight self - attention mechanisms. For example, the Linear Spatial Reduction Attention of PyramidVision Transformer (PVT) reduces the computational amount and memory requirements by linearly reducing the spatial dimensions of the feature map. However, although lightweight self - attention mechanisms have advantages in capturing global dependencies, they still have deficiencies in capturing local detailed information.

[0006] In practical applications, super-resolution reconstruction of face images requires simultaneous attention to local detail information and global dependency relationships. Using only convolutional neural networks or self-attention mechanisms alone may not be able to meet the requirements of both aspects simultaneously. Therefore, current technologies face challenges in improving image reconstruction quality and reducing computational complexity. There is a need for a method that can effectively combine convolutional operations with self-attention mechanisms to achieve efficient and high-quality super-resolution reconstruction of face images. Summary of the Invention

[0007] This application aims to solve at least one of the technical problems in the related art to some extent.

[0008] To this end, the purpose of this application is to propose a lightweight face super-resolution reconstruction method based on hybrid residual attention, which is used to solve the problems of neglecting local and global dimension feature fusion and high computational complexity in current face super-resolution reconstruction methods based on self-attention mechanisms. By constructing a lightweight hybrid residual self-attention module and an inverted fusion module, the face image reconstruction ability can be effectively improved, and the parameter complexity and computational amount can be reduced.

[0009] To achieve the above object, an embodiment of this application proposes a lightweight face super-resolution reconstruction method based on hybrid residual attention, including the following steps:

[0010] Construct a local convolution module and a lightweight self-attention module. The local convolution module and the lightweight self-attention module are respectively used to capture local detail information and global dependency relationships of the shallow feature map, forming local features and global features.

[0011] Integrate the local convolution module and the lightweight self-attention module to construct a hybrid residual attention module, and perform feature fusion on the local features and the global features through the hybrid residual attention module.

[0012] Extract the initial features of the low-resolution face image, concatenate multiple hybrid residual attention modules in series, use the initial features as the input, and fuse multi-layer features through an inverted fusion module to obtain the target deep comprehensive features.

[0013] Input the target deep comprehensive features into an image upsampling module to generate a high-resolution face reconstruction image.

[0014] Optionally, the capture process of the local convolution model includes:

[0015]

[0016] Wherein, is the initial feature of the low-resolution face image, represents the convolutional layer function with a convolution kernel size of 1×1, Represents the channel attention mechanism function, Represents the per-channel convolution function with a convolution kernel size of 3×3, is the local feature.

[0017] Optionally, the capture process of the lightweight self-attention module includes:

[0018] Construct a lightweight linear space compression layer to reduce the number of parameters and computational complexity. Its expression is as follows:

[0019]

[0020] Among them, is the global average maximum pooling layer function with a convolution kernel size of 7×7, Represents the convolution layer function with a convolution kernel size of 1 and a stride of 1. LN(∙) represents the layer normalization layer function, which normalizes in the feature dimension, Represents the GELU activation function, is the initial feature, is the output feature map for subsequent key and value generation, is the number of channels of the feature map, is the length of the feature map, is the width of the feature map;

[0021] Input the initial feature into the lightweight linear space compression layer to generate the lightweight key feature map and value feature map, and use the linear layer function to generate the key and value , and input the initial feature into the linear layer function to generate the query . Its expression is as follows:

[0022]

[0023]

[0024] Among them, and represent the generated key and value and query , and are the dimensions of the key and value respectively, and represent the linear layer function, represents the function of the layer normalization layer, Represents a lightweight linear space compression layer function for reducing the number of parameters and computational complexity;

[0025] Generate keys , values and queries are divided into h heads, and the dimension of each head is , . Self-attention calculation is performed on each head respectively, and the expression is:

[0026]

[0027]

[0028] Among them, , represent the query and key matrices of the i-th head respectively, is the attention score matrix, represents the value matrix of the i-th head, is the output of the i-th head;

[0029] Concatenate the output matrices of all heads along the feature dimension to obtain the concatenated matrix , and the expression is:

[0030]

[0031] Among them, represents the concatenation function;

[0032] Map the concatenated matrix back to the original input dimension through a linear transformation to obtain , then the expression of the final output is:

[0033]

[0034] Among them, is the output weight matrix;

[0035] Apply residual connection and layer normalization to the output to generate the final global feature, and the expression is:

[0036]

[0037]

[0038] Among them, represents the intermediate calculation result, Denotes a fully connected layer unit, which consists of a fully connected layer, a GELU activation function, a Dropout regularization layer, and a fully connected layer. For the global feature.

[0039] Optionally, the constructed hybrid residual attention module includes a main path and a residual path. Integrating the local convolutional module and the lightweight self-attention module to construct the hybrid residual attention module, and performing feature fusion on the local feature and the global feature through the hybrid residual attention module, including:

[0040] The main path captures the global dependence of the shallow feature map through the lightweight self-attention module to form a global feature ; The residual path captures the local detail information of the shallow feature map through the local convolutional module to form a local feature , and the expression is:

[0041]

[0042]

[0043] Where Represents the lightweight self-attention module, Represents the local convolutional module;

[0044] Fusing the global feature of the main path and the local feature of the residual path to obtain the deep comprehensive feature of a single hybrid residual attention module, and the expression is:

[0045]

[0046] Where Is the deep comprehensive feature of a single hybrid residual attention module.

[0047] The extraction of the initial feature of the low-resolution face image includes:

[0048] Extracting the initial feature of the low-resolution face image through the initial feature extraction module, and the expression is as follows:

[0049]

[0050] Where Represents the input low-resolution face image, Represents a convolutional layer function with a convolution kernel of 3×3, a stride of 1, a padding of 1, and an output channel of 64, Is the extracted initial feature.

[0051] Optionally, the fusion process of the reverse fusion module includes:

[0052] If there are four cascaded hybrid residual attention modules, starting from the output of the last cascaded hybrid residual attention module, directly use the features of this layer, and then layer by layer forward, concatenate the features of the current layer with the fusion result of the previous layer, and then perform feature fusion through convolution operations to obtain richer context information by accumulating multi-level feature information. The expression is as follows:

[0053]

[0054] Among them, represents the feature map extracted from the i-th layer hybrid residual attention module, represents the feature fusion result at the i-th layer, represents the 1×1 convolutional layer function, represents the concatenation operation of feature maps.

[0055] Optionally, cascading and concatenating multiple hybrid residual attention modules, using the initial feature as the input, and fusing multi-layer features through a reverse fusion module to obtain the target deep comprehensive feature, including:

[0056] If there are four cascaded hybrid residual attention modules, input the initial feature into four consecutive hybrid residual attention modules, and send the output result to the reverse fusion module to extract the target deep comprehensive feature. The expression is as follows:

[0057]

[0058] Among them, represents the hybrid residual attention module, represents the reverse fusion module, is the finally extracted target deep comprehensive feature.

[0059] Optionally, inputting the target deep comprehensive feature into an image upsampling module to generate a high-resolution face reconstruction image, including:

[0060]

[0061] Among them, represents the convolutional layer function with a convolution kernel size of 3×3, represents the upsampling convolutional sub-element function, which can determine the magnification factor by adjusting the number of output channels, is the high-resolution face reconstruction image with the corresponding magnification factor.

[0062] The technical solutions provided by the embodiments of the present application at least bring the following beneficial effects:

[0063] The local convolutional module and the lightweight self-attention module are used to capture the local detailed information and global dependency relationships of the shallow feature map respectively, forming local features and global features. The local convolutional module and the lightweight self-attention module are integrated to construct a hybrid residual attention module, and the local features and global features are fused through the residual structure. The initial features of the low-resolution face image are extracted, multiple hybrid residual attention modules are concatenated in series, and the multi-layer features are fused through the reverse fusion module to obtain the target deep comprehensive features. The target deep comprehensive features are input into the image upsampling module to generate a high-resolution face reconstruction image. In summary, combining global and local features improves the reconstruction quality of face images. At the same time, the design of the lightweight self-attention module effectively reduces the computational complexity, making it suitable for resource-constrained devices and real-time scenarios.

[0064] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, in which:

[0066] Figure 1 is a flowchart of a lightweight face super-resolution reconstruction method based on hybrid residual attention according to an embodiment of the present application;

[0067] Figure 2 is a schematic diagram of a local convolutional module according to an embodiment of the present application;

[0068] Figure 3 is a schematic diagram of a lightweight self-attention module according to an embodiment of the present application;

[0069] Figure 4 is a flowchart of a reverse fusion module according to an embodiment of the present application;

[0070] Figure 5 is a schematic diagram of a face super-resolution reconstruction network based on hybrid residual attention according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present application, and should not be construed as limiting the present application.

[0072] The object of the present invention is to solve the technical problems in the existing face super-resolution technology that it is difficult to effectively integrate global information and local information, and the computational complexity is relatively high, and to provide a face super-resolution reconstruction method based on a hybrid residual attention mechanism. This method can make full use of the interaction of cross-layer features, better integrate global and local information, significantly improve the quality of face image reconstruction, and at the same time effectively reduce the computational complexity, and is applicable to resource-constrained devices and real-time scenarios.

[0073] The following describes a lightweight face super-resolution reconstruction method based on hybrid residual attention according to an embodiment of the present application with reference to the accompanying drawings.

[0074] Figure 1 is a flowchart of a lightweight face super-resolution reconstruction method based on hybrid residual attention shown according to an embodiment of the present application, as Figure 1 shown, this method includes the following steps:

[0075] Step 101, construct a local convolution module and a lightweight self-attention module. The local convolution module and the lightweight self-attention module are respectively used to capture the local detail information and global dependency relationship of the shallow feature map, and form local features and global features.

[0076] In the embodiment of the present application, for extracting initial features from a low-resolution face image, a local convolution module and a lightweight self-attention module are constructed to capture the local detail information and global dependency relationship of the shallow feature map, and form local features and global features.

[0077] Figure 2 is a schematic diagram of the local convolution module shown according to an embodiment of the present application.

[0078] Refer to Figure 2 , the specific calculation process of the capture process of this local convolution model is as follows:

[0079]

[0080] Among them, is the initial feature of the low-resolution face image, represents a convolution layer function with a convolution kernel size of 1×1, represents a channel attention mechanism function, represents a per-channel convolution function with a convolution kernel size of 3×3, is the final output local feature.

[0081] Figure 3 is a schematic diagram of the lightweight self-attention module shown according to an embodiment of the present application.

[0082] Refer to Figure 3, the specific calculation process of the capture process of this lightweight self-attention module is as follows:

[0083] First, construct a lightweight linear spatial compression layer to reduce the number of parameters and computational complexity. Its expression is as follows:

[0084]

[0085] Among them, is the global average maximum pooling layer function with a convolution kernel size of 7×7, represents the convolution layer function with a convolution kernel size of 1 and a stride of 1. LN(∙) represents the layer normalization layer function, which is normalized in the feature dimension. represents the GELU activation function, is the initial feature, is the output feature map, which is used for the subsequent generation of keys and values . is the number of channels of the feature map, is the length of the feature map, is the width of the feature map.

[0086] Next, input the initial feature into the lightweight linear spatial compression layer to generate the lightweight key feature map and the value feature map, and use the linear layer function to generate the key , value . At the same time, input the initial feature into the linear layer function to generate the query . Its expression is as follows:

[0087]

[0088]

[0089] Among them, , represent the generated key , value and query . and are the dimensions of the key and value respectively. , represent the linear layer function. represents the function of the layer normalization layer. represents the lightweight linear spatial compression layer function, which is used to reduce the number of parameters and computational complexity.

[0090] Again, generate the key , value and query Divided into h heads, with the dimension of each head being 、 , and then self-attention calculation is performed on each head separately.

[0091] Specifically, first calculate the scaled dot-product attention score for each head, and the expression is:

[0092]

[0093] Among them, 、 represent the query and key matrices of the i-th head respectively, is the attention score matrix.

[0094] Next, calculate the output of the i-th head, and the expression is:

[0095]

[0096] Among them, represents the value matrix of the i-th head, is the output of the i-th head.

[0097] After performing self-attention calculation on each head separately, the output matrices of all heads are concatenated along the feature dimension to obtain the concatenated matrix , and the expression is:

[0098]

[0099] Among them, represents the concatenation function.

[0100] Next, map the concatenated matrix back to the original input dimension through a linear transformation to obtain , then the final output has the following expression:

[0101]

[0102] Among them, is the output weight matrix;

[0103] Finally, apply residual connection and layer normalization to the output to generate the final global feature, and the expression is:

[0104]

[0105]

[0106] Among them, represents the intermediate calculation result, Represents a fully connected layer unit, which is composed of a fully connected layer, a GELU activation function, a Dropout regularization layer, and a fully connected layer, and is the global feature of the final output.

[0107] Step 102: Integrate the local convolutional module and the lightweight self-attention module to construct a hybrid residual attention module, and perform feature fusion on the local feature and the global feature through the hybrid residual attention module.

[0108] This step involves the construction process of a single-layer hybrid residual attention module.

[0109] In the embodiment of the present application, the constructed hybrid residual attention module includes a main path and a residual path.

[0110] Among them, the main path captures the global dependency relationship of the shallow feature map through the lightweight self-attention module to form a global feature ; the residual path captures the local detailed information of the shallow feature map through the local convolutional module to form a local feature , and the expression is:

[0111]

[0112]

[0113] Among them, represents the lightweight self-attention module, represents the local convolutional module.

[0114] Then, perform feature fusion on the global feature of the main path and the local feature of the residual path to obtain the deep comprehensive feature of a single hybrid residual attention module. The expression is:

[0115]

[0116] Among them, is the deep comprehensive feature of a single hybrid residual attention module.

[0117] Step 103: Extract the initial features of the low-resolution face image, concatenate multiple hybrid residual attention modules in series, take the initial features as the input, and fuse the multi-layer features through the reverse fusion module to obtain the target deep comprehensive feature.

[0118] This step involves the construction process of a multi-layer hybrid residual attention module, and realizes the fusion of multi-layer features through the reverse fusion module.

[0119] Figure 4 is the flowchart of the reverse fusion module shown according to the embodiment of the present application.

[0120] Refer toFigure 4 If there are four cascaded hybrid residual attention modules, start from the output of the last cascaded hybrid residual attention module, directly use the features of this layer, and then move forward layer by layer. Concatenate the features of the current layer with the fusion result of the previous layer, and then perform feature fusion through convolution operations to obtain richer context information through the accumulation of multi-level feature information. The expression is as follows:

[0121]

[0122] Among them, represents the feature map extracted from the i-th layer hybrid residual attention module, represents the feature fusion result at the i-th layer, represents the 1×1 convolutional layer function, represents the concatenation operation of the feature maps.

[0123] It can be understood that referring to the above formula, for i = 4, directly use the feature map of this layer, that is, , for i = 3, 2, 1, first concatenate the current feature map and the fusion result of the previous layer to obtain the concatenated feature map , and then perform convolution operations on the concatenated feature map to obtain the final fusion result .

[0124] Thus, the full construction introduction of the lightweight self-attention module, local convolution module, hybrid residual attention module, and reverse fusion module is realized, and the subsequent application stage is carried out.

[0125] First, for low-resolution face images, extract their initial features through the initial feature extraction module. The expression is as follows:

[0126]

[0127] Among them, represents the input low-resolution face image, represents the convolutional layer function with a 3×3 convolutional kernel, a stride of 1, a padding of 1, and an output channel of 64, is the extracted initial feature.

[0128] Still taking the above embodiment as an example, as Figure 5 shown, if there are four cascaded hybrid residual attention modules, input the initial feature into four consecutive hybrid residual attention modules, and send the output result to the reverse fusion module to extract the target deep comprehensive feature. The expression is as follows:

[0129]

[0130] Among them, represents the hybrid residual attention module, represents the reverse fusion module, is the finally extracted target deep comprehensive feature.

[0131] It should be noted that Figure 5 is a schematic diagram of the face super-resolution reconstruction network based on hybrid residual attention shown in the embodiments of the present application. In the present application, the overall system composed of the hybrid residual attention module, the reverse fusion module and the image upsampling module is called the face super-resolution reconstruction network based on hybrid residual attention.

[0132] Step 104, input the target deep comprehensive feature into the image upsampling module to generate a high-resolution face reconstruction image.

[0133] Referring to Figure 5 the face super-resolution reconstruction network based on hybrid residual attention shown, this network consists of three parts: initial feature extraction, target deep comprehensive feature extraction and final upsampling reconstruction.

[0134] After inputting the initial feature into four consecutive hybrid residual attention modules and sending the output result to the reverse fusion module to extract the target deep comprehensive feature, in this step, finally, the generation of the high-resolution face reconstruction image is realized through the image upsampling module.

[0135] Specifically, the image upsampling module is represented as follows:

[0136]

[0137] Among them, represents the convolutional layer function with a convolutional kernel size of 3×3, represents the upsampling convolutional sub-element function, which can determine the magnification factor by adjusting the number of output channels of is the high-resolution face reconstruction image with the corresponding magnification factor.

[0138] It can be understood that in the present application, by adjusting the number of output channels of

[0139] It should be understood that various forms of the above processes can be used, reordering, adding or deleting steps. For example, the steps described in the present application can be executed in parallel, sequentially or in a different order, as long as the results expected by the technical solution of the present application can be achieved, and no limitation is made herein.

[0140] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of this application shall be included within the protection scope of this application.

Claims

1. A lightweight face super-resolution reconstruction method based on hybrid residual attention, characterized in that Including: Construct a local convolution module and a lightweight self-attention module. The local convolution module and the lightweight self-attention module are respectively used to capture the local detail information and global dependency relationship of the shallow feature map, and form local features and global features. Integrate the local convolution module and the lightweight self-attention module to construct a hybrid residual attention module, and perform feature fusion on the local features and the global features through the hybrid residual attention module. The constructed hybrid residual attention module includes a main path and a residual path. The integrating the local convolution module and the lightweight self-attention module to construct a hybrid residual attention module, and performing feature fusion on the local features and the global features through the hybrid residual attention module includes: The main path captures the global dependencies of the shallow feature map through the lightweight self-attention module to form global features ; the residual path captures the local detail information of the shallow feature map through the local convolution module to form local features , and the expression is: Among them, represents the lightweight self-attention module, represents the local convolution module; Fuse the global features of the main path and the local features of the residual path to obtain the deep comprehensive features of a single hybrid residual attention module. The expression is: Among them, is the deep comprehensive feature of a single said hybrid residual attention module; Extract the initial features of the low-resolution face image, concatenate multiple hybrid residual attention modules in series, take the initial features as the input, and fuse the multi-layer features through a reverse fusion module to obtain the target deep comprehensive features. Input the target deep comprehensive features into an image upsampling module to generate a high-resolution face reconstruction image.

2. The method according to claim 1, characterized in that, The capture process of the local convolution model includes: Among them, is the initial feature of the low-resolution face image, represents the convolutional layer function with a convolutional kernel size of 1×1, represents the channel attention mechanism function, represents the per-channel convolution function with a convolutional kernel size of 3×3, is the local feature.

3. The method according to claim 2, wherein The capture process of the lightweight self-attention module includes: Construct a lightweight linear space compression layer to reduce the number of parameters and computational complexity. The expression is as follows: Among them, is the global average max pooling layer function with a convolution kernel size of 7×7. represents the convolution layer function with a convolution kernel size of 1 and a stride of 1. LN(∙) represents the layer normalization layer function, which performs normalization in the feature dimension. represents the GELU activation function. is the initial feature. is the output feature map, which is used for the subsequent generation of keys and values. is the number of channels of the feature map. is the length of the feature map. is the width of the feature map. Input the initial feature into the lightweight linear space compression layer to generate a lightweight key feature map and value feature map, and generate a key using the linear layer function , value , and input the initial feature into the linear layer function to generate a query , and its expression is as follows: Among them, and represent the generated keys and values and queries , and are the dimensions of the key and value respectively, and represent the linear layer function, represents the function of the layer normalization layer, represents the lightweight linear space compression layer function, which is used to reduce the number of parameters and computational complexity; Generate keys values and queries into h heads, each with a dimension of , , and perform self-attention calculations on each head separately. The expression is as follows: Among them, , respectively represent the query and key matrices of the i-th head, is the attention score matrix, represents the value matrix of the i-th head, is the output of the i-th head; Output of all headers The matrices are concatenated along the feature dimension to obtain a concatenated matrix , and the expression is: Among them, represents a splicing function; The concatenation matrix is mapped back to the original input dimension through a linear transformation to obtain , and then the final output has the following expression: ​ Among them, is the output weight matrix; The output Applying residual connections and layer normalization to generate the final global features, the expression is: Among them, represents an intermediate calculation result, represents a fully connected layer unit, which is composed of a fully connected layer, a GELU activation function, a Dropout regularization layer, and a fully connected layer, is the global feature.

4. The method according to claim 3, characterized in that, The extracting the initial features of the low-resolution face image includes: Extract the initial features of the low-resolution face image through an initial feature extraction module. The expression is as follows: Among them, represents the input low-resolution face image, represents a convolutional layer function with a convolutional kernel of 3×3, a stride of 1, a padding of 1, and an output channel of 64, is the initial feature extracted.

5. The method according to claim 4, wherein The fusion process of the reverse fusion module includes: If there are four hybrid residual attention modules connected in series, start from the output of the last hybrid residual attention module connected in series, directly use the features of this layer, and then layer by layer forward. Concatenate the features of the current layer with the fusion result of the previous layer, and then perform feature fusion through convolution operations. Obtain richer context information by accumulating multi-level feature information. The expression is as follows: Among them, represents the feature map extracted from the i-th layer of the hybrid residual attention module, represents the feature fusion result at the i-th layer, represents the 1×1 convolutional layer function, represents the concatenation operation of the feature maps.

6. The method according to claim 5, characterized in that The concatenating multiple hybrid residual attention modules in series, taking the initial features as the input, and fusing the multi-layer features through a reverse fusion module to obtain the target deep comprehensive features includes: If there are four cascaded hybrid residual attention modules, the initial feature is input into four consecutive hybrid residual attention modules, and the output result is sent to the reverse fusion module to extract the target deep comprehensive feature. The expression is as follows: Among them, represents the hybrid residual attention module, represents the reverse fusion module, is the finally extracted target deep comprehensive feature.

7. The method according to claim 6, wherein The inputting the target deep comprehensive features into an image upsampling module to generate a high-resolution face reconstruction image includes: Among them, represents a convolutional layer function with a convolution kernel size of 3×3, represents an upsampling convolutional sub-element function, which can determine the magnification factor by adjusting the number of output channels, and is a high-resolution face reconstruction image with the corresponding magnification factor.

Citation Information

Patent Citations

  • Lightweight double-branch convolutional neural network for image target detection and detection method thereof

    CN114648684A

  • Transform and convolutional neural network combination-based face super-resolution reconstruction method and system

    CN115953296A