Image Super-Resolution Reconstruction Model and Method Based on Residual Hybrid Attention Network

By combining the residual hybrid attention network of convolutional neural network and Transformer, the problem of difficult training and insufficient feature utilization of image super-resolution reconstruction models in the prior art is solved, and high-precision image super-resolution reconstruction is achieved, and more texture details are restored.

CN115222601BActive Publication Date: 2025-07-29FUZHOU UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210940743.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-06
Publication Date
2025-07-29
Estimated Expiration
2042-08-06

AI Technical Summary

Technical Problem

The existing image super-resolution reconstruction methods have excessive network depth, resulting in high training difficulties and high training skills requirements, and the convolutional neural network fails to make full use of the internal self-similarity and multi-scale features of the image, resulting in poor reconstruction results.

Method used

The image super-resolution reconstruction model based on residual hybrid attention network is adopted, combined with convolutional neural network and Transformer, and the residual separation hybrid attention module and efficient Swin Transformer module are used to extract local and global features, and the feature map is processed in parallel using channel separation technology to fuse high and low frequency information.

Benefits of technology

Higher precision image super-resolution reconstruction is achieved, more texture details are restored, image quality is improved, and extremely deep network structure is not required.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115222601B_ABST
    Figure CN115222601B_ABST
Patent Text Reader

Abstract

The present invention provides an image super-resolution reconstruction model and method based on a residual hybrid attention network, including: a shallow feature extraction module, a deep feature extraction module, and a reconstruction module. The shallow feature extraction module extracts shallow features from a low-resolution image; the deep feature extraction module consists of multiple cascaded residual separation hybrid attention groups and a global residual connection, extracts and fuses features from the shallow features to obtain deep features; the reconstruction module uses sub-pixel convolution to upsample the deep features to obtain an image with a higher resolution. The residual separation hybrid attention module uses channel separation technology to split the feature map and feeds it into two branch modules in parallel for processing, fuses the local features extracted by the residual triple attention module and the global features extracted by the efficient Swin Transformer module to obtain rich high- and low-frequency information. Through the method of the present invention, an image with richer details can be obtained, and higher-precision super-resolution reconstruction can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and image processing, and in particular to an image super-resolution reconstruction model and method based on a residual hybrid attention network. Background Art

[0002] In the field of electronic image applications, people often expect to obtain high-resolution images. High resolution means a high pixel density in the image, which can provide more details, and these details are indispensable in many practical applications. For example, high-resolution medical images are very helpful for doctors to make correct diagnoses; it is easy to distinguish similar objects from similar ones using high-resolution satellite images; if high-resolution images can be provided, the performance of pattern recognition in computer vision will be greatly improved.

[0003] Limited by funds and technology, the acquired pictures are difficult to reach the expected resolution in most cases. However, the image is restricted in many aspects during the acquisition and processing process, such as optical blurring caused by defocusing, diffraction, etc. in the digital imaging process, motion blurring caused by a limited shutter speed, the density of the sensing unit will affect the aliasing effect, random noise in the image sensor or during image transmission, etc. These factors will all affect the generation quality of the image. Therefore, it is very necessary to find a method to enhance the current resolution level.

[0004] Image super-resolution reconstruction technology, as a post-processing means, can enhance the resolution of images without increasing hardware costs. Image super-resolution reconstruction aims to restore a high-resolution image from a given low-resolution image through an algorithm. Currently, the methods of image super-resolution reconstruction can be mainly divided into three categories: interpolation-based super-resolution reconstruction, reconstruction-based super-resolution reconstruction, and learning-based super-resolution reconstruction. The interpolation method means increasing the size of the image by inserting new pixels around the original pixels of the image, and after inserting the pixels, values also need to be assigned to these pixels to restore the image content and achieve the effect of improving the image resolution. This kind of method is relatively simple in calculation, easy to understand and implement, but there will be problems such as ringing effect and serious loss of high-frequency information in the reconstruction result. The reconstruction method establishes an observation model for the image acquisition process, and then realizes super-resolution reconstruction by solving the inverse problem of the observation model. The reconstruction method has improved in restoring details, but its performance decreases as the scale factor increases, and this method is very time-consuming. The super-resolution reconstruction based on deep learning uses to establish an "end-to-end" mapping between low-resolution image patches and high-resolution image patches to restore high-frequency information and has achieved good reconstruction results.

[0005] Currently, mainstream algorithms often design very deep network architectures, which require long training time. Moreover, as the network depth increases, the training difficulty becomes greater, and the required training techniques also increase. At the same time, low-resolution inputs contain rich low-frequency information, which is equally treated among channels, hindering the convolutional neural network from learning more high-frequency information. In addition, the current convolutional neural networks for super-resolution do not fully utilize features at multiple scales, limiting the learning ability of the network. Meanwhile, simply using a convolutional neural network to build a model cannot fully utilize the self-similarity within the image and capture the long-range dependencies within the image. Therefore, it is very necessary to solve the existing problems and reconstruct high-quality images. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide an image super-resolution reconstruction model and method based on a residual hybrid attention network, so as to restore more texture details and improve the accuracy of image super-resolution reconstruction.

[0007] To achieve the above purpose, the present invention adopts the following technical solutions: An image super-resolution reconstruction model based on a residual hybrid attention network, including a shallow feature extraction module, a deep feature extraction module, and a reconstruction module: The shallow feature extraction module consists of a 3×3 convolutional layer. By utilizing the characteristic of the convolutional layer being good at extracting features, it extracts the shallow features of the low-resolution input image. The specific operation is as follows:

[0008] M0 = H SF (I LR )

[0009] Where H SF (·) represents the shallow feature extraction module, I LR represents the input low-resolution image, and M0 represents the shallow feature map;

[0010] The deep feature extraction module consists of multiple residual separation hybrid attention groups and a 3×3 convolution, and extracts high-level features from the shallow features. This process is expressed as follows:

[0011]

[0012] M DF = f 3×3 (M n )

[0013] Where represents the i-th multiple residual separation hybrid attention group, n represents the number of multiple hybrid attention residual groups, M i-1 , M i , M n represent the intermediate feature maps of the multiple residual separation hybrid attention groups, and f 3×3(·) represents a convolution operation with a convolution kernel size of 3×3, and M DF represents the deep feature map.

[0014] The reconstruction module consists of a sub-pixel convolution layer and a 3×3 convolution. A sub-pixel convolution layer is used to upsample the deep features extracted by the deep feature extraction module, and the information flow is reorganized into a feature map with a specified upsampling ratio. This process is described as:

[0015] I SR = H UP (M DF + M0) + Bicubic(I LR )

[0016] where I SR represents the high-resolution image after super-resolution reconstruction, H UP (·) represents the reconstruction module, and Bicubic(·) represents bicubic interpolation of the low-resolution image to the target resolution.

[0017] In a preferred embodiment, the multiple residual separation hybrid attention group consists of multiple residual separation hybrid attention modules and a 3×3 convolution layer.

[0018] In a preferred embodiment, the residual separation hybrid attention module consists of two 1×1 convolution layers, a residual triple attention module, an efficient Swin Transformer module, residual connection, channel splitting and splicing operations. The specific calculation formula expression is as follows:

[0019] X1, X2 = H SPL (f 1×1 (X))

[0020] Y1 = RTAB(X1)

[0021] Y2 = ESTB(X2)

[0022] Z = f 1×1 (H CAT (Y1, Y2)) + X

[0023] where X represents the input feature of the residual separation hybrid attention module, f 1×1 (·) represents a 1×1 convolution layer, H SPL (·) represents the operation of channel splitting, X1 and X2 represent the split feature maps, RTAB represents the operation of the residual triple attention module, Y1 represents the output feature of the residual triple attention module, ESTB represents the operation of the efficient Swin Transformer module, Y2 represents the output feature of the efficient Swin Transformer module, and H CAT(·) represents the operation of channel splicing, and Z represents the output feature of the residual separation and mixed attention module.

[0024] In a preferred embodiment, the residual triple attention module consists of two 3×3 convolutional layers, a RELU activation function, a residual connection, and a triple attention module. The specific calculation formula is as follows:

[0025] X O =F TAM (f 3×3 (RELU(f 3×3 (X I ))))+X

[0026] Among them, X I represents the input feature of the residual triple attention module, X o represents the output feature of the residual triple attention module, F TAM (·) represents the operation of the triple attention module, and RELU(·) represents the RELU activation function; among them, the triple attention module consists of two cross-dimensional interaction modules and a spatial attention module. The specific calculation formula is as follows:

[0027]

[0028] Among them represents the output features of the two cross-dimensional interaction modules, represents the output feature of the spatial attention module, and Y represents the output feature of the triple attention module.

[0029] In a preferred embodiment, the cross-dimensional interaction module includes a 7×7 convolutional layer, a dimension permutation operation, a Z-Pool layer, a channel connection operation, a max pooling operation, and an average pooling operation; the specific calculation formula is as follows:

[0030] X′1=H PER (X1)

[0031] Z-Pool(X)=H CAT (H MP (X), H AP (X))

[0032] X″1=Z-Pool(X′1)

[0033]

[0034]

[0035] Among them, X′1 represents the result of the dimension permutation operation, X″1 represents the result of the operation through the Z-Pool layer, and H PER(·) represents the dimension permutation operation, and Z-Pool(·) represents the operation of the Z-Pool layer, H CAT (·) represents the operation of concatenating along a specific dimension in the given input sequence feature map, H MP (·) and H AP (·) represent the max pooling operation and the average pooling operation along a specific dimension respectively, f 7×7 (·) represents the convolution operation with a convolution kernel size of 7×7, H IN (·) represents the instance normalization operation, δ(·) represents the Sigmoid function, and · represents the operation of multiplying by channels.

[0036] In a preferred embodiment, the spatial attention module includes a 7×7 convolutional layer, a Z-Pool layer, and an instance normalization operation. The specific calculation formula is as follows:

[0037]

[0038] where X3, represent the input feature and output feature of the spatial attention module.

[0039] In a preferred embodiment, the efficient Swin Transformer module is composed of two layer normalization operations, a self-attention calculation module based on moving windows, a local feature extraction feed-forward network, and a residual connection;

[0040] The calculation formula of the efficient Swin Transformer module is as follows:

[0041] Q = LN(W Q X)

[0042] K = LN(W K X)

[0043] V = LN(W V X)

[0044]

[0045]

[0046]

[0047] where, W Q 、W K 、W VThe transformation matrix for calculating Q, K, and V is denoted as, LN(·) represents the layer normalization operation, Q, K, and V respectively represent the query, key, and value matrices, SoftMax(·) represents the SoftMax function, SW-MSA(·) represents the shifted window self-attention module, and LeFF(·) represents the local enhanced multi-layer perceptron module. represents the output feature of the efficient Swin Transformer module, d represents the dimension of the K matrix, and Attention represents the self-attention calculation.

[0048] The present invention also provides an image super-resolution reconstruction method based on a residual hybrid attention network, which adopts the image super-resolution reconstruction model based on the residual hybrid attention network, and includes the following steps:

[0049] Step S1: Establish a training set according to the image degradation model to obtain N low-resolution images I LR and the corresponding real high-resolution images I LR corresponding to the N low-resolution images I HR ; where N is an integer greater than 1.

[0050] Step S2: Input the low-resolution images into the shallow feature extraction module to extract the shallow features of the images.

[0051] Step S3: Input the shallow features into the deep feature extraction module to extract deep features.

[0052] Step S4: Input the deep features into the reconstruction module, perform sub-pixel convolution to complete the upsampling process, and reconstruct the final high-resolution image.

[0053] Step S5: Optimize the image super-resolution reconstruction model through a loss function. The loss function uses the average L1 error between N reconstructed high-resolution images and the corresponding real high-resolution images, and the expression is:

[0054]

[0055] where L1 represents the L1 loss function.

[0056] Compared with the prior art, the present invention has the following beneficial effects: The present invention combines a convolutional neural network and a Transformer, and adopts a channel separation technique to split the feature map and feed it into two branch modules in parallel for processing. Through the residual separation hybrid attention module, the local features extracted by the triple attention module based on the convolutional neural network and the global features extracted by the efficient Swin Transformer module based on the Transformer are fused to obtain rich high- and low-frequency information. Through the method of the present invention, an image with richer details can be obtained, and higher-precision super-resolution reconstruction can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 is the structural diagram of the residual separation hybrid attention network in the preferred embodiment of the present invention;

[0058] Figure 2 is the structural diagram of the residual separation hybrid attention group in the preferred embodiment of the present invention;

[0059] Figure 3 is the structural diagram of the residual separation hybrid attention module in the preferred embodiment of the present invention;

[0060] Figure 4 is the structural diagram of the residual triple attention module in the preferred embodiment of the present invention;

[0061] Figure 5 is the structural diagram of the efficient Swin Transformer module in the preferred embodiment of the present invention;

[0062] Figure 6 is the schematic flow diagram of the image super-resolution reconstruction method in the preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0064] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0065] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0066] AsFigures 1 to 6 As shown in the figure, this embodiment provides an image super-resolution reconstruction model based on a residual separation hybrid attention network, where the model includes: a shallow feature extraction module, a deep feature extraction module, and a reconstruction module. The shallow feature extraction module extracts shallow features from the low-resolution image; the deep feature extraction module is composed of multiple cascaded residual separation hybrid attention groups and a global residual connection, extracts and fuses features from the shallow features to obtain deep features; the reconstruction module uses a sub-pixel convolutional layer to upsample the deep features to obtain a higher-resolution image. The residual separation hybrid attention module uses channel separation technology to split the feature map and feeds it into two branch modules in parallel for processing, fuses the local features extracted by the residual triple attention module and the global features extracted by the efficient Swin Transformer module to obtain rich high- and low-frequency information.

[0067] Step 1: Extract shallow features from the input image. Specifically:

[0068] Utilizing the characteristic that the convolutional layer is good at extracting features, a 3×3 convolutional layer is used to extract the shallow features of the low-resolution input image. The specific operation is as follows:

[0069] M0 = H SF (I LR )

[0070] where H SF (·) represents the shallow feature extraction module, I LR represents the input low-resolution image, and M0 represents the shallow feature map;

[0071] Step 2: Pass the shallow features through the deep feature extraction module composed of multiple cascaded residual separation hybrid attention groups and a global residual connection to perform feature extraction and feature fusion to obtain deep features:

[0072] Step 2.1: As Figure 2 shown, by inputting the shallow features into the deep feature extraction module composed of n cascaded residual separation hybrid attention groups and a global residual connection, feature extraction and feature fusion can be performed on the shallow features to obtain richer and deeper features. Among them, the cascaded residual separation hybrid attention group contains m cascaded residual separation hybrid attention modules. In this example, n = 6 and m = 6 are taken. The specific calculation is represented by the following formula:

[0073]

[0074] M DF = f 3×3 (M n )

[0075] where Denote the i-th multi-residual separation hybrid attention group, n denotes the number of multi-hybrid attention residual groups, M i-1 、M i 、M n represent the intermediate feature maps of the multi-residual separation hybrid attention group, f 3×3 (·) represents a convolution operation with a convolution kernel size of 3×3, M DF represents the deep feature map.

[0076] Figure 3 is the structural diagram of the residual separation hybrid attention module. This module contains 1 residual triple attention module, 1 efficient Swin Transformer module, two 1×1 convolutional layers, and channel splitting and connection operations. Its calculation is expressed by the following formula:

[0077] X1, X2 = H SPL (f 1×1 (X))

[0078] Y1 = RTAB(X1)

[0079] Y2 = ESTB(X2)

[0080] z = f 1×1 (H CAT (Y1, Y2)) + X

[0081] Among them, X represents the input feature of the residual separation hybrid attention module, f 1×1 (·) represents a 1×1 convolutional layer, H SPL (·) represents the operation of channel splitting, X1 and X2 represent the split feature maps, RTAB represents the operation of the residual triple attention module, Y1 represents the output feature of the residual triple attention module, ESTB represents the operation of the efficient Swin Transformer module, Y2 represents the output feature of the efficient Swin Transformer module, H CAT (·) represents the operation of channel concatenation, and Z represents the output feature of the residual separation hybrid attention module.

[0082] Figure 4 is the structural diagram of the residual triple attention module. This module contains two 3×3 convolutional layers, a RELU activation function, a triple attention module, an average pooling operation, and a residual connection. The residual triple attention module makes full use of the advantages of CNN in extracting inherent features by leveraging the correlation between different dimensions of the feature map, enabling the network to learn richer image information. Its calculation is expressed by the following formula:

[0083] X O = F TAM (f 3×3 (RELU(f3×3 (X I ))))+X

[0084] Among them, X I represents the input feature of the residual triple attention module, and X O represents the output feature of the residual triple attention module. F TAM (·) represents the operation of the triple attention module, and RELU(·) represents the RELU activation function. The triple attention module is composed of two cross-dimensional interaction modules and a spatial attention module. The specific calculation formula is as follows:

[0085]

[0086] Among them represents the output features of the two cross-dimensional interaction modules, represents the output features of the spatial attention module, and Y represents the output features of the triple attention module. The cross-dimensional interaction module includes a 7×7 convolutional layer, a dimension permutation operation, a Z-Pool layer, a channel connection operation, a max pooling operation, and an average pooling operation. The specific calculation formula is as follows:

[0087] X′1 = H PER (X1)

[0088] Z-Pool(X) = H CAT (H MP (X), H AO (X))

[0089] X″1 = Z-Pool(X′1)

[0090]

[0091]

[0092] Among them, X′1 represents the result of the dimension permutation operation, and X″1 represents the result of the operation of the Z-Pool layer. H PER (·) represents the dimension permutation operation, Z-Pool(·) represents the operation of the Z-Pool layer, H CAT (·) represents the operation of connecting along a specific dimension in the given input sequence feature map, H MP (·) and H AP (·) respectively represent the max pooling operation and the average pooling operation along a specific dimension. f 7×7 (·) represents the convolution operation with a convolution kernel size of 7×7, H IN (·) represents the instance normalization operation, δ(·) represents the Sigmoid function, and · represents the operation of multiplying by channels, Represents the output features of the cross-dimensional interaction module.

[0093] The spatial attention module includes a 7×7 convolutional layer, a Z-Pool layer, and an instance normalization operation. The specific calculation formula is as follows:

[0094]

[0095] Where X3, Represents the input and output features of the spatial attention module.

[0096] Figure 5 Figure Figure 5 is the structural diagram of the efficient Swin Transformer module. This module includes two layer normalization operations, a shifted window self-attention module, a local enhanced multi-layer perceptron module, and a residual connection operation. Its calculation is represented by the following formula:

[0097] Q = LN(W Q X)

[0098] K = LN(W K X)

[0099] V = LN(W V X)

[0100]

[0101]

[0102]

[0103] Where, W Q , W K , W V Represent the transformation matrices for calculating Q, K, and V. d represents the dimension of the K matrix. Attention represents the self-attention calculation. LN(·) represents the layer normalization operation. Q, K, and V represent the query, key, and value matrices. SoftMax(·) represents the SoftMax function. SW-MSA(·) represents the shifted window self-attention module. LeFF(·) represents the local enhanced multi-layer perceptron module. Represents the output features of the efficient Swin Transformer module.

[0104] Step 3: Reconstruct the deep feature map. The predicted feature map is upsampled to reconstruct a high-resolution image. The expression is as follows:

[0105] I SR = H UP (M DF + M0) + Bicubic(I LR )

[0106] Among them, I SR represents the high-resolution image after super-resolution reconstruction, and H UP (·) represents the reconstruction module, and Bicubic(·) represents bicubic interpolation of the low-resolution image to the target resolution.

[0107] An image super-resolution reconstruction method applied to the above image super-resolution reconstruction model is described in detail as follows.

[0108] S1. Establish a training set according to the image degradation model to obtain N low-resolution images I LR and the corresponding real high-resolution images I LR for the N low-resolution images I HR ; where N is an integer greater than 1;

[0109] S2. Input the low-resolution image into the shallow feature extraction module to extract the shallow features of the image;

[0110] S3. Input the shallow features into the deep feature extraction module to extract deep features;

[0111] S4. Input the deep features into the reconstruction module, perform sub-pixel convolution to complete the upsampling process, and reconstruct the final high-resolution image;

[0112] S5. Optimize the image super-resolution reconstruction model through the loss function. The loss function uses the average L1 error between the N reconstructed high-resolution images and the corresponding real high-resolution images, and the expression is:

[0113]

[0114] where L1 represents the L1 loss function.

[0115] To better illustrate the effectiveness of the present invention, the embodiments of the present invention also use a comparative experiment to compare the reconstruction effects.

[0116] Specifically, the embodiments of the present invention use 800 high-resolution images in DIV2K as the training set, and use Set5, Set14, B100, Urban100, and Manga109 as the test sets respectively. Perform bicubic downsampling on the original high-resolution images to obtain the corresponding low-resolution images.

[0117] After constructing the training set, the model is trained and tested on the PyTorch framework. The low-resolution images in the training set are cropped into 64×64 image patches, and 48 groups of 64×64 image patches are randomly input each time for 500 epochs of training. The Adam gradient descent method is used to optimize the network parameters, where the parameters of the Adam optimizer are set as β1 = 0.9, β2 = 0.999, and ε = 10 -8 . The initial learning rate is set to 2×10 -4 , and it is halved after the {250, 400, 425, 450, 475}th epochs. The number of RSHABs is set to 36, and the number of channels is set to 180. The peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) are used to evaluate the model performance. The present invention uses five benchmark datasets, namely Set5, Set14, B100, Urban100, and Manga109, to test the model performance. For the comparative experiment, 11 representative image super-resolution reconstruction methods are selected to compare with the experimental results of the present invention. The experimental results are shown in Table 1, where RSHAN is the method proposed in the present invention.

[0118] Table 1. Comparison of average PSNR and SSIM values on 5 test sets

[0119]

[0120] In summary, the present invention has the following advantages and effects compared with the prior art:

[0121] (1) In the embodiment of the present invention, by adopting the residual separation hybrid attention module, the local modeling ability of the residual triple attention module based on the convolutional neural network and the non-local modeling ability of the efficient Swin Transformer module are combined. On the premise of ensuring that the number of parameters is similar to that of the currently best-performing reconstruction model, the information relationship between different dimensions of the image is effectively utilized, significantly improving the ability of the super-resolution reconstruction model to extract high-frequency information. At the same time, by using the captured long-distance dependence relationship, the extracted features are made more abundant.

[0122] (2) In the embodiment of the present invention, by adopting the structure of global residual embedded cascaded residuals, the network can bypass the low-frequency information in the low-resolution input, learn more high-frequency residual information, and obtain rich detail features. And without building a very deep network, high-resolution reconstructed images with good effects can also be obtained.

Claims

1. An image super-resolution reconstruction system based on a residual hybrid attention network, characterized in that It includes a shallow feature extraction module, a deep feature extraction module, and a reconstruction module: The shallow feature extraction module consists of a 3×3 convolutional layer. By leveraging the feature extraction ability of the convolutional layer, it extracts the shallow features of the low-resolution input image. The specific operations are as follows: M0 = H SF (I LR ) Among them, H SF (·) represents the shallow feature extraction module, I LR represents the input low-resolution image, and M0 represents the shallow feature map; The deep feature extraction module consists of multiple residual separation hybrid attention groups and a 3×3 convolution, and extracts high-level features from the shallow features. This process is expressed as follows: M DF = f 3×3 (M n ) Among them represents the i-th multiple residual separation mixed attention group, n represents the number of multiple mixed attention residual groups, M i-1 、M i 、M n represent the intermediate feature maps of the multiple residual separation mixed attention group, f 3×3 (·) represents a convolution operation with a convolution kernel size of 3×3, M DF represents the deep feature map; The reconstruction module consists of a sub-pixel convolutional layer and a 3×3 convolution. It uses a sub-pixel convolutional layer to upsample the deep features extracted by the deep feature extraction module and reorganizes the information flow into a feature map with a specified upsampling ratio. This process is described as: I SR = H UP (M DF + M0)+ Bicubic(I LR ) where I SR represents the high-resolution image after super-resolution reconstruction, HUP(·) represents the reconstruction module, and Bicubic(·) represents bicubic interpolation of the low-resolution image to the target resolution; The multiple residual separation hybrid attention group consists of multiple residual separation hybrid attention modules and a 3×3 convolutional layer; The residual separation hybrid attention module consists of two 1×1 convolutional layers, a residual triple attention module, an efficient Swin Transformer module, residual connections, channel splitting, and splicing operations. The specific calculation formula expression is as follows: X1,X2 = H SPL (f 1×1 (X)) Y1 = RTAB(X1) Y2 = ESTB(X2) Z = f 1×1 (H CAT (Y1, Y2)) + X Among them, X represents the input feature of the residual separation and hybrid attention module, and f 1×1 (·) represents a 1×1 convolutional layer, and H SPL (·) represents the operation of channel splitting. X1 and X2 represent the split feature maps. RTAB represents the operation of the residual triple attention module. Y1 represents the output feature of the residual triple attention module. ESTB represents the operation of the efficient Swin Transformer module. Y2 represents the output feature of the efficient Swin Transformer module, and H CAT (·) represents the operation of channel concatenation. Z represents the output feature of the residual separation and hybrid attention module; The residual triple attention module consists of two 3×3 convolutional layers, a RELU activation function, residual connections, and a triple attention module. The specific calculation formula expression is as follows: X O = F TAM (f 3×3 (RELU(f 3×3 (X I )))) + X Among them, X I represents the input feature of the residual triple attention module, and X O represents the output feature of the residual triple attention module. F TAM (·) represents the operation of the triple attention module, and RELU(·) represents the RELU activation function. The triple attention module is composed of two cross-dimensional interaction modules and a spatial attention module. The specific calculation formula is as follows: Among them represents the output features of two cross-dimensional interaction modules represents the output features of the spatial attention module, and Y represents the output features of the triple attention module; The efficient Swin Transformer module consists of two layer normalization operations, a self-attention calculation module based on moving windows, a local feature extraction feed-forward network, and residual connections; The calculation formula of the efficient Swin Transformer module is as follows: Q = LN(W Q X) K = LN(W K X) V = LN(W V X) Among them, W Q , W K , W V represent the transformation matrices for calculating Q, K, and V. LN(·) represents the layer normalization operation. Q, K, and V represent the query, key, and value matrices respectively. SoftMax(·) represents the SoftMax function. SW-MSA(·) represents the shifted window multi-head self-attention module. LeFF(·) represents the local enhanced multi-layer perceptron module. represents the output feature of the efficient Swin Transformer module. d represents the dimension of the K matrix. Attention represents the self-attention calculation.

2. The image super-resolution reconstruction system based on the residual hybrid attention network according to claim 1, wherein The cross-dimensional interaction module includes a 7×7 convolutional layer, a dimension permutation operation, a Z-Pool layer, a channel connection operation, a max pooling operation, and an average pooling operation; The specific calculation formula is as follows: X1′ = H PER (X1) Z-Pool(X) = H CAT (H MP (X), H AP (X)) X1″ = Z-Pool(X1′) Among them, x‘1 represents the result of the dimension permutation operation, X″1 represents the result of the operation through the Z-Pool layer, H PER (·) represents the dimension permutation operation, Z-Pool(·) represents the operation of the Z-Pool layer, H CAT (·) represents the operation of concatenating along a specific dimension in the feature map of the given input sequence, H MP (·) and H AP (·) represent the max pooling operation and the average pooling operation along a specific dimension respectively, f 7×7 (·) represents the convolution operation with a convolution kernel size of 7×7, H IN (·) represents the instance normalization operation, δ(·) represents the Sigmoid function, and · represents the operation of multiplying by channels.

3. The image super-resolution reconstruction system based on the residual hybrid attention network according to claim 1, wherein, The spatial attention module includes a 7×7 convolutional layer, a Z-Pool layer, and an instance normalization operation. The specific calculation formula is as follows: Among them, X3, represent the input feature and output feature of the spatial attention module.

4. Image super-resolution reconstruction method based on residual hybrid attention network, characterized in that, The image super-resolution reconstruction system based on the residual hybrid attention network described in any one of claims 1-3 is adopted, including the following steps: Step S1: Establish a training set according to the image degradation model to obtain N low-resolution images I LR and the corresponding true high-resolution images I LR for the N low-resolution images I HR ; where N is an integer greater than 1; Step S2: Input the low-resolution image into the shallow feature extraction module to extract the shallow features of the image; Step S3: Input the shallow features into the deep feature extraction module to extract the deep features; Step S4: Input the deep features into the reconstruction module, perform sub-pixel convolution to complete the upsampling process, and reconstruct the final high-resolution image; Step S5: Optimize the image super-resolution reconstruction model through a loss function. The loss function uses the average L1 error between N reconstructed high-resolution images and the corresponding real high-resolution images. The expression is: Where L1 represents the L1 loss function.

Citation Information

Patent Citations

  • Mine image super-resolution reconstruction method and system based on multi-scale residual network

    CN113592718A

  • Image super-resolution reconstruction method and system based on residual channel attention network

    CN114429422A