Lightweight Image Super-Resolution Reconstruction Method Based on Dual Attention Mechanism
By introducing a lightweight design of the dual attention mechanism in image super-resolution network, the high consumption problems existing in storage and computing of existing large-scale network models are solved, and a more efficient image super-resolution reconstruction effect is achieved.
Patent Information
- Application Number
- CN202211168904.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-09-25
AI Technical Summary
Existing large-scale image super-resolution network models have huge consumption in storage space and computing costs, and are difficult to use on mobile devices.
A lightweight image super-resolution reconstruction method based on the dual attention mechanism is proposed. By constructing a lightweight network model, it uses enhanced local feature extraction blocks, simple channel attention modules and enhanced spatial attention modules to extract rich hierarchical features and important channel information and spatial features.
It achieves better image super-resolution reconstruction effect under smaller parameters and calculation amounts, improves the reconstruction performance of network models, and obtains better results on image quality indicators PSNR and SSIM.
Smart Images

Figure CN115496658B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more particularly, to a lightweight image super-resolution reconstruction method based on a dual attention mechanism in the field of image super-resolution reconstruction technology. Background Art
[0002] The goal of image super-resolution (SR) reconstruction is to recover the corresponding high-resolution image from a given low-resolution image. Image super-resolution reconstruction, as a typical low-level computer vision task, plays an important role in fields such as satellite remote sensing, biometric recognition, image analysis, and video surveillance. However, SR is an ill-posed problem, that is, a low-resolution (LR) image can be degraded from different high-resolution (HR) images. The rise of deep learning has provided a powerful tool for solving this problem. Dong et al. first applied the CNN-based method to image super-resolution reconstruction and developed an SRCNN network with only three convolutional layers. Kim et al. proposed a VDSR network with 20 convolutional layers, achieving better performance than SRCNN. The results show that the reconstruction performance of image super-resolution can be improved by deepening the depth of the network. Subsequently, the RDN network and RCAN network proposed by Zhang et al. increased the network depth to 100 layers and 400 layers respectively.
[0003] However, these large-scale network models consume huge storage space and a large amount of computational cost. In practical applications, it is difficult to use these network models on mobile devices. For this reason, scholars have successively proposed a series of strategies to reduce the number of parameters or computational amount. The DRRN network proposed by Tai et al. uses a recursive mechanism to reduce the number of network parameters. The CARN-M network proposed by Ahn uses grouped convolution to reduce the number of network parameters. The IMDN network proposed by Hui et al. uses information multi-distillation blocks to gradually extract hierarchical features and aggregates them according to the importance of candidate features. Although these super-resolution network models consider reducing the number of parameters when performing lightweight design, there are still many redundant computations, and a more efficient super-resolution network model can be further constructed by reducing redundant computations and using more effective modules. Summary of the Invention
[0004] The present invention proposes a lightweight image super-resolution reconstruction method based on a dual attention mechanism.
[0005] In order to solve the above technical problems, the technical solution of the present invention is as follows:
[0006] A lightweight image super-resolution reconstruction method based on a dual attention mechanism, comprising the following steps:
[0007] Step 1, constructing a training data set:
[0008] (1a) First, perform data augmentation on 900 high-resolution images in the DIV2K dataset, and then perform bicubic interpolation downsampling on the high-resolution images after data augmentation to obtain corresponding low-resolution images;
[0009] (1b) Combine the high-resolution images and the corresponding low-resolution images to form a pair of training samples, thereby obtaining a training dataset;
[0010] Step 2, construct a super-resolution network model:
[0011] The network model includes four parts: a shallow feature extraction module, an enhanced local feature extraction group, a deep feature fusion module, and an upsampling module. The shallow feature extraction module consists of a 3×3 convolutional layer for extracting shallow features; the enhanced local feature extraction group is stacked by 8 enhanced local feature extraction blocks. After inputting the shallow features into the enhanced local feature extraction group, 8 enhanced local feature extraction blocks sequentially extract deep-level features; the deep feature fusion module fuses and filters the deep-level features to obtain deep fusion features; finally, add the shallow features and the deep fusion features, and obtain the reconstructed high-resolution image through the upsampling module;
[0012] Step 3, train the super-resolution network model to obtain a trained super-resolution network model:
[0013] Input all the training samples in the training dataset into the super-resolution network model, and use the gradient descent method to iteratively update the parameters of the network until the loss function converges, thereby obtaining a trained super-resolution network model;
[0014] Step 4, perform super-resolution reconstruction on the low-resolution image:
[0015] Input the low-resolution image in the natural scene into the trained super-resolution network model, and obtain the super-resolution reconstructed high-resolution image after processing.
[0016] The data augmentation in step (1a) refers to: performing random horizontal flipping, random vertical flipping, and random rotation by 90°, 180°, and 270° on the high-resolution images of the DIV2K dataset.
[0017] The bicubic interpolation downsampling in step (1a) refers to: I LR = I HR ↓ S , I LR represents the low-resolution image, I HR represents the high-resolution image, ↓ S represents the bicubic interpolation downsampling operation, and S represents the downsampling factor.
[0018] The enhanced local feature extraction block in step 2 consists of two enhanced spatial attention modules and two residual local feature layers. First, important spatial features are selected from the input features of the first enhanced spatial attention module, and then intermediate features are extracted through two stacked residual local feature layers. Finally, another enhanced spatial attention module is used to obtain more important feature regions. The enhanced spatial attention module consists of a 1×1 convolutional layer, a 3×3 convolutional layer, a max pooling layer, a convolutional group (two 3×3 convolutional layers in series), an upsampling layer, and a Sigmoid function. The residual local feature layer consists of an LN layer, two 1×1 convolutional layers, a 3×3 convolutional layer, a GELU activation function, and a simple channel attention module. The simple channel attention module consists of a global average pooling layer, a 1×1 convolutional layer, and a Sigmoid activation function.
[0019] The depth feature fusion module in step 2 consists of two stacked convolutional layers (a 1×1 convolutional layer and a 3×3 convolutional layer) and a GELU activation function, which fuses and filters the depth hierarchical features.
[0020] The upsampling module in step 2 consists of a 3×3 convolutional layer and an efficient sub-pixel convolutional layer. The 3×3 convolutional layer is used to compress the channels of the feature map, and the sub-pixel convolutional layer reorganizes the low-resolution feature map through multiple channels to obtain a high-resolution image.
[0021] In step 3, in the training of the super-resolution network model of the present invention, the L1 loss function is adopted. The mathematical expression of the L1 loss function is:
[0022]
[0023] L (θ) represents the L1 loss function, θ represents the parameters to be learned by the super-resolution network model, N represents the number of images used for training, and ‖·‖ 1 represents the mean absolute error, represents the i-th high-resolution original image, represents the i-th image after super-resolution reconstruction.
[0024] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention proposes a lightweight image super-resolution reconstruction method based on a dual attention mechanism. This method constructs a lightweight network model, extracts rich hierarchical features through an efficient enhanced local feature extraction block, and introduces a simple channel attention module and an enhanced spatial attention module to obtain important channel information and spatial feature information, thereby further improving the reconstruction performance of the network model. Experimental results show that this method can achieve better image super-resolution reconstruction results with fewer parameters and computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 Flowchart of the lightweight image super-resolution reconstruction method based on the dual attention mechanism in Embodiment 1
[0026] Figure 2 Schematic diagram of the super-resolution network model in Embodiment 1
[0027] Figure 3 Schematic diagram of the enhanced local feature extraction block in Embodiment 1
[0028] Figure 4 Schematic diagram of the residual local feature layer in Embodiment 1
[0029] Figure 5 Schematic diagram of the simple channel attention module in Embodiment 1
[0030] Figure 6 Schematic diagram of the enhanced spatial attention module in Embodiment 1
[0031] Figure 7 Schematic diagram of the random super-resolution reconstruction effect in Embodiment 1 DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0033] Refer to Figure 1 to further describe the implementation steps of the present invention in detail.
[0034] Embodiment 1
[0035] Step 1, construct a training dataset:
[0036] (1a) First, perform data augmentation on 900 high-resolution images in the DIV2K dataset, and then perform bicubic interpolation downsampling on the data-augmented high-resolution images to obtain corresponding low-resolution images.
[0037] (1b) Combine the high-resolution images and the corresponding low-resolution images into a pair of training samples to obtain a training dataset.
[0038] The data augmentation described in step (1a) refers to: randomly horizontally flipping, vertically flipping, randomly rotating by 90°, 180°, and 270° the high-resolution images in the DIV2K dataset;
[0039] The bicubic interpolation downsampling described in step (1a) refers to: I LR = I HR ↓ S where I LR represents the low-resolution image, I HR represents the high-resolution image, ↓ S represents the bicubic interpolation downsampling operation, and S represents the downsampling factor.
[0040] Step 2, construct a super-resolution network model:
[0041] The structure of the super-resolution network model is as Figure 2 shown. The super-resolution network model includes a shallow feature extraction module, an enhanced local feature extraction group, a deep feature fusion module, and an upsampling module.
[0042] (2a) Shallow feature extraction module:
[0043] The shallow feature extraction module consists of a 3×3 convolutional layer. The low-resolution image is input into the shallow feature extraction module to obtain the shallow feature F 0 , and its expression formula is as follows:
[0044] F 0 = W 0 * I LR
[0045] where W 0 represents the weight of the 3×3 convolutional layer, I LR represents the low-resolution image, and * represents the convolution operation.
[0046] (2b) Enhanced local feature extraction group:
[0047] The enhanced local feature extraction group contains 8 enhanced local feature extraction blocks (ELFB). Each enhanced local feature extraction block consists of two enhanced spatial attention modules (ESA) and two residual local feature layers (RLFL). The shallow feature F 0 is input into the enhanced local feature extraction group to extract deep-level features F d , where d = 1, 2, 3... D, and D is the number of enhanced local feature extraction blocks;
[0048] Specifically, 8 enhanced local feature extraction blocks are used for deep hierarchical feature extraction. The deep hierarchical feature F of the output of the d-th enhanced local feature extraction block d (1 ≤ d ≤ D) has the following expression formula:
[0049]
[0050] In the formula, represents the d-th enhanced local feature extraction block.
[0051] The structure of the enhanced local feature extraction block (ELFB) is as Figure 3 shown. Given the input feature F d-1 , the d-th enhanced local feature extraction block first selects important spatial features from the input of the first enhanced spatial attention module (ESA), then extracts intermediate features through two residual local feature layers (RLFL), and finally uses another enhanced spatial attention module (ESA) to obtain more important spatial features. Its expression formula is as follows:
[0052] F d = H ESA (H RLFL (H RLFL (H ESA (F d-1 ))))
[0053] In the formula, H RLFL (·) represents the residual local feature layer, and H ESA (·) represents the enhanced spatial attention module.
[0054] The structure of the residual local feature layer (RLFL) is as Figure 4 shown. The residual local feature layer first normalizes the input feature through layer normalization (LN), then expands the number of channels using the first 1×1 convolutional layer, then uses a 3×3 grouped convolutional layer to obtain rich local features, then adds non-linearity to the network model through the GELU activation function, then uses a simple channel attention module (SCA) to obtain important channel features, then refines the features through a 1×1 convolutional layer, and finally adds the refined features to the original input feature to obtain the output feature Its expression formula is as follows:
[0055]
[0056] In the formula, W 1 represents the weight of the first 1×1 convolutional layer, and W 2 represents the weight of the 3×3 grouped convolutional layer, and W 3Denote the weights of the last 1×1 convolutional layer, H SCA (·) represents the simple channel attention module.
[0057] The structure of the simple channel attention module (SCA) is as Figure 5 shown. In the simple channel attention module, first, the input feature is subjected to global average pooling operation, then passed through a 1×1 convolutional layer, and finally, the output feature is normalized using the Sigmoid activation function. Its expression formula is as follows:
[0058]
[0059] In the formula, Pool(·) represents the global average pooling layer, W 1 denotes the weights of the 1×1 convolutional layer, denotes the output feature of the simple channel attention module.
[0060] The structure of the enhanced spatial attention module (ESA) is as Figure 6 shown. In the enhanced spatial attention module, the input feature first passes through a 1×1 convolutional layer that reduces the input channels to obtain a compact feature then through a 3×3 convolutional layer with a stride of 2, followed by a max-pooling layer H pool (·) performs max-pooling operation, and then passes through a convolutional group (two 3×3 convolutional layers in series) H g (·), followed by an upsampling layer H up (·) implemented by bilinear interpolation to restore the feature spatial size, and added to the feature to obtain the intermediate feature Finally, a 1×1 convolutional layer is used to restore the number of input channels, and the Sigmoid function is used to obtain the important spatial features Its expression formula is as follows:
[0061]
[0062]
[0063]
[0064] In the formula, denotes the weights of the first 1×1 convolutional layer, denotes the weights of the 3×3 convolutional layer, denotes the weights of the last 1×1 convolutional layer.
[0065] (2c) Depth feature fusion module:
[0066] The depth feature fusion module fuses and filters the depth-level features through two stacked convolutional layers (a 1×1 convolutional layer and a 3×3 convolutional layer) and a GELU activation function, and the depth-level feature F d (d = 8) is input into the depth feature fusion module to obtain the depth fusion feature F DFF , and its expression formula is as follows:
[0067] F DFF = W 2 *(GELU(W 1 *F d ))
[0068] In the formula, W 1 , W 2 respectively represent the weights of the 1×1 convolutional layer and the 3×3 convolutional layer.
[0069] (2d) Upsampling module:
[0070] The upsampling module consists of a 3×3 convolutional layer and an efficient sub-pixel convolutional layer. First, the depth fusion feature F DFF is added to the shallow feature F 0 , and then input into the upsampling module to obtain the final super-resolution reconstructed image I SR , and its expression formula is as follows:
[0071] I SR = F up (W 3 *(F DFF + F 0 ))
[0072] In the formula, F up represents the channel rearrangement operation of the sub-pixel convolutional layer, and W 3 represents the weight of the 3×3 convolutional layer.
[0073] Step 3, train the super-resolution network model to obtain the trained super-resolution network model:
[0074] Input all the training samples in the training dataset into the super-resolution network model, and use the gradient descent method to iteratively update the parameters of the network until the loss function converges, so as to obtain the trained super-resolution network model;
[0075] In step 3, in the training of the super-resolution network model of the present invention, the L1 loss function is adopted, and the mathematical expression of the L1 loss function is:
[0076]
[0077] L (θ)Let \(L_1\) denote the loss function, \(\theta\) denote the parameters to be learned by the super-resolution network model, \(N\) denote the number of images used for training, and \(\|\cdot\|\) 1 denote the mean absolute error, denote the \(i\)-th high-resolution original image, denote the \(i\)-th image after super-resolution reconstruction.
[0078] Step 4, perform super-resolution reconstruction on the low-resolution image:
[0079] Input the low-resolution image in the natural scene into the trained super-resolution network model, and after processing, obtain the high-resolution image of super-resolution reconstruction.
[0080] To better illustrate the technical effects of the present invention, in this embodiment, the method provided by the present invention and the existing Bicubic algorithm, SRCNN algorithm, FSRCNN algorithm, VDSR algorithm, LapSRN algorithm, DRRN algorithm, IDN algorithm, CARN algorithm, IMDN algorithm are respectively used to conduct 4-fold super-resolution experiments on five benchmark datasets for image super-resolution: Set5, Set14, B100, Urban100, Manga109, and the obtained results are shown in Table 1
[0081] Table 1 is the PSNR / SSIM index analysis of the 4-fold super-resolution results
[0082]
[0083] Among them, the evaluation metrics are respectively the peak signal-to-noise ratio (PSNR) and the structural similarity (SSIM). The higher these two values are, the more similar the details and results of the obtained high-resolution image are to the original real image, and the better the super-resolution reconstruction effect. The bold indicates that this result is the best result.
[0084] Meanwhile, the superiority of the present invention can also be verified from the visualization results.
[0085] Specifically, as Figure 7 shown, the present invention has the best super-resolution reconstruction effect on randomly selected pictures. In the case of 4-fold super-resolution of the Barbara image in the Set14 dataset, the method of the present invention can more correctly and clearly restore the texture information of the book. These visualization results further demonstrate the effectiveness of the present invention.
[0086] In summary, compared with the mainstream algorithms, the method proposed in the present invention not only achieves better results in terms of the image quality metrics PSNR and SSIM, but also has a lower number of parameters and computational complexity of the network model, realizing a good trade-off between the model size and the image reconstruction quality. It can be seen therefrom that a lightweight image super-resolution reconstruction method based on a dual attention mechanism proposed in the present invention is a lightweight and efficient super-resolution reconstruction method.
[0087] The above description is only a specific embodiment of the present invention and does not limit the implementation manners of the present invention. However, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. cannot be made without departing from the spirit and principle of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A lightweight image super-resolution reconstruction method based on a dual attention mechanism, characterized in that, it specifically includes the following steps: Step 1, construct a training data set: (1a) First, perform data augmentation on 900 high-resolution images in the DIV2K data set, and then perform bicubic interpolation downsampling on the data-augmented high-resolution images to obtain corresponding low-resolution images; (1b) Combine the high-resolution images and the corresponding low-resolution images into a pair of training samples to obtain a training data set; Step 2, construct a super-resolution network model: The network model includes four parts: a shallow feature extraction module, an enhanced local feature extraction group, a deep feature fusion module, and an upsampling module: The shallow feature extraction module consists of a 3×3 convolutional layer for extracting shallow features; The enhanced local feature extraction group consists of two enhanced spatial attention modules and two residual local feature layers; The enhanced spatial attention module consists of a 1×1 convolutional layer, a 3×3 convolutional layer, a max pooling layer, a convolutional group composed of two 3×3 convolutional layers in series, an upsampling layer, and a Sigmoid function. This module can select important spatial features; The residual local feature layer consists of an LN layer, two 1×1 convolutional layers, a 3×3 convolutional layer, a GELU activation function, and a simple channel attention module; The simple channel attention module consists of a global average pooling layer, a 1×1 convolutional layer, and a Sigmoid activation function. This module can select important channel features; After inputting the shallow features into the enhanced local feature extraction group, 8 enhanced local feature extraction blocks are used to extract deep hierarchical features in sequence; The deep feature fusion module fuses and filters the deep hierarchical features to obtain deep fusion features; Finally, add the shallow features and the deep fusion features, and obtain the reconstructed high-resolution image through the upsampling module; Step 3, train the super-resolution network model to obtain a trained super-resolution network model: Input all the training samples in the training data set into the super-resolution network model, and use the gradient descent method to iteratively update the parameters of the network until the loss function converges, thereby obtaining a trained super-resolution network model; Step 4, perform super-resolution reconstruction on the low-resolution image: Input the low-resolution image in the natural scene into the trained super-resolution network model, and obtain the super-resolution reconstructed high-resolution image after processing.
2. The lightweight image super-resolution reconstruction method based on a dual attention mechanism according to claim 1, characterized in that, The data augmentation in the step (1a) refers to randomly horizontally flipping, randomly vertically flipping, and randomly rotating the high-resolution images in the DIV2K dataset by 90°, 180°, and 270°; the bicubic interpolation downsampling in the step (1a) means that: I LR = I HR ↓ S , where I LR represents the low-resolution image, I HR represents the high-resolution image, ↓ S represents the bicubic interpolation downsampling operation, and S represents the downsampling factor.
3. The lightweight image super-resolution reconstruction method based on a dual attention mechanism according to claim 1, characterized in that, the deep feature fusion module in step 2 consists of a 1×1 and a 3×3 stacked convolutional layer and a GELU activation function, and this module fuses and filters the deep hierarchical features.
4. A lightweight image super-resolution reconstruction method based on a dual attention mechanism according to claim 1, characterized in that, the upsampling module in step 2 is composed of a 3×3 convolutional layer and an efficient sub-pixel convolutional layer, where the 3×3 convolutional layer is used to compress the feature map channels, and the sub-pixel convolutional layer obtains a high-resolution image by reorganizing between multiple channels of the low-resolution feature map.
5. A lightweight image super-resolution reconstruction method based on a dual attention mechanism according to claim 1, characterized in that, in step 3, in the training of the super-resolution network model, the L1 loss function is adopted, and the mathematical expression of the L1 loss function is: L (θ) represents the L1 loss function, θ represents the parameters to be learned by the super-resolution network model, N represents the number of images used for training, and ‖·‖ 1 represents the mean absolute error, represents the i-th original high-resolution image, represents the i-th image after super-resolution reconstruction.