A lightweight image super-resolution method based on serial high-frequency attention
By designing a lightweight method based on serial high-frequency attention in the image super-resolution task, the existing attention mechanism has large memory usage, slow inference speed and no obvious guidance of high-frequency features, and the effect of reducing memory usage, improving inference speed and improving reconstruction quality is achieved.
Patent Information
- Application Number
- CN202210466344.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-04-29
AI Technical Summary
The existing attention mechanisms take up too much video memory in image super-resolution tasks, the inference speed is too slow, and the guidance of high-frequency features is not obvious.
A lightweight image super-resolution method based on serial high-frequency attention is designed, and a serial ERB+HFAB structure is adopted. The high-frequency attention module is composed of dimension reduction convolution, edge detection convolution, dimension up convolution, batch normalization layer and Sigmoid layer. The recovery of image high-frequency edge information is enhanced by learning a weight of 0 to 1 for each pixel.
Video memory usage was reduced by 72%, inference speed increased by 38%, and significantly improved reconstruction quality, especially in the recovery of image edge information.
Smart Images

Figure CN114897690B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and relates to deep learning, computer vision, and computer image understanding, especially image super-resolution technology, which is used to guide deep learning neural networks to focus on high-frequency features so as to enhance the reconstruction quality, and is a lightweight image super-resolution method based on serial high-frequency attention. Background Art
[0002] Lightweight image super-resolution refers to the technology of using a model with fast running speed and less video memory occupancy to restore a low-resolution image into a clear high-resolution image. This technology can not only be directly used for enhancing the quality of real-life images, but also provide an effective preprocessing means for downstream tasks such as small object target detection, segmentation, and human key point detection. In actual application scenarios, many times we hope that the model can reconstruct pictures faster, such as super-resolution of magnetic resonance images for diagnosing patients' conditions, super-resolution of Design Ideas in Microsoft 365, super-resolution of underwater cruise environment images, and preprocessing for target detection. These actual needs have made lightweight image super-resolution a hot research topic.
[0003] Since the birth of deep learning, methods based on convolutional neural networks (CNNs) have made great progress in the field of image super-resolution. SRCNN creatively designed a three-layer CNN to learn the mapping from low resolution to high resolution, achieving a significant improvement compared with traditional methods. After that, more and more inspiring ideas have been introduced, such as residual learning, feature fusion, cleverly designed loss functions, and the recently popular attention mechanism, which have promoted the development of the field of image super-resolution.
[0004] The attention mechanism has been proven effective in various computer vision tasks. Its goal is to guide the network to focus on important signals while reducing the attention to unimportant signals. Since SENet achieved great success in image classification tasks, researchers in image super-resolution have proposed various variants of the attention mechanism. The Residual Channel Attention Networks (RCAN) first incorporated channel attention into residual blocks. The Residual Non-Local Attention Networks (RNAN) introduced local and global attention mechanisms to scale intermediate features. The Second-order Attention Network (SAN) designed a channel attention mechanism using second-order statistical information of features, achieving better results than the first-order channel attention mechanism. The Residual Feature Aggregation Network (RFA) proposed to enhance the spatial attention mechanism to obtain feature maps with a larger receptive field. The Holistic Attention Network (HAN) proposed layer attention mechanism and channel-spatial attention mechanism to model elements in different convolutional layers, different channels, and different spaces. Although these attention mechanism methods have made great progress, multi-branch structures and inefficient operators, such as 7x7 convolutions, are suboptimal in lightweight image super-resolution tasks. More powerful and efficient attention mechanism modules need to be further studied in depth. Summary of the Invention
[0005] The problem to be solved by the present invention is that the currently commonly used attention mechanisms have excessive video memory occupancy, too slow inference speed, and insignificant guidance for high-frequency features.
[0006] The technical solution of the present invention is as follows: A lightweight image super-resolution method based on serial high-frequency attention. An image super-resolution model based on high-frequency attention is built. After the low-resolution image is subjected to feature extraction, high-frequency learning is performed, and then it is reconstructed into a high-resolution image. The high-frequency learning module includes a serial ERB+HFAB structure. The ERB+HFAB structure is connected with a high-frequency attention module HFAB after each enhanced residual block ERB. The high-frequency attention module HFAB is composed of a dimensionality reduction convolution, an edge detection convolution, a dimensionality increase convolution, a batch normalization layer, and a Sigmoid layer. By learning a weight from 0 to 1 for each pixel, the recovery of the high-frequency edge information of the image by the convolutional neural network is strengthened. First, the input feature map is convolved for dimensionality reduction, then a rough edge map is obtained through the edge detection convolution, then the edge map is refined through the enhanced residual block, the dimension is transformed back to the input space through the dimensionality increase convolution, and finally, after passing through the batch normalization layer BN to reach the non-saturation point of the Sigmoid function, it is sent to the Sigmoid function to learn a weight from 0 to 1 for each pixel to obtain an attention map, and the attention map and the input feature map are multiplied pixel by pixel to achieve feature correction; when training the image super-resolution model, the L1 loss function is used to calculate the distance between the reconstructed high-resolution image and the high-definition image positive sample, and the gradient of the parameters of each layer of the network is deduced therefrom, and the Adam optimizer is used for supervised training.
[0007] When designing the high-frequency attention module of the present invention, first, the inference time of the meta-operator is tested, and the operator with the highest efficiency is selected to build the attention module; secondly, through the analysis of the video memory, a serial structure is adopted to reduce the video memory occupancy instead of a parallel structure; then, by explicitly introducing a learnable Laplace edge detection operator, the learning of high-frequency features is enhanced; finally, a batch normalization layer is introduced to accelerate the convergence of the module.
[0008] The present invention has the following outstanding innovation points: (1) The present invention discovers that the 3x3 convolution is more efficient and can bring a larger receptive field, and uses the 3x3 convolution for dimensionality increase and reduction instead of the 1x1 convolution used in the existing method; (2) The method of the present invention adopts a completely serial structure to reduce the video memory occupancy and improve the inference speed, instead of the parallel structure adopted by the existing attention module; (3) The method of the present invention applies a learnable Laplace edge detection operator in the attention module to enhance the features in the high-frequency region; (4) The method of the present invention discovers that the batch normalization layer helps the network to converge in the attention module, while the attention modules of the existing attention mechanisms do not adopt the batch normalization layer.
[0009] The beneficial effects of the present invention are as follows:
[0010] 1. The attention module of the present invention is more efficient. The maximum video memory occupancy is reduced by 72% compared with the existing method, and the inference speed is increased by 38%, which can well meet the lightweight image super-resolution task.
[0011] 2. The reconstruction quality of the present invention is higher, which can improve the problem of blurred edge information of the reconstructed image and achieve better reconstruction quality. Judged by the peak signal-to-noise ratio (PSNR), the method of the present invention has achieved the highest reconstruction quality on five datasets, namely Set5, Set14, B100, Urban100, and Manga109. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a flowchart of the present invention.
[0013] Figure 2 It is a test of the inference time of the meta-operator of the present invention on a GTX 1080Ti.
[0014] Figure 3 It is a video memory analysis of the serial and parallel modules of the present invention.
[0015] Figure 4 It is a network structure based on high-frequency attention proposed by the present invention.
[0016] Figure 5 It is a comparison diagram of the enhanced residual block (ERB) and the ordinary residual block (RB) adopted by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The present invention proposes a lightweight image super-resolution method based on serial high-frequency attention, which can greatly improve the reconstruction quality while ensuring the inference efficiency of the model. The method of the present invention constructs an image super-resolution model based on high-frequency attention. After the low-resolution image is subjected to feature extraction, high-frequency learning is carried out, and then it is reconstructed into a high-resolution image. The core of the present invention lies in constructing a serial high-frequency attention module, which is composed of a dimensionality reduction convolution, an edge detection convolution, a dimensionality increase convolution, a batch normalization layer, and a Sigmoid layer, and strengthens the recovery of the high-frequency edge information of the image by the convolutional neural network by learning a weight from 0 to 1 for each pixel.
[0018] Figure 1 It shows the main process of constructing an image super-resolution model based on high-frequency attention, and further describes the implementation of the present invention in combination with the drawings and specific embodiments.
[0019] Step 1: Meta-operator time test. A convolutional neural network model consists of basic components such as convolutional layers, activation layers, up / downsampling layers, normalization layers, etc. and four arithmetic operations. The operation time of meta-operators is an important part of the model inference time. Mainstream deep learning frameworks have different optimization efforts for different operators. For example, due to the generality of 3x3 convolutions, they have been highly optimized for parallelism on hardware devices, while the computational density of larger convolutional kernels such as 5x5 is much lower than that of 3x3 convolutions. Therefore, it is very necessary to select efficient operators for building a lightweight image super-resolution model, which can achieve the lowest inference time overhead under the same model capacity. Based on the super-resolution model EDSR-baseline, this invention adds multiple repeated operators to obtain a new model, and then obtains the approximate inference time of each operator by taking the difference. Set the input image size to 3x480x480, run 5 inferences on a GTX 1080Ti and take the average time to reduce errors, and get Figure 2 . The following conclusions can be drawn from the figure:
[0020] (1) It is best to use convolutional layers with a 3x3 convolutional kernel size. The number of parameters of a lightweight image super-resolution model is generally only a few hundred K. 3x3 convolutions can bring a larger receptive field compared to 1x1 convolutions. More importantly, the computational amount of 3x3 convolutions is 9 times that of 1x1 convolutions, yet the time overhead is only 2 times, which is more efficient;
[0021] (2) Select LeakyReLU as the activation function. LeakyReLU has a larger non-linear mapping space compared to ReLU, and the inference time is significantly better than that of PReLU;
[0022] (3) Avoid using bound operators. Although there are many methods suggesting to downsample first and then upsample to obtain a larger receptive field, the combination of a single downsampling and a single upsampling is equivalent to half of the inference time of a 3x3 convolution. Image super-resolution models all repeat multiple sub-structures, and combining upsampling and downsampling will cause a significant additional overhead in the final model. In addition, channel separation and channel concatenation are also bound operators, which are common in the current best models IMDN and RFDN, but this invention finds that they are the main factors restricting performance;
[0023] (4) Use sub-pixel convolution for the upsampling module. Sub-pixel convolution is significantly better than other upsampling methods.
[0024] Finally, this invention selects a total of 6 most efficient operators, namely 3x3 convolution, LeakyReLU, element-wise addition, multiplication, Sigmoid without dependence, and sub-pixel convolution, to build the module.
[0025] Step 2: Maximum VRAM Occupancy Analysis. In addition to the commonly used metrics of parameter quantity, computational complexity, and inference time, the AIM, an international lightweight image super-resolution competition, also includes the maximum VRAM as a performance metric. The proposal of the maximum VRAM metric is reasonable: the VRAM of a graphics card is limited. In actual applications, multiple models often collaborate to complete tasks. It is desired that each model can occupy as little VRAM as possible. Otherwise, the program may abort abnormally due to a model occupying too much VRAM at a certain moment. Additionally, when deployed to mobile devices, excessive VRAM occupancy can also cause problems such as overheating. Let M i represent the VRAM occupied when passing through the i-th node during inference, n represent the total number of nodes in the model, and M represent the maximum VRAM occupied during inference, which is defined as:
[0026] M = max(M1, M2, …, M n )
[0027] M i consists of four parts: the VRAM occupied by the input features the VRAM occupied by the output features the VRAM occupied by the temporarily stored features that will be accessed in the future and the VRAM occupied by the model parameters which is defined as:
[0028]
[0029] In the lightweight image super-resolution task, the VRAM occupied by the model can be ignored compared to the VRAM occupied by the features. With the same input resolution and parameters, the VRAM occupancy of this layer mainly depends on the size of the temporarily stored feature maps. To analyze the reason behind the doubling of VRAM caused by the feature fusion layer in the current best model, consider the simple serial structure and the fusion parallel topology structure adopted by RFDN, as Figure 3 shown. For ease of description, assume that the size of the feature map remains unchanged in each dimension after passing through a 3x3 convolution. Multiple intermediate features are concatenated along the channel dimension, and a 1x1 convolution is used to reduce the number of feature maps to the same as the number of channels of the input image after concatenation. Since the ReLU function can be directly performed on the original feature map, which means that the input and output of the ReLU layer share VRAM, the activation layer is omitted in the figure. For a convolution with a kernel size of C in ×C out ×K×K, the input and output features cannot share VRAM because each pixel point needs to be accessed C out ×K×K times and the Winograd algorithm is used.
[0030] First, consider Figure 3The serial structure in (a). In the serial structure, each node is only related to the current input and output feature maps. The feature maps of the previous layers other than this can be not saved after forward propagation. Therefore, the video memory occupancy of each convolutional node is approximately M input +M output , and stacking the same serial structures only increases the number of network parameters, and the resulting video memory can be ignored. When the size of the feature map remains unchanged, the maximum video memory M serial is 2×M input .
[0031] Then consider Figure 3 the parallel structure based on feature fusion in (b). The feature maps related to the 1x1 convolutional fusion layer will be saved in the video memory after the initial calculation, so it will cause M kept to increase significantly. Taking the second 3x3 convolutional layer as an example, there will be three features occupying the video memory: the input of the convolutional layer, the output of the convolutional layer, and the input of the first convolutional layer that will be used for fusion. Therefore, the video memory occupancy of this node will be 3 times that of the input feature map. Similarly, in the feature splicing layer, the video memory occupied by splicing 3 features along the channel dimension will be 6×M input . Suppose there are N features of the same size participating in the fusion. Then the video memory occupied by the splicing node, with index i, is 2×N×M input . Thus, the ratio of the video memory M parallel occupied by the fusion parallel structure to the video memory M serial occupied by the simple serial structure is as follows:
[0032]
[0033] That is, at least N times the relationship. The above analysis has been well verified on the RFDN model: the 400K serial structure occupies about 30M of video memory, and the RFDN parallel structure with global fusion occupies about 200M of video memory, about 7 times the maximum video memory occupancy. In order to minimize the maximum video memory occupancy as much as possible, the present invention should first consider the serial structure when designing the network block and avoid using multiple parallel connections at a certain node.
[0034] Step 3: Model design based on the serial high-frequency attention module. After the analysis of the first 2 steps, according to the determined operators and feature fusion structures, build an image super-resolution model based on high-frequency attention, especially the high-frequency attention module HFAB among them.
[0035] The present invention constructs a network based on the enhanced residual block ERB (enhanced residual block), as Figure 5 shown, Figure 5 (a) is the residual block RB,Figure 5 (b) shows the enhanced residual block ERB. The reconstruction accuracy of ERB and RB is comparable, but the two skip connections in ERB can be merged with the parallel convolutions during the inference stage, reducing the memory access overhead. In RB, however, there is a non-linear operation in the middle of the skip connection and they cannot be merged. ERB can improve the inference speed by 10% compared to RB. The overall structure of the image super-resolution model of the present invention is basically the same as that of the existing methods. The difference lies in the high-frequency learning part. The present invention designs a serial ERB+HFAB structure, and a high-frequency attention module HFAB is applied after each ERB to enhance high-frequency information.
[0036] The task of the high-frequency attention module HFAB is to assign a weight between 0 and 1 to each pixel point on the feature map to represent their importance during the model learning process. The goal that HFAB hopes to achieve is that the edge detail pixels can be more finely restored during the recovery process. To achieve this goal, the Laplacian operator is introduced in the HFAB module to guide the branch to pay attention to the edge details. The convolution template of the Laplacian operator is defined as:
[0037]
[0038] Initialize the 3x3 convolution with this weight, and it can be continuously updated during the subsequent learning process, so that the template has a stronger correction ability.
[0039] The overall structure of the image super-resolution model of the present invention and the structure of the high-frequency attention module are as Figure 4 shown. Denote the input of the k-th HFAB as F k-1 , and the output as F k . The learning process of HFAB can be formally described. To reduce the parameter overhead brought by the attention branch, first reduce the dimension of the feature map with a 3x3 convolution. The feature after dimensionality reduction by Conv squeeze (·) is:
[0040]
[0041] where LReLU(·) is the LeakyReLU non-linear activation function. Then obtain a rough edge map through the edge detection layer
[0042]
[0043] Then refine the edge map through the enhanced residual block E(·) to obtain
[0044]
[0045] Finally, transform it to the input space through a 3x3 convolution to obtain
[0046]
[0047] Next, it passes through the batch normalization layer to reach the non-saturation point of the Sigmoid function, and is fed into the Sigmoid function to learn a weight between 0 and 1 for each pixel, obtaining the attention map Attention k :
[0048]
[0049] Finally, the attention map and the input features are multiplied pixel by pixel to achieve feature correction:
[0050] F k = Attention k × F k-1
[0051] Batch normalization is a linear operation. In the HFAB designed in the present invention, the batch normalization layer BN can be fused with the previous 3x3 convolution during inference, thereby accelerating inference. Let the mean, variance, and numerical stability parameter of BN be μ, σ, and ∈, the learned scale factor and offset be γ and β, the weights and offsets of the 3x3 convolution be W3 and b3, and the input be X. The fusion process of the BN layer and the previous convolutional layer is described as:
[0052]
[0053] After the reparameterization of the combined linear operation, the entire network structure only contains the following 6 efficient operators: 3x3 convolution, ReLU activation function, element-wise addition, element-wise multiplication, Sigmoid, and sub-pixel convolution. The efficient operators and serial modules ensure the efficiency of the model operation, and the high-frequency attention module for adaptively scaling features ensures the reconstruction performance of the model.
[0054] Step 4: Training of the model. The batch size for each iteration is set to 16, and the learning rate is initially set to 1×10 -5 and becomes half every 200,000 iterations, and a total of 1,000,000 iterations are trained. The L1 loss function is used to calculate the distance between the reconstructed high-resolution image and the high-definition image positive sample, for regularization, and the gradients of the parameters of each layer of the network are derived therefrom. The Adam optimizer is used to optimize the model.
[0055] Step 5: Testing of the model. After training, save the network weights and reconstruct the images in the test set. Through testing and verification, compared with the RFDN (Residual Feature Distillation Network), the champion solution of the lightweight image super-resolution model in the AIM2020-ESR competition in the inference stage, the present invention reduces the video memory occupancy by 72% and improves the inference speed by 38%.
Claims
1. A lightweight image super-resolution method based on serial high-frequency attention, characterized in that Build an image super-resolution model based on high-frequency attention. After feature extraction, the low-resolution image undergoes high-frequency learning and then is reconstructed into a high-resolution image. The high-frequency learning module includes a serial ERB+HFAB structure. The ERB+HFAB structure is connected to a high-frequency attention module HFAB after each enhanced residual block ERB. The high-frequency attention module HFAB consists of dimensionality reduction convolution, edge detection convolution, dimensionality increase convolution, batch normalization layer and Sigmoid layer. By learning a 0 to 1 weight for each pixel, the convolutional neural network is strengthened to restore the high-frequency edge information of the image. First, the input feature map is convolved to reduce the dimensionality, and then the edge detection is performed. The rough edge map is obtained by convolution, and then the edge map is refined by the enhanced residual block. The dimension is transformed back to the input space by the dimensionality-raising convolution. Finally, the unsaturated point of the Sigmoid function is reached by the batch normalization layer BN, and the global information of the image is introduced. The Sigmoid function is sent to learn a weight from 0 to 1 for each pixel to obtain the attention map. The attention map and the input feature map are multiplied pixel by pixel to achieve feature correction. When training the image super-resolution model, the L1 loss function is used to calculate the distance between the reconstructed high-resolution image and the high-definition image positive sample, and the gradient of the parameters of each layer of the network is derived from this, and the Adam optimizer is used for supervised training. The high-frequency attention module HFAB is specifically: Denote the input of the k-th HFAB as F k-1 , and the output as F k . First, reduce the dimension of the feature map using a 3x3 convolution. After passing through Conv squeeze (·), the reduced feature is: where LReLU(·) is a non-linear activation function, and then a rough edge map is obtained through the edge detection layer W Laplacian is the Laplacian operator, which is used to guide the branch to focus on edge details; Then, it passes through the enhanced residual block E k (·) to refine the edge map to obtain Then, it is upsampled to the input space through a 3x3 convolution to obtain Next, it passes through the batch normalization layer to reach the non-saturation point of the Sigmoid function, and then is fed into the Sigmoid function to learn a weight between 0 and 1 for each pixel, obtaining the attention map Attention k : Finally, the attention map and the input feature map are multiplied pixel by pixel to achieve feature correction: F k = Attention k × F k-1 The batch normalization layer BN in HFAB is fused with the previous 3x3 convolution in the inference stage. Let the mean, variance and numerical stability parameters of BN be μ, σ and ∈ respectively, the learned scale factor and offset be γ and β, the weight and offset of 3x3 convolution be W3 and b3, let the input be X, and the convolution layer and BN layer parameter fusion process is:
2. The lightweight image super-resolution method based on serial high-frequency attention according to claim 1, characterized in that Analyze the meta-operators, test the inference time of the meta-operators, and select the most efficient operators to build the attention module, including: 1) Use a convolutional layer with a kernel size of 3x3; 2) Select LeakyReLU as the activation function; 3) Avoid using bound operators; 4) The upsampling module uses sub-pixel convolution.
3. The lightweight image super-resolution method based on serial high-frequency attention according to claim 2, characterized in that The high-frequency attention module HFAB is built using six operators: 3x3 convolution, LeakyReLU, independent element-by-element addition, element-by-element multiplication, Sigmoid, and sub-pixel convolution.
Citation Information
Patent Citations
Remote sensing image super-resolution reconstruction method based on lightweight generative model
CN113538234A
Lightweight super-resolution reconstruction method based on adaptive weight learning
CN113538244A