A Single-Image Super-Resolution Algorithm Based on a Feedback Mechanism
By constructing feedback subnetwork and optimization algorithms, the problem of single image super-resolution algorithm in deep-level feature extraction and receptive field fixation is solved, and efficient image super-resolution generation is achieved, improving image quality and model performance.
Patent Information
- Application Number
- CN202210678755.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-15
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-06-15
AI Technical Summary
The existing single-image super-resolution algorithm based on convolutional neural networks has limitations in the problems of deep feature extraction and receptive field fixation, resulting in partial loss of information flow and insufficient ability to acquire high-dimensional feature.
A single image super-resolution algorithm based on feedback mechanism is adopted to construct a feedback subnet containing coordinate attention module and densely connected receptive field module, and optimize model parameters with Charbonnier loss function and Adam backpropagation algorithm to achieve efficient fusion and extraction of features.
It improves the peak signal-to-noise ratio of the image, reduces the complexity of model training, and generates high-quality high-resolution images, which have great performance advantages.
Smart Images

Figure CN114897704B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a single-image super-resolution technology and belongs to the technical field of digital image processing. Background Art
[0002] When comparing a high-resolution image (HR) with a low-resolution image (LR), the former has clearer textures. In most computer vision tasks, the size of the image resolution often has a very important impact on the final result of this task. Various computer vision tasks have very high requirements for the image resolution, and the size of the image resolution often determines whether the task can be successful.
[0003] The super-resolution reconstruction technology, that is, through an algorithm, the low-resolution image is enlarged to a high-resolution image, and then the details of the image are enhanced and optimized, which can effectively improve the quality of the picture, achieve efficient data transmission, and reduce costs. Reconstructing a corresponding clear high-resolution image from a degraded low-resolution image is a typical low-level computer vision task and also a research hotspot in the current field of image processing. With the wide popularity of taking pictures, the super-resolution technology has once again attracted extensive attention.
[0004] Image super-resolution can be classified into single-frame image super-resolution (SISR), multi-frame image super-resolution (MISR), and video super-resolution (VSR) according to the type of input. Compared with MISR and VSR, SISR can obtain less information and has a higher processing complexity. After simple processing of the results obtained by the SISR method, it can be applied to MISR and VSR.
[0005] In recent years, scholars have proposed a variety of single-image super-resolution methods based on convolutional neural networks and have exceeded the performance of traditional algorithms. These methods are all composed of a series of convolutional layers with non-linear activation functions. Their superior performance has prompted them to be gradually applied to more and more industrial productions. These algorithms based on convolutional neural networks have achieved great success compared with classical algorithms, but these algorithms also have some limitations. As the number of network layers deepens, the forward form of the structure will cause partial loss of information flow, and useful deep-level features cannot be obtained. Moreover, using a convolutional kernel with a fixed size will result in a fixed receptive field, unable to sense changes in a larger range and unable to improve the ability to obtain high-dimensional features. Summary of the Invention
[0006] Objective of the Invention: To solve the above problems, the present invention provides a single-image super-resolution algorithm based on a feedback mechanism, which can improve the peak signal-to-noise ratio of an image and reduce the training complexity of a model.
[0007] The technical solution adopted by the present invention is as follows: A single-image super-resolution algorithm based on a feedback mechanism, comprising the following steps:
[0008] (1) Obtain the DIV2K dataset as the training set for training a neural network, and use bicubic downsampling to obtain image pairs;
[0009] (2) Utilize the feedback mechanism to construct a single-image super-resolution generation network model;
[0010] (3) Input the image pairs into the single-image super-resolution generation network model, use the Charbonnier loss as the loss function, and optimize the model parameters through the Adam backpropagation algorithm to obtain a trained super-resolution network model.
[0011] Specifically, the step (2) includes:
[0012] 1) Extract the initial feature f0 from the given input through two convolutional layers;
[0013] 2) Input the initial feature f0 into the feedback sub-network, and represent the feedback sub-network as FB. Then, the output feature of the i-th feedback network is
[0014]
[0015] where, FB represents the operation process of the feedback sub-network, and the output of the i-th feedback sub-network is F i ;
[0016] 3) Concatenate f0 and the outputs of the N-th feedback sub-network, and add them to f0 to obtain the high-dimensional feature F fuse ;
[0017] 4) Obtain the high-resolution feature by passing F fuse through a convolutional layer and a sub-pixel convolutional layer;
[0018] 5) Obtain the high-resolution output by passing the high-resolution feature through a convolutional layer.
[0019] Specifically, the feedback sub-network takes the output f0 of the shallow feature extraction network as the feature input. The shallow feature input will participate in the recursive process of the feedback sub-network as the input each time. The input of the feedback sub-network consists of the output of the previous feedback sub-network and the output of the shallow feature extraction.
[0020] Specifically, the feedback sub-network includes a coordinate attention module and a dense connection receptive field module. The two outputs of the feedback sub-network are subjected to attention allocation through the coordinate attention module, and then features are extracted through the dense connection receptive field module.
[0021] Specifically, the coordinate attention module consists of a pooling layer, a convolutional layer, and an activation function layer. First, pooling operations are performed on the input feature map along the X-axis and Y-axis to obtain two features, which are concatenated along the channel dimension; then, channel compression is performed through the convolutional layer, and then split and restored to the original number of channels through the convolutional layer. Finally, the output range is compressed to [0,1] through the activation function to obtain the features on the X-axis and Y-axis. Multiplying the input by these two attentions can obtain the final output.
[0022] Specifically, the dense connection receptive field module consists of a receptive field module and a convolutional layer. The output of each layer serves as the input of all subsequent layers.
[0023] Specifically, the receptive field module has a multi-branch convolutional layer and a dilated convolutional layer to obtain receptive fields of different scales, reduce the number of network parameters, and obtain richer features.
[0024] Specifically, step (3) includes the following steps:
[0025] 1) Use the Charbonnier loss as the loss function and input f0 into the feedback sub-network
[0026]
[0027] where C is the number of channels of the image; H is the height of the image; W is the width of the image; is each pixel point of the generated high-resolution image; y i,j,k is each pixel point of the real image;
[0028] 2) Use the Adam backpropagation algorithm to optimize the model parameters.
[0029] Beneficial effects:
[0030] The present invention proposes a new recursive sub-network, which makes full use of the feature information extracted at different stages through the attention mechanism and the receptive field module. Coordinate attention allocates the outputs at different stages and the initial input. The receptive field module fully increases the receptive field size of the module, and can extract more effective features. The present invention can generate high-resolution pictures and has great performance advantages. Description of the Drawings
[0031] Figure 1 It is the framework diagram of this method.
[0032] Figure 2 Schematic diagram of the expansion of the feedback sub-network in this method.
[0033] Figure 3 Schematic diagram of the feedback block in this method.
[0034] Figure 4 Schematic diagram of the coordinate attention module in this method.
[0035] Figure 5 Schematic diagram of the dense connection receptive field module in this method.
[0036] Figure 6 Schematic diagram of the receptive field module in this method. Detailed implementation method
[0037] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation methods.
[0038] A single-image super-resolution algorithm based on a feedback mechanism of the present invention Figure 1 The specific network of the present invention includes the following steps:
[0039] (1) Obtain the DIV2K dataset and use bicubic downsampling to obtain the training dataset pair.
[0040] (2) The low-resolution input extracts the shallow feature f0 through two convolutional layers.
[0041] (3) As Figure 2 shown, the shallow feature f0 and the output of the previous feedback sub-network are continuously iteratively output through the feedback sub-network, and the two outputs of the first time are both f0.
[0042] (4) As Figure 3 shown, the feedback sub-network includes two parts: a coordinate attention and a dense connection receptive field module. The output of the feedback sub-network first performs attention distribution through coordinate attention, and the coordinate attention module is as Figure 4 shown. The input feature first performs pooling operations on the X-axis and Y-axis
[0043]
[0044]
[0045] Next, the pooling output of the X-axis and the pooling output of the Y-axis are concatenated to obtain coordinate information, and the channels are compressed through a convolutional layer and mapped through a non-linear activation function.
[0046] f = δ(F1[z h ,z w )
[0047] Among them, [·] represents the cascading operation, δ represents the non-linear activation mapping, and F1 represents the convolution layer operation with a convolution kernel size of 1×1, and the channel compression rate of the input is r.
[0048] Next, the feature f is split, and the split features pass through a convolution layer and a Sigmoid activation function to obtain an output with the same number of channels as the input feature. Figure 1 Same number of channels as the input feature.
[0049] h = σ(F h (f h ))
[0050] g w = σ(F w (f w ))
[0051] Among them, F(·) represents a convolution layer with a convolution kernel size of 1×1, the channel magnification rate is r, and the channel size of the output feature map is C. σ represents the Sigmoid activation function.
[0052] Multiplying the input by the two attentions can obtain the final output:
[0053]
[0054] (5) The second part of the feedback sub-network is the densely connected receptive field module, as Figure 5 shown, using a densely connected structure. Figure 5 The receptive field module in Figure 6 is shown as i , using a multi-branch convolution layer and dilated convolution. The input passes through a convolution layer to obtain the initial feature M0, and the output of the i-th receptive field module is denoted as M MRFB .
[0055]
[0056] Finally, the initial feature M0 and the outputs of all receptive field modules are fused through a cascading operation, and the channel dimension is compressed through a convolution layer to obtain the final output.
[0057] F = Conv 1×1 ([M0,..., M N )
[0058] Among them, [·] represents the cascading operation, Conv 1×1 represents the convolution operation with a convolution kernel size of 1×1, and F is the output feature.
[0059] (6) The output F of the feedback sub-network is obtained through the densely connected receptive field module. i, and is input into the feedback sub-network together with the initial feature f0 for iteration.
[0060]
[0061] (7) The outputs of the feedback blocks for N iterations are subjected to feature fusion in a cascaded manner. The cascaded features are compressed in channels through a 1×1 convolutional layer, and the compressed features are added to the initial feature f0.
[0062] F fuse =Conv 1×1 ([f0,F1,...,F N )+f0
[0063] (8) F fuse continuously extracts features through a 1×1 convolutional layer, then performs upsampling through sub-pixel convolution to map the low-resolution feature map to a high-resolution feature map, and finally reconstructs through a 3×3 convolutional layer to output the final super-resolution image.
[0064] (9) The Charbonnier loss function is used, and the Adam backpropagation algorithm is used for model optimization training. β1 and β2 are set to 0.9 and 0.99 respectively, and the initial learning rate is set to 1e-4. Finally, the trained super-resolution network model is obtained.
[0065] (10) When the magnification factor is 2, the test PSNR results based on the Set5, Set14, BSD100, and Urban100 datasets are shown in Table 1, with the unit of dB.
[0066] Table 1. Comparison of PSNR Results
[0067] Method Set5 Set14 BSD100 Urban100 Bicubic 33.66 30.24 29.56 26.88 SRCNN 36.66 32.45 31.36 29.50 DRCN 37.63 33.04 31.85 30.75 CARN 37.76 33.52 32.09 31.92 Our Model 37.94 33.69 32.37 31.94
Claims
1. A single-image super-resolution algorithm based on a feedback mechanism, characterized in that: It includes the following steps: (1) Obtain the DIV2K dataset as the training set for training the neural network, and use bicubic downsampling to obtain image pairs; (2) Utilize a feedback mechanism to construct a single-image super-resolution generation network model; (3) Input the image pairs into the single-image super-resolution generation network model, use the Charbonnier loss as the loss function, and optimize the model parameters through the Adam backpropagation algorithm to obtain a trained super-resolution network model; The step (2) includes: 1) Extract the initial feature f0 from the given input through two convolutional layers; 2) Input the initial feature f0 into the feedback sub-network. Denote the feedback sub-network as FB. Then the output feature of the i-th feedback network is Among them, F FB represents the operation process of the feedback sub-network, and the output of the feedback sub-network at the i-th time is F i ; 3) Concatenate the output of f0 and the N - time feedback sub - network and add it to f0 to obtain the high - dimensional feature F fuse ; 4) Obtain high-resolution features of F fuse through a convolutional layer and a sub-pixel convolutional layer; 5) Obtain the high-resolution output through a convolutional layer for the high-resolution feature; The step (3) includes the following steps: 1) Use the Charbonnier loss as the loss function and input f0 into the feedback sub-network Among them, C is the number of channels of the image; H is the height of the image; W is the width of the image; is each pixel of the generated high-resolution image; y i,j,k is each pixel of the real image; 2) Optimize the model parameters through the Adam backpropagation algorithm.
2. The single-image super-resolution algorithm based on a feedback mechanism according to claim 1, wherein: The feedback sub-network takes the output f0 of the shallow feature extraction network as the feature input. The shallow feature input will participate in the recursive process of the feedback sub-network as the input each time. The input of the feedback sub-network consists of the output of the previous feedback sub-network and the output of the shallow feature extraction.
3. The single-image super-resolution algorithm based on a feedback mechanism according to claim 2, characterized in that: The feedback sub-network includes a coordinate attention module and a densely connected receptive field module. The two-way output of the feedback sub-network performs attention allocation through the coordinate attention module, and then extracts features through the densely connected receptive field module.
4. A single-image super-resolution algorithm based on a feedback mechanism according to claim 3, characterized in that: The coordinate attention module consists of a pooling layer, a convolutional layer, and an activation function layer. First, perform pooling operations on the input feature map along the X-axis and Y-axis to obtain two features and concatenate them along the channel dimension; then perform channel compression through a convolutional layer, then split and restore the original number of channels through a convolutional layer, and finally compress the output range to [0,1] through the activation function to obtain the features on the X-axis and Y-axis. Multiply the input by these two attentions to obtain the final output.
5. A single-image super-resolution algorithm based on a feedback mechanism according to claim 4, characterized in that: The densely connected receptive field module consists of a receptive field module and a convolutional layer. The output of each layer serves as the input for all subsequent layers.
6. The single-image super-resolution algorithm based on a feedback mechanism according to claim 5, wherein: The receptive field module has a multi-branch convolutional layer and a dilated convolutional layer to obtain receptive fields of different scales, reduce the number of network parameters, and obtain richer features.
Citation Information
Patent Citations
Lightweight image super-division method and system based on attention feedback mechanism
CN113409191A
Super-resolution reconstruction method based on attention mechanism
CN113706386A