A lightweight image super-resolution method based on full point-by-point convolution
By introducing spatial shift operations and full point-by-point convolution into the image super-resolution network, a lightweight shift convolution network SCNet is built, which solves the problems of resource limitation and insufficient feature aggregation capabilities in the prior art, and achieves efficient image super-resolution.
Patent Information
- Application Number
- CN202310000876.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2043-01-03
AI Technical Summary
The existing lightweight image super-resolution network based on 3×3 convolution is difficult to deploy on resource-constrained platforms, and it is difficult to build an effective image super-resolution network based on point-by-point convolution.
A lightweight image super-resolution method based on all point-by-point convolution is proposed. By introducing spatial shift operations, features are split along the channel direction and moved along different spatial directions, local aggregation of features is realized and shifted convolution is formed. This method replaces the 3×3 convolution in the standard residual structure and stacks the shift residual units to construct a lightweight image super-resolution network SCNet.
The possibility of deployment on resource-constrained platforms is realized, the amount of parameters and calculations is reduced, and the ability to aggregate local features is improved, achieving better results than the existing lightweight super-resolution model.
Smart Images

Figure CN116051375B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a lightweight image super-resolution method based on full point-by-point convolution. Background Art
[0002] Image super-resolution (SISR) aims to reconstruct a high-resolution (HR) image from its corresponding degraded low-resolution (LR) image. With the rapid development of deep learning, it has received much attention in the research community and has made breakthrough progress in image super-resolution tasks. Researchers have proposed many methods based on convolutional neural networks (CNNs), which continuously improve the upper limit of image super-resolution tasks by carefully designing network modules and introducing more reasonable structural priors. However, these methods usually introduce very complex model architectures. Although they improve the performance of image super-resolution, the corresponding number of parameters and computational complexity make these methods difficult to deploy on resource-constrained platforms. Accordingly, structurally efficient and lightweight super-resolution methods are crucial for real application scenarios. Many researchers have explored many lightweight image super-resolution models from the perspective of reducing the amount of computation and parameters. Among them, 3×3 convolution is widely used, and larger convolution kernels can further improve model performance, but the number of parameters and computational cost increase rapidly. The small kernel of point-by-point convolution (1×1 convolution) can reduce the number of parameters, but due to the lack of fusion and integration with the local features of neighboring pixels, it is difficult to build an effective image super-resolution network using point-by-point convolution directly. This paper proposes a method that has the best of both worlds: a lightweight image super-resolution model is realized only through 1×1 point-by-point convolution. Summary of the invention
[0003] The purpose of the present invention is to solve the problem of performance degradation caused by further compressing the existing lightweight image super-resolution network based on 3×3 convolution, and proposes a lightweight image super-resolution method based on full point-by-point convolution.
[0004] The present invention is realized by the following technical scheme. The present invention proposes a lightweight image super-resolution method based on full point-by-point convolution. The method is specifically as follows: a spatial shift operation is introduced, and the features extracted by point-by-point convolution are split into different groups along the channel direction, and different groups of features are moved along different spatial directions, and the moved features are integrated with neighboring pixel features in the channel direction; point-by-point convolution and spatial shift operations are collectively referred to as shifted convolution, and shifted convolution has extremely low parameter amount and calculation amount in point-by-point convolution, and can also perform effective local feature aggregation; a residual structure unit is obtained based on the shifted convolution, and the residual structure unit is further stacked to obtain a lightweight image super-resolution network architecture: a shifted convolution network SCNet; the SCNet includes shallow feature extraction, deep feature extraction and high-resolution image reconstruction modules; given a low-resolution image I LR , first use the shallow feature extractor f s Map it to the specified hidden layer feature space to obtain the feature map F s =f s (I LR ), then, the shallow feature map is passed through the deep feature extractor f d , extract the deep feature map F d =f d (F s ), finally, the high-resolution image reconstruction module f u Upsample the deep features to obtain the final super-resolution result I SR =f u (F (d) ).
[0005] Furthermore, the spatial shift operation is specifically as follows: firstly, the input feature map is split into N groups evenly through channels, where N represents the number of shifted neighbor features; then, different groups are moved in different directions by specified step sizes, and after moving in different directions, the neighbor features at corresponding positions are aggregated.
[0006] Furthermore, based on the shifted convolution layer, all 3×3 convolutions in the standard residual structure are replaced by point-by-point convolutions, in which the spatial shift operation is embedded. The improved shifted residual structure unit includes a shifted convolution and a point-by-point convolution and an activation layer.
[0007] Beneficial effects of the present invention:
[0008] (i) The present invention designs a shift convolution, which combines the spatial shift operation with point-by-point convolution, and expands the local feature aggregation capability of point-by-point convolution.
[0009] (ii) The present invention proposes a shifted residual unit based on shifted convolution, and uses the shifted residual unit to implement a lightweight shifted super-resolution network SCNet.
[0010] (iii) This paper verifies and analyzes the effectiveness of the proposed SCNet, and shows that SCNet achieves better results than existing lightweight super-resolution models on public benchmark test sets. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 This is the overall structure diagram of SCNet;
[0012] Figure 2 Implement a flow chart for the spatial shift operation;
[0013] Figure 3 The conventional residual unit and the shifted residual unit proposed in the present invention, wherein (a) is the conventional residual unit, and (b) is the shifted residual unit proposed in the present invention;
[0014] Figure 4 Comparison chart of subjective results between SCNet and other methods;
[0015] Figure 5 It is the objective indicator result diagram;
[0016] Figure 6 Feature selection result graphs for various shift operations;
[0017] Figure 7 The calculation amount comparison chart of SCNet and several other cutting-edge methods;
[0018] Figure 8 This is the ablation analysis result diagram of the shift operation;
[0019] Fig. 9 The following is a graph of SCNet ablation analysis results for different model sizes. DETAILED DESCRIPTION
[0020] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] On various resource-limited devices, such as mobile devices, lightweight network architectures are critical for deploying image super-resolution deep models. In order to balance model capacity and computational complexity, 3×3 convolution is widely used in current super-resolution models based on convolutional neural networks. Compared with 3×3 convolution, point-by-point convolution (1×1 convolution) has a smaller computational complexity, but lacks the aggregation of spatial features, so it is difficult to obtain an effective super-resolution model using only point-by-point convolution. This invention rethinks how to use point-by-point convolution in a lightweight super-resolution network. Using the spatial shift operation, the neighbor feature information is manually aggregated along the channel direction, and a lightweight image super-resolution model implemented by full point-by-point convolution combined with spatial shift operation is proposed, named SCNet. A large number of experiments have shown that methods based on full point-by-point convolution can achieve and exceed existing lightweight super-resolution models.
[0022] In order to solve the problem that point-by-point convolution cannot achieve feature aggregation, a spatial shift operation (Spatial-Shift) is introduced. The features extracted by point-by-point convolution are split into different groups along the channel direction, and different groups of features are moved along different spatial directions. The moved features are integrated with neighboring pixel features in the channel direction. Through the spatial shift operation, the shortcomings of point-by-point convolution feature extraction are effectively filled. Here, point-by-point convolution and spatial shift operations are collectively referred to as shift convolution (Shift-Conv). Shift convolution has extremely low parameter amount and calculation amount in point-by-point convolution, and can also perform effective local feature aggregation. On this basis, the present invention proposes a residual structure unit based on shift convolution, and further stacks the residual structure unit to propose a lightweight image super-resolution network architecture: shift convolution network (SCNet). The shift network only contains one operator, point-by-point convolution, and the calculation process is extremely simple. Compared with the super-resolution network based on 3×3 convolution, the shift network designed with the same architecture can reduce the parameter amount and calculation amount to one-ninth of the original model. On the other hand, thanks to the lightweight features of the shifted convolution, the present invention further expands the depth and width (feature dimension) of the shifted network, and achieves further improvement on the basis of less computational complexity than the existing methods.
[0023] Image super-resolution aims to transform the LR image I LR Converted to the corresponding HR image I HR , thus generating the SR result I SR The present invention proposes a lightweight image super-resolution network SCNet implemented by full point-by-point convolution. Referring to the general architecture design based on existing convolutional neural networks, SCNet mainly consists of three parts: shallow feature extraction, deep feature extraction and HR image reconstruction module. The specific implementation is as follows Figure 1 shown.
[0024] Given a low-resolution image I LR , first use the shallow feature extractor f s Map it to the specified hidden layer feature space to obtain the feature map F s =f s (I LR ). Then, the shallow feature map is passed through the deep feature extractor f d , extract the deep feature map F d =f d (F s ). Finally, the high-resolution image reconstruction module f u Upsample the deep features to obtain the final super-resolution result I SR =f u (F (d) ).
[0025] The objective function of learning is to minimize the difference between the super-resolution result and the target high-resolution image:
[0026] L=|I SR -I HR |1.
[0027] The overall framework and training process of SCNet proposed in the present invention are as described above. The specific implementation details of the shifted convolution and shifted residual unit are introduced below.
[0028] Shifted convolution includes point-by-point convolution and spatial shift operations. The spatial shift operation is used to align neighboring features along the channel direction, thereby achieving local feature aggregation. The specific spatial shift operation is implemented as follows: Figure 2 As shown. The input feature map is first split into N groups, where N represents the number of shifted neighbor features. To keep consistent with the 3×3 convolution, the present invention uses eight groups by default. Then different groups are moved in different directions by specified step sizes. Figure 2 As shown in the figure, after moving in different directions, the aggregation of neighboring pixel features at corresponding positions is achieved. Here, in order to be consistent with the 3×3 convolution, the present invention adopts 8 directions and a step size of 1 as the default setting. It is worth noting that compared with the 3×3 convolution, the shifted convolution not only achieves local feature aggregation, but also can be further extended to long-range feature relationship extraction by controlling the selection of shifted feature points.
[0029] Based on the above shifted convolution layer, the present invention replaces all 3×3 convolutions in the standard residual structure with point-by-point convolutions, in which the spatial shift operation is embedded. The improved shifted residual unit includes a shifted convolution and a point-by-point convolution and an activation layer. The specific implementation details are shown in Figure 3Based on the shifted residual unit, the present invention realizes SCNet of different scales by stacking different shifted residual blocks.
[0030] In order to further verify the beneficial effects of the method of the present invention, the present invention is described in detail through the following embodiments.
[0031] The model described in the present invention is trained on DIV2K and Flickr2K, and the training data contains 3450 high-resolution images in total. The present invention verifies the 2x, 3x and 4x super-resolution models. During training, the input low-resolution image is cropped to a block of 64×64 size. During testing, the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) indicators are introduced as evaluation indicators. They are calculated in the Y channel of the super-resolution result converted YCbCr space.
[0032] Stacking shifted residual units of different sizes realizes lightweight SCNets of different scales. The smallest model SCNet-T is implemented by 16 64-channel shifted residual blocks, the basic model SCNet-B is implemented by 64 64-channel shifted residual blocks, and the largest SCNet-L is implemented by 32 128-dimensional shifted residual blocks.
[0033] The proposed SCNet in this invention is compared with various image super-resolution methods based on convolutional neural networks, including SRCNN (Image super-resolution using deep convolutional networks), VDSR (Accurate image super-resolution using very deep convolutional networks), DRRN (Image super-resolution via deep recursive residual network.), LapNet (DeepLaplacian pyramid networks for fast and accurate super-resolution.), ECBSR (Edge-oriented convolution block for real-time super resolution on mobiledevices.), LPARA (LAPAR:linearly-assembled pixel-adaptive regression networkfor single image super-resolution and beyond.), CARN (Fast,accurate,andlightweight super-resolution with cascading residual network.), IMDN (Lightweight image super-resolution with information multi-distillationnetwork.), FDIWN (Feature distillation interaction weighting network forlightweight image super-resolution.), LBNet (Lightweight bimodal network forsingle-image super-resolution via symmetric CNN and recursive transformer.) and ShuffleMixer (Shufflemixer:An efficient convnet for image super-resolution.).
[0034] Subjective results Figure 4 The super-resolution results of several images selected from Urban100 are shown. It can be seen that the SCNet super-resolution results have a clearer subjective effect than other CNN methods. At the same time, SCNet can also better reconstruct some edges and texture parts.
[0035] Objective results Figure 5 The objective performance of different super-resolution methods is listed using the above metrics. Figure 5 The best results are underlined, and the results of SCNet proposed in this invention are bolded. SCNet-L achieves the best quantitative performance. Compared with IMDN and SRResNet, which are widely used as benchmarks, SCNet-L improves by 0.26dB and 0.28dB respectively. In addition, the quantitative objective results are divided according to the size of the method model. Taking 4x super-resolution as an example, it can be seen that SCNet has achieved very good results in the comparison of methods of different scales. In particular, SCNet-B, with less than 700K parameters, has surpassed the existing CNN methods, except that it is lower than LBNet on Set5, but LBNet contains more parameters. In addition, the specific computational complexity comparison is summarized in Figure 7 In the figure, we can see that SCNets of different sizes have achieved a better balance between performance and computational complexity. Compared with LAPAR-C with a small number of parameters, SCNet-T has a lower computational complexity when it contains more parameters. Compared with the larger SRResNet, SCNet-L has fewer parameters and computational complexity and achieves better results. This is because the amount of point-by-point convolution calculations and parameters is one-ninth of that of 3×3 convolution, which expands to a deeper SCNet with stronger fitting capabilities, especially SCNet-B, which stacks 64 shifted residual units and achieves better super-resolution results with only one-third of the computational effort of SRResNet.
[0036] In order to verify the role of each component in the proposed SCNet, the present invention further conducts a series of ablation studies.
[0037] Shift convolution: First, the influence of different shift strategies in shift convolution is analyzed. This paper designs 5 different shift strategies such as Figure 6 Here, four points are selected and divided into two groups in different directions, as shown in Figure 6 (a) and Figure 6 (b) The default 8 points are as follows: Figure 6 As shown in (c), the moving step length is further expanded to obtain Figure 6 (d). Finally, Figure 6 (c) and Figure 6The points in (d) are merged to obtain Figure 6 (e) to verify the influence of the number of points. The specific results are shown in Figure 8 It can be seen that different shift operation settings have a significant impact on the final performance of the model. When the shift operation only selects 4 points, the SCNet index decreases overall. This is because it is difficult to effectively model neighbor relationships by selecting features from only 4 points. Compared with the default 8 points, the shift setting of the empty 8 points achieved better results. This is mainly because the empty shift introduces a larger receptive field. This also verifies that the shift operation can be designed by different parameter priors, and even obtain long-distance relationship modeling, which has better flexibility than general convolution. Finally, the overall effect of SCNet using 16 points has decreased, mainly because 16 points need to divide the features into 16 groups, and the number of features in each group is too small to perform effective relationship modeling.
[0038] The model capacity is due to the fact that shifted convolution requires only a very small number of parameters and computation, so SCNet can be easily expanded. Here, the scalability of SCNet is mainly discussed. This paper designs SCNets of different sizes and verifies them on the 4x super-resolution task. The objective indicators are shown in Fig. 9 It can be found that SCNet, which is implemented only by point-by-point convolution, has good scalability and can be effectively expanded to different scales of parameters. Here, ablation verification is performed from the two aspects of depth and width (feature dimension). In general, the benefits brought by expanding the depth are better than using a larger feature dimension.
[0039] The present invention proposes a lightweight image super-resolution network SCNet that is fully implemented using point-by-point convolution. Compared with general 3×3 convolution, point-by-point convolution contains fewer parameters and lower computational cost, but lacks the key feature of local feature fusion. In order to solve this problem, the present invention extends the point-by-point shifted convolution through spatial shift operation, enables it to have the ability of feature aggregation through manual feature aggregation, and the spatial shift operation has no additional computational cost. Based on the shifted convolution, the present invention replaces the 3×3 convolution in the standard residual structure and proposes a shifted residual unit. SCNet of different model sizes is realized by stacking shifted residual units of different sizes. Finally, the SCNet proposed in the present invention achieved the best results on multiple public test data sets. In addition, the present invention also verifies the effectiveness of the different modules proposed in the present invention through detailed ablation analysis.
[0040] The above is a detailed introduction to the lightweight image super-resolution method based on full point-by-point convolution proposed in the present invention. In this article, specific examples are used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A lightweight image super-resolution method based on full point-by-point convolution, characterized by: The method specifically includes: introducing a spatial shift operation, splitting the features extracted by point-by-point convolution into different groups along the channel direction, and moving the features of different groups along different spatial directions, so that the moved features are integrated with the features of neighboring pixels in the channel direction; the point-by-point convolution and spatial shift operations are collectively referred to as shifted convolution, and the shifted convolution has extremely low parameter and computational complexity in the point-by-point convolution, and can also perform effective local feature aggregation; based on the shifted convolution, a residual structure unit is obtained, and the residual structure unit is further stacked to obtain a lightweight image super-resolution network architecture: a shifted convolution network SCNet; the SCNet includes shallow feature extraction, deep feature extraction and high-resolution image reconstruction modules; given a low-resolution image I LR , first use the shallow feature extractor f s Map it to the specified hidden layer feature space to obtain the feature map F s =f s (I LR ), then, the shallow feature map is passed through the deep feature extractor f d , extract the deep feature map F d =f d (F s ), finally, the high-resolution image reconstruction module f u Upsample the deep features to obtain the final super-resolution result I SR =f u (F (d) ).
2. The method according to claim 1, characterized in that The spatial shift operation is specifically as follows: firstly, the input feature map is split into N groups evenly through channels, where N represents the number of shifted neighbor features; then, different groups are moved in different directions by specified step sizes, and after moving in different directions, the neighbor features at corresponding positions are aggregated.
3. The method according to claim 2, characterized in that Based on the shifted convolution layer, all 3×3 convolutions in the standard residual structure are replaced by point-by-point convolutions, in which the spatial shift operation is embedded. The improved shifted residual structure unit contains a shifted convolution and a point-by-point convolution and an activation layer.
Citation Information
Patent Citations
Content-guide Residual Network for Image Super-Resolution
AU2020100200A4
Image super-resolution reconstruction method based on cascade residual convolutional neural network
CN110276721A