Lightweight image super-resolution method and device based on spatial domain and frequency domain
Through the lightweight image super-resolution method based on the airspace and frequency domain, the problem of single image features and time-consuming calculation in the prior art is solved, and the efficient image super-resolution effect is achieved, and the computing efficiency and image quality of the model are improved.
Patent Information
- Application Number
- CN202510130560.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-05
AI Technical Summary
Although the super-resolution images reconstructed by neural networks in the prior art have high quantitative indexes PSNR and SSIM, there are problems such as single image features in the network, insufficient attention information and time-consuming to calculate the model.
The lightweight image super-resolution method based on the airspace and frequency domain is adopted to obtain shallow features through dynamic convolution, and combined with the multi-scale perception of the airspace feature calculation module and the high-frequency enhanced frequency domain feature calculation module, the middle and deep features are gradually extracted, and finally the super-resolution image is generated through feature fusion.
Effectively capture texture information on a larger spatial structure, emphasize high-frequency image information such as texture details, and realize the super-resolution effect of lightweight calculations, improving the computing efficiency and image quality of the model.
Smart Images

Figure CN120070180A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image denoising, and specifically relates to a lightweight image super-resolution method and device. Background Art
[0002] With the rapid development of digital image processing technology, image super-resolution (SR) technology plays an increasingly important role in fields such as multimedia, medical imaging, and satellite image processing. Image super-resolution technology aims to improve the resolution of images through algorithms and restore the detailed information of images to meet the needs of high-resolution display devices.
[0003] However, with the popularization of mobile devices and embedded systems, the demand for lightweight image processing models is increasing day by day. Traditional image super-resolution models often have numerous parameters and high computational complexity, making it difficult to run in real time on resource-constrained devices. Therefore, the research on lightweight models has become particularly important. Lightweight models aim to reduce the number of parameters and computational complexity of the model while maintaining or improving the performance of image processing. Such models can run efficiently on mobile devices and embedded systems to meet the needs of real-time processing.
[0004] Research shows that existing deep learning-based research has achieved good results in quantitative metrics such as PSNR and structural similarity index SSIM, but these models contain a large number of parameters, slow computational efficiency, and it is difficult to achieve real-time image super-resolution on mobile devices and embedded systems. Some existing lightweight super-resolution networks need to be improved in terms of model structure and can more effectively combine attention mechanism modules. In recent years, many studies on image super-resolution have shown that neural networks trained using the attention mechanism can enhance tissue edge textures, but usually only use a single attention mechanism, and there is no obvious improvement in obtaining attention information.
[0005] In summary, although the super-resolution images reconstructed by neural networks in the prior art have high quantitative metrics PSNR and SSIM, there are problems such as single image feature acquisition by the network, insufficient attention information, and large model calculation time consumption. Summary of the Invention
[0006] To solve the above technical problems, the present invention provides a lightweight image super-resolution method and device based on the spatial domain and frequency domain, which can provide better quality image super-resolution services for mobile devices.
[0007] Specifically, the method includes the following steps:
[0008] S1: Collect pictures, perform dynamic convolution, and obtain shallow features;
[0009] S2: Input the shallow features into the spatial feature calculation module with multi-scale perception to obtain middle-level features;
[0010] S3: Input the middle-level features into the frequency domain feature calculation module with high-frequency enhancement to obtain deep features;
[0011] S4: Convolve the deep features to obtain enhanced features;
[0012] S5: Fuse the enhanced features with the original input features and finally output the super-resolution image.
[0013] Preferably, S1 includes the following steps:
[0014] S1.1: Collect a three-channel image x in the RGB color space;
[0015] S1.2: Perform dynamic convolution calculation to obtain shallow features:
[0016] First, perform convolution:
[0017] f = Conv 3×3 (x)
[0018] In the above formula, Conv n×n (.) represents a convolution kernel of size n;
[0019] Then, perform global feature extraction to obtain the global feature vector G:
[0020] G = adaptive average pool(f)
[0021] In the above formula, adaptive averagepool(.) represents adaptive average pooling;
[0022] Next, perform dynamic weight K generation:
[0023] K = softmax(W 2 ·ReLU(W 1 ·G))
[0024] In the above formula, W 1 and W 2 represent learnable weight matrices, ReLU represents the activation function, and the softmax formula is: N is the number of feature channels;
[0025] Finally, perform dynamic convolution operation to obtain the shallow feature f 1 :
[0026] f 1 = Conv k (f)
[0027] In the above formula, substitute Conv k (.) represents a convolution kernel with convolution weight K.
[0028] Preferably, S2 includes the following steps:
[0029] S2.1: Perform two downsampling operations on the shallow features:
[0030] f 1D2 = Down 2 (f 1 )
[0031] f 1D4 = Down 4 (f 1 )
[0032] In the above formula, Down n (.) represents an image downsampling operation with a multiple of n;
[0033] S2.2: Calculate channel attention for features at 3 different sizes:
[0034] f cab1 = CAB(f 1 )
[0035] f cab2 = CAB(f 1D2 )
[0036] f cab3 = CAB(f 1D4 )
[0037] In the above formula, CAB (Channel Attention Block) is a channel attention module, and its specific calculation method is:
[0038]
[0039] In the above formula, max pooling(.) represents max pooling, represents matrix multiplication calculation;
[0040] S2.3: Perform image upsampling operations on 2 of the 3 calculation results in the previous step:
[0041] f cab2U2 = Up 2 (f cab2 )
[0042] f cab3U4 = Up 4 (f cab3 )
[0043] In the above formula, Up n(.) represents an image upsampling operation with a multiple of n;
[0044] S2.4: Concatenate the 3 image features to obtain the middle-level features:
[0045]
[0046] In the above formula, represents the concatenation operation;
[0047] Preferably, S3 includes the following steps:
[0048] S3.1: Calculate the mean feature map of f 2 :
[0049] f m = Mean(f 2 )
[0050] In the above formula, Mean(.) represents calculating the mean of each channel in the graph;
[0051] S3.2: Obtain the high-frequency feature map through difference calculation, and combine f 2 to output the deep features:
[0052]
[0053] Preferably, S4 includes the following steps:
[0054] S4.1: Perform convolution calculation on f 3 to obtain the enhanced features:
[0055] f 4 = Conv 3×3 (f 3 )
[0056] Preferably, S5 includes the following steps:
[0057] S5.1: Fuse the features of f 4 and x, and finally output the super-resolution image X:
[0058] X = x + f 4
[0059] S5.2: Obtain a lightweight image super-resolution model based on the spatial and frequency domains;
[0060] S5.3: Evaluate the super-resolution effect of the model image according to the calculated quantitative indicators PSNR and SSIM.
[0061] The second technical solution adopted by the present invention is: A lightweight image super-resolution device based on the spatial and frequency domains, including:
[0062] Multi-scale perception spatial domain module: used to capture texture information on a larger spatial structure;
[0063] High-frequency enhancement frequency domain module: used to emphasize high-frequency information of images such as texture details;
[0064] Feature fusion module: used to merge the enhanced features and the original image to generate the reconstructed super-resolution image.
[0065] The beneficial effects of the method and device of the present invention are as follows: First, through convolution calculation, shallow features are obtained and input into the multi-scale perception spatial domain feature calculation module to obtain middle-level features. Then, the middle-level features are input into the high-frequency enhancement frequency domain feature calculation module to obtain deep features. Finally, convolution calculation is performed to obtain enhanced features, and the enhanced features are fused with the original input features to finally output a super-resolution image. The present invention can effectively address the challenges commonly existing in the image super-resolution of existing neural network models. On the one hand, the multi-scale perception spatial domain module is used to effectively capture texture information on a larger spatial structure. On the other hand, the high-frequency enhancement frequency domain module is used to emphasize high-frequency information of images such as texture details. At the same time, the current module and process are both lightweight calculations, and finally a lightweight image super-resolution effect is achieved. Brief Description of the Drawings
[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the prior art and the embodiments. The following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0067] Figure 1 It is a schematic flowchart of a lightweight image super-resolution method based on the spatial domain and the frequency domain of the present invention;
[0068] Figure 2 It is a schematic diagram of a specific embodiment of a lightweight image super-resolution method based on the spatial domain and the frequency domain of the present invention;
[0069] Figure 3 It is a schematic diagram of the structure of the spatial domain feature calculation module of a lightweight image super-resolution method based on the spatial domain and the frequency domain of the present invention;
[0070] Figure 4 It is a schematic diagram of the structure of the frequency domain feature calculation module of a lightweight image super-resolution method based on the spatial domain and the frequency domain of the present invention;
[0071] Figure 5 It is the super-resolution effect diagram after the specific implementation of the present invention. Specific Implementation Scheme
[0072] To make the objectives, features, and advantages of the present invention more apparent and understandable, the following will describe the technical solutions in the embodiments of the present invention clearly and completely in conjunction with the accompanying drawings in the embodiments of the present invention. It should be noted that the following detailed descriptions are all illustrative and are intended to provide further explanations for the present application. Unless otherwise specified, all other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0073] The embodiments of the present invention provide a lightweight image super-resolution method and device based on the spatial domain and frequency domain, which are used to achieve high-quality image super-resolution effects by using lightweight computing methods at the algorithm level.
[0074] A typical embodiment of the present invention refers to Figure 1 、 Figure 2 , and the method includes the following steps:
[0075] S1: Collect pictures, perform dynamic convolution, and obtain shallow features;
[0076] S2: Input the shallow features into the spatial domain feature calculation module with multi-scale perception to obtain middle-level features;
[0077] S3: Input the middle-level features into the frequency domain feature calculation module with high-frequency enhancement to obtain deep features;
[0078] S4: Convolve the deep features to obtain enhanced features;
[0079] S5: Fuse the enhanced features with the original input features and finally output the super-resolution image.
[0080] Furthermore, as a preferred embodiment of this method, S1 includes the following steps:
[0081] S1.1: Collect a three-channel image x in the RGB color space;
[0082] Furthermore, the collection of the three-channel image x in the RGB color space specifically includes:
[0083] Collect the original clear image dataset at high resolution, marked with the "hr" label. And perform image downsampling operations. The downsampling operations can use image interpolation operations such as bilinear interpolation and bicubic interpolation to obtain the three-channel image x in the RGB color space, marked with the "lr" label.
[0084] S1.2: Perform dynamic convolution calculation to obtain shallow features:
[0085] First, perform convolution:
[0086] f = Conv 3×3(x)
[0087] In the above formula, Conv n×n (.) represents a convolution kernel of size n.
[0088] Then, global feature extraction is performed to obtain the global feature vector G:
[0089] G = adaptive averagepool(f)
[0090] In the above formula, adaptive averagepool(.) represents adaptive average pooling.
[0091] Next, dynamic weight K generation is performed:
[0092] K = softmax(W 2 · ReLU(W 1 · G))
[0093] In the above formula, W 1 and W 2 represent learnable weight matrices, ReLU represents an activation function, and the softmax formula is: N is the number of feature channels.
[0094] Finally, dynamic convolution operation is performed to obtain the shallow feature f 1 :
[0095] f 1 = Conv k (f)
[0096] In the above formula, the Conv k (.) represents a convolution kernel with convolution weight K.
[0097] Furthermore, S2 includes the following steps:
[0098] S2.1: Perform two downsampling operations on the shallow feature:
[0099] f 1D2 = Down 2 (f 1 )
[0100] f 1D4 = Down 4 (f 1 )
[0101] In the above formula, Down n (.) represents an image downsampling operation with a multiple of n.
[0102] S2.2: Perform channel attention calculation on the features at three different sizes:
[0103] f cab1 = CAB(f 1 )
[0104] f cab2 = CAB(f 1D2 )
[0105] f cab3 = CAB(f 1D4 )
[0106] In the above formula, CAB (Channel Attention Block) is the channel attention module, and its specific calculation method is as follows:
[0107]
[0108] In the above formula, max pooling(.) represents max pooling, represents matrix multiplication calculation.
[0109] S2.3: Perform image upsampling operations on two of the three calculation results in the previous step:
[0110] f cab2U2 = Up 2 (f cab2 )
[0111] f cab3U4 = Up 4 (f cab3 )
[0112] In the above formula, Up n (.) represents an image upsampling operation with a multiple of n.
[0113] S2.4: Concatenate the three image features to obtain the middle-level features:
[0114]
[0115] In the above formula, represents the concatenation operation.
[0116] Furthermore, S3 includes the following steps:
[0117] S3.1: Calculate the mean feature map of f 2 :
[0118] f m = Mean(f 2 )
[0119] In the above formula, Mean(.) represents calculating the mean of each channel in the graph.
[0120] S3.2: Obtain the high-frequency feature map through difference calculation and combine it with f 2 Output deep features:
[0121]
[0122] Furthermore, S4 includes the following steps:
[0123] Perform convolution calculation on f 3 to obtain enhanced features:
[0124] f 4 = Conv 3×3 (f 3 )
[0125] Furthermore, S5 includes the following steps:
[0126] S5.1: Fuse the features of f 4 and x, and finally output the super-resolution image X:
[0127] X = x + f 4
[0128] S5.2: Obtain a lightweight image super-resolution model based on the spatial and frequency domains;
[0129] S5.3: Evaluate the super-resolution effect of the model image according to the calculated quantitative indicators PSNR and SSIM.
[0130] The above is a specific description of the preferred embodiment of the present invention, but the present invention is not limited to the described embodiment. Those skilled in the art can make other equivalent deformations or substitutions without departing from the spirit of the invention, and these equivalent deformations or substitutions are included within the scope defined by the claims of the application.
Claims
1. A lightweight image super-resolution method based on spatial domain and frequency domain, characterized in that: The method comprises the following steps: S1: Collect pictures, perform dynamic convolution, and obtain shallow features; S2: Input shallow features into the multi-scale perception spatial feature calculation module to obtain mid-level features; S3: Input the middle-level features into the high-frequency enhanced frequency domain feature calculation module to obtain the deep-level features; S4: Convolve the deep features to obtain enhanced features; S5: Fusion the enhanced features with the original input features, and finally output a super-resolution image.
2. The lightweight image super-resolution method based on spatial domain and frequency domain according to claim 1, characterized in that: The S1 comprises the following steps: S1.1: Collect a three-channel image x in RGB color space; S1.2: Perform dynamic convolution calculation to obtain shallow features: First, perform the convolution: f=Conv 3×3 (x) In the above formula, Conv n×n (.) represents the convolution kernel of size n; Then perform global feature extraction to obtain the global feature vector G: G = adaptive average pool (f) In the above formula, adaptive averagepool(.) represents adaptive average pooling; Then generate the dynamic weight K: K = softmax(W2·ReLU(W1·G)) In the above formula, W1 and W2 represent learnable weight matrices, ReLU represents the activation function, and the softmax formula is: N is the number of feature channels; Finally, a dynamic convolution operation is performed to obtain the shallow feature f1: f1=Conv k (f) In the above formula, substitute Conv k (.) represents the convolution kernel with convolution weight K.
3. The lightweight image super-resolution method based on spatial domain and frequency domain according to claim 1, characterized in that: S2 includes the following steps: S2.1: Perform two downsampling operations on shallow features: f 1D2 =Down2(f1) in 1D4 =Down4(f1) In the above formula, Down n (.) represents the image downsampling operation with a multiple of n; S2.2: Channel attention calculation for features at three different scales: f cab1 =CAB(f1) f cab2 =CAB(f 1D2 ) f cab3 =CAB(f 1D4 ) In the above formula, CAB (Channel Attention Block) is the channel attention module, and its specific calculation method is: In the above formula, max pooling(.) represents the maximum pooling, Represents matrix multiplication calculation; S2.3: Perform image upsampling operation on 2 of the 3 calculation results in the previous step: f cab2U2 =Up2(f cab2 ) f cab3U4 =Up4(f cab3 ) In the above formula, Up n (.) represents an image upsampling operation with a multiple of n; S2.4: Concatenate the three image features to obtain the middle-level features: In the above formula, Represents a splicing operation.
4. The lightweight image super-resolution method based on spatial domain and frequency domain according to claim 1, characterized in that: S3 includes the following steps: S3.1: Calculate the mean feature map of f2: f m =Mean(f2) In the above formula, Mean(.) represents the mean of each channel in the calculation graph; S3.2: Obtain high-frequency feature maps through difference calculation, and output deep features in combination with f2:
5. The lightweight image super-resolution method based on spatial domain and frequency domain according to claim 1, characterized in that: S4 includes the following steps: S4.1: Perform convolution calculation on f3 to obtain enhanced features: f4=Conv 3×3 (f3)。 6. The lightweight image super-resolution method based on spatial domain and frequency domain according to claim 1, characterized in that: S5 includes the following steps: S5.1: Fuse the features of f4 and x, and finally output the super-resolution image X: X=x+f4 S5.2: Obtain a lightweight image super-resolution model based on spatial domain and frequency domain; S5.3: Evaluate the super-resolution effect of the model image by calculating the quantitative indicators PSNR and SSIM.
7. A lightweight image super-resolution device based on spatial domain and frequency domain, characterized in that: include: Multi-scale-aware spatial domain module: used to capture texture information on larger spatial structures; Frequency domain module for high-frequency enhancement: used to emphasize high-frequency information of images such as texture details; Feature fusion module: used to merge the enhanced features with the original image to generate a reconstructed super-resolution image.
Citation Information
Patent Citations
Image super-resolution method based on frequency domain and spatial domain
CN114463183A
Remote sensing image super-resolution reconstruction method based on feature information distillation network
CN115601236A
Light-weight image super-resolution reconstruction method and device based on frequency domain separation network
CN115775206A
Image super-resolution method and system combining multi-scale local and global information
CN117726516A
Depth image enhancement method based on super-resolution
CN117788297A