Image super-resolution reconstruction method and device based on residual pyramid and frequency spectrum fusion
The image super-resolution reconstruction method of residual pyramid and spectrum fusion solves the problems of insufficient utilization of frequency domain information and large model size, achieves efficient image reconstruction and lightweight deployment, and improves image reconstruction accuracy and multi-scale information modeling capabilities.
Patent Information
- Application Number
- CN202510975107.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-06-18
- Filing Date
- 2025-07-15
- Publication Date
- 2025-09-26
AI Technical Summary
Existing image super-resolution reconstruction technology has problems such as insufficient utilization of frequency domain information, insufficient restoration of edge details, large model size, difficulty in deployment, and lack of multi-scale fusion mechanism, resulting in poor image reconstruction effect and difficulty in model deployment.
An image super-resolution reconstruction method based on residual pyramid and spectrum fusion is adopted. Through four stages of shallow feature extraction, deep feature extraction in spatial domain, deep feature extraction in frequency domain and image reconstruction, combined with lightweight Transformer submodule and spectrum fusion module, multi-scale feature extraction and dynamic weighting of spectrum attention are realized, and a lightweight network is designed.
Significantly improve image reconstruction accuracy, enhance multi-scale information modeling capabilities, achieve coordinated enhancement in the frequency domain and spatial domain, reduce the number of model parameters, and is suitable for edge devices.
Smart Images

Figure CN120707389A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning computer vision and image processing technology, and specifically relates to an image super-resolution reconstruction method and device based on residual pyramid and spectrum fusion. Background Art
[0002] Image super-resolution reconstruction technology aims to convert low-resolution images into high-resolution images through reconstruction methods. It is widely used in medical imaging, remote sensing image processing, security monitoring, video enhancement, and other fields. In recent years, deep learning has made significant progress in the field of super-resolution, with convolutional neural networks becoming the mainstream technology due to their efficient feature extraction capabilities.
[0003] In recent years, deep learning methods, especially the combination of Transformers and CNNs, have made breakthrough progress. However, there are still many problems to be solved: First, the frequency domain information is insufficiently utilized. Most traditional methods only focus on the contextual relationship between spatial pixels, ignoring the high-frequency structural information contained in the image in the frequency domain. Secondly, edge details are not fully restored. During the magnification process, the original texture edges in the image often become blurred or distorted due to incomplete information. The model is large and relatively difficult to deploy. High-performance super-resolution models often contain a large number of parameters and high computational complexity, which is not conducive to deployment on edge devices. The final problem is that the current model lacks a multi-scale fusion mechanism. The single-scale feature extraction method cannot effectively capture the multi-level and cross-scale semantic information in the image.
[0004] The invention is titled "Lightweight image super-resolution method and device based on efficient frequency domain Transformer", with publication number CN119180752B. This invention proposes an image super-resolution reconstruction model based on the combination of Transformer and frequency domain, which overcomes the problem of high computational complexity of Transformer, but does not fully combine frequency domain information with spatial domain information.
[0005] Therefore, how to efficiently fuse spatial and frequency domain information to achieve a balance between local details and global structure while lightweighting the model structure has become an important topic in current image super-resolution research. Summary of the Invention
[0006] To overcome the shortcomings of the aforementioned prior art, the present invention aims to provide a method and device for image super-resolution reconstruction based on residual pyramid and spectral fusion. This method features multi-scale feature extraction, dual-domain fusion of spatial and spectral domains, dynamic spectral attention weighting, lightweight network design, and end-to-end trainability. This method significantly improves image reconstruction accuracy, enhances multi-scale information modeling capabilities, and achieves synergistic enhancement in the frequency and spatial domains, thereby achieving better image reconstruction while reducing the number of model parameters.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is:
[0008] An image super-resolution reconstruction method based on residual pyramid and spectrum fusion, comprising the following steps:
[0009] Step 1: Cropping the images in the dataset to obtain an image block training dataset; the data in the dataset is an image pair consisting of a low-resolution image and its corresponding high-resolution image; the low-resolution image is cropped and its corresponding high-resolution image constitutes the data in the image block training dataset;
[0010] Step 2: Based on the image block training set, an image super-resolution reconstruction model based on residual pyramid and spectrum fusion is constructed to synthesize a high-resolution image;
[0011] Step 3: After setting the hyperparameters, loss function, optimizer, and number of iterations of the image super-resolution reconstruction model based on residual pyramid and spectrum fusion, the model is trained to obtain a trained lightweight image super-resolution model;
[0012] Step 4: Input the low-resolution image to be reconstructed into the trained lightweight image super-resolution model to obtain a high-resolution image.
[0013] In step 1, DF2K is used as a training data set. Each high-resolution image in the data set is downsampled to generate a corresponding low-resolution image. The two constitute a pair of training samples. During the training process, a low-resolution image and its corresponding high-resolution image are selected, and image blocks of 64×64 in size are randomly cropped from the low-resolution image as input. The corresponding high-resolution image with the same magnification is used as the "true value" to calculate the reconstruction error, guiding the network to continuously adjust the parameters to restore richer details and more accurate texture information. In step 2, the image super-resolution reconstruction model based on residual pyramid and spectrum fusion includes four stages: shallow feature extraction, deep feature extraction based on spatial domain, deep feature extraction based on frequency domain, and image reconstruction;
[0014] In the shallow feature extraction stage, the input low-resolution image is subjected to preliminary convolution processing to obtain a shallow feature map. The obtained shallow feature map is first subjected to deep feature extraction based on the spatial domain to capture spatial detail information, and then subjected to deep feature extraction based on the frequency domain to extract high-frequency details, and finally subjected to an upsampling-based image reconstruction module to synthesize a high-resolution image.
[0015] The specific process is:
[0016] Step 2.1: The shallow feature extraction module uses a 3×3 convolutional layer to extract shallow features from the low-resolution image;
[0017] F0=H conv3×3 (I LR )
[0018] Among them, I LR Represents the input low-resolution image, H conv3×3 Represents a shallow convolution extraction network constructed using a 3×3 convolution kernel, which is used to extract the basic feature F0, that is, the extracted shallow feature;
[0019] Step 2.2: Spatial-based deep feature extraction: Input the basic feature F0 to the spatial-based deep feature extraction module for deep feature extraction. This module consists of two parts: a lightweight Transformer submodule and a residual pyramid submodule.
[0020] Step 2.3: Input the extracted spatial deep features into the frequency domain-based feature extraction module for spectrum fusion. The frequency domain-based feature extraction module is divided into two parts: Fourier domain and wavelet domain to extract frequency domain deep features;
[0021] Step 2.4: The image reconstruction module uses a dynamic upsampling module to reconstruct the image using deep features in the frequency domain.
[0022] The step 2.2 is specifically as follows:
[0023] Step 2.2.1: The basic feature F0 is first compressed and dimensionally reduced through the lightweight Transformer submodule. The input basic feature F0 is compressed through linear transformation to reduce its dimension and generate the intermediate feature F r ;
[0024] F r =W r reshape (F0)
[0025] Among them, W r It is a learnable weight matrix with a dimension of C / 2, which is used to compress the feature dimension;
[0026] Step 2.2.2: By inputting the intermediate feature F r , calculate the query (Q), key (K) and value (V) and window attention calculation, the formula is as follows:
[0027] Q,K,V=reshape(Linear(F r ))→split(3)
[0028]
[0029] Among them, d k is the dimension of the key, F rIt is the intermediate feature passed through the lightweight Transformer submodule. Linear() is a fully connected transformation. Q, K, and V represent query (Q), key (K), and value (V) vectors respectively. They are obtained by reshaping the input features after linear transformation and dividing them along the channel dimension. Split (3) means dividing the linearly mapped vector into three parts along the channel direction and assigning them to Q, K, V, and Q respectively. i ,K i ,V i Represents the query, key and value vectors at the i-th position respectively. Softmax() normalizes the attention scores of all positions to make them a probability distribution. i Represents the attention output of the i-th position;
[0030] Step 2.2.3: Perform attention result splicing. After splicing the attention results of all windows, restore the original channel dimension to obtain the enhanced feature F a ;
[0031] Step 2.2.4: After the function enhancement, the multi-layer perception unit will enhance the feature F a The input is sent to the function-enhanced multi-layer perception unit, where channel expansion is first performed and Ghost features are generated. The formula is as follows:
[0032] F mlp =Conv2(DWConv1(A·B))
[0033] Among them, A·B are two features obtained by convolution operation and channel segmentation, and DWConv1 is a 7×7 depth-wise separable convolution;
[0034] Step 2.2.5: Combine the residual structure with the learnable scaling factor α and output the final feature map F1.
[0035] F1=F mlp α+F1
[0036] After the lightweight Transformer submodule extracts features, F1 inputs the second module of the deep feature extraction spatial domain part, that is, through multi-scale spatial shift operation (Shift) and depth-wise separable dilated convolution (DilatedConv) operation;
[0037] Step 2.2.6: First, extract information at different scales. By shifting the feature map at different scales (such as Shift4, Shift8, Shift16), the model's ability to perceive local features at different scales is enhanced. The formula is as follows:
[0038] Shift i =Shifti (H DilatedConv (Shift i (F1)))
[0039] Among them, Shift i H represents the spatial shift operation of the i-th scale (e.g., Shift4, Shift8, Shift16). DilatedConv (·) represents the depthwise separable dilated convolution operation, which expands the receptive field and extracts features at different scales;
[0040] Step 2.2.7: The feature maps of different scales are weighted fused by learnable weights to obtain the fused feature map x fused , the formula is as follows:
[0041]
[0042] Among them, Weight(Shift i ) is a learnable weight that controls the importance of each scale feature in the fusion;
[0043] Step 2.2.8: Fused feature map x fused It is further refined through 1×1 convolution operation and the feature expression ability is enhanced through ReLU activation function. The formula is as follows:
[0044] x out =ReLU(H conv1×1 (x fused ))
[0045] Among them, H conv1×1 is a 1×1 convolution operation, which is used to further refine the fused feature map. ReLU(·) is the ReLU activation function, which enhances the nonlinear expression ability of the model.
[0046] Step 2.2.9: Finally, output feature x out Feature map x after weighted fusion with multi-scale features fused The residuals are added and enhanced by the gated spatial attention module (GSAU);
[0047] F2=GSAU(x fused +x out )
[0048] After a complete deep feature extraction module based on spatial domain, the spatial feature F2 is generated, and the gated spatial attention module can enhance the important regional features and improve the expressive ability of the model.
[0049] The step 2.3 is specifically as follows:
[0050] Step 2.3.1: Convert the spatial domain feature F2 to the frequency domain through two-dimensional fast Fourier transform (FFT) to extract the frequency information;
[0051] F f =FFT(F2)
[0052] Among them, F2 is the spatial domain feature, FFT is Fourier transform, F f is the feature after Fourier transform;
[0053] Step 2.3.2: Extract local texture details through wavelet convolution operation;
[0054] F w =WaveConv(F2)
[0055] Among them, WaveConv is wavelet convolution, F w is the feature extracted after wavelet convolution;
[0056] Step 2.3.3: Concatenate the Fourier transformed features with the features extracted by wavelet convolution and connect them with the residual of the original input features to generate the final fused features;
[0057] F fusion =Conv(1×1)(Concat(F f ,F w ))+F2
[0058] Among them, Conv(1×1) is 1×1 convolution, F f is the feature after Fourier transform, F w is the feature extracted after wavelet convolution, F2 is the spatial feature, and F fusion is the fused feature, and Concat is the concatenation operation.
[0059] The step 2.4 is specifically as follows:
[0060] The image reconstruction module uses a dynamic upsampling module to perform image reconstruction, and the deep features are fused in the dynamic upsampling module to generate a final high-resolution image;
[0061] I HR =Dysample(F WFFB )
[0062] Among them, F WFFB It is the feature map after deep feature extraction, I HR The reconstructed high-resolution image.
[0063] The step 3 is specifically as follows: setting the hyper parameters of the model (such as the number of feature channels 64, the initial learning rate 2×10 -4), define the loss function (such as Charbonnier Loss) and optimizer (Adam), and train the constructed super-resolution model under the condition of 500,000 iterations until a trained lightweight model is obtained.
[0064] The step 4 is specifically as follows: inputting the low-resolution image to be reconstructed into a trained lightweight super-resolution model, and directly outputting the corresponding high-resolution image through shallow layer extraction within the model, deep feature fusion in the spatial and frequency domains, and dynamic upsampling. An image super-resolution reconstruction device based on residual pyramid and spectrum fusion includes:
[0065] A memory for storing a computer program for implementing an image super-resolution reconstruction method based on residual pyramid and spectrum fusion;
[0066] The processor is configured to implement an image super-resolution reconstruction method based on residual pyramid and spectrum fusion when executing the computer program.
[0067] A computer-readable storage medium stores a computer program, which, when executed by a processor, can implement the image super-resolution reconstruction method based on residual pyramid and spectrum fusion according to the present invention.
[0068] Beneficial effects of the present invention:
[0069] Multi-domain fusion enhances high-frequency restoration capabilities: For the first time, the wavelet transform and Fourier transform are fused in parallel in the image super-resolution task, modeling spectral information from both global and local perspectives, significantly improving the fidelity of edges and details.
[0070] Multi-scale pyramid module structure efficiently extracts features: It extracts features of different granularities by stacking multi-scale residuals, enhancing the model's generalization ability in complex scenarios.
[0071] The lightweight Transformer is suitable for edge devices: while ensuring performance, it greatly reduces the model's computational complexity and parameters, making it suitable for resource-constrained platforms such as mobile terminals and embedded systems.
[0072] The space-frequency synergy mechanism improves reconstruction quality: it enables effective reverse guidance of frequency domain information on spatial features, enhancing the image's perceived contrast and local texture restoration capabilities.
[0073] The present invention significantly improves the image reconstruction accuracy, enhances the multi-scale information modeling capability, and realizes the coordinated enhancement of the frequency domain and the spatial domain, thereby better performing image reconstruction while reducing the number of model parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 It is an overall flow chart of the method of the present invention.
[0075] Figure 2 This is the overall model structure diagram of image super-resolution based on residual pyramid and spectrum fusion.
[0076] Figure 3 Schematic diagram of the specific structure of the residual pyramid module.
[0077] Figure 4 Schematic diagram of the specific structure of the spectrum fusion module.
[0078] Figure 5 The figure shows a comparison of the number of parameters and PSNR indicators of the recent models on the Manga109 dataset.
[0079] Figure 6 This is a visual comparison of the model with other models on the Set14 dataset.
[0080] Figure 7 The following is a visual comparison of the model with other models on the Manga109 dataset. DETAILED DESCRIPTION
[0081] The present invention will be further described in detail below with reference to the accompanying drawings.
[0082] like Figure 1 As shown in FIG. 1 , as a specific embodiment of the present invention, an image super-resolution reconstruction method based on residual pyramid and spectrum fusion specifically includes the following steps:
[0083] Step 1: Use DF2K as the training dataset, which consists of 800 images from DIV2K and 2,650 images from Flickr2K. Each high-resolution image is downsampled to a corresponding low-resolution image, which together form a training pair. During training, 64×64 image patches are randomly cropped from the low-resolution images as input to the model in Step 2. The corresponding high-resolution image patches are also captured as supervisory labels for subsequent feature extraction, model training, and reconstruction optimization.
[0084] Step 2: Construct an image super-resolution reconstruction model based on residual pyramid and spectrum fusion.
[0085] Specifically, the image super-resolution reconstruction model structure based on residual pyramid and spectrum fusion in the present invention is as follows: Figure 2 shown.
[0086] In step 2, an image super-resolution reconstruction model based on residual pyramid and spectrum fusion was constructed. The specific process consists of four stages: shallow feature extraction, deep feature extraction based on the spatial domain, deep feature extraction based on the frequency domain, and image reconstruction. The purpose of this step is to fully extract and reconstruct potential high-frequency details in low-resolution images by constructing a deep neural network that fuses spatial and frequency domain information. The model introduces a residual pyramid in its structure to capture multi-scale spatial features, and uses a spectrum fusion module to enhance the expression of high-frequency details. The overall process establishes a basic feature representation through shallow feature extraction, then extracts structural texture and frequency information in the spatial and frequency domains respectively, and finally achieves high-quality image reconstruction through dynamic upsampling. This step effectively improves the accuracy and clarity of image super-resolution reconstruction, reduces artifact generation, and has good computational efficiency and generalization capabilities.
[0087] The specific steps are as follows:
[0088] Step 2.1: Shallow feature extraction and basic feature construction.
[0089] First, input a low-resolution image I LR After normalization and convolution operations, the basic feature F0 is obtained, and the formula is as follows:
[0090] F0=H conv3×3 (LayerNorm(I LR ))
[0091] Among them, LayerNorm(·) represents the layer normalization operation, H conv3×3 Indicates a convolution operation using a 3×3 convolution kernel;
[0092] Through this operation, the low-resolution input image is mapped to a high-dimensional feature space, ready for further deep feature extraction.
[0093] Step 2.2: Deep feature extraction based on spatial domain.
[0094] Input the basic feature F0 to the deep feature extraction module based on the spatial domain, which consists of two parts: a lightweight Transformer sub-module and a residual pyramid sub-module.
[0095] Step 2.2.1: The basic feature F0 is first compressed and dimensionally reduced through the lightweight Transformer submodule. The input basic feature F0 is compressed through linear transformation to reduce its dimension and generate the intermediate feature F r ;
[0096] F r =W r reshape (F0)
[0097] Among them, W r It is a learnable weight matrix with dimension C / 2, which is used to compress the feature dimension.
[0098] Step 2.2.2: By inputting the intermediate feature F r , calculate the query (Q), key (K) and value (V) and window attention calculation, the formula is as follows:
[0099] Q,K,V=reshape(Linear(F r ))→split(3)
[0100]
[0101] Among them, d k is the dimension of the key, F r It is the intermediate feature passed through the lightweight Transformer submodule. Linear() is a fully connected transformation. Q, K, and V represent query (Q), key (K), and value (V) vectors respectively. They are obtained by reshaping the input features after linear transformation and dividing them along the channel dimension. Split (3) means dividing the linearly mapped vector into three parts along the channel direction and assigning them to Q, K, V, and Q respectively. i ,K i ,V i Represents the query, key and value vectors at the i-th position respectively. Softmax() normalizes the attention scores of all positions to make them a probability distribution. i represents the attention output of the i-th position.
[0102] Step 2.2.3: Perform attention result splicing. After splicing the attention results of all windows, restore the original channel dimension to obtain the enhanced feature F a ;
[0103] Step 2.2.4: After the function enhancement, the multi-layer perception unit will enhance the feature F a The input is sent to the function-enhanced multi-layer perception unit, where channel expansion is first performed and Ghost features are generated. The formula is as follows:
[0104] F mlp =Conv2(DWConv1(A·B))
[0105] Among them, A·B are two features obtained by convolution operation and channel segmentation, and DWConv1 is a 7×7 depth-wise separable convolution;
[0106] Step 2.2.5: Combine the residual structure with the learnable scaling factor α and output the final feature map F1.
[0107] F1=F mlp α+F1
[0108] After the lightweight Transformer submodule extracts features, F1 inputs the second module of the deep feature extraction spatial domain part, that is, through multi-scale spatial shift operation (Shift) and depth-wise separable dilated convolution (DilatedConv) operation;
[0109] Step 2.2.6: First, extract information at different scales. By shifting the feature map at different scales (such as Shift4, Shift8, Shift16), the model's ability to perceive local features at different scales is enhanced. The formula is as follows:
[0110] Shift i =Shift i (H DilatedConv (Shift i (F1)))
[0111] Among them, Shift i H represents the spatial shift operation of the i-th scale (e.g., Shift4, Shift8, Shift16). DilatedConv (·) represents the depthwise separable dilated convolution operation, which expands the receptive field and extracts features at different scales;
[0112] Step 2.2.7: Multi-scale feature weighted fusion.
[0113] The feature maps of different scales are weighted and fused through learnable weights to obtain the fused feature map x fused This weighted fusion can dynamically adjust the contribution of different scale features according to the importance of each scale feature and optimize the image reconstruction effect. The formula is as follows:
[0114]
[0115] Among them, Weight(Shift i ) are learnable weights that control the importance of each scale feature in the fusion. Through weighted fusion, information at different scales is effectively integrated, which helps improve the model's ability to express multi-scale features;
[0116] Step 2.2.8: 1×1 convolution and activation function refine features.
[0117] The fused feature map x fused It is further refined through 1×1 convolution operation and the feature expression ability is enhanced through ReLU activation function. The formula is as follows:
[0118] xout =ReLU(H conv1×1 (x fused ))
[0119] Among them, H conv1×1 is a 1×1 convolution operation, which is used to further refine the fused feature map. ReLU(·) is the ReLU activation function, which enhances the nonlinear expression ability of the model. Through this refining process, the obtained x out The feature maps have stronger expressive power and are ready for further processing.
[0120] Step 2.2.9: Combine the feature maps after weighted fusion of multi-scale features and input them into the gated spatial attention module.
[0121] Finally, the output feature x out Feature map x after weighted fusion with multi-scale features fused The residuals are added and enhanced by the gated spatial attention module (GSAU).
[0122] F2=GSAU(x fused +x out )
[0123] After a complete deep spatial feature extraction module, the spatial feature F2 is generated. The gated spatial attention module can enhance the features of important regions and improve the expressiveness of the model.
[0124] Step 2.3: The feature map after spatial domain processing is input into the frequency domain-based feature extraction module for spectrum fusion, which is mainly divided into two parts: Fourier domain and wavelet domain.
[0125] Step 2.3.1: Convert the feature map to the frequency domain through two-dimensional fast Fourier transform (FFT) to extract the frequency information.
[0126] F f =FFT(F2)
[0127] Among them, F2 is the spatial domain feature, FFT is Fourier transform, F f is the feature after Fourier transform.
[0128] Step 2.3.2: Extract local texture details through wavelet convolution operation.
[0129] F w =WaveConv(F2)
[0130] Among them, WaveConv is wavelet convolution, F w is the feature extracted after wavelet convolution.
[0131] Step 2.3.3: Concatenate the Fourier transformed features with the features extracted by wavelet convolution and connect them with the residual of the original input features to generate the final fused features.
[0132] F fusion =Conv(1×1)(Concat(F f ,F w ))+F2
[0133] Among them, Conv(1×1) is 1×1 convolution, F f is the feature after Fourier transform, F w is the feature extracted after wavelet convolution, F2 is the spatial feature, and F fusion is the fused feature, and Concat is the concatenation operation.
[0134] Step 2.4: Image reconstruction module.
[0135] Image reconstruction is based on the upsampling operation. The upsampling module adopts an adaptive upsampling structure (Dysample). Compared with traditional pixel rearrangement or interpolation methods, this structure has stronger context modeling capabilities and edge preservation capabilities, and can achieve content-aware high-quality detail restoration in image reconstruction. This module receives the fused feature map from step 2.2 (feature extraction based on the spatial domain) and step 2.3 (feature extraction based on the frequency domain) as input, and adds it to the shallow feature map obtained in step 2.1 through a residual connection, thereby retaining both shallow detail information and deep semantic features. The fusion result is then input into the image reconstruction module for dynamic upsampling, and finally generates a high-resolution image output. This process realizes a closed loop from low-resolution input to high-quality image reconstruction.
[0136] The deep feature extraction module consists of multiple residual pyramids, each of which contains an efficient attention module, a residual pyramid module and a wavelet Fourier fusion module;
[0137] The shallow features are serially passed through the efficient attention module and the residual pyramid module to extract and output deep features;
[0138] The residual pyramid module uses multi-scale displacement operations and depth-wise separable convolution to extract multi-scale features;
[0139]
[0140] Among them, Shift i Represents spatial shift operations at different scales, H dilatedConv Denotes depth-separable convolution operation, Weight(Shift i ) are learnable weights that control the importance of features at each scale.
[0141] Step 3: Set the hyperparameters of the image super-resolution reconstruction based on residual pyramid and spectrum fusion. The specific parameters are as follows: use 64 feature channels; use Adam optimizer for training; and the initial learning rate is 2×10 -4 Charbonnier Loss was used to improve training stability. The number of iterations was 500,000. Note that this model did not use data augmentation (such as Mixup or RGB channel shuffling) or additional training strategies (such as pre-training or cosine learning rate scheduling). The model was obtained by training with parameter settings.
[0142] Step 4: Input the low-resolution image into the image super-resolution reconstruction model based on residual pyramid and spectrum fusion to obtain a high-resolution image. First, the low-resolution image to be processed is normalized and resized to ensure consistency with the model input format; then the pre-processed image is input into the trained super-resolution model, and sequentially passes through modules such as shallow feature extraction, deep feature extraction in the spatial and frequency domains, feature fusion and residual connection, and dynamic upsampling reconstruction; finally, the model outputs the reconstructed high-resolution image, and can perform post-processing operations such as denormalization and color space restoration according to actual application needs. This step realizes super-resolution inference reconstruction of any low-resolution image, with the advantages of fast inference speed, high image quality, and clear detail restoration, and is suitable for actual image enhancement scenarios.
[0143] This paper proposes two models: lightweight model: the backbone network contains 5 residual pyramids and spectrum fusion blocks; basic model: the backbone network contains 6 residual pyramids and spectrum fusion blocks.
[0144] To verify the effectiveness of the invented architecture, comparisons are made with multiple advanced lightweight SR methods (with scaling factors of 3× and 4×), including: CNN-based methods: EDSR, SRMDNF, CARN, AWSRN-M, MADNet, SMSR; Transformer-based methods: SwinIR, ESRT, ELAN-light, DiVANet, SwinIR-NG, SCNet, IPG-Tiny.
[0145] Some of the implementation results are shown in Table 1 and Table 2, where ours represents the two models designed by the present invention.
[0146]
[0147]
[0148] Table 2: Accuracy comparison of different methods on multiple datasets at X4 magnification
[0149]
[0150] See also Figure 3 : The residual pyramid module combines features of different dimensions and enhances the feature extraction capability of the model in different dimensions.
[0151] See also Figure 4 :The spectrum fusion module of wavelet Fourier transform combines the frequency domain characteristics of the spectrum fusion module with the spatial domain characteristics of the residual pyramid module. The model performance is significantly improved compared with other models, and the model is lightweight.
[0152] See also Figure 5 : On the manga109 dataset, the PSNR and parameter comparison chart with other models shows that the model of the present invention has significant improvements in both lightweight and performance.
[0153] See also Figure 6 and Figure 7 : In the visualization results of the two data sets, the model of the present invention has significant improvement in the restoration of details and textures compared with other models.
[0154] An image super-resolution reconstruction device based on residual pyramid and spectrum fusion, comprising
[0155] A memory for storing a computer program for implementing an image super-resolution reconstruction method based on residual pyramid and spectrum fusion;
[0156] The processor is configured to implement an image super-resolution reconstruction method based on residual pyramid and spectrum fusion when executing the computer program.
[0157] A computer-readable storage medium stores a computer program, which, when executed by a processor, can implement the image super-resolution reconstruction method based on residual pyramid and spectrum fusion according to the present invention.
Claims
1. An image super-resolution reconstruction method based on residual pyramid and spectrum fusion, characterized in that: The following steps are included: Step 1: Cropping the images in the dataset to obtain an image block training dataset; the data in the dataset is an image pair consisting of a low-resolution image and its corresponding high-resolution image; the low-resolution image is cropped and its corresponding high-resolution image constitutes the data in the image block training dataset; Step 2: Based on the image block training set, an image super-resolution reconstruction model based on residual pyramid and spectrum fusion is constructed to synthesize a high-resolution image; Step 3: After setting the hyperparameters, loss function, optimizer, and number of iterations of the image super-resolution reconstruction model based on residual pyramid and spectrum fusion, the model is trained to obtain a trained lightweight image super-resolution model; Step 4: Input the low-resolution image to be reconstructed into the trained lightweight image super-resolution model to obtain a high-resolution image.
2. The image super-resolution reconstruction method based on residual pyramid and spectrum fusion according to claim 1, characterized in that: In step 1, a 64×64 image block is randomly cropped from the low-resolution image as input, and the corresponding high-resolution image with the same magnification is used to calculate the reconstruction error, guiding the network to continuously adjust parameters to restore richer details and more accurate texture information.
3. The image super-resolution reconstruction method based on residual pyramid and spectrum fusion according to claim 1, characterized in that: In step 2, the image super-resolution reconstruction model based on residual pyramid and spectrum fusion includes four stages: shallow feature extraction, deep feature extraction based on spatial domain, deep feature extraction based on frequency domain, and image reconstruction; In the shallow feature extraction stage, the input low-resolution image is subjected to preliminary convolution processing to obtain a shallow feature map. The obtained shallow feature map is first subjected to deep feature extraction based on the spatial domain to capture spatial detail information, and then subjected to deep feature extraction based on the frequency domain to extract high-frequency details, and finally subjected to an upsampling-based image reconstruction module to synthesize a high-resolution image.
4. The image super-resolution reconstruction method based on residual pyramid and spectrum fusion according to claim 3, characterized in that: The specific process is: Step 2.1: The shallow feature extraction module uses a 3×3 convolutional layer to extract shallow features from the low-resolution image; F0=H conv3×3 (I LR ) Among them, I LR Represents the input low-resolution image, H conv3×3 Represents a shallow convolution extraction network constructed using a 3×3 convolution kernel, which is used to extract the basic feature F0, that is, the extracted shallow feature; Step 2.2: Input the basic feature F0 to the deep feature extraction module based on spatial domain for deep feature extraction. The module consists of two parts: lightweight Transformer submodule and residual pyramid submodule; Step 2.3: Input the extracted spatial deep features into the frequency domain-based feature extraction module for spectrum fusion. The frequency domain-based feature extraction module is divided into two parts: Fourier domain and wavelet domain to extract frequency domain deep features; Step 2.4: The image reconstruction module uses a dynamic upsampling module to reconstruct the image using deep features in the frequency domain.
5. The image super-resolution reconstruction method based on residual pyramid and spectrum fusion according to claim 4, characterized in that: The step 2.2 is specifically as follows: Step 2.2.1: The basic feature F0 is first compressed and dimensionally reduced through the lightweight Transformer submodule. The input basic feature F0 is compressed through linear transformation to generate the intermediate feature F r ; F r =W r ·reshape(F0) Among them, W r It is a learnable weight matrix with a dimension of C / 2, which is used to compress the feature dimension; Step 2.2.2: By inputting the intermediate feature F r , calculate the query (Q), key (K) and value (V) and window attention calculation, the formula is as follows: Q,K,V=reshape(Linear(F r ))→split(3) Among them, d k is the dimension of the key, F r It is the intermediate feature passed through the lightweight Transformer submodule. Linear() is a fully connected transformation. Q, K, and V represent query (Q), key (K), and value (V) vectors respectively. They are obtained by reshaping the input features after linear transformation and dividing them along the channel dimension. Split (3) means dividing the linearly mapped vector into three parts along the channel direction and assigning them to Q, K, V, and Q respectively. i ,K i ,V i Represents the query, key and value vectors at the i-th position respectively. Softmax() normalizes the attention scores of all positions to make them a probability distribution. i Represents the attention output of the i-th position; Step 2.2.3: After concatenating the attention results of all windows, restore the original channel dimension to obtain the enhanced feature F a ; Step 2.2.4: Enhance the feature F a The input is sent to the function-enhanced multi-layer perception unit, where channel expansion is first performed and Ghost features are generated. The formula is as follows: F mlp =Conv2(DWConv1(A·B)) Among them, A·B are two features obtained by convolution operation and channel segmentation, and DWConv1 is a 7×7 depth-wise separable convolution; Step 2.2.5: Combine the residual structure with the learnable scaling factor α to output the final feature map F1; F1=F mlp ·α+F1 After the lightweight Transformer submodule extracts features, F1 will be input into the second module of the deep feature extraction spatial domain part, that is, through multi-scale spatial shift operation and depth-wise separable dilated convolution operation; Step 2.2.6: First, extract information of different scales by shifting the feature map at different scales. The formula is as follows: Shift i =Shift i (H DilatedConv (Shift i (F1))) Among them, Shift i H represents the spatial shift operation of the i-th scale (e.g., Shift4, Shift8, Shift16). DilatedConv (·) represents the depthwise separable dilated convolution operation, which expands the receptive field and extracts features at different scales; Step 2.2.7: Perform weighted fusion of feature maps of different scales using learnable weights to obtain the fused feature map x fused , the formula is as follows: Among them, Weight(Shift i ) is a learnable weight that controls the importance of each scale feature in the fusion; Step 2.2.8: Fused feature map x fused It is further refined through 1×1 convolution operation and the feature expression ability is enhanced through ReLU activation function. The formula is as follows: x out =ReLU(H conv1×1 ( x fused )) Among them, H conv1×1 is a 1×1 convolution operation, and ReLU(·) is the ReLU activation function; Step 2.2.9: Finally, output feature x out Feature map x after weighted fusion with multi-scale features fused The residuals are added and enhanced by the gated spatial attention module; F2=GSA(x fused +x out ) After a complete deep feature extraction module based on spatial domain, the spatial domain feature F2 is generated.
6. The image super-resolution reconstruction method based on residual pyramid and spectrum fusion according to claim 4, characterized in that: The step 2.3 is specifically as follows: Step 2.3.1: Convert the spatial domain feature F2 to the frequency domain through two-dimensional fast Fourier transform to extract the frequency information; F f =FFT(F2) Among them, F2 is the spatial domain feature, FFT is Fourier transform, F f is the feature after Fourier transform; Step 2.3.2: Extract local texture details through wavelet convolution operation; F w =WaveConv(F2) Among them, WaveConv is wavelet convolution, F w is the feature extracted after wavelet convolution; Step 2.3.3: Concatenate the Fourier transformed features with the features extracted by wavelet convolution and connect them with the residual of the original input features to generate the final fused features; F fusion =Conv(1×1)(Concat(F f ,F w ))+F2 Among them, Conv(1×1) is 1×1 convolution, F f is the feature after Fourier transform, F w is the feature extracted after wavelet convolution, F2 is the spatial feature, and F fusion is the fused feature, and Concat is the concatenation operation.
7. The image super-resolution reconstruction method based on residual pyramid and spectrum fusion according to claim 4, characterized in that: The step 2.4 is specifically as follows: The image reconstruction module uses a dynamic upsampling module to perform image reconstruction, and the deep features are fused in the dynamic upsampling module to generate a final high-resolution image; I HR =Dysample(F WFFB ) Among them, F WFFB It is the feature map after deep feature extraction, I HR The reconstructed high-resolution image.
8. An image super-resolution reconstruction device based on residual pyramid and spectrum fusion, characterized in that: include: A memory for storing a computer program for implementing the image super-resolution reconstruction method based on residual pyramid and spectrum fusion according to any one of claims 1 to 7; A processor, configured to implement the image super-resolution reconstruction method based on residual pyramid and spectrum fusion according to any one of claims 1 to 7 when executing the computer program.
9. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the image super-resolution reconstruction method based on residual pyramid and spectrum fusion according to any one of claims 1 to 7 can be implemented.
Citation Information
Patent Citations
Lightweight image super-resolution method and device based on efficient frequency domain Transformer
CN119180752B
Cited By
Business expansion work order handwritten date identification method and system
CN121033870A
Emotion and fatigue integrated monitoring system and method based on electroencephalogram signals
CN121154180A
Wafer image super-resolution method and related equipment
CN121481846A
Super-resolution remote sensing image reconstruction method, system and equipment based on frequency domain enhancement
CN121481853A