Lightweight Image Super-Resolution Reconstruction Method Based on Frequency Domain-Spatial Domain Assisted Mamba

By introducing the frequency domain-spatial domain Mamba layer and the transpose-based and self-attention feature interaction layer into the image super-resolution model, the problems of large amount of parameters and high computational complexity of the existing model are solved, and an efficient image super-resolution effect is achieved.

CN119251051BActive Publication Date: 2025-06-27SOUTHWEST UNIVERSITY FOR NATIONALITIES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411309230.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2025-06-27
Estimated Expiration
2044-09-19

AI Technical Summary

Technical Problem

The existing image super-resolution model has a large amount of parameters and high computational complexity, making it difficult to efficiently capture global context information and local features.

Method used

A lightweight image super-resolution reconstruction method based on frequency domain-space domain assisted Mamba is adopted to construct a lightweight image super-resolution reconstruction model by introducing the frequency domain-space domain Mamba layer and a transpose and self-attention feature interaction layer. This model enhances feature interaction through frequency domain and spatial domain to achieve efficient reconstruction of images.

Benefits of technology

The calculation complexity and inference time of the model are reduced, the efficiency and quality of image reconstruction are improved, and the lightweight image super-resolution effect is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119251051B_ABST
    Figure CN119251051B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing technology, and discloses a lightweight image super-resolution reconstruction method based on frequency domain-spatial domain assisted Mamba, including the steps of: obtaining an original image dataset and performing data preprocessing on it to obtain a training dataset and a test dataset; introducing a frequency domain-spatial domain Mamba layer and a transpose and self-attention feature interaction layer to construct a lightweight image super-resolution reconstruction model based on frequency domain-spatial domain assisted Mamba; inputting the training dataset into the lightweight image super-resolution reconstruction model based on frequency domain-spatial domain assisted Mamba for training to obtain a trained lightweight image super-resolution reconstruction model with weights, which is used to reconstruct a low-resolution image into a high-resolution image; inputting the test dataset into the trained lightweight image super-resolution reconstruction model with weights to obtain a reconstructed high-resolution visualization image; this method realizes lightweight image super-resolution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically relates to a lightweight image super-resolution reconstruction method based on frequency domain-spatial domain assisted Mamba. Background Art

[0002] In the context of the digital age, image quality has become one of the key indicators for measuring technological progress and user experience. With the continuous iteration of technology and the increasing improvement of user requirements, high-resolution images have become the standard configuration and common expectation in many fields. With their excellent resolution, rich detail performance, and better visual experience, they bring users a more vivid and realistic visual enjoyment. However, there are still a large number of low-resolution images in real life, which makes image super-resolution technology crucial for improving image resolution.

[0003] Traditionally, methods for improving image super-resolution mainly focused on physical-level improvements, such as enhancing the quality and pixel density of camera sensors. Although this method can improve the resolution and quality of images, it also has significant drawbacks and limitations. Specifically, physical-level improvements often require expensive hardware upgrade costs and are not applicable to all scenarios. Therefore, due to technical and cost factors, a comprehensive physical upgrade of the device may be unrealistic.

[0004] In recent years, with the rapid development of deep learning technology, a series of deep learning-based image super-resolution methods have emerged, including methods based on convolutional neural networks (CNNs) and methods based on Transformers. However, these methods all face challenges of different problems. First, in existing CNN-based models, as the model depth increases, the computational complexity increases significantly. At the same time, the fixed size of the convolutional kernel limits the receptive field of the model, making it difficult to efficiently capture global context information. Second, compared with CNN-based models, traditional Transformer-based methods can capture context information better, but their computational complexity is proportional to the square of the input token length, resulting in high computational overhead and inference time. Recently, although the Mamba-based image super-resolution model has achieved linear computational complexity, it lacks consideration in capturing local features. Summary of the Invention

[0005] In view of the above deficiencies in the prior art, the present invention provides a lightweight image super-resolution reconstruction method based on frequency domain-spatial domain assisted Mamba to solve the technical problems of large parameter quantities and high computational complexity in current image super-resolution models.

[0006] To achieve the above invention objective, the technical solution adopted by the present invention is as follows:

[0007] A lightweight image super-resolution reconstruction method based on frequency-domain and spatial-domain assisted Mamba, comprising the following steps:

[0008] S1. Obtain the original image dataset and perform data preprocessing on it to obtain the training dataset and the test dataset;

[0009] S2. Introduce the frequency-domain and spatial-domain Mamba layer and the transpose and self-attention feature interaction layer to construct a lightweight image super-resolution reconstruction model based on frequency-domain and spatial-domain assisted Mamba;

[0010] S3. Input the training dataset into the lightweight image super-resolution reconstruction model based on frequency-domain and spatial-domain assisted Mamba for training to obtain a trained lightweight image super-resolution reconstruction model with weights, which is used to reconstruct low-resolution images into high-resolution images;

[0011] S4. Input the test dataset into the trained lightweight image super-resolution reconstruction model with weights to obtain the reconstructed high-resolution visualization image.

[0012] Further, step S1 specifically includes:

[0013] S11. Obtain the publicly available DIV2K dataset, use the DIV2K dataset as the HR image dataset, and perform bicubic interpolation on each image in the HR image dataset to downsample it by 2, 3, and 4 times to obtain images with multiple downsampling multiples;

[0014] S12. Perform random flipping and rotation operations on the images with multiple downsampling multiples to obtain enhanced low-resolution LR images;

[0015] S13. Use the HR image dataset as the first label data, use the enhanced low-resolution LR images as the first original data, and use the first label data and the first original data as the training dataset;

[0016] S14. Obtain the publicly available Set5 dataset, Set14 dataset, B100 dataset, Urban100 dataset, and Manga109 dataset and use them as the HR images of the test set. Respectively perform bicubic interpolation on each dataset to downsample it by 2, 3, and 4 times to obtain the LR images of the test set;

[0017] S15. Use the HR images of the test set as the second label data, use the LR images of the test set as the second original data, and use the second label data and the second original data as the test dataset.

[0018] Further, the lightweight image super-resolution reconstruction model based on frequency domain-spatial domain assisted Mamba in step S2 includes a shallow feature extraction module, a deep feature extraction module, and a high-resolution image reconstruction module;

[0019] The deep feature extraction module includes M feature mixing group modules and a second 3×3 convolutional layer;

[0020] The feature mixing group module includes N multi-level feature interaction layers and a first 3×3 convolutional layer;

[0021] The multi-level feature interaction layer includes a first LayerNorm layer, a frequency domain-spatial domain Mamba layer, a second LayerNorm layer, and a transpose and self-attention feature interaction layer;

[0022] The frequency domain-spatial domain Mamba layer includes a visual Mamba layer, a frequency domain-spatial domain enhancement layer, and a first 1×1 convolutional layer;

[0023] The visual Mamba layer includes an upper branch module, a lower branch module, and a third Linear layer;

[0024] The upper branch module includes a first Linear layer, a first 3×3 depth convolutional layer, a first SiLU layer, a 2D selective scanning module, and a third LayerNorm layer;

[0025] The lower branch module includes a second Linear layer and a second SiLU layer;

[0026] The frequency domain-spatial domain enhancement layer includes a frequency domain enhancement module and a spatial domain enhancement module;

[0027] The frequency domain enhancement module includes a second 3×3 depth convolutional layer, a Fourier transform layer, a second 1×1 convolutional layer, a first ReLU layer, and an inverse Fourier transform layer;

[0028] The spatial domain enhancement module includes a third 3×3 depth convolutional layer and a third 1×1 convolutional layer;

[0029] The transpose and self-attention feature interaction layer includes a third 3×3 convolutional layer, a spatial channel permutation layer, a spatial self-attention layer, a channel self-attention layer, and a fifth 1×1 convolutional layer;

[0030] The spatial channel permutation layer includes a first permutation layer, a fifth 3×3 depth convolutional layer, a second permutation layer, and a sixth 3×3 depth convolutional layer;

[0031] The channel self-attention layer includes an adaptive average pooling layer, a fourth 3×3 depth convolutional layer, and a second sigmoid layer;

[0032] The spatial self-attention layer includes a second ReLU layer, a fourth 1×1 convolutional layer, and a first sigmoid layer.

[0033] Further, step S3 specifically includes:

[0034] S31. Input the training data set into the shallow feature extraction module of the lightweight image super-resolution reconstruction model based on frequency-domain and spatial-domain assisted Mamba, and perform convolution processing using a 3×3 convolutional layer to obtain shallow features F SH ;

[0035] S32. Input the shallow features F SH into the deep feature extraction module for feature mixing extraction to obtain deep features

[0036] S33. Add the deep features to the shallow features F SH to implement residual connection, and then input them into the high-resolution image reconstruction module to use convolutional layers and PixelShuffle upsampling to implement image reconstruction, obtaining reconstructed features I SR , and use the L1 loss function to optimize the lightweight image super-resolution reconstruction model based on frequency-domain and spatial-domain assisted Mamba during training to obtain a trained lightweight image super-resolution reconstruction model with weights based on frequency-domain and spatial-domain assisted Mamba.

[0037] Further, step S32 specifically includes:

[0038] S321. Input the shallow features F SH into the first LayerNorm layer for normalization operation to obtain the first normalized features X, and average-divide the first normalized features X from the channel dimension to obtain feature X1 and feature X2, that is:

[0039] X1, X2 = Split(X, dim = C)

[0040] where Split represents the feature splitting function, dim represents the splitting dimension, and C represents the number of channels;

[0041] S322. Input feature X1 and feature X2 into the visual Mamba layer and the frequency-domain and spatial-domain enhancement layer of the frequency-domain and spatial-domain Mamba layer respectively for visual Mamba and frequency-domain and spatial-domain enhancement processing, then splice them in the channel dimension, and input them into the first 1×1 convolutional layer for preliminary feature interaction to obtain preliminary global-local features F g_l , specifically including:

[0042] S3221. Input feature X1 into the visual Mamba layer for visual Mamba processing to capture context information and obtain global features F m ;

[0043] S3222. Input the feature X2 into the frequency-domain and spatial-domain enhancement layer for frequency-domain and spatial-domain enhancement processing to enhance the context information and local information, and obtain the frequency-domain and spatial-domain enhanced feature F f_s ;

[0044] S3223. Concatenate the global feature F m and the frequency-domain and spatial-domain enhanced feature F f_s in the channel dimension, and input them into the first 1×1 convolutional layer for preliminary feature interaction to obtain the preliminary global-local feature F g_l ;

[0045] S323. Add the preliminary global-local feature F g_l and the shallow feature F SH to achieve local residual connection and obtain the final global-local feature F G_L ;

[0046] S324. Input the final global-local feature F G_L into the second LayerNorm layer for normalization operation to obtain the second normalized feature

[0047] S325. After inputting the second normalized feature into the transposed and self-attention based feature interaction layer for feature interaction and fusion, add it to the final global-local feature F G_L to achieve local residual connection and obtain the interaction feature

[0048] S326. After repeating steps S321 - S325 N times for the interaction feature , input it into the first 3×3 convolutional layer to obtain the mixed feature

[0049] S327. After repeating step S326 M times for the mixed feature , input it into the second 3×3 convolutional layer to obtain the deep feature

[0050] Furthermore, step S3221 specifically includes:

[0051] S32211. Input the feature X1 into the first Linear layer for linear transformation operation to obtain the upper-branch transformation feature F u_l , and input it into the first 3×3 depth convolutional layer for encoding operation to obtain the upper-branch encoding feature F u_c , and input it into the first SiLU layer for activation operation to obtain the upper-branch activation feature F u_s, and input it into the 2D selective scanning module for global feature extraction, and then input it into the third LayerNorm layer for normalization operation to obtain the first feature F u ;

[0052] S32212. Input the feature X1 into the second Linear layer for linear transformation operation to obtain the lower branch transformed feature F d_l , and input it into the second SiLU layer for activation operation to obtain the second feature F d ;

[0053] S32213. Multiply the first feature F u with the second feature F d , and input it into the third Linear layer for linear activation to obtain the global feature F m .

[0054] Furthermore, step S3222 specifically includes:

[0055] S32221. Input the feature X2 into the second 3×3 depth convolution layer for depth convolution operation to extract local features, obtain the first local feature of the feature X2, and input it into the Fourier transform layer for Fourier transform, which is used to transform the first local feature of the feature X2 from the spatial domain to the frequency domain, and then input it into the second 1×1 convolution layer for convolution operation and then into the first ReLU layer for activation operation to obtain the activated frequency domain feature, and then input it into the inverse Fourier transform layer for inverse transformation, which is used to transform the activated frequency domain feature from the frequency domain to the spatial domain to obtain the frequency domain enhanced feature F f ;

[0056] S32222. Input the feature X2 into the third 3×3 depth convolution layer for depth convolution operation to extract local features, obtain the second local feature of the feature X2, and input it into the third 1×1 convolution layer for convolution operation to obtain the spatial domain enhanced feature F s ;

[0057] S32223. Add the frequency domain enhanced feature F f to the spatial domain enhanced feature F s to obtain the frequency domain - spatial domain enhanced feature F f_s .

[0058] Furthermore, step S325 specifically includes:

[0059] S3251. Input the second normalized feature into the third 3×3 convolution layer for preliminary feature fusion and channel compression, and compress it to C / 4 to obtain the initialized feature F i ;

[0060] S3252. Input the initialized feature F iThe first permutation layer of the input space channel permutation layer performs a permutation operation by permuting the spatial features of the initialized feature F i into channel dimension features to obtain the first permutation feature F i_p1 , and inputs it into the fifth 3×3 depth convolution layer for fineness feature extraction, and inputs it into the second permutation layer for permutation operation to convert the fineness feature from the channel dimension to the spatial feature, obtaining the second permutation feature F i_p2 , and inputs it into the sixth 3×3 depth convolution layer for spatial feature encoding to obtain the permutation fusion feature F i_p ;

[0061] S3253. Input the initialized feature F i into the second ReLU layer of the spatial self-attention layer for non-linear activation, input it into the fourth 1×1 convolution layer to reduce the channel dimension to 1, and input it into the first sigmoid layer for activation operation to obtain the first activation feature F i_s1 . Multiply the first activation feature F i_s1 with the initialized feature F i to obtain the spatial self-attention fusion feature F i_s ;

[0062] S3254. Input the initialized feature F i into the adaptive average pooling layer of the channel self-attention layer for pooling operation to obtain the initialized feature with a spatial size of 1×1, input it into the fourth 3×3 depth convolution layer for encoding, and input it into the second sigmoid layer for non-linear activation operation to obtain the second activation feature F i_s2 . Multiply the second activation feature F i_s2 with the initialized feature F i to obtain the channel self-attention fusion feature F i_c ;

[0063] S3255. Add the permutation fusion feature F i_p , the spatial self-attention fusion feature F i_s and the channel self-attention fusion feature F i_c , input it into the fifth 1×1 convolution layer to expand the number of channels to C, and then add it to the final global-local feature F G_L to obtain the interaction feature

[0064] Furthermore, the calculation formula for optimizing the lightweight image super-resolution reconstruction model based on frequency domain-spatial domain assisted Mamba during training using the L1 loss function in step S33 is:

[0065] L1 = ‖I HR - I SR ‖

[0066] where IHR Represents the high-resolution image features corresponding to the HR image dataset.

[0067] Furthermore, step S4 specifically includes:

[0068] According to the trained lightweight image super-resolution reconstruction model based on frequency-domain-spatial-domain assisted Mamba with weights, the test dataset is respectively input into the trained lightweight image super-resolution reconstruction model based on frequency-domain-spatial-domain assisted Mamba with ×2, ×3, and ×4 weights for testing, to obtain quantitative results and reconstructed high-resolution visualization images;

[0069] Among them, the quantitative results include peak signal-to-noise ratio and structural similarity;

[0070] The calculation formula for the peak signal-to-noise ratio is:

[0071]

[0072] Among them, MSE represents the mean square error between image Y and image Z in the test dataset, a represents the row index of the image position, b represents the column index of the image position, H represents the height of the image, W represents the width of the image, PSNR represents the peak signal-to-noise ratio, n represents the number of bits per pixel in the image, log 10 Represents the logarithmic function;

[0073] The calculation formula for the structural similarity is:

[0074]

[0075] Among them, SSIM(Y,Z) represents the structural similarity value between image Y and image Z, μ Y Represents the average value of image Y, μ Z Represents the average value of image Z, σ YZ Represents the covariance between image Y and image Z, σ Y Represents the variance of image Y, σ Z Represents the variance of image Z, c1 and c2 respectively represent the first constant and the second constant used to prevent the denominator from being 0, k1 and k2 both represent constants, and L represents the dynamic range of pixel values in the image.

[0076] The present invention has the following beneficial effects:

[0077] The lightweight image super-resolution reconstruction method based on frequency-domain and spatial-domain assisted Mamba proposed by the present invention introduces frequency domain and spatial domain enhancements on the basis of Mamba to supplement local information and enhance context information. At the same time, in order to fully interact the above three different levels of information, a feature interaction layer based on transpose and self-attention is proposed to achieve feature interaction and fusion. And the mitigation of feature redundancy promotes the efficiency of image reconstruction, speeds up model inference and reduces computational complexity, and finally realizes lightweight image super-resolution. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 It is a schematic flow chart of the lightweight image super-resolution reconstruction method based on frequency-domain and spatial-domain assisted Mamba proposed by the present invention;

[0079] Figure 2 It is a schematic structural diagram of a lightweight image super-resolution reconstruction model based on frequency-domain and spatial-domain assisted Mamba;

[0080] Figure 3 It is a schematic structural diagram of the Visual Mamba layer;

[0081] Figure 4 It is a schematic structural diagram of the frequency-domain and spatial-domain enhancement layer;

[0082] Figure 5 It is a schematic structural diagram of the feature interaction layer based on transpose and self-attention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0083] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0084] As Figure 1 shown, the lightweight image super-resolution reconstruction method based on frequency-domain and spatial-domain assisted Mamba includes the following steps S1-S4:

[0085] S1. Obtain the original image dataset and perform data preprocessing on it to obtain the training dataset and the test dataset.

[0086] In this embodiment, the purpose of performing data preprocessing on the original image dataset is to obtain high-resolution images and low-resolution images, and use the high-resolution images as the labels of the low-resolution images for the training of the lightweight image super-resolution reconstruction model based on frequency-domain and spatial-domain assisted Mamba in the subsequent steps.

[0087] Specifically, step S1 specifically includes S11 - S15:

[0088] S11. Obtain the publicly available DIV2K dataset, use the DIV2K dataset as the HR image dataset, and perform downsampling on each image in the HR image dataset by a factor of 2, 3, and 4 using bicubic interpolation to obtain images with multiple downsampling factors.

[0089] S12. Perform random flipping and rotation operations on the images with multiple downsampling factors to obtain enhanced low - resolution LR images.

[0090] S13. Use the HR image dataset as the first label data, the enhanced low - resolution LR images as the first original data, and the first label data and the first original data as the training dataset.

[0091] S14. Obtain the publicly available Set5 dataset, Set14 dataset, B100 dataset, Urban100 dataset, and Manga109 dataset and use them as the HR images of the test set. Respectively perform downsampling on each dataset by a factor of 2, 3, and 4 using bicubic interpolation to obtain the LR images of the test set.

[0092] S15. Use the HR images of the test set as the second label data, the LR images of the test set as the second original data, and the second label data and the second original data as the test dataset.

[0093] S2. Introduce the frequency - spatial domain Mamba layer and the transpose - and self - attention feature interaction layer to construct a lightweight image super - resolution reconstruction model based on the frequency - spatial domain assisted Mamba.

[0094] In this embodiment, the frequency - spatial domain Mamba layer and the transpose - and self - attention feature interaction layer are introduced to construct a lightweight image super - resolution reconstruction model based on the frequency - spatial domain assisted Mamba. Its connection relationship and structure are specifically as Figure 2 shown. The lightweight image super - resolution reconstruction model based on the frequency - spatial domain assisted Mamba includes a shallow feature extraction module, a deep feature extraction module, and a high - resolution image reconstruction module. The shallow feature extraction module is used to extract shallow features, the deep feature extraction module is used to extract deep features, and the high - resolution image reconstruction module is used for image reconstruction.

[0095] Specifically, the deep feature extraction module includes M feature mixing group modules and a second 3×3 convolutional layer; the feature mixing group module includes N multi - level feature interaction layers and a first 3×3 convolutional layer; the multi - level feature interaction layer includes a first LayerNorm layer, a frequency - spatial domain Mamba layer, a second LayerNorm layer, and a transpose - and self - attention feature interaction layer.

[0096] Among them, the frequency-domain spatial-domain Mamba layer includes a visual Mamba layer, a frequency-domain spatial-domain enhancement layer, and a first 1×1 convolutional layer.

[0097] Among them, the structure and connection relationship of the visual Mamba layer are as Figure 3 shown, including an upper branch module, a lower branch module, and a third Linear layer; the upper branch module includes a first Linear layer, a first 3×3 depth convolutional layer, a first SiLU layer, a 2D selective scanning module, and a third LayerNorm layer; the lower branch module includes a second Linear layer and a second SiLU layer.

[0098] Among them, the structure and connection relationship of the frequency-domain spatial-domain enhancement layer are as Figure 4 shown, including a frequency-domain enhancement module and a spatial-domain enhancement module; the frequency-domain enhancement module includes a second 3×3 depth convolutional layer, a Fourier transform layer, a second 1×1 convolutional layer, a first ReLU layer, and an inverse Fourier transform layer; the spatial-domain enhancement module includes a third 3×3 depth convolutional layer and a third 1×1 convolutional layer.

[0099] Among them, the structure and connection relationship of the transpose and self-attention feature interaction layer are as Figure 5 shown, including a third 3×3 convolutional layer, a spatial channel permutation layer, a spatial self-attention layer, a channel self-attention layer, and a fifth 1×1 convolutional layer; the spatial channel permutation layer includes a first permutation layer, a fifth 3×3 depth convolutional layer, a second permutation layer, and a sixth 3×3 depth convolutional layer; the channel self-attention layer includes an adaptive average pooling layer, a fourth 3×3 depth convolutional layer, and a second sigmoid layer; the spatial self-attention layer includes a second ReLU layer, a fourth 1×1 convolutional layer, and a first sigmoid layer.

[0100] In addition, it can also be seen from Figure 2 that there is a residual connection structure in the first summation symbol ⊕ between the shallow feature extraction module and the deep feature extraction module, and there is also a residual connection structure between the first summation symbol ⊕ and the second summation symbol ⊕ of the deep feature extraction module, and there is a residual structure connected between the third summation symbol ⊕ after the deep feature extraction module and the shallow feature extraction module to maintain stable training.

[0101] S3. Input the training data set into the lightweight image super-resolution reconstruction model based on the frequency-domain spatial-domain assisted Mamba for training to obtain a trained lightweight image super-resolution reconstruction model with weights based on the frequency-domain spatial-domain assisted Mamba, which is used to reconstruct low-resolution images into high-resolution images.

[0102] In this implementation, when training the lightweight image super-resolution reconstruction model based on frequency-domain and spatial-domain assisted Mamba, the feature mixing group is set to 4, the multi-level feature interaction layer is set to [4, 6, 4, 4], the number of channels is set to 72, and the L1 loss function is used to optimize the model to train the lightweight image super-resolution reconstruction model, so as to obtain a trained lightweight image super-resolution reconstruction model with weights for reconstructing low-resolution images into high-resolution images. In addition, during the training process, the number of iterations is 500K, the initial learning rate is set to 2e -4 , and when the number of iterations is 250K, 350K, 400K, and 450K respectively, the learning rate is continuously halved to maintain the stability of training.

[0103] Specifically, step S3 specifically includes S31 - S33:

[0104] S31. Input the training dataset into the shallow feature extraction module of the lightweight image super-resolution reconstruction model based on frequency-domain and spatial-domain assisted Mamba, and perform convolution processing using a 3×3 convolutional layer to obtain shallow features F SH .

[0105] In this embodiment, the training dataset is used as the input feature I LR ∈R H×W×3 and input into the shallow feature extraction module of the lightweight image super-resolution reconstruction model based on frequency-domain and spatial-domain assisted Mamba for shallow feature extraction. Among them, F SH ∈R H ×W×C , where R represents a matrix, H is the height, W is the width, and C represents the number of channels.

[0106] S32. Input the shallow features F SH into the deep feature extraction module for feature mixing extraction to obtain deep features

[0107] Specifically, step S32 specifically includes S321 - S327:

[0108] S321. Input the shallow features F SH into the first LayerNorm layer for normalization operation to obtain the first normalized feature X, and average split the first normalized feature X from the channel dimension to obtain feature X1 and feature X2, that is:

[0109] X1, X2 = Split(X, dim = C)

[0110] where Split represents the feature splitting function, dim represents the splitting dimension, and C represents the number of channels.

[0111] In this embodiment, X ∈ R H×W×C , and after the first normalized feature X is evenly divided from the channel dimension to obtain features X1 and X2, X1, X2 ∈ R H×W×C′ , That is, C′ is half of the number of channels C.

[0112] S322. Input features X1 and X2 into the visual Mamba layer and the frequency-domain spatial-domain enhancement layer of the frequency-domain spatial-domain Mamba layer respectively for visual Mamba and frequency-domain spatial-domain enhancement processing, then splice them in the channel dimension, and input them into the first 1×1 convolutional layer for preliminary feature interaction to obtain the preliminary global-local feature F g_l , specifically including:

[0113] S3221. Input feature X1 into the visual Mamba layer for visual Mamba processing to capture context information and obtain the global feature F m .

[0114] In this embodiment, F m ∈ R H×W×C′ .

[0115] Specifically, step S3221 specifically includes S32211 - S32213:

[0116] S32211. Input feature X1 into the first Linear layer for linear transformation operation to obtain the upper-branch transformed feature F u_l , and input it into the first 3×3 depth convolutional layer for encoding operation to obtain the upper-branch encoded feature F u_c , and input it into the first SiLU layer for activation operation to obtain the upper-branch activated feature F u_s , and input it into the 2D selective scanning module for global feature extraction, and input it into the third LayerNorm layer for normalization operation to obtain the first feature F u .

[0117] In this embodiment, when feature X1 is input into the first Linear layer for linear transformation operation, the number of channels is halved, so

[0118] S32212. Input feature X1 into the second Linear layer for linear transformation operation to obtain the lower-branch transformed feature F d_l , and input it into the second SiLU layer for activation operation to obtain the second feature F d .

[0119] In this embodiment, when feature X1 is input into the second Linear layer for linear transformation operation, the number of channels is halved, so

[0120] S32213. Multiply the first feature F u by the second feature F d and input it into the third Linear layer for linear activation to obtain the global feature F m .

[0121] In this embodiment, according to steps S32211 - S32213, the expression of the global feature F m is as follows:

[0122]

[0123] S3222. Input the feature X2 into the frequency-domain to spatial-domain enhancement layer for frequency-domain to spatial-domain enhancement processing to enhance the context information and local information, and obtain the frequency-domain to spatial-domain enhanced feature F f_s .

[0124] In this embodiment, F f_s ∈R H×W×C′ .

[0125] Specifically, step S3222 specifically includes S32221 - S32223:

[0126] S32221. Input the feature X2 into the second 3×3 depth convolution layer for depth convolution operation and extract the local feature to obtain the first local feature of the feature X2, and input it into the Fourier transform layer for Fourier transform to convert the first local feature of the feature X2 from the spatial domain to the frequency domain, and then input it into the second 1×1 convolution layer for convolution operation and then input it into the first ReLU layer for activation operation to obtain the activated frequency-domain feature, and then input it into the inverse Fourier transform layer for inverse transformation to convert the activated frequency-domain feature from the frequency domain to the spatial domain to obtain the frequency-domain enhanced feature F f .

[0127] In this embodiment, F f ∈R H×W×C′ , and the expression of the frequency-domain enhanced feature F f is as follows:

[0128] F f =RFFT(ReLU1(Conv2 1×1 (FFT(DConv2 3×3 (X2)))))

[0129] where FFT represents the Fourier transform and RFFT represents the inverse Fourier transform.

[0130] S32222. Input the feature X2 into the third 3×3 depth convolution layer for depth convolution operation to extract local features, obtaining the second local feature of the feature X2, and then input it into the third 1×1 convolution layer for convolution operation to obtain the spatially enhanced feature F s .

[0131] In this embodiment, F s ∈R H×W×C′ , and the expression of the spatially enhanced feature F s is:

[0132] F s =Conv3 1×1 (DConv3 3×3 (X2)).

[0133] S32223. Add the frequency-domain enhanced feature F f to the spatially enhanced feature F s to obtain the frequency-domain and spatially enhanced feature F f_s .

[0134] In this embodiment, F f_s ∈R H×W×C′ .

[0135] S3223. Concatenate the global feature F m and the frequency-domain and spatially enhanced feature F f_s in the channel dimension, and then input it into the first 1×1 convolution layer for preliminary feature interaction to obtain the preliminary global-local feature F g_l .

[0136] In this embodiment, F g_l ∈R H×W×C , and the expression of the preliminary global-local feature F g_l is:

[0137] F g_l =Conv1 1×1 (Cat((F m ,F f_s ),dim = C))

[0138] where Cat represents the concatenation operation, and dim = C represents the concatenation dimension.

[0139] S323. Add the preliminary global-local feature F g_l to the shallow feature F SH to achieve local residual connection, obtaining the final global-local feature F G_L .

[0140] In this embodiment, F G_L ∈R H×W×C .

[0141] S324. Input the final global-local feature F G_L into the second LayerNorm layer for normalization operation to obtain the second normalized feature

[0142] In this embodiment,

[0143] S325. After inputting the second normalized feature into the transpose and self-attention feature interaction layer for feature interaction and fusion, add it to the final global-local feature F G_L to achieve local residual connection and obtain the interaction feature

[0144] In this embodiment,

[0145] Specifically, step S325 specifically includes S3251 - S3255:

[0146] S3251. Input the second normalized feature into the third 3×3 convolutional layer for preliminary feature fusion and channel compression, and compress it to C / 4 to obtain the initialization feature F i .

[0147] In this embodiment, and the expression of the initialization feature F i is:

[0148]

[0149] S3252. Input the initialization feature F i into the first permutation layer of the spatial channel permutation layer for permutation operation. By permuting the spatial feature of the initialization feature F i into the channel dimension feature, obtain the first permutation feature F i_p1 , and input it into the fifth 3×3 depth convolutional layer for fineness feature extraction, and then input it into the second permutation layer for permutation operation to convert the fineness feature from the channel dimension to the spatial feature, obtain the second permutation feature F i_p2 , and input it into the sixth 3×3 depth convolutional layer for spatial feature encoding to obtain the permutation fusion feature F i_p .

[0150] In this embodiment, the spatial channel permutation layer performs cross-permutation through space and channel; among them, and the expression of the permutation fusion feature F i_p is:

[0151] F i_p = DConv63×3 (δ2(DConv5 3×3 (δ1(F i ))))

[0152] Among them, δ1 represents the first permutation layer, and δ2 represents the second permutation layer.

[0153] S3253. Input the initialized feature F into the second ReLU layer of the spatial self-attention layer for non-linear activation, input it into the fourth 1×1 convolutional layer to reduce the channel dimension to 1, and input it into the first sigmoid layer for activation operation to obtain the first activation feature F i_s1 , multiply the first activation feature F i_s1 with the initialized feature F i to obtain the spatial self-attention fusion feature F i_s .

[0154] In this embodiment, the spatial self-attention layer fuses different hierarchical information into a single channel; among them, and the expression of the spatial self-attention fusion feature F i_s is: F i_s = Sigmoid1(Conv4 1×1 (ReLU2(F i )))*F i .

[0155] S3254. Input the initialized feature F i into the adaptive average pooling layer of the channel self-attention layer for pooling operation to obtain the initialized feature with a spatial size of 1×1, input it into the fourth 3×3 depth convolutional layer for encoding, and input it into the second sigmoid layer for non-linear activation operation to obtain the second activation feature F i_s2 , multiply the second activation feature F i_s2 with the initialized feature F i to obtain the channel self-attention fusion feature F i_c .

[0156] In this embodiment, the channel self-attention layer fuses different hierarchical information into a single space; among them, and the expression of the channel self-attention fusion feature F i_c is:

[0157] F i_c = Sigmoid2(DConv4 3×3 (AdaptiveAvgPool2d(F i )))*F i

[0158] Among them, AdaptiveAvgPool2d is adaptive average pooling.

[0159] S3255. Add the permutation fusion feature F i_p , the spatial self-attention fusion feature F i_s and the channel self-attention fusion feature F i_c . Then input them into the fifth 1×1 convolutional layer to expand the number of channels to C, and add them to the final global-local feature F G_L to obtain the interaction feature

[0160] In this embodiment, and the interaction feature has the following expression:

[0161] S326. After repeating steps S321 - S325 N times for the interaction feature , input it into the first 3×3 convolutional layer to obtain the mixed feature

[0162] In this embodiment,

[0163] S327. After repeating step S326 M times for the mixed feature , input it into the second 3×3 convolutional layer to obtain the deep feature

[0164] In this embodiment,

[0165] S33. Add the deep feature to the shallow feature F SH . After implementing residual connection, input it into the high-resolution image reconstruction module to perform image reconstruction using convolutional layers and PixelShuffle upsampling to obtain the reconstructed feature I SR . Then use the L1 loss function to optimize the lightweight image super-resolution reconstruction model based on frequency-domain-spatial-domain assisted Mamba during training, and obtain the trained lightweight image super-resolution reconstruction model with weights based on frequency-domain-spatial-domain assisted Mamba.

[0166] In this embodiment, the reconstructed feature I SR has the following expression:

[0167]

[0168] Among them, P represents PixelShuffle upsampling, γ represents the upsampling multiple, Conv 3×3 performs convolution processing through a 3×3 convolutional layer, and ISR ∈R H×W×3 。

[0169] Specifically, the calculation formula for optimizing the lightweight image super-resolution reconstruction model based on frequency-domain-spatial-domain assisted Mamba during training using the L1 loss function in step S33 is as follows:

[0170] L1 = ‖I HR -I SR ‖

[0171] where I HR represents the high-resolution image features corresponding to the HR image dataset.

[0172] S4. Input the test dataset into the trained lightweight image super-resolution reconstruction model based on frequency-domain-spatial-domain assisted Mamba with weights to obtain the reconstructed high-resolution visualization image.

[0173] Specifically, step S4 specifically includes:

[0174] According to the trained lightweight image super-resolution reconstruction model based on frequency-domain-spatial-domain assisted Mamba with weights, input the test dataset into the trained lightweight image super-resolution reconstruction model based on frequency-domain-spatial-domain assisted Mamba with ×2, ×3, and ×4 weights respectively for testing to obtain the quantitative results and the reconstructed high-resolution visualization image; among them, the quantitative results include the peak signal-to-noise ratio and the structural similarity.

[0175] The calculation formula for the peak signal-to-noise ratio is:

[0176]

[0177] where MSE represents the mean square error between image Y and image Z in the test dataset, a represents the row index of the image position, b represents the column index of the image position, H represents the height of the image, W represents the width of the image, PSNR represents the peak signal-to-noise ratio, n represents the number of bits per pixel in the image, and log 10 represents the logarithmic function.

[0178] The calculation formula for the structural similarity is:

[0179]

[0180] where SSIM(Y,Z) represents the structural similarity value between image Y and image Z, μ Y represents the average value of image Y, μ Z represents the average value of image Z, σ YZ represents the covariance between image Y and image Z, σ Y represents the variance of image Y, σZ represents the variance of image Z, c1 and c2 respectively represent the first constant and the second constant used to prevent the denominator from being zero, k1 and k2 both represent constants, and L represents the dynamic range of pixel values in the image.

[0181] In this embodiment, a higher peak signal-to-noise ratio (PSNR) indicates that the difference between the reconstructed high-resolution visualization image and the HR (high-resolution) image is smaller, while a structural similarity value (SSIM) close to 1 indicates a higher structural similarity between the reconstructed high-resolution visualization image and the HR (high-resolution) image.

[0182] In the present invention, specific embodiments are used to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0183] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A lightweight image super-resolution reconstruction method based on frequency domain-spatial domain assisted Mamba, characterized in that: The following steps are involved: S1. Obtain the original image data set and perform data preprocessing on it to obtain the training data set and the test data set; S2, introduce the frequency-spatial domain Mamba layer and the transposition and self-attention feature interaction layer to build a lightweight image super-resolution reconstruction model based on the frequency-spatial domain assisted Mamba; The lightweight image super-resolution reconstruction model based on frequency-spatial domain assisted Mamba includes a shallow feature extraction module, a deep feature extraction module, and a high-resolution image reconstruction module; The deep feature extraction module includes M feature mixing group modules and the second 3×3 convolutional layer; The feature mixing group module includes N multi-level feature interaction layers and the first 3×3 convolutional layer; The multi-level feature interaction layer includes the first LayerNorm layer, the frequency-spatial domain Mamba layer, the second LayerNorm layer, and the transposition and self-attention based feature interaction layer; The frequency-spatial domain Mamba layer includes a visual Mamba layer, a frequency-spatial domain enhancement layer, and the first 1×1 convolutional layer; The visual Mamba layer includes the upper branch module, the lower branch module, and the third Linear layer; The upper branch module includes the first Linear layer, the first 3×3 depth convolution layer, the first SiLU layer, the 2D selective scanning module, and the third LayerNorm layer; The lower branch module includes the second Linear layer and the second SiLU layer; The frequency domain-spatial domain enhancement layer includes a frequency domain enhancement module and a spatial domain enhancement module; The frequency domain enhancement module includes a second 3×3 depth convolution layer, a Fourier transform layer, a second 1×1 convolution layer, a first ReLU layer, and an inverse Fourier transform layer; The spatial domain enhancement module includes a third 3×3 depth convolution layer and a third 1×1 convolution layer; The feature interaction layer based on transposition and self-attention includes the third 3×3 convolution layer, spatial channel permutation layer, spatial self-attention layer, channel self-attention layer, and the fifth 1×1 convolution layer; The spatial channel permutation layer includes a first permutation layer, a fifth 3×3 depth convolution layer, a second permutation layer, and a sixth 3×3 depth convolution layer; The channel self-attention layer includes an adaptive average pooling layer, a fourth 3×3 depth convolution layer, and a second sigmoid layer; The spatial self-attention layer includes the second ReLU layer, the fourth 1×1 convolution layer, and the first sigmoid layer; S3, inputting the training data set into the lightweight image super-resolution reconstruction model based on the frequency domain-spatial domain assisted Mamba for training, and obtaining a trained lightweight image super-resolution reconstruction model based on the frequency domain-spatial domain assisted Mamba with weights, which is used to reconstruct the low-resolution image into a high-resolution image; S4. Input the test data set into the trained lightweight image super-resolution reconstruction model based on frequency domain-spatial domain assisted Mamba with weights to obtain a reconstructed high-resolution visualization image.

2. The lightweight image super-resolution reconstruction method based on frequency domain-spatial domain assisted Mamba according to claim 1, characterized in that: Step S1 specifically includes: S11, obtaining a public DIV2K dataset, using the DIV2K dataset as a HR image dataset, and downsampling each image in the HR image dataset by 2, 3, and 4 times using a bicubic interpolation method to obtain images with multiple downsampling multiples; S12, performing random flipping and rotation operations on the images with multiple downsampling times to obtain enhanced low-resolution LR images; S13, using the HR image data set as the first label data, using the enhanced low-resolution LR image as the first original data, and using the first label data and the first original data as the training data set; S14, obtain the public Set5 dataset, Set14 dataset, B100 dataset, Urban100 dataset and Manga109 dataset and use them as HR images of the test set, and use the bicubic interpolation method to downsample each dataset by 2, 3 and 4 times respectively to obtain the LR images of the test set; S15. Use the HR image of the test set as the second label data, use the LR image of the test set as the second original data, and use the second label data and the second original data as the test data set.

3. The lightweight image super-resolution reconstruction method based on frequency domain-spatial domain assisted Mamba according to claim 2 is characterized in that: Step S3 specifically includes: S31, input the training data set into the shallow feature extraction module of the lightweight image super-resolution reconstruction model based on frequency domain-spatial domain assisted Mamba, and use 3×3 convolution layer for convolution processing to obtain shallow features ; S32, shallow features Input the deep feature extraction module to extract mixed features and obtain deep features ; S33, deep features With shallow features After adding, the residual connection is realized, and then input into the high-resolution image reconstruction module to realize image reconstruction using convolution layer and PixelShuffle upsampling to obtain the reconstructed features , and adopt The loss function optimizes the frequency-domain-spatial-domain assisted Mamba-based lightweight image super-resolution reconstruction model in training, and obtains a trained weighted frequency-domain-spatial-domain assisted Mamba-based lightweight image super-resolution reconstruction model.

4. The lightweight image super-resolution reconstruction method based on frequency domain-spatial domain assisted Mamba according to claim 3 is characterized in that: Step S32 specifically includes: S321, shallow features Input the first LayerNorm layer for normalization operation to obtain the first normalized feature , and the first normalized feature Perform average segmentation from the channel dimension to obtain features and Features ,Right now: in, represents the feature segmentation function, represents the split dimension, Indicates the number of channels; S322, the characteristics and Features The visual Mamba layer and the frequency-spatial domain enhancement layer of the frequency-spatial domain Mamba layer are respectively input for visual Mamba and frequency-spatial domain enhancement processing, and then spliced ​​in the channel dimension, and input into the first 1×1 convolutional layer for preliminary feature interaction to obtain preliminary global-local features. , including: S3221, the characteristics Input the visual Mamba layer for visual Mamba processing to capture context information and obtain global features ; S3222, the characteristics The input frequency domain-spatial domain enhancement layer performs frequency domain-spatial domain enhancement processing to enhance context information and local information and obtain frequency domain-spatial domain enhancement features. ; S3223, global features Frequency-spatial domain enhancement features Splicing is performed in the channel dimension and input into the first 1×1 convolutional layer for preliminary feature interaction to obtain preliminary global-local features ; S323, the preliminary global-local features With shallow features Add together to achieve local residual connection and obtain the final global-local features ; S324, the final global-local feature Input the second LayerNorm layer for normalization operation to obtain the second normalized feature ; S325, the second normalized feature After the input is subjected to feature interaction fusion based on the transposition and self-attention feature interaction layers, it is combined with the final global-local feature Add, realize local residual connection, and obtain interactive features ; S326, interactive features After repeating steps S321-S325 N times, the first 3×3 convolutional layer is input to obtain the mixed features. ; S327, Mix the features After repeating step S326 M times, the second 3×3 convolutional layer is input to obtain the deep features. .

5. The lightweight image super-resolution reconstruction method based on frequency domain-spatial domain assisted Mamba according to claim 4 is characterized in that: Step S3221 specifically includes: S32211、Characteristics Input the first Linear layer for linear transformation operation to obtain the upper branch transformation features , and input into the first 3×3 deep convolutional layer for encoding operation to obtain the upper branch encoding feature , and input into the first SiLU layer for activation operation to obtain the upper branch activation feature , and input into the 2D selective scanning module for global feature extraction, and input into the third LayerNorm layer for normalization operation to obtain the first feature ; S32212, the characteristics Input the second Linear layer for linear transformation operation to obtain the lower branch transformation features , and input into the second SiLU layer for activation operation to obtain the second feature ; S32213, the first feature With the second feature Multiply and input into the third Linear layer for linear activation to obtain global features .

6. The lightweight image super-resolution reconstruction method based on frequency domain-spatial domain assisted Mamba according to claim 5, characterized in that: Step S3222 specifically includes: S32221、Characteristics Input the second 3×3 depth convolution layer to perform depth convolution operation and extract local features to obtain features The first local feature of the feature is input into the Fourier transform layer for Fourier transform. The first local feature is converted from the spatial domain to the frequency domain, and is input into the second 1×1 convolution layer for convolution operation, and then input into the first ReLU layer for activation operation to obtain the activated frequency domain feature, and then input into the inverse Fourier transform layer for inverse transformation, which is used to convert the activated frequency domain feature from the frequency domain to the spatial domain to obtain the frequency domain enhanced feature ; S32222、Characteristics Input the third 3×3 deep convolution layer to perform deep convolution operation to extract local features and obtain features The second local feature is input into the third 1×1 convolution layer for convolution operation to obtain the spatial domain enhanced feature ; S32223, frequency domain enhancement features Enhanced features in spatial domain Add together to get the frequency domain-spatial domain enhancement features .

7. The lightweight image super-resolution reconstruction method based on frequency domain-spatial domain assisted Mamba according to claim 6, characterized in that: Step S325 specifically includes: S3251, the second normalized feature Input the third 3×3 convolutional layer for preliminary feature fusion and channel compression, and compress it into , get the initialization feature ; S3252, initialize the feature The first permutation layer of the input spatial channel permutation layer performs a permutation operation by initializing the feature The spatial features are replaced by channel dimension features to obtain the first replacement feature , and input the fifth 3×3 deep convolutional layer to extract the fineness feature, and input the second permutation layer to perform the permutation operation, converting the fineness feature from the channel dimension to the spatial feature to obtain the second permutation feature , and input the sixth 3×3 deep convolutional layer for spatial feature encoding to obtain the permutation fusion feature ; S3253, initialize the feature The input is the second ReLU layer of the spatial self-attention layer for nonlinear activation, and the input is the fourth 1×1 convolution layer to reduce the channel dimension to 1, and the input is the first sigmoid layer for activation operation to obtain the first activation feature , the first activated feature With initialization feature Multiply them together to get the spatial self-attention fusion feature ; S3254, initialize the feature The adaptive average pooling layer of the input channel self-attention layer performs a pooling operation to obtain an initialization feature with a spatial size of 1×1, and enters the fourth 3×3 deep convolutional layer for encoding, and enters the second sigmoid layer for nonlinear activation operation to obtain the second activation feature , activate the second feature With initialization feature Multiply them together to get the channel self-attention fusion feature ; S3255, replace the fusion feature , spatial self-attention fusion features And channel self-attention fusion features Add and input the fifth 1×1 convolutional layer to expand the number of channels to , and then with the final global-local feature Add together to get the interaction feature .

8. The lightweight image super-resolution reconstruction method based on frequency domain-spatial domain assisted Mamba according to claim 7, characterized in that: In step S33, The loss function is used to optimize the lightweight image super-resolution reconstruction model based on frequency domain-spatial domain assisted Mamba during training as follows: in, Represents the high-resolution image features corresponding to the HR image dataset.

9. The lightweight image super-resolution reconstruction method based on frequency domain-spatial domain assisted Mamba according to claim 8, characterized in that: Step S4 specifically includes: According to the trained lightweight image super-resolution reconstruction model based on frequency domain-spatial domain assisted Mamba with weights, the test data set is input into the trained lightweight image super-resolution reconstruction model based on frequency domain-spatial domain assisted Mamba with ×2, ×3 and ×4 weights for testing, and the quantitative results and reconstructed high-resolution visual images are obtained; Among them, quantitative results include peak signal-to-noise ratio and structural similarity; The peak signal-to-noise ratio is calculated as: in, Represents the images in the test dataset With image The mean square error between The row index representing the image position, The column index representing the image position, Indicates the height of the image. Indicates the width of the image. represents the peak signal-to-noise ratio, The number of bits per pixel in the image, represents the logarithmic function; The calculation formula of structural similarity is: in, Representing images With image The structural similarity value between Representing images The average value of Representing images The average value of Representing images With image The covariance between Representing images The variance of Representing images The variance of , They represent the first constant and the second constant used to prevent the denominator from being 0, , All represent constants, Represents the dynamic range of pixel values ​​in an image.

Citation Information

Patent Citations

  • Method for realizing super-resolution for real-world text image through double-branch network capable of sensing multiple features

    CN116703725A

  • Self-adaptive efficient image super-resolution method

    CN118350995A