A remote sensing image super-resolution reconstruction method based on a cascaded sparse Mamba

CN122335555BActive Publication Date: 2026-09-25KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610785812.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-09-25
Estimated Expiration
2046-06-02

AI Technical Summary

Technical Problem

[0006]本发明的目的在于提供一种基于级联稀疏Mamba的遥感图像超分辨率重建方法,旨在解决现有技术在高倍率(×4、×8、×16)重建任务中存在模型训练困难、重建图像边缘模糊与纹理丢失、感知质量不佳以及全局建模与计算效率难以平衡的技术问题

Benefits of technology

[0058]1、本发明通过构建两级级联重建架构,将高倍率上采样任务分解为逐步优化的子任务并引入中间监督信号,有效缓解了单阶段大倍率重建的病态性与误差累积问题,显著降低了模型在高倍率任务中的训练难度;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122335555B_ABST
    Figure CN122335555B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of remote sensing image super-resolution reconstruction method based on cascade sparse Mamba, belong to remote sensing image super-resolution reconstruction technical field.The method includes: low-resolution remote sensing image is input into remote sensing image super-resolution reconstruction network, in shallow feature extraction part, the shallow feature of low-resolution remote sensing image is extracted, and shallow feature map is obtained;In first level reconstruction part, shallow feature map is input into first level sparse visual Mamba group, and first level deep feature map is obtained;In second level reconstruction part, first level deep feature map is input into second level sparse visual Mamba group, and second level deep feature map is obtained;Shallow feature map, first level deep feature map and second level deep feature map are fused, and high-resolution reconstruction image is obtained.The present application is aimed at solving the technical problems that the existing technology has blurred edge of reconstruction image, texture loss, poor perception quality and the difficulty in balancing global modeling and computational efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for super-resolution reconstruction of remote sensing images based on cascaded sparse Mamba, belonging to the field of remote sensing image super-resolution reconstruction technology. Background Technology

[0002] Remote Sensing Image Super-Resolution (RSISR) aims to reconstruct high-resolution images from low-resolution images and is a key technology for improving the usability of remote sensing data. Existing deep learning-based methods are mainly divided into three categories: methods based on Convolutional Neural Networks (CNNs), methods based on Transformers, and linear time series modeling methods based on selective state space (Mamba). CNN-based methods improve the network's representation ability through residual learning and attention mechanisms, but they are limited by local receptive fields and struggle to effectively capture long-range dependencies in remote sensing images. Transformer-based methods achieve global modeling through self-attention mechanisms, but their computational complexity increases quadratically with image size, resulting in excessive computational overhead in high-magnification (e.g., ×8, ×16) reconstruction tasks. While Mamba-based methods combine linear computational complexity and a global receptive field, they homogenize all features, failing to distinguish between smooth backgrounds (low-frequency) and fine ground structures (high-frequency) in remote sensing images.

[0003] To address the aforementioned issues, existing research has proposed various improvement schemes. For example, some methods enhance global information fusion by mining multi-scale feature associations or strengthen hierarchical feature transfer by constructing extremely deep networks; other methods employ multi-stage cascaded or progressive reconstruction frameworks, decomposing high-magnification upsampling into multiple sub-tasks, effectively mitigating the ill-conditioning and error accumulation of single-stage reconstruction. Furthermore, some researchers have introduced Mamba into the field of image super-resolution, designing a Mamba reconstruction method based on multi-directional selective scanning, and exploring the application of state-space models in feature extraction.

[0004] However, the above methods still have the following shortcomings: First, existing cascaded methods are mostly based on CNN or Transformer architectures. CNNs are difficult to adaptively model, and Transformers have excessively high computational complexity, failing to balance efficiency and performance. Second, the standard Mamba module treats all features equally, resulting in computational redundancy in smooth regions. Furthermore, uneven resource allocation weakens the ability to recover high-frequency details, a limitation particularly pronounced when high-magnification reconstructions suffer from significant information loss. Third, existing methods often rely on pixel-wise loss functions, which can easily lead to overly smooth reconstruction results, blurred edges and textures, and poor visual realism, making it difficult to meet the urgent need for high-quality images in remote sensing image interpretation tasks.

[0005] Therefore, in order to address the core problems in high-magnification remote sensing image super-resolution reconstruction tasks, such as the difficulty of model training, poor perception quality of reconstruction results, and the need to optimize computational efficiency, there is an urgent need for a reconstruction method that can effectively integrate global context, differentiate high and low frequency features, and simultaneously take into account pixel fidelity and visual realism. Summary of the Invention

[0006] The purpose of this invention is to provide a method for super-resolution reconstruction of remote sensing images based on cascaded sparse Mamba, aiming to solve the technical problems of existing technologies in high-magnification (×4, ×8, ×16) reconstruction tasks, such as difficulty in model training, blurred edges and loss of texture in reconstructed images, poor perceptual quality, and difficulty in balancing global modeling and computational efficiency.

[0007] To achieve the above objectives, the technical solution of this invention is: a remote sensing image super-resolution reconstruction method based on cascaded sparse Mamba. This method constructs a two-stage cascaded reconstruction architecture to decompose high-magnification upsampling into progressively optimized sub-tasks and introduces intermediate supervision signals. In each stage of reconstruction, a sparse visual Mamba module is used to adaptively divide high-frequency and low-frequency features through a channel routing mechanism, and these features are differentiated by Mamba branches and lightweight CNN branches, respectively. Simultaneously, perceptual loss is introduced to construct multi-dimensional training constraints to improve the visual realism of the reconstructed image, including the following steps:

[0008] Step 1: Construct a remote sensing image super-resolution reconstruction network; wherein, the remote sensing image super-resolution reconstruction network includes a shallow feature extraction part, a first-level reconstruction part, and a second-level reconstruction part.

[0009] Step 2: Input the low-resolution remote sensing image into the remote sensing image super-resolution reconstruction network for reconstruction, specifically:

[0010] In the shallow feature extraction section, shallow features of the low-resolution remote sensing image are extracted to obtain a shallow feature map;

[0011] In the first-level reconstruction section, the shallow feature map is input into the first-level sparse visual Mamba group to obtain the first-level deep feature map;

[0012] In the second-level reconstruction part, the first-level deep feature map is input into the second-level sparse visual Mamba group to obtain the second-level deep feature map;

[0013] Step 3: Fuse the shallow feature map, the first-level deep feature map, and the second-level deep feature map, and then upsample them through subpixel convolution to obtain a high-resolution reconstructed image.

[0014] Optionally, obtaining the shallow feature map specifically involves:

[0015] Define low-resolution remote sensing images as Shallow features are generated through 3×3 convolutional layers. ,in, , These represent the height and width of the low-resolution remote sensing image, respectively. The number of channels for the feature is expressed as:

[0016]

[0017] in, This indicates that convolutional layers are used to process low-resolution remote sensing images.

[0018] Optionally, obtaining the first-level deep feature map specifically involves:

[0019] shallow features Deep features are generated through first-level sparse visual Mamba group processing, and then optimized by a 3×3 convolution to obtain the first-level deep feature map. The expression is:

[0020]

[0021] in, This indicates that sparse visual Mamba groups are used to process the features.

[0022] Optionally, obtaining the second-level deep feature map specifically involves:

[0023] The first-level deep feature map The input is fed into the second-level sparse visual Mamba group, and processed through three layers of sparse visual Mamba groups and one 3×3 convolutional layer to obtain the second-level deep feature map, expressed as:

[0024]

[0025] in, This is the second-level deep feature map.

[0026] Optionally, both the first-level sparse visual Mamba group and the second-level sparse visual Mamba group are composed of several stacked sparse visual Mamba modules, and feature propagation is optimized through convolutional layers and skip connections. Each sparse visual Mamba module is specifically:

[0027] First, the input features of the sparse visual Mamba module are defined as follows: ,in, The length of the feature sequence. The number of channels is a feature. After layer normalization, low-frequency feature maps are obtained by partitioning the channel routing module (CRM). and high-frequency feature maps ,in, and These are the channel dimensions of the low-frequency feature map and the high-frequency feature map, respectively;

[0028] Secondly, the low-frequency feature map Features are extracted by successively using depthwise separable convolutional layers and 1×1 convolutional layers to obtain the features. ;

[0029] Then, the high-frequency feature map First, the feature is obtained by passing through a fully connected layer and SiLU activation. Then the features After depthwise separable convolution, 2D Mamba selective scanning, and layer normalization, and combined with features Perform element-wise multiplication to obtain the features. The expression is:

[0030]

[0031]

[0032] in, It is a fully connected layer. It is a two-dimensional selective scanning module. This indicates that depthwise separable convolutional layers are used to process the features. This indicates element-wise multiplication. Presentation layer normalization processing;

[0033] Finally, and After splicing along the channel dimension, with Perform skip connections to obtain features Then the features After layer normalization, convolutional layers, and channel attention blocks, and with features Residual connections are performed to obtain the output features of the sparse visual Mamba module. .

[0034] Optionally, the channel routing module CRM specifically comprises:

[0035] First, define the input characteristics of the channel routing module CRM as follows: ,right The model undergoes deformation processing, sequentially passing through depthwise separable convolution, batch normalization, and GELU activation to obtain features. ;

[0036] Then, the features Three parallel feed paths:

[0037] The first path generates a routing weight map through convolution and the Sigmoid function, and then the low-frequency weight map is obtained by dividing the path into blocks. and high-frequency weighting diagram The expression is:

[0038]

[0039] in, This indicates a block-based operation along the channel dimension;

[0040] The second path, through average pooling, convolution, and batch normalization, yields low-frequency content features. The expression is:

[0041]

[0042] in, This indicates batch normalization processing. This indicates average pooling.

[0043] The third path, through convolution, batch normalization, GELU, convolution, and batch normalization, yields high-frequency content features. The expression is:

[0044]

[0045] Finally, the characteristics of low-frequency content Low-frequency weighted graph Element-wise multiplication yields the low-frequency feature map output by the channel routing module CRM. High-frequency content features With high-frequency weighting graph Element-wise multiplication yields the high-frequency feature map output by the channel routing module CRM. .

[0046] Optionally, the remote sensing image super-resolution reconstruction network is trained by constructing a total loss, specifically as follows:

[0047] The first-level deep feature map shallow features After element-wise addition, the image is fed into a subpixel convolutional layer to obtain the intermediate reconstructed image. and the intermediate truth image Calculate the first-level reconstruction loss ,in, The scaling factor for the first-level reconstruction is expressed as:

[0048]

[0049]

[0050] in, This indicates a subpixel convolution operation. This indicates the calculation of the L1 loss function. This represents element-wise addition.

[0051] Reconstruct images using high resolution Compared to true high-resolution images Calculate the second-level reconstruction loss and compare it with the first-level reconstruction loss. The total pixel-level loss is obtained after weighted summation. The expression is:

[0052]

[0053] in, and These represent the weights of the first-level reconstruction loss and the second-level reconstruction loss in the total loss, respectively.

[0054] Acquire high-resolution reconstructed images Compared to true high-resolution images The LPIPS values ​​between the two values ​​are used as the perceptual loss. The total loss is obtained by weighted summing of the total pixel-level loss and the perceptual loss. The expression is:

[0055]

[0056] in, To perceive loss weights, This indicates the operation of calculating LPIPS values.

[0057] The beneficial effects of this invention are:

[0058] 1. This invention constructs a two-level cascaded reconstruction architecture, decomposes the high-magnification upsampling task into progressively optimized sub-tasks and introduces intermediate supervision signals, effectively alleviating the ill-conditioning and error accumulation problems of single-stage high-magnification reconstruction, and significantly reducing the training difficulty of the model in high-magnification tasks.

[0059] 2. This invention proposes a sparse visual Mamba module, which uses a channel routing mechanism to adaptively divide high-frequency and low-frequency features. For key high-frequency details, a Mamba branch is used to perform long-range semantic association modeling, and a lightweight CNN branch is used to process smooth low-frequency backgrounds. While maintaining the global receptive field and linear computational complexity, this invention significantly enhances the model's ability to recover key details such as edges and textures in remote sensing images, and effectively controls computational overhead.

[0060] 3. This invention introduces LPIPS-based perceptual loss as a training constraint to guide the model to approximate the distribution of real images in the high-level semantic feature space, effectively overcoming the oversmoothing problem that is easily caused by traditional pixel loss functions, and greatly improving the visual realism and semantic consistency of the reconstructed images. Attached Figure Description

[0061] Figure 1 This is a schematic diagram of the overall structure and process of the remote sensing image super-resolution reconstruction network of the present invention;

[0062] Figure 2 This is a schematic diagram of the sparse visual Mamba module and channel routing module of the present invention;

[0063] Figure 3 This is a comparison chart of the visual effects of various methods on the UCMerced dataset when the magnification is ×8, as described in the embodiments of the present invention.

[0064] Figure 4 The LAM visualization analysis diagrams for the various methods involved in the embodiments of the present invention are shown below;

[0065] Figure 5 This is a comparison chart of the visual effects of reconstruction results of different model variants involved in the embodiments of the present invention. Detailed Implementation

[0066] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0067] Example 1: A method for super-resolution reconstruction of remote sensing images based on cascaded sparse Mamba, comprising the following steps:

[0068] Step 1: As Figure 1 As shown, a remote sensing image super-resolution reconstruction network is constructed; wherein, the remote sensing image super-resolution reconstruction network includes a shallow feature extraction part, a first-level reconstruction part, and a second-level reconstruction part.

[0069] Step 2: Input the low-resolution remote sensing image into the remote sensing image super-resolution reconstruction network for reconstruction, specifically:

[0070] In the shallow feature extraction section, shallow features of the low-resolution remote sensing image are extracted to obtain a shallow feature map, specifically as follows:

[0071] Define low-resolution remote sensing images as Shallow features are generated through 3×3 convolutional layers. ,in, , These represent the height and width of the low-resolution remote sensing image, respectively. The number of channels for the feature is expressed as:

[0072]

[0073] in, This indicates that convolutional layers are used to process low-resolution remote sensing images; in this embodiment, It is 64. It is 64. It is 180.

[0074] In the first-level reconstruction section, the shallow feature map is input into the first-level sparse visual Mamba group to obtain the first-level deep feature map, specifically:

[0075] shallow features Deep features are generated through first-level sparse visual Mamba group processing, and then optimized by a 3×3 convolution to obtain the first-level deep feature map. The expression is:

[0076]

[0077] in, This indicates that the features are processed using the Sparse Vision Mamba Group (SVMG).

[0078] In the second-level reconstruction section, the first-level deep feature map is input into the second-level sparse visual Mamba group to obtain the second-level deep feature map, specifically:

[0079] The first-level deep feature map The input is fed into the second-level sparse visual Mamba group, and processed through 3 layers of SVMG and one 3×3 convolutional layer to obtain the second-level deep feature map, expressed as:

[0080]

[0081] in, This is the second-level deep feature map.

[0082] Optionally, such as Figure 2 As shown, both the first-level sparse visual Mamba group and the second-level sparse visual Mamba group are composed of several stacked sparse visual Mamba modules, and feature propagation is optimized through convolutional layers and skip connections. Each sparse visual Mamba module is specifically as follows:

[0083] First, the input features of the Sparse Visual Mamba Module (SVMM) are defined as follows: ,in, The length of the feature sequence is , and , The number of channels is a feature. After layer normalization, low-frequency feature maps are obtained by partitioning the data using the Channel Routing Module (CRM). and high-frequency feature maps ,in, and Let be the channel dimensions of the low-frequency feature map and the high-frequency feature map, respectively, expressed as:

[0084]

[0085] in, This indicates that CRM is used to process the features. The representation layer is normalized. In this embodiment, in the first-level reconstruction part, the channel dimensions of the low-frequency feature map and the high-frequency feature map are 105 and 75, respectively. In the second-level reconstruction part, the channel dimensions of the low-frequency feature map and the high-frequency feature map are 75 and 105, respectively.

[0086] Secondly, the low-frequency feature map Features are extracted by successively using depthwise separable convolutional layers and 1×1 convolutional layers to obtain the features. The expression is:

[0087]

[0088] in, This indicates that depthwise separable convolutional layers are used to process the features.

[0089] Then, the high-frequency feature map First, the feature is obtained by passing through a fully connected layer and SiLU activation. Then the features After depthwise separable convolution, 2D Mamba selective scanning, and layer normalization, and combined with features Perform element-wise multiplication to obtain the features. The expression is:

[0090]

[0091]

[0092] in, It is a fully connected layer. It is a two-dimensional selective scanning module. This represents element-wise multiplication.

[0093] Finally, and After splicing along the channel dimension, with Perform skip connections to obtain features Then the features After layer normalization, convolutional layers, and channel attention blocks to optimize the feature channels, the feature channels are finally optimized. Residual connections are performed to obtain the output features of the sparse visual Mamba module. The expression is:

[0094]

[0095]

[0096] in, This indicates a splicing operation along the channel dimension. This indicates that channel attention blocks are used to process the features. This indicates an element-wise addition operation.

[0097] Optionally, such as Figure 2 As shown, the channel routing module specifically comprises:

[0098] First, define the input characteristics of the channel routing module CRM as follows: ,right The model undergoes deformation processing, sequentially passing through depthwise separable convolution, batch normalization, and GELU activation to obtain features. The expression is:

[0099]

[0100] in, This indicates batch normalization processing;

[0101] Then, the features Three parallel feed paths:

[0102] The first path generates a routing weight map through convolution and the Sigmoid function, and then the low-frequency weight map is obtained by dividing the path into blocks. and high-frequency weighting diagram The expression is:

[0103]

[0104] in, This indicates a block-based operation along the channel dimension;

[0105] The second path, through average pooling, convolution, and batch normalization, yields low-frequency content features. The expression is:

[0106]

[0107] in, This indicates average pooling.

[0108] The third path, through convolution, batch normalization, GELU, convolution, and batch normalization, yields high-frequency content features. The expression is:

[0109]

[0110] Finally, the characteristics of low-frequency content Low-frequency weighting graph Element-wise multiplication yields the low-frequency feature map output by the channel routing module CRM. ,and High-frequency content features With high-frequency weighting graph Element-wise multiplication yields the high-frequency feature map output by the channel routing module CRM. ,and .

[0111] It is understood that, through the channel routing module constructed in this embodiment, the feature map is dynamically divided into high-frequency component feature map and low-frequency component feature map according to the feature response. The high-frequency branch relies on local convolution to respond to details and edges, while the low-frequency branch relies on the global average pooling low-pass filtering effect to suppress local details and noise and retain the overall energy of the channel.

[0112] Optionally, the remote sensing image super-resolution reconstruction network is trained by constructing a total loss, specifically as follows:

[0113] The first-level deep feature map shallow features After element-wise addition, the image is fed into a subpixel convolutional layer to obtain the intermediate reconstructed image. and the intermediate truth image Calculate the first-level reconstruction loss ,in, The scaling factor for the first-level reconstruction is expressed as:

[0114]

[0115]

[0116] in, This indicates the calculation of the L1 loss function; in this embodiment, the reconstruction ratio of the first level is... Take 4. The dimensions are 256×256×3;

[0117] Reconstruct images using high resolution Compared to true high-resolution images Calculate the second-level reconstruction loss and compare it with the first-level reconstruction loss. The total pixel-level loss is obtained after weighted summation. The expression is:

[0118]

[0119] in, and These represent the weights of the first-level reconstruction loss and the second-level reconstruction loss in the total loss, respectively.

[0120] Acquire high-resolution reconstructed images Compared to true high-resolution images The LPIPS values ​​between the two values ​​are used as the perceptual loss. The total loss is obtained by weighted summing of the total pixel-level loss and the perceptual loss. The expression is:

[0121]

[0122] in, To perceive loss weights, This indicates the operation of calculating LPIPS values.

[0123] Optionally, in this embodiment, and All values ​​are 0.5. The value is set to 0.02. Furthermore, this embodiment uses the Adam optimizer, with parameters set to... , =0.99, The initial learning rate is The batch size is set to 4, the total number of training epochs is 500, and the learning rate is halved every 125 iterations.

[0124] Step 3: Fuse the shallow feature map, the first-level deep feature map, and the second-level deep feature map, and then upsample them through subpixel convolution to obtain a high-resolution reconstructed image.

[0125] Optionally, the expression for obtaining the high-resolution reconstructed image is:

[0126]

[0127] in, To reconstruct high-resolution images, This indicates a subpixel convolution operation. This represents the total magnification of the super-resolution reconstruction.

[0128] Based on the specific implementation details, the effectiveness of the technical solution of the present invention will be demonstrated through experiments.

[0129] Specifically, this experiment was validated on the UCMerced and AID public datasets. The hardware environment was an NVIDIA GeForce RTX 4090 GPU (24GB VRAM), and the software environment was the PyTorch framework. The following evaluation metrics were used: PSNR (Peak Signal-to-Noise Ratio), a measure of pixel similarity, with higher values ​​being better; SSIM (Structural Similarity Index), a measure of overall image structure and texture similarity, ranging from 0 to 1, with values ​​closer to 1 indicating better performance and closer alignment with human visual perception; LPIPS (Learned Perceptual Patch Similarity), representing human visual perception error and simulating subjective human visual perception, with lower values ​​being better; and CLIP-IQA (Clip-Based Image Quality Assessment), used to directly evaluate the overall aesthetics and naturalness of an image, ranging from 0 to 1, with values ​​closer to 1 indicating better image quality. The experimental results are shown in Table 1.

[0130] Table 1 Comparison results of ×8 tasks on the UCMerced dataset

[0131]

[0132] As shown in Table 1, in the UCMerced dataset ×8 scaling task, the CSPPG-SR of this invention achieves a peak signal-to-noise ratio (PSNR) of 23.32 dB, a structural similarity index (SSIM) of 0.5925, a learned perceptual patch similarity (LPIPS) of 0.3219, and a CLIP-IQA of 0.8592, all of which outperform existing comparative methods. In particular, compared to the basic MambaIR, CSPPG-SR improves PSNR, SSIM, and CLIP-IQA by 0.06 dB, 0.0028, and 0.0338, respectively, while decreasing LPIPS by 0.1322.

[0133] Furthermore, such as Figure 3 As shown, the visual advantages of CSMamba-SR in high-magnification reconstruction tasks are intuitively verified. In the UCMerced dataset ×8 magnification reconstruction results, the aircraft images reconstructed by CNN-based methods (LGCNet, HSENet) exhibit obvious edge blurring and texture homogenization problems, with loose fuselage outlines and severe loss of details. Transformer-based methods (HAT, DAT) can preserve the overall structure of the aircraft, but the color transitions of the fuselage are abrupt, and artifacts and texture breaks appear in some areas. MambaIR and CSMambaSR improve structural consistency through long-range dependency modeling, but there is still room for improvement in details such as the tail fin outline. The method of this invention not only clearly restores the outline edges of the aircraft, but also accurately restores the color changes and structural lines of the fuselage, and significantly reduces artifacts and oversmoothing phenomena.

[0134] Furthermore, such as Figure 4 As shown, the Local Attribution Map (LAM) visualization analysis of each method is presented. The LAM coverage and diffusion index (DI) reflect the model's ability to utilize global spatial information. A higher DI value indicates that the model can more fully integrate the correlation information of distant pixels. LAM analysis shows that the diffusion index DI of this invention reaches 21.21, significantly higher than SwinIR (10.80), TTST (11.28), and MambaIR (18.58). This demonstrates that the sparse visual Mamba module of this invention, through differentiated modeling of high-frequency and low-frequency features, can more efficiently capture long-distance ground feature relationships in remote sensing images, thereby achieving more accurate structural reconstruction and detail supplementation.

[0135] Furthermore, an ablation experiment was conducted in this embodiment to verify the effectiveness of each component. The experimental results are shown in Table 2.

[0136] Table 2. Ablation experiment results for the UCMerced dataset ×8 task (bold indicates best results).

[0137]

[0138] In Table 2, the two-stage cascade represents the first and second stage reconstruction parts. As shown in Table 2, the baseline model (without two-stage cascade, without SVMM, and without perceptual loss) has a PSNR of 23.26 dB and LPIPS of 0.4541. After introducing the two-stage cascade architecture, the PSNR increases to 23.30 dB, and LPIPS decreases to 0.4494. After introducing SVMM, the PSNR increases to 23.27 dB, and SSIM increases to 0.5902. After introducing perceptual loss, LPIPS decreases significantly to 0.3463, and CLIP-IQA increases to 0.8551. The complete model (i.e., the present invention, two-stage cascade + SVMM + perceptual loss) achieves the best performance, with a PSNR of 23.32 dB, SSIM of 0.5925, LPIPS of 0.3219, and CLIP-IQA of 0.8592.

[0139] Furthermore, such as Figure 5 The image shown is a comparison of the visual effects of reconstruction results from different model variants. Figure 5 The reconstruction quality of each variant of the remote sensing image is visually demonstrated and compared with the high-resolution remote sensing image (HR). Among them, the reconstruction result of variant 1 has obvious edge blurring; variant 2 has improved structural integrity but still has artifacts; variant 4 has enhanced detail recovery ability but lacks visual realism; variant 7 presents a clearer dividing line but the body structure is somewhat distorted; variant 9 (complete model) not only clearly restores color and texture, but also maintains good visual naturalness.

[0140] Based on the experimental results above, it can be seen that, as verified by the UCMerced and AID public datasets, the present invention outperforms existing CNN-based, Transformer-based, and Mamba-based methods in quantitative metrics such as PSNR, SSIM, LPIPS, and CLIP-IQA in high-magnification reconstruction tasks such as ×4, ×8, and ×16.

[0141] In summary, this invention first extracts shallow features from low-resolution remote sensing images using a 3×3 convolutional layer to generate a shallow feature map. Further, it constructs a two-stage cascaded reconstruction architecture, decomposing the high-magnification upsampling task into progressively optimized sub-tasks. The first-stage reconstruction extracts first-level deep features through stacked sparse visual Mamba groups and convolutional layers, and introduces intermediate supervision signals during the training phase to calculate the first-stage reconstruction loss, mitigating error accumulation and reducing model training difficulty. The first-level deep features are then input into the second-stage reconstruction, where another set of sparse visual Mamba groups and convolutional layers extract second-level deep features. Finally, the shallow features, first-level deep features, and second-level deep features are fused and upsampled via sub-pixel convolution to generate the final high-resolution reconstructed image. The sparse visual Mamba group consists of multiple stacked sparse visual Mamba modules. Each module adaptively divides high-frequency and low-frequency features through a channel routing mechanism. The high-frequency branch uses two-dimensional Mamba selective scanning for long-range semantic association modeling, while the low-frequency branch uses lightweight depthwise separable convolution processing. This enhances detail recovery capabilities while maintaining the global receptive field and linear computational complexity. Furthermore, an LPIPS-based perceptual loss is introduced as a training constraint to improve the visual realism of the reconstructed image. This invention effectively improves the pixel fidelity and perceptual quality of high-magnification (×4, ×8, ×16) remote sensing image super-resolution reconstruction, and is suitable for downstream interpretation tasks such as fine segmentation of remote sensing images, target detection, and land cover classification.

[0142] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. A method for super-resolution reconstruction of remote sensing images based on cascaded sparse Mamba, characterized in that, The method includes the following steps: Step 1: Construct a remote sensing image super-resolution reconstruction network; wherein, the remote sensing image super-resolution reconstruction network includes a shallow feature extraction part, a first-level reconstruction part, and a second-level reconstruction part. Step 2: Input the low-resolution remote sensing image into the remote sensing image super-resolution reconstruction network for reconstruction, specifically: In the shallow feature extraction section, shallow features of the low-resolution remote sensing image are extracted to obtain a shallow feature map; In the first-level reconstruction section, the shallow feature map is input into the first-level sparse visual Mamba group to obtain the first-level deep feature map; In the second-level reconstruction part, the first-level deep feature map is input into the second-level sparse visual Mamba group to obtain the second-level deep feature map; Step 3: Fuse the shallow feature map, the first-level deep feature map, and the second-level deep feature map, and then perform upsampling through subpixel convolution to obtain a high-resolution reconstructed image; Both the first-level sparse visual Mamba group and the second-level sparse visual Mamba group are composed of several stacked sparse visual Mamba modules, and feature propagation is optimized through convolutional layers and skip connections. Each sparse visual Mamba module is specifically as follows: First, the input features of the sparse visual Mamba module are defined as follows: ,in, The length of the feature sequence. The number of channels is a feature. After layer normalization, low-frequency feature maps are obtained by partitioning the channel routing module (CRM). and high-frequency feature maps ,in, and These are the channel dimensions of the low-frequency feature map and the high-frequency feature map, respectively; Secondly, the low-frequency feature map Features are extracted by successively using depthwise separable convolutional layers and 1×1 convolutional layers to obtain the features. ; Then, the high-frequency feature map First, the feature is obtained by passing through a fully connected layer and SiLU activation. Then the features After depthwise separable convolution, 2D Mamba selective scanning, and layer normalization, and combined with features Perform element-wise multiplication to obtain the features. ; Finally, and After splicing along the channel dimension, with Perform skip connections to obtain features Then the features After layer normalization, convolutional layers, and channel attention blocks, and with features Residual connections are performed to obtain the output features of the sparse visual Mamba module. ; The channel routing module CRM is specifically as follows: First, define the input characteristics of the channel routing module CRM as follows: ,right The model undergoes deformation processing, sequentially passing through depthwise separable convolution, batch normalization, and GELU activation to obtain features. ; Then, the features Three parallel feed paths: The first path generates a routing weight map through convolution and the Sigmoid function, and then the low-frequency weight map is obtained by dividing the path into blocks. and high-frequency weighting diagram ; The second path, through average pooling, convolution, and batch normalization, yields low-frequency content features. ; The third path, through convolution, batch normalization, GELU, convolution, and batch normalization, yields high-frequency content features. ; Finally, the characteristics of low-frequency content Low-frequency weighted graph Element-wise multiplication yields the low-frequency feature map output by the channel routing module CRM. High-frequency content features With high-frequency weighting graph Element-wise multiplication yields the high-frequency feature map output by the channel routing module CRM. .

2. The remote sensing image super-resolution reconstruction method based on cascaded sparse Mamba as described in claim 1, characterized in that, The obtained shallow feature map is specifically as follows: Define low-resolution remote sensing images as Shallow features are generated through 3×3 convolutional layers. ,in, , Here, the height and width of the low-resolution remote sensing image are respectively expressed as: ; in, This indicates that convolutional layers are used to process low-resolution remote sensing images.

3. The method for super-resolution reconstruction of remote sensing images based on cascaded sparse Mamba as described in claim 2, characterized in that, The specific steps for obtaining the first-level deep feature map are as follows: shallow features Deep features are generated through first-level sparse visual Mamba group processing, and then optimized by a 3×3 convolution to obtain the first-level deep feature map. The expression is: ; in, This indicates that sparse visual Mamba groups are used to process the features.

4. The method for super-resolution reconstruction of remote sensing images based on cascaded sparse Mamba as described in claim 3, characterized in that, The specific steps for obtaining the second-level deep feature map are as follows: The first-level deep feature map The input is fed into the second-level sparse visual Mamba group, and processed through three layers of sparse visual Mamba groups and one 3×3 convolutional layer to obtain the second-level deep feature map, expressed as: ; in, This is the second-level deep feature map.

5. The method for super-resolution reconstruction of remote sensing images based on cascaded sparse Mamba as described in claim 1, characterized in that, The remote sensing image super-resolution reconstruction network is trained by constructing a total loss, specifically as follows: The first-level deep feature map shallow features After element-wise addition, the image is fed into a subpixel convolutional layer to obtain the intermediate reconstructed image. and the intermediate truth image Calculate the first-level reconstruction loss ,in, The scaling factor for the first-level reconstruction is expressed as: ; ; in, This indicates a subpixel convolution operation. This indicates the calculation of the L1 loss function. This represents element-wise addition. Reconstruct images using high resolution Compared to true high-resolution images Calculate the second-level reconstruction loss and compare it with the first-level reconstruction loss. The total pixel-level loss is obtained after weighted summation. The expression is: ; in, and These represent the weights of the first-level reconstruction loss and the second-level reconstruction loss in the total loss, respectively. Acquire high-resolution reconstructed images Compared to true high-resolution images The LPIPS values ​​between the two values ​​are used as the perceptual loss. The total loss is obtained by weighted summing of the total pixel-level loss and the perceptual loss. The expression is: ; in, To perceive loss weights, This indicates the operation of calculating LPIPS values.

Citation Information

Patent Citations

  • Remote sensing image super-resolution reconstruction method based on cross-scale Mama

    CN120598784A

  • Image super-resolution system and method based on high and low frequency separation sensing Mama

    CN121639473A