Remote sensing image panchromatic sharpening method and system fusing Mama and CNN under detail enhancement guidance
Through the Mamba and CNN fusion method under the guidance of detailed enhancement, the problem of image quality degradation in full color sharpening is solved, high-quality HRMS image generation is achieved, and objective and subjective evaluation of the fusion effect is improved.
Patent Information
- Application Number
- CN202510324405.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-25
AI Technical Summary
Existing full-color sharpening methods usually improve the spatial resolution of multispectral images through direct upsampling operations, resulting in problems such as image quality degradation, blurring and incomplete spectral information recovery.
The Mamba and CNN fusion method under detailed enhancement guidance is adopted to enhance the LRMS texture characteristics through PAN image Qualcomm information, and local and global feature information is extracted in combination with CNN and Mamba, and integrated into the fusion branch step by step, and finally reconstruct high-quality HRMS images.
It realizes effective retention of rich texture information and spectral characteristics, and improves the subjective visual effect and objective evaluation indicators of the fusion results, such as SSIM, PSNR, SAM, RMSE, ERGAS and CC.
Smart Images

Figure CN120374447A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a panchromatic sharpening method and system for remote sensing images that integrates Mamba and CNN under the guidance of detail enhancement. Background Art
[0002] With the progress of technology and the development of society, the demand for high-resolution multispectral remote sensing images in fields such as urban planning, resource exploration, and climate monitoring is increasing. However, due to the limitations of the hardware functions of remote sensing satellite sensors and related costs, existing satellite imaging systems cannot produce multispectral images (Multispectral, MS) with high spatial resolution (High Resolution, HR). In other words, the multispectral images obtained by multispectral remote sensing sensor imaging devices have rich spectral information, but their spatial resolution is low, that is, they are actually low-resolution multispectral images (Low Resolution Multispectral, LRMS); relatively, the panchromatic images (Panchromatic, PAN) obtained by panchromatic remote sensing sensor imaging devices have high spatial resolution and strong texture information, but lack rich spectral information. Therefore, through image fusion technology, the multispectral image with rich spectral information and the panchromatic image with high spatial resolution are fused (also known as panchromatic sharpening) to produce a fused image with rich spectral information and high spatial resolution, that is, a high-resolution multispectral image (High Resolution Multispectral, HRMS), which is beneficial to the completion of downstream vision tasks.
[0003] In recent years, due to the excellent performance of deep learning methods in many computer vision tasks, researchers have actively applied them to the field of panchromatic and multispectral remote sensing image fusion. Inspired by the Super-Resolution Convolutional Neural Network (SRCNN), Masi et al. first introduced deep learning methods into the field of panchromatic and multispectral remote sensing image fusion, proposed a panchromatic sharpening method called PNN, and creatively regarded the fusion problem of panchromatic and multispectral remote sensing images as a special case of the multispectral remote sensing image super-resolution problem. By adopting a relatively simplified network architecture, PNN successfully achieved the expected results. Since then, researchers have fully utilized the latest progress of deep learning in this field and introduced more complex network models. Yuan et al. introduced multi-scale feature extraction and residual learning into the basic Convolutional Neural Network (CNN) architecture according to the characteristics of multi-scale ground object information in remote sensing images, and proposed the MSDCNN network. This network can capture more complex features and significantly improve the processing ability of ground object information at different scales. Xiang et al. proposed a weighted fusion model MC-JAFN. This model enhances important feature information, reduces unnecessary information interference, and significantly improves the spatial resolution and fusion efficiency. Liu et al. first attempted to apply the Generative Adversarial Network (GAN) to the fusion of panchromatic and multispectral remote sensing images and proposed the PSGAN model. Without real values, this model uses the adversarial mechanism between the generator and the discriminator to further improve the quality of panchromatic sharpening. Ma et al. transformed the panchromatic sharpening problem into a multi-task problem and designed a dual discriminator network structure with coexisting spectral discriminator and spatial discriminator, Pan-GAN. Pan-GAN aims to simultaneously preserve the rich spectral information in the multispectral image and the high-resolution spatial details in the panchromatic image. Zhou et al. first attempted to introduce Transformer into the field of panchromatic and multispectral remote sensing image fusion and proposed a Transformer-based image fusion network model PanFormer. This model uses Transformer to extract modal features from PAN and MS images respectively, and captures the redundant and complementary information between these two modalities through a unique cross-attention module. This fusion framework promotes the information interaction between PAN and MS images, thus achieving a significant improvement in the fusion effect.
[0004] The CNN-based method extracts features from the input image through convolution operations. However, the convolution operation only focuses on extracting local feature information in the image and is difficult to capture the context global feature information in the input image. The GAN based on the adversarial mechanism transforms the image fusion problem into the problem of the generator generating an image that conforms to the expected data distribution, and continuously improves the performance of the generator through the adversarial game between the generator and the discriminator. The GAN-based method has achieved good panchromatic and multispectral remote sensing image fusion effects. However, the training process of the GAN model is complex, unstable, and prone to mode collapse, etc. Compared with the CNN that only focuses on extracting local features in the input image, the Transformer-based method can extract global feature information from the input image data and has achieved good image fusion effects. However, the network model of the Transformer network is often large, has many training parameters, high computational complexity, and long training time, etc. In contrast, in the face of these challenges, the Mamba model provides a novel and efficient solution. The Mamba model can achieve effective sequence modeling with linear complexity and shows great potential in the field of long sequence modeling and generation such as long texts, high-resolution images, videos, etc. He et al. first introduced the Mamba model into the field of panchromatic and multispectral remote sensing image fusion and proposed the Pan-Mamba model. This model further processes the features extracted by Mamba through the channel-swapping Mamba block and the cross-modal Mamba block, thus realizing the effective interaction and fusion of cross-modal information, and further bringing significant performance improvement and new development directions to the problem of panchromatic and multispectral remote sensing image fusion.
[0005] Many existing panchromatic sharpening methods usually first perform a single upsampling operation on the LRMS to match its spatial resolution with that of the PAN; then output the fused image through feature extraction, fusion, and reconstruction. Such a fusion strategy will lead to insufficient extraction of the spatial and spectral information of the source LRMS, resulting in problems such as blurred fusion results and incomplete restoration of spectral information. Summary of the Invention
[0006] To solve the problems existing in the background technology, the purpose of the present invention is to propose a panchromatic sharpening method and system for remote sensing images that integrates Mamba and CNN under the guidance of detail enhancement, that is, a panchromatic sharpening method for remote sensing images that integrates Mamba and Convolutional Neural Network (CNN) under the guidance of detail enhancement. Specifically, first, the texture feature information of the LRMS is enhanced by the high-pass information of the PAN image; subsequently, the local and global feature information of the PAN and the enhanced LRMS images are extracted by CNN and Mamba, and the extracted feature information is integrated into the CNN-based fusion branch level by level for sufficient interaction and fusion of local and global information in the two modalities; finally, a high-quality HRMS image is reconstructed and generated.
[0007] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0008] A panchromatic sharpening method for remote sensing images that integrates Mamba and CNN under the guidance of detail enhancement, characterized in that the method is as follows:
[0009] S1: First, enhance the texture feature information of the low-spatial-resolution multispectral image LRMS through the high-pass information of the panchromatic image PAN to obtain the enhanced LRMS image;
[0010] S2: Subsequently, extract the local and global feature information of the PAN and the enhanced LRMS image through the convolutional neural network CNN and Mamba, and integrate the extracted feature information into the CNN-based fusion branch level by level for sufficient interaction and fusion of local and global information in the two modalities;
[0011] S3: Finally, reconstruct and generate a high-quality high-spatial-resolution multispectral image HRMS.
[0012] Furthermore, the specific method in S1 is divided into two stages:
[0013] The first stage: First, downsample the PAN image by 4 times and concatenate it with the original LRMS image in the channel dimension to obtain the initial input F1; then, extract features from F1 through a 3×3 convolutional layer, and use the PixelShuffle operation to improve the spatial resolution to obtain the feature map F2; subsequently, input the obtained feature map F2 into the residual connection block and the dense connection block in parallel to further capture and enhance the feature information, and obtain the feature maps F3 and F4 respectively; then, concatenate and output the feature maps F3 and F4 enhanced by the residual connection and the dense connection, and restore to the original LRMS image channel number after passing through a convolutional layer; finally, add this output to the upsampled original LRMS image by 2 times and add the high-pass information of the downsampled PAN image by 2 times to obtain the output feature map F of the first stageout1 Specifically, the above process is expressed by the formula:
[0014]
[0015] where I M ∈R h×w×B represents an LRMS image of size h×w with B bands, and I P ∈R H×W×1 represents a PAN image of size H×W×1, where W / w = H / h = r, W and H are equal to 128, and r is equal to 4; the superscript ↑ and subscript ↓ represent upsampling and downsampling operations respectively; concat(*) represents the concatenation operation of features in the channel dimension, conv(*) represents the convolution operation, Pix(*) represents the PixelShuffle operation of sub-pixel convolution, ResNet(*) represents the residual connection operation, DenseNet(*) represents the dense connection operation, and H(*) represents the operation of obtaining high-pass information;
[0016] Second stage: The output feature map F out1 of the first stage is concatenated with the PAN image downsampled by a factor of 2 in the channel dimension to obtain the feature map F5; then, similar to the first stage, the feature map F5 is processed, that is, feature extraction by a 3×3 convolutional layer and the PixelShuffle operation is used to improve the spatial resolution to obtain the feature map F6; then, the feature map F6 is respectively subjected to residual and dense connection processing to obtain the feature maps F7 and F8; subsequently, the feature maps F7 and F8 are concatenated and added to the upsampled LRMS image by a factor of 4 and the high-pass filtered information of the PAN image, so as to obtain the enhanced multi-spectral image, denoted as MS+; specifically, the above process is expressed by the formula:
[0017]
[0018] Furthermore, the specific method of S2 is as follows: The pre-fused MS+ and PAN images are first respectively passed through a shallow feature extraction block composed of 3×3 convolutions to capture their respective basic features, and the feature maps F MS and F PAN are respectively obtained; subsequently, the feature maps F MS and F PAN extracted by the shallow feature extraction blockGlobal feature extraction is performed through the Mamba block to enhance the understanding of complex scenes. Next, the output of the Mamba block is concatenated with the original MS+ and PAN images in the channel dimension and then input into the first feature fusion layer. Then, the output of the fusion layer is concatenated with the outputs of the PAN branch and the MS branch after being processed by the next-level Mamba block respectively, and then input into the next-level feature fusion layer. And so on, this process is carried out step by step to enable full information interaction and fusion between the MS+ image and the PAN image. Finally, the output of the last-level feature fusion layer is further concatenated with the output of the Mamba block to complete the task of the entire feature extraction and interaction module and obtain the final feature map F out2 Specifically, the above process is expressed by the formula as follows:
[0019]
[0020] where F MS represents the shallow features extracted from the MS+ image, F PAN represents the shallow features extracted from the PAN image, D n represents the output of the nth fusion layer, and n takes values of 1, 2, 3 respectively; Shallow(*) represents the shallow feature extraction operation, Fu[*] represents the fusion layer operation, and M h (*) represents the operation of the hth Mamba block, and h takes values of 1, 2, 3, 4 respectively.
[0021] Furthermore, S3 specifically is that the reconstruction module receives the feature map processed by the feature extraction and interaction module and then reconstructs and generates the final high-resolution multispectral image. Specifically, the operation process of the S3 method is expressed by the formula as follows:
[0022] I HRMS = Sig(conv(σ(conv(F out2 )))) (4)
[0023] where F out2 represents the output of the feature extraction and interaction module, σ(*) represents the activation function PReLU, and Sig(*) represents the Sigmoid activation function.
[0024] A pan-sharpening system for remote sensing images that fuses Mamba and CNN under the guidance of detail enhancement, characterized by including:
[0025] Pre-fusion module: The pre-fusion module generates an enhanced MS image with the same resolution as the PAN image;
[0026] Feature extraction and interaction module: The enhanced MS and PAN are input into this module to promote the learning of complementary features and at the same time suppress redundant information;
[0027] Reconstruction module: Generate the finally fused high-resolution multispectral image through the reconstruction module.
[0028] Furthermore, the feature extraction and interaction module consists of three branches, namely the multispectral branch, the panchromatic branch, and the intermediate fusion branch. The intermediate fusion branch realizes the complementarity and integration of local and global feature information.
[0029] The following beneficial effects can be obtained through the above technical solutions:
[0030] The method proposed in the present invention effectively combines the CNN focusing on local feature extraction and the Mamba focusing on global feature modeling, thereby achieving good panchromatic sharpening effects for remote sensing images. Experiments are carried out using the publicly available dataset. The results show that, compared with the existing methods, the fused results obtained by the method proposed in this patent have richer texture information and better subjective visual effects. In addition, the objective evaluation indexes obtained by the proposed method are generally better than those of other comparison methods. Specifically: the structural similarity index (SSIM) is about 15.04% better than the average value of the comparison methods, the peak signal-to-noise ratio (PSNR) is about 19.66% better than the average value of the comparison methods, the spectral angle mapper (SAM) is about 21.84% better than the average value of the comparison methods, the root mean square error (RMSE) is about 42.20% better than the average value of the comparison methods, the relative dimensionless global error (ERGAS) is about 27.53% better than the average value of the comparison methods, the correlation coefficient (CC) is about 3.90% better than the average value of the comparison methods, and the universal image quality index (UIQI) is about 9.52% better than the average value of the comparison methods. This further illustrates that the fused results obtained by the method proposed in the present invention not only effectively retain the spatial information of the panchromatic image, but also effectively maintain the spectral characteristics of the multispectral image. Description of the Drawings
[0031] Figure 1 is the overall fusion framework diagram of the method proposed in the present invention.
[0032] Figure 2 is the Mamba network structure diagram.
[0033] Figure 3 is the Fusion network structure diagram.
[0034] Figure 4 is the Reconstruction network structure diagram.
[0035] Figure 5 is the Shallow network structure diagram. Detailed Embodiments
[0036] The following further describes the present invention with reference to the drawings:
[0037] Many existing pansharpening methods usually adopt a single upsampling operation to enhance the spatial resolution of low-resolution multispectral images (LRMS) to a level matching that of panchromatic images (PAN). However, such methods have certain deficiencies: they often overlook the image quality degradation problem caused by directly upsampling LRMS images. Specifically, although the direct upsampling operation can quickly adjust the spatial resolution of LRMS, due to the lack of effective protection of image details and spectral information, it is prone to introducing problems such as blurring, distortion, or spectral aberration, thus reducing the quality of the fusion result. Therefore, a pansharpening method and system for remote sensing images that fuse Mamba and CNN under the guidance of detail enhancement are proposed.
[0038] This embodiment proposes a pansharpening system for remote sensing images that fuses Mamba and CNN under the guidance of detail enhancement. The system consists of three core modules: (1) Pre-fusion module: An enhanced MS image with the same resolution as the PAN image is generated through the pre-fusion module; (2) Feature extraction and interaction module: The enhanced MS and PAN are input into this module to promote the learning of complementary features while suppressing redundant information; (3) Reconstruction module: The final fused high-resolution multispectral image is generated through the reconstruction module.
[0039] Based on the above embodiment, a pansharpening method for remote sensing images that fuses Mamba and CNN under the guidance of detail enhancement is proposed, that is, a pansharpening method for remote sensing images that fuses Mamba and Convolutional Neural Network (CNN) under the guidance of detail enhancement. Specifically, first, the texture feature information of LRMS is enhanced through the high-pass information of the PAN image; subsequently, local and global feature information of the PAN and the enhanced LRMS images are extracted by CNN and Mamba, and the extracted feature information is integrated into the CNN-based fusion branch level by level for full interaction and fusion of local and global information of the two modalities; finally, a high-quality HRMS image is reconstructed.
[0040] As Figures 1-5 shown, the embodiment is further elaborated as follows:
[0041] Overall framework
[0042] Multispectral images have rich spectral information, while panchromatic images have high-resolution spatial texture information. In view of the complementarity of information between multispectral images and panchromatic images, the present invention proposes a pansharpening method for remote sensing images that fuses Mamba and CNN under the guidance of detail enhancement. The method consists of three parts: a pre-fusion module, a feature extraction and interaction module, and a reconstruction module. The overall framework is as Figure 1 shown. Each module is elaborated in detail below:
[0043] (1) Pre - fusion module: A significant shortcoming of existing fusion methods is that they often overlook the image quality degradation problem caused by directly upsampling the LRMS image. To address this issue, the present invention introduces the information of the PAN image before fusion to enhance the LRMS image; then, by integrating the rich details in the high - resolution PAN image, the fine structures and texture information lost during the upsampling process of the LRMS image can be effectively compensated. In the pre - fusion module, a two - stage fusion strategy is adopted to gradually improve the quality of the LRMS image. First, the PAN image is downsampled by a factor of 4 and concatenated with the original LRMS image in the channel dimension to obtain the initial input F1; then, after extracting features from F1 through a 3×3 convolutional layer, the PixelShuffle operation is used to increase the spatial resolution, resulting in the feature map F2; subsequently, the obtained feature map F2 is fed into the residual connection block and the dense connection block in parallel to further capture and enhance the feature information, obtaining the feature maps F3 and F4 respectively; then, the feature maps F3 and F4 enhanced by the residual connection and the dense connection are concatenated, output, and restored to the number of channels of the original LRMS image after passing through a convolutional layer; finally, this output is added to the original LRMS image upsampled by a factor of 2, and the high - pass information after downsampling the PAN image by a factor of 2 is added to obtain the output feature map F of the first stage. out1 Specifically, the above process is expressed by the formula:
[0044]
[0045] where I M ∈R h×w×B represents the LRMS image of size h×w with B bands, and I P ∈R H×W×1 represents the PAN image of size H×W×1, where W / w = H / h = r, W and H are equal to 128, and r is equal to 4; the superscript ↑ and the subscript ↓ represent the upsampling and downsampling operations respectively; concat(*) represents the concatenation operation of features in the channel dimension, conv(*) represents the convolution operation, Pix(*) represents the sub - pixel convolution PixelShuffle operation, ResNet(*) represents the residual connection operation, DenseNet(*) represents the dense connection operation, and H(*) represents the operation of obtaining the high - pass information.
[0046] In the second - stage fusion, the output feature map F of the first stage out1It is concatenated with the PAN image downsampled by a factor of 2 in the channel dimension to obtain the feature map F5. Then, similar to the first stage, the feature map F5 is processed, i.e., feature extraction by a 3×3 convolutional layer and PixelShuffle operation to improve the spatial resolution, resulting in the feature map F6. Next, residual and dense connection processing are respectively performed on the feature map F6 to obtain the feature maps F7 and F8. Subsequently, after being concatenated, the feature maps F7 and F8 are added to the upsampled by a factor of 4 LRMS image and the high-pass filtered information of the PAN image, thereby obtaining the enhanced multispectral image, denoted as MS+. Specifically, the above process is expressed by the formula as follows:
[0047]
[0048] (2) Feature Extraction and Interaction Module: This module mainly adopts a hybrid model combining Mamba and CNN, rather than using Mamba or CNN alone. This hybrid structure fully exploits the advantages of the two models: CNN is good at extracting local features from the original image, while Mamba has a strong global feature extraction ability. In addition, the adaptive mechanism of Mamba also allows the model to dynamically adjust its parameters according to the characteristics of the input data, ensuring stability and flexibility in different application scenarios. Specifically, this module consists of three branches, namely the multispectral branch, the panchromatic branch, and the intermediate fusion branch. Briefly, by introducing the intermediate fusion branch, this module further realizes the complementarity and integration of local and global feature information. More specifically, the pre-fused MS+ and PAN images are first respectively passed through a shallow feature extraction block composed of 3×3 convolutions to capture their respective basic features, obtaining the feature maps F MS and F PAN ; Subsequently, the feature maps F MS and F PAN extracted by the shallow feature extraction block are subjected to global feature extraction through the Mamba block, thereby enhancing the understanding of complex scenes; Next, the outputs of the Mamba block are concatenated with the original MS+ and PAN images in the channel dimension and then input into the first feature fusion layer; Then, the output of the fusion layer is concatenated again with the outputs of the PAN branch and the MS branch respectively after being processed by the next-level Mamba block and then input into the next-level feature fusion layer; And so on, this process is carried out step by step to enable full information interaction and fusion between the MS+ image and the PAN image; Finally, the output of the last-level feature fusion layer is further concatenated with the output of the Mamba block, thereby completing the task of the entire feature extraction and interaction module and obtaining the final feature map F out2 . Specifically, the above process is expressed by the formula as follows:
[0049]
[0050] where F MSRepresents the shallow features extracted from the MS+ image, F PAN Represents the shallow features extracted from the PAN image, D n Represents the output of the nth fusion layer, where n takes values 1, 2, 3; Shallow(*) represents the shallow feature extraction operation, Fu[*] represents the fusion layer operation, M h (*) represents the hth Mamba block operation, where h takes values 1, 2, 3, 4.
[0051] (3) Reconstruction module: The reconstruction module receives the feature map F processed by the feature extraction and interaction module out2 , and thus reconstructs and generates the final high-resolution multispectral image I HRMS . This process is expressed by the formula:
[0052] I HRMS = Sig(conv(σ(conv(F out2 )))) (10)
[0053] where F out2 represents the output of the feature extraction and interaction module, σ(*) represents the activation function PReLU, and Sig(*) represents the Sigmoid activation function.
[0054] Loss function
[0055] In the method proposed in this patent, the total loss consists of three parts, namely content loss, gradient loss, and structural similarity loss. Among them, the content loss is expressed by the formula:
[0056] L int = ||I GT - M net (I M , I P )||1 (11)
[0057] where I GT represents the ground truth, Mnet(*) represents the fusion result of the LRMS and PAN images obtained by the method proposed in this patent, and ||*||1 represents the l1 norm.
[0058] The gradient loss is expressed by the formula:
[0059]
[0060] where mean(*) represents calculating the average value in the channel dimension. Represents the Sobel gradient operator.
[0061] The structural similarity loss is expressed by the formula:
[0062] L ssim = 1 - ssim(I GT , M net (I M , I P )) (13)
[0063] In summary, the total loss function of the method proposed in this patent is defined as:
[0064] L total = λ1L int + λ2L grad + λ3L ssim (14)
[0065] Among them, λ1, λ2, and λ3 are balance parameters that control each loss term. Specifically, in the method proposed in this patent, these parameters are set to: λ1 = 10, λ2 = 1, and λ3 = 10.
[0066] The present invention introduces the high-frequency information of the PAN image before fusion to enhance the LRMS image, and effectively compensates for the fine structure and texture information lost in the upsampling process of the LRMS image by integrating the rich details in the high-resolution PAN image.
[0067] The hybrid structure designed by the present invention makes full use of the global feature modeling ability of Mamba and the local feature extraction ability of CNN, and promotes the full interaction and fusion of spatial detail features and spectral information features through the intermediate fusion branch, so as to finally reconstruct and generate a high-quality HRMS image.
[0068] The above are all preferred embodiments of the present invention. For those of ordinary skill in the art, without departing from the principle of the present invention, the modifications of various equivalent forms of the present invention all fall within the protection scope of the appended claims of this application.
Claims
1. A panchromatic sharpening method for remote sensing images that integrates Mamba and CNN under the guidance of detail enhancement, characterized in that: The method is as follows: S1: First, enhance the texture feature information of the low-spatial-resolution multispectral image LRMS through the high-pass information of the panchromatic image PAN to obtain the enhanced LRMS image; S2: Subsequently, extract the local and global feature information of PAN and the enhanced LRMS image through the convolutional neural network CNN and Mamba, and gradually integrate the extracted feature information into the CNN-based fusion branch to facilitate the full interaction and fusion of local and global information of the two modalities; S3: Finally, reconstruct and generate a high-quality high-spatial-resolution multispectral image HRMS.
2. A panchromatic sharpening method for remote sensing images that fuses Mamba and CNN under the guidance of detail enhancement according to claim 1, characterized in that: The specific method in S1 is divided into two stages: The first stage: First, downsample the PAN image by a factor of 4 and concatenate it with the original LRMS image in the channel dimension to obtain the initial input F1; then, extract features from F1 through a 3×3 convolutional layer and use the PixelShuffle operation to enhance the spatial resolution to obtain the feature map F2; subsequently, input the obtained feature map F2 into the residual connection block and the dense connection block in parallel to further capture and enhance the feature information, obtaining the feature maps F3 and F4 respectively; then, concatenate and output the feature maps F3 and F4 enhanced by the residual connection and the dense connection, and restore to the number of channels of the original LRMS image after passing through a convolutional layer; finally, add this output to the original LRMS image upsampled by a factor of 2 and add the high-pass information after downsampling the PAN image by a factor of 2 to obtain the output feature map F of the first stage out1 Specifically, the above process is expressed by the formula as follows: where, I M ∈ R h×w×B represents an LRMS image of size h × w with B bands, and I P ∈ R H×W×1 represents a PAN image of size H × W × 1, where Ww = Hh = r, W and H are equal to 128, and r is equal to 4; the superscript ↑ and subscript ↓ represent upsampling and downsampling operations respectively; concat(*) represents the concatenation operation of features in the channel dimension, conv(*) represents the convolution operation, Pix(*) represents the PixelShuffle operation of sub-pixel convolution, ResNet(*) represents the residual connection operation, DenseNet(*) represents the dense connection operation, and H(*) represents the operation of obtaining high-pass information; The second stage: The output feature map F of the first stage out1 is concatenated with the PAN image downsampled by a factor of 2 in the channel dimension to obtain the feature map F5; then, similar to the first stage, the feature map F5 is processed, that is, feature extraction by a 3×3 convolutional layer and PixelShuffle operation to improve the spatial resolution to obtain the feature map F6; then, the residual and dense connection processing are respectively performed on the feature map F6 to obtain the feature maps F7 and F8; subsequently, after the feature maps F7 and F8 are concatenated, they are added to the high-pass filtered information of the LRMS image upsampled by a factor of 4 and the PAN image, so as to obtain the enhanced multispectral image, denoted as MS+; specifically, the above process is expressed by the formula as follows:
3. A panchromatic sharpening method for remote sensing images that fuses Mamba and CNN under the guidance of detail enhancement according to claim 1 or 2, characterized in that: The specific method of S2 is as follows: The pre-fused MS+ and PAN images are first respectively passed through a shallow feature extraction block composed of 3×3 convolutions to capture their respective basic features, and the feature maps F MS and F PAN are obtained respectively; Subsequently, the feature maps F MS and F PAN extracted by the shallow feature extraction block are subjected to global feature extraction through the Mamba block, so as to enhance the understanding of complex scenes; Next, the output of the Mamba block is concatenated with the original MS+ and PAN images in the channel dimension, and then input into the first feature fusion layer; Then, the output of the fusion layer is concatenated again with the outputs of the PAN branch and the MS branch respectively after being processed by the next-level Mamba block, and then input into the next-level feature fusion layer; And so on, this process is carried out step by step to achieve sufficient information interaction and fusion between the MS+ image and the PAN image; finally, the output of the last-level feature fusion layer is further connected to the output of the Mamba block, thereby completing the task of the entire feature extraction and interaction module and obtaining the final feature map F out2 Specifically, the above process is expressed by the formula as follows: Among them, F MS represents the shallow features extracted from the MS+ image, and F PAN represents the shallow features extracted from the PAN image. D n represents the output of the nth fusion layer, where n takes values of 1, 2, and 3 respectively; Shallow(*) represents the shallow feature extraction operation, Fu[*] represents the fusion layer operation, and M h (*) represents the operation of the hth Mamba block, where h takes values of 1, 2, 3, and 4 respectively.
4. A panchromatic sharpening method for remote sensing images that integrates Mamba and CNN under the guidance of detail enhancement according to claim 1, characterized in that: Specifically, in S3, the reconstruction module receives the feature map processed by the feature extraction and interaction module, thereby reconstructing and generating the final high-resolution multispectral image. Specifically, the operation process of the S3 method is expressed by the formula: I HRMS = Sig(conv(σ(conv(F out2 )))) (4) Among them, F out2 represents the output of the feature extraction and interaction module, σ(*) represents the activation function PReLU, and Sig(*) represents the Sigmoid activation function.
5. A panchromatic sharpening system for remote sensing images that integrates Mamba and CNN under the guidance of detail enhancement, characterized in that: Including: Pre-fusion module: Generate an enhanced MS image with the same resolution as the PAN image through the pre-fusion module; Feature extraction and interaction module: Input the enhanced MS and PAN into this module to promote the learning of complementary features while suppressing redundant information; Reconstruction module: Generate the final fused high-resolution multispectral image through the reconstruction module.
6. A panchromatic sharpening system for remote sensing images that integrates Mamba and CNN under the guidance of detail enhancement, as claimed in claim 5, wherein: The feature extraction and interaction module consists of three branches, namely the multispectral branch, the panchromatic branch, and the intermediate fusion branch. The intermediate fusion branch realizes the complementarity and integration of local and global feature information.
Citation Information
Cited By
Multispectral image panchromatic sharpening method based on plug-and-play gradient feature guidance fusion
CN120976059A
A multispectral image panchromatic sharpening method based on plug-and-play gradient feature-guided fusion
CN120976059B
Dual-domain fusion panchromatic sharpening method and system based on wavelet transform and Mama
CN121582099A