Three-source remote sensing image fusion classification method based on hybrid Mama network

By adopting a hybrid Mamba network in the multi-source remote sensing data fusion classification, a variety of Mamba module extraction and fusion features are designed, the problems of information redundancy and feature loss in the three-source data processing are solved, and high-precision land coverage classification is achieved.

CN120014367APending Publication Date: 2025-05-16LIAONING NORMAL UNIVERSITY
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510400708.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing multi-source remote sensing data fusion classification method has problems of information redundancy and key feature loss when processing three-source data, and it is difficult to make full use of the complementarity between multi-source data, resulting in limited classification performance.

Method used

Using a three-source remote sensing image fusion classification method based on hybrid Mamba network, the advantageous features of high-spectral, multi-spectral and radar data are extracted, and the deep interactive correlation between multi-modal features is explored.

Benefits of technology

It effectively reduces the redundant information in the feature fusion process, focuses on key variables, and significantly improves the accuracy and reliability of land cover classification. The overall accuracy of experimental results on Houston 2013 and Augsburg City datasets reached 92.13% and 63.02% respectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014367A_ABST
    Figure CN120014367A_ABST
Patent Text Reader

Abstract

The invention discloses a three-source remote sensing image fusion classification method based on a hybrid Mama network, and the method comprises the steps: extracting dominant features from different modes through a convolutional neural network and a plurality of Mama models, and synchronously modeling the spectrum sequence information of a hyperspectrum from the forward direction and the reverse direction through a bidirectional spectrum Mama module, the global spectrum dependency relationship and the local detail features in the hyperspectral image are captured; meanwhile, a central sensing Mama module is adopted to extract global spatial features of the multispectral image, and the core position of a central pixel in a remote sensing classification task is highlighted; finally, deep interaction and adaptive fusion of three-source data features are realized through a three-modal fusion Mamba module, the discrimination capability of joint features is significantly improved, reliable feature representation is provided for high-precision land cover classification, and the final classification effect is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of three-source remote sensing image processing, and in particular to a three-source remote sensing image fusion classification method based on a hybrid Mamba network, which can effectively mine and extract advantageous features between different modalities and explore deep interactive correlations between multimodal features, thereby achieving high classification accuracy. Background Art

[0002] In recent years, with the rapid progress of remote sensing technology and the rapid development of sensor technology, it has become possible to obtain multi-type and multi-temporal remote sensing data. Multi-source remote sensing data have different spatial resolutions, spectral characteristics and sensitivity to surface features, which enable them to provide multi-angle and multi-level information, providing more accurate and comprehensive information support for surface cover classification. Hyperspectral images provide spectral curves of each pixel in multiple bands, and their spectral resolution can usually reach the nanometer level, so that hyperspectral data can identify and classify objects with subtle spectral differences. Multispectral images provide spectral information of a few wide bands, and have weak ability to distinguish objects with subtle differences, but they have high spatial resolution and can provide clearer spatial details of objects. Radar remote sensing, as an active remote sensing technology, can penetrate clouds and other atmospheric interference and has the characteristics of working all day and all weather. Therefore, through multi-source data fusion, the complementary advantages of different data sources can be fully utilized, thereby significantly improving the accuracy and reliability of surface cover classification. Multi-source remote sensing data fusion can not only overcome the shortcomings of a single data source, but also provide a more comprehensive solution for object identification and dynamic monitoring in complex environments. It has become a prominent research focus in the field of land use / land cover classification research.

[0003] In the early research stage, multi-source remote sensing data fusion classification mainly relied on traditional algorithms. Gu et al. proposed a method for LiDAR and optical image fusion classification based on a multi-kernel learning model, integrating the designed multi-kernel learning model into feature fusion, seeking the best combination kernel, and then guiding the optimization of the classifier to achieve better classification performance. Xia et al. used random subspace to divide the extracted hyperspectral and LiDAR features into several non-overlapping subsets, and then independently transformed these subsets to provide more reliable feature support for the final classification. Rasti et al. used sparse low-rank technology to generate low-rank fusion features from the acquired hyperspectral and LiDAR spatial and elevation features, and sent them to the classifier to obtain classification results. However, the performance of these traditional methods is often limited by the ability to manually design features, and they perform poorly when faced with complex nonlinear data, making it difficult to fully explore the complex relationships between multi-source data.

[0004] In recent years, the rapid development of deep learning technology has brought new ideas to the fusion classification of multi-source remote sensing images. Roy et al. proposed a joint learning mechanism based on CNN and morphological features for the fusion classification of HS and LiDAR images. 3D CNN and morphological dilation and erosion were used to extract the spatial-spectral joint features of hyperspectral and the elevation features of LiDAR data respectively. Xu et al. developed a three-source fusion classification framework (called Fusion-CNN). The network has three branches corresponding to HS, LiDAR and ultra-high resolution images, and finally adds the extracted three-source features to generate fusion features. However, the inherent convolution kernel size of CNN makes it have certain limitations in modeling long-range dependencies. Therefore, the introduction of Transformer can capture the contextual information between distant pixels through its self-attention mechanism. Gao et al. developed a Transformer based on cross-scale mixed attention for the fusion classification of HS and MS. This method uses a spectral spatial mixer to extract the features of HS and MS, and integrates scale feature calibration in the process. Yang et al. integrated the cross-modal attention mechanism and spectral self-attention mechanism in the modal fusion visual Transformer, reducing the alignment requirements of heterogeneous features of HS and LiDAR. Song et al. designed a multi-head complementary attention mechanism in the multi-source Transformer complementor to extract local texture features and global features of the image and achieve comprehensive feature fusion. However, the computational complexity of the Transformer is proportional to the square of the input sequence length. This high computational complexity makes it face a significant efficiency bottleneck when processing large-scale image data. Recently, a new deep learning model Mamba has begun to attract widespread attention. It significantly reduces the computational complexity by introducing a state space model and a selective scanning mechanism, while maintaining a strong sequence modeling capability. Liao et al. proposed a Mamba-based network for HS and LiDAR fusion classification, and explored the relationship between the two modalities by constructing a multimodal Mamba fusion module. Zhang et al. designed Cross-SSM in the spatial Mamba and spectral Mamba modules and combined the reverse bottleneck structure to fuse bimodal HS and LiDAR features. Ma et al. developed a cross-modal spatial-spectral interaction Mamba network for MS and panchromatic image fusion classification, and captured global spatial-spectral features from different dimensions by designing a multi-path selective scanning mechanism. However, the Mamba model has just emerged, and its potential in multi-source remote sensing image fusion classification tasks needs to be further explored.

[0005] In general, although the multi-source remote sensing image fusion classification technology has been developed to a certain extent in recent years, the existing methods still have the following two problems:

[0006] (1) Faced with various types of available remote sensing data, most of the current fusion classification research focuses on two sources of remote sensing data, and there are very few fusion classification studies on more remote sensing data, which greatly limits the full mining and utilization of potential complementary information between multi-source remote sensing data. Especially when facing surface scenes with high heterogeneity and complex environmental characteristics, the fusion method of single or two-source data often cannot meet application requirements.

[0007] (2) The existing dual-source and very few triple-source remote sensing data fusion classification methods often face the problems of information redundancy and loss of key features in the process of extracting joint features. Especially when triple-source data is involved, the dimension of the features is further expanded. While the information richness increases, it also brings additional redundancy, resulting in the complementarity between features not being fully utilized, which weakens the classification performance to a certain extent. Summary of the invention

[0008] The present invention aims to solve the above-mentioned technical problems existing in the prior art and to provide a three-source remote sensing image fusion classification method which can effectively mine and extract advantageous features between different modalities and explore the deep interactive correlations between multimodal features, thereby achieving high classification accuracy.

[0009] The technical solution of the present invention is: a three-source remote sensing image fusion classification method based on a hybrid Mamba network, which is carried out in the following steps:

[0010] Step 1. Create and initialize the hybrid Mamba network N HMF , the N HMF Contains 4 sub-networks N for feature extraction HSPA 、N HSPE 、N MSPA and N RSPA , 1 sub-network N for feature fusion TF and 1 subnetwork N for classification CLS ;

[0011] Step 2. Input the training set H of hyperspectral images, the training set M of multispectral images, the training set R of radar images, the manually annotated pixel point coordinate set and label set, and HMF Conduct training;

[0012] Step 3 Input the hyperspectral image H′, multispectral image M′ and radar image R¢ to be tested, and use the trained network N HMF Complete pixel classification.

[0013] The step 1 is specifically as follows:

[0014] Step 1.1 Create and initialize subnetwork N HSPA , the sub-network NHSPA It consists of three groups of 2d convolution modules, namely Conv1_0, Conv1_1, and Conv1_2;

[0015] Conv1_0 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and a nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation;

[0016] Conv1_1 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation;

[0017] The Conv1_2 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation, 1 layer of activation operation and 1 layer of Dropout operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and selects the nonlinear activation function LeakyReLU with a parameter of 0.01 as the activation function and the Dropout with a parameter of 0.5 for operation;

[0018] Step 1.2 Create and initialize subnetwork N HSPE , the sub-network N HSPE It consists of 1 bidirectional spectrum Mamba module BSM;

[0019] The BSM performs the following 7 steps:

[0020] (a) The original hyperspectral image is input into the global average pooling layer GAP to obtain the feature Among them, B is the batch size of the input network, C1 is the number of channels of the hyperspectral image;

[0021] (b) Use reshape operation to change feature F gap The shape of

[0022] (c) For feature F g ' ap Perform projection transformation operation, and use 2d convolution layer Conv2_0 to perform feature mapping to obtain feature F pro , where the size of the convolution kernel in the convolution layer is 1×1, and each convolution kernel performs convolution operation with a step size of 1 pixel;

[0023] (d) For feature F pro Perform position encoding and then flip to get feature F p ' ro ;

[0024] (e) The features F pro and F p ' ro The input parameter is shared in the spectrum Mamba module to obtain the feature F s and F s ′, where the spectral Mamba module consists of 1 Linear2_0 layer, 1 1d convolution layer Conv2_1, 1 activation function layer using nonlinear activation function SiLU, 1 S6_2_0 layer, 1 Linear2_1 layer, 1 activation function layer using nonlinear activation function SiLU, 1 element-wise multiplication operation and 1 Linear2_2 layer;

[0025] (f) The feature F s and F s ′ is added element by element, and the average operation is performed along the second dimension to obtain the feature F m ;

[0026] (g) The feature F m The input is processed in the 2d convolution layer Conv2_2 to obtain the output feature F of the BSM module bsm , where the size of the convolution kernel in the convolution layer is 1×1, and each convolution kernel performs convolution operation with a step size of 1 pixel;

[0027] Step 1.3 Create and initialize subnetwork N MSPA , the sub-network N MSPA It consists of 2 center-aware Mamba modules, 1 residual connection, 3 groups of 2d convolution modules and 1 cross-network residual connection. The 2 center-aware Mamba modules are CAM_1 and CAM_2, and the 3 groups of 2d convolution modules are Conv3_0, Conv3_1, and Conv3_2.

[0028] The CAM_1 performs the following five steps:

[0029] (a1) The original multispectral image MS is input into the projection transformation layer to obtain the feature F l , where the projection transformation layer contains one 2d convolution layer Conv3_3 and one normalization layer LayerNorm3_0. The convolution kernel size of Conv3_3 is 1×1, and each convolution kernel performs convolution operation with a step size of 1 pixel;

[0030] (b1) For feature F lPerform dual position encoding, then perform a Dropout operation with a parameter of 0.01, and then input it into the normalization layer LayerNorm3_1 to obtain the feature F n ;

[0031] (c1) The feature F n Input into the linear layer Linear3_0, and use the nonlinear activation function SiLU to obtain the gated branch feature F gate ;

[0032] (d1) The feature F n The input is sent to the linear layer Linear3_1, and then processed by the depthwise separable convolution layer DWConv3_0 and the nonlinear activation function SiLU to obtain the feature F dw , where the size of the DWConv3_0 convolution kernel is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the feature F dw It is sent to the four-way central spiral scanning FDCSS_0 for feature extraction, where FDCSS_0 starts from the upper left, lower left, lower right and upper right pixels of the current patch, and scans the entire image patch from the outside to the inside in a clockwise spiral order, and forms four Token sequences S1, S2, S3 and S4 with different arrangement orders. Then the four different Token sequences are input into the S6_3_0, S6_3_1, S6_3_2 and S6_3_3 models for processing, and the sequences are reorganized into image features F according to the original arrangement order. s1 、F s2 、F s3 and F s4 , the output feature F of the FDCSS_0 module is calculated according to formula (1) fdc ,in and is the parameter automatically learned by the network, and then the feature F fdc The input normalization layer LayerNorm3_2 is processed to obtain the main branch feature F main ;

[0033]

[0034] (e1) According to formula (2), the gated branch feature F is integrated gate and the main branch feature F main Get feature F fu , then the feature F fu The input is sent to the linear layer Linear3_2 and processed by a 2D convolution layer Conv3_4 to obtain the output feature F of the CAM_1 module. c1 , where the convolution kernel size of Conv3_4 is 1×1, and each convolution kernel performs convolution operation with a step size of 1 pixel;

[0035] F fu =F gate ×F main (2)

[0036] The CAM_2 performs the following five steps:

[0037] (a2) The output feature F of CAM_1 module c1 Input projection transformation layer processing to obtain feature F l ′, where the projection transformation layer contains one 2D convolution layer Conv3_5 and one normalization layer LayerNorm3_3. The convolution kernel size of Conv3_5 is 1×1, and each convolution kernel performs convolution operation with a step size of 1 pixel.

[0038] (b2) For feature F l ′ performs dual position encoding, then performs a Dropout operation with a parameter of 0.01, and then inputs the normalization layer LayerNorm3_4 to obtain the feature F n ′;

[0039] (c2) The feature F n ′ is input into the linear layer Linear3_3, and processed by the nonlinear activation function SiLU to obtain the gated branch feature F g ' ate ;

[0040] (d2) The feature F n ′ is input into the linear layer Linear3_4, and then processed by the depthwise separable convolution layer DWConv3_1 and the nonlinear activation function SiLU to obtain the feature F d ' w , where the size of the DWConv3_1 convolution kernel is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the feature F d ' w It is sent to the four-way central spiral scanning FDCSS_1 for feature extraction, where FDCSS_1 starts from the upper left, lower left, lower right and upper right pixels of the current patch, scans the entire image patch from outside to inside in a clockwise spiral order, and forms four Token sequences S1′, S2′, S3′ and S4′ with different arrangement orders. Then the four different Token sequences are input into the S6_3_4, S6_3_5, S6_3_6 and S6_3_7 models for processing, and the sequences are reorganized into image features F according to the original arrangement order. s ′1, F s ′2, F s ′3 and F s′4, according to formula (3), the output feature F of the FDCSS_1 module is calculated f ' dc ,in and is the parameter automatically learned by the network, and then the feature F f ' dc The input normalization layer LayerNorm3_5 is processed to obtain the main branch feature F m ' ain ;

[0041]

[0042] (e2) According to formula (4), the gated branch feature F is fused g ' ate and the main branch feature F m ' ain Get feature F f ' u , then the feature F f ' u The input is sent to the linear layer Linear3_5 and processed by a 2D convolution layer Conv3_6 to obtain the output feature F of the CAM_2 module. c2 , where the convolution kernel size of Conv3_6 is 1×1, and each convolution kernel performs convolution operation with a step size of 1 pixel;

[0043] F f ' u =F g ' ate ×F m ' ain (4)

[0044] The residual connection is calculated according to formula (5), where θ1 and θ2 are parameters automatically learned by the network;

[0045] F cam =θ1×MS+θ2×F c2 (5)

[0046] The Conv3_0 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation;

[0047] The Conv3_1 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation;

[0048] The Conv3_2 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation, 1 layer of activation operation and 1 layer of Dropout operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and selects the nonlinear activation function LeakyReLU with a parameter of 0.01 as the activation function and the Dropout with a parameter of 0.5 for operation;

[0049] The cross-network residual connection is calculated according to formula (6), where θ3 and θ4 are parameters automatically learned by the network, and F conv3_2 It is the output feature of convolution module Conv3_2;

[0050] F cres =θ3×F bsm +θ4×F conv3_2 (6)

[0051] Step 1.4 Create and initialize subnetwork N RSPA , the sub-network N RSPA It consists of three groups of 2d convolution modules, namely Conv4_0, Conv4_1, and Conv4_2;

[0052] The Conv4_0 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation;

[0053] The Conv4_1 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation;

[0054] The Conv4_2 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation, 1 layer of activation operation and 1 layer of Dropout operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and selects the nonlinear activation function LeakyReLU with a parameter of 0.01 as the activation function and the Dropout with a parameter of 0.5 for operation;

[0055] Step 1.5 Establish and initialize the sub-network N for feature fusion TF , the sub-network N TF It contains a tri-modal fusion Mamba module TFM, which consists of a projection transformation layer, 3 visual Mamba modules, an element-by-element summation operation and a 2d convolution layer Conv5_0. The projection transformation layer contains a 2d convolution layer Conv5_1 and a normalization layer LayerNorm5_0. The convolution kernel sizes of Conv5_0 and Conv5_ are both 1×1, and each convolution kernel performs convolution operations with a step size of 1 pixel. The three visual Mamba modules are ViM_1, ViM_2 and ViM_3.

[0056] The sub-network N TF Follow the 3 steps below:

[0057] (a) The hyperspectral feature F hspa , multi-spectral feature F mspa and radar signature F rspa The input projection transformation layer processes the feature F h ' spa 、F m ' spa and F r ' spa ;

[0058] (b) The feature F h ' spa 、F m ' spa and F r ' spa The input is processed in three visual Mamba modules ViM_1, ViM_2 and ViM_3;

[0059] The ViM_1, ViM_2 and ViM_3 simultaneously perform the following 8 steps:

[0060] (b1) The feature F h ' spa 、F m ' spa and F r ' spaThe normalization layers LayerNorm5_1, LayerNorm5_2, and LayerNorm5_3 are input for processing, and then the position encoding is added to obtain the feature F. h 、F m and F r ;

[0061] (b2) The feature F h 、F m and F r Input into normalization layers LayerNorm5_4, LayerNorm5_5 and LayerNorm5_6 respectively to obtain feature F hl 、F ml and F rl ;

[0062] (b3) The feature F hl Input into the linear layer Linear5_0, and use the nonlinear activation function SiLU to obtain the gated branch feature F of ViM_1 h_gate , the feature F ml Input into the linear layer Linear5_1, and use the nonlinear activation function SiLU to obtain the gated branch feature F of ViM_2 m_gate , the feature F rl Input into the linear layer Linear5_2, and use the nonlinear activation function SiLU to obtain the gated branch feature F of ViM_3 r_gate ;

[0063] (b4) The feature F hl The input is sent to the linear layer Linear5_3, and then processed by the depthwise separable convolution layer DWConv5_0 and the nonlinear activation function SiLU to obtain the feature F hdw , the feature F ml The input is sent to the linear layer Linear5_4, and then processed by the depthwise separable convolution layer DWConv5_1 and the nonlinear activation function SiLU to obtain the feature F mdw , the feature F rl The input is sent to the linear layer Linear5_5, and then processed by the depthwise separable convolution layer DWConv5_2 and the nonlinear activation function SiLU to obtain the feature F rdw , where the convolution kernel size of DWConv5_0, DWConv5_1 and DWConv5_2 is 3×3, and each convolution kernel performs convolution operation with a step size of 1 pixel;

[0064] (b5) The feature F hdw 、F mdw and F rdwThe input is sent to the cross-modal selective scanning module CMSS for deep modeling. CMSS converts the feature F hdw 、F mdw and F rdw The three modal pixels at each position {(i,j)|i=(0,1,···,H-1),j=(0,1,···,W-1)} are regarded as a scan path. When the image size is H×W, there are H×W scan paths. According to formula (7), H×W token sequences are generated, and each sequence contains only three pixels. and They come from the same position of different pixels, where D is the feature dimension after projection transformation; then the H×W Token sequences are input into the corresponding S6 k=0,1,...,H×W The feature modeling is performed in the output sequence, and finally the image reorganization operation is performed on the output sequence, and all pixels in the same order in the output sequence are combined into an image to obtain the output feature F of CMSS. hc 、F mc and F rc ;

[0065]

[0066] (b6) The feature F hc Input normalization layer LayerNorm5_7 to obtain the main branch feature F of ViM_1 h_main , the feature F mc Input normalization layer LayerNorm5_8 to obtain the main branch feature F of ViM_2 m_main , the feature F rc Input normalization layer LayerNorm5_9 to obtain the main branch feature F of ViM_3 r_main ;

[0067] (b7) Calculate the fusion features F of ViM_1, ViM_2 and ViM_3 according to formulas (8), (9) and (10) respectively V1 、F V2 and F V3 ;

[0068] F V1 =F h_gate ×F h_main (8)

[0069] F V2 =F m_gate ×F m_main (9)

[0070] F V3 =F r_gate ×F r_main(10)

[0071] (b8) The feature F V1 、F V2 and F V3 Input into linear layers Linear5_6, Linear5_7 and Linear5_8 respectively to obtain feature F V ′1, F V ′2 and F V ′3, and finally the output features F of ViM_1, ViM_2 and ViM_3 are calculated according to formulas (11), (12), (13) V1_out 、F V2_out and F V3_out ;

[0072] F V1_out =F V ′1+F h (11)

[0073] F V2_out =F V ′2+F m (12)

[0074] F V3_out =F V ′3+F r (13)

[0075] (c) According to formula (14), the fusion feature F V1_out 、F V2_out and F V3_out And get the sub-network N TF Output:

[0076]

[0077] Step 1.6 Create and initialize subnetwork N CLS , the sub-network N CLS It consists of two groups of convolution modules, namely Conv6_0 and Conv6_1;

[0078] The Conv6_0 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 1×1, each convolution kernel performs convolution operation with a step size of 1 pixel, and the nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation;

[0079] The Conv6_1 includes 1 AdaptiveAvgPool pooling operation, 1 layer of 2d convolution operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 1×1, each convolution kernel performs convolution operation with a step size of 1 pixel, and selects Softmax as the activation function for operation.

[0080] The step 2 is specifically as follows:

[0081] Step 2.1. Extract all pixel points with labels from the training set H of the hyperspectral image based on the manually labeled pixel point coordinates. Extract all pixel points with labels from the training set M of multispectral images Extract all pixel points with labels from the radar image training set R in, Represents X H The i1th pixel point in Represents X M The i1th pixel point in Represents X R The i1th pixel point in , N represents the total number of pixel points with labels;

[0082] Step 2.2. Use X H Each pixel point is taken as the center to divide H into a series of hyperspectral pixel blocks of size P′P X M Each pixel point is taken as the center to divide M into a series of multispectral pixel blocks of size P′P X R Each pixel point is the center and R is divided into a series of radar pixel blocks of size P′P.

[0083] Will and As the training set of the fusion classification neural network, the samples in the training set are integrated into tuples The form of network input is Represents the pixel pairs composed of hyperspectral images, multispectral images and radar images in the training set, which have the same spatial coordinates. express and Corresponding true category label, set the number of iterations iter←1, and execute steps 2.3 to 2.7;

[0084] Step 2.3. Use subnetwork N HSPA 、N HSPE 、N MSPA and N RSPA Extract features of the training set;

[0085] Step 2.3.1 Use subnetwork N HSPA Training set for hyperspectral images Perform feature extraction to obtain the spatial features F of the hyperspectral image hspa ;

[0086] Step 2.3.2 Use subnetwork N HSPE Training set for hyperspectral images Perform feature extraction to obtain the spectral feature F of the hyperspectral image hspe ;

[0087] Step 2.3.3 Use subnetwork N MSPA Training set for multispectral images Perform feature extraction to obtain the spatial features F of the multispectral image mspa ;

[0088] Step 2.3.4 Use subnetwork N RSPA Training set for radar images Perform feature extraction to obtain the spatial feature F of the radar image rspa ;

[0089] Step 2.4. Use subnetwork N TF Fusion feature F hspa 、F mspa and F rspa Get joint features

[0090] Step 2.5. Use subnetwork N CLS Joint Features Perform classification and calculate the classification prediction result TR pred ;

[0091] Step 2.6. Calculate the cross entropy loss function according to the definition of formula (14);

[0092]

[0093] Among them, C represents the number of categories, p n,c is the predicted probability, y n,c is the probability vector corresponding to the label;

[0094] In step 2.7, if all pixel blocks in the training set have been processed, proceed to step 2.8; otherwise, take a group of unprocessed pixel blocks from the training set and return to step 2.3;

[0095] Step 2.8 Let iter←iter+1. If the number of iterations iter>Total_iter, the trained hybrid Mamba network N is obtained.HMF , go to step 3, otherwise, use the reverse error propagation algorithm based on stochastic gradient descent and the prediction loss L CE Update N HMF , go to step 2.3 to reprocess all pixel blocks in the training set, and Total_iter represents the preset number of iterations.

[0096] The step 3 is as follows:

[0097] Step 3.1. Extract all pixel points in H′ to form a set Extract all pixel points in M′ to form a set Extract all pixel points in R¢ to form a set in, Indicates T H The i2th pixel of Indicates T M The i2th pixel of Indicates T R The i2th pixel of , U represents the total number of all pixels;

[0098] Step 3.2. Take T H Each pixel point is taken as the center to divide H′ into a series of hyperspectral pixel blocks of size P′P, forming a hyperspectral image set to be tested. T M Each pixel point is taken as the center to divide M′ into a series of multispectral pixel blocks of size P′P, forming a multispectral image set to be tested T R Each pixel point of is taken as the center to divide R¢ into a series of radar pixel blocks of size P′P, forming the radar image set to be tested

[0099] Step 3.3. Use the fully trained sub-network N HSPA 、N HSPE 、N MSPA and N RSPA Extract the features of the image to be tested;

[0100] Step 3.3.1 Use the fully trained sub-network N HSPA The hyperspectral image collection to be tested Perform feature extraction to obtain the spatial features F of the hyperspectral image h ' spa ;

[0101] Step 3.3.2 Use the fully trained sub-network N HSPE The hyperspectral image collection to be tested Perform feature extraction to obtain the spectral feature F of the hyperspectral image h' spe ;

[0102] Step 3.3.3 Use the fully trained sub-network N MSPA The multispectral image collection to be tested Perform feature extraction to obtain the spatial features F of the multispectral image m ' spa ;

[0103] Step 3.3.4 Use the fully trained sub-network N RSPA The radar image collection to be tested Perform feature extraction to obtain the spatial feature F of the radar image r ' spa ;

[0104] Step 3.4. Use the fully trained sub-network N TF Fusion feature F h ' spa 、F m ' spa and F r ' spa Get joint features

[0105] Step 3.5. Use the fully trained sub-network N CLS Joint Features Classify and calculate the classification prediction result TE pred .

[0106] Compared with the prior art, the present invention improves the accuracy of multi-source remote sensing data fusion classification at the two levels of feature extraction and feature fusion, which is reflected in the following three aspects: First, a bidirectional spectral Mamba module is proposed for the dominant spectral features of hyperspectral images. By synchronously modeling the spectral sequence information of the hyperspectral images from the forward and reverse directions, the integrity of the spectral information and the complementarity of the bidirectional features are ensured, and the long-range dependencies and spectral variation laws in the hyperspectral data are fully captured; Second, a center-aware Mamba module is proposed to extract the dominant spatial features of multispectral images. The core position of the center pixel in the remote sensing classification task is highlighted through the designed four-way center spiral scanning mechanism, so that the model can more effectively capture the correlation between the center pixel and its surrounding context and enhance the feature expression ability of the center pixel; Third, a trimodal fusion Mamba module is proposed. The scanning path is extended from a single modality to multiple modalities through the designed cross-modal selective scanning mechanism, so that the multimodal information is incorporated into the framework of joint modeling in the generation stage, and the efficient complementarity and enhancement of multimodal information are achieved. Therefore, the present invention can effectively extract the advantageous features of different data sources, optimize the feature fusion strategy, reduce the redundant information of the fusion process, and focus on key variables, and has the characteristics of good feature extraction and fusion quality and high detection accuracy. The experimental results show that the overall accuracy of the present invention on the Houston2013 and Augsburg City data sets reached 92.13% and 63.02% respectively, effectively improving the accuracy of fusion classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0107] Figure 1 is a flow chart of an embodiment of the present invention.

[0108] Figure 2 Flow chart of the bidirectional spectral Mamba module in an embodiment of the present invention.

[0109] Figure 3 This is a flow chart of the central perception Mamba module CAM_1 in an embodiment of the present invention.

[0110] Figure 4 This is a flow chart of the central perception Mamba module CAM_2 in an embodiment of the present invention.

[0111] Figure 5 Flow chart of the tri-modal fusion Mamba module in an embodiment of the present invention.

[0112] Figure 6 This is a comparison chart of the classification results of the Houston2013 dataset between the embodiment of the present invention and the AM3Net method, TBCNN method, MFT method, MDAS method, Fusion-CNN method and S2FL method.

[0113] Figure 7This is a comparison chart of the classification results of the Augsburg City data set between the embodiment of the present invention and the AM3Net method, TBCNN method, MFT method, MDAS method, Fusion-CNN method and S2FL method. DETAILED DESCRIPTION

[0114] A three-source remote sensing image fusion classification method based on a hybrid Mamba network of the present invention is as follows Figure 1 As shown, proceed as follows:

[0115] Step 1. Create and initialize the hybrid Mamba network N HMF , the N HMF Contains 4 sub-networks N for feature extraction HSPA 、N HSPE 、N MSPA and N RSPA , 1 sub-network N for feature fusion TF and 1 subnetwork N for classification CLS ;

[0116] Step 1.1 Create and initialize subnetwork N HSPA , the sub-network N HSPA It consists of three groups of 2d convolution modules, namely Conv1_0, Conv1_1, and Conv1_2;

[0117] Conv1_0 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and a nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation;

[0118] Conv1_1 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation;

[0119] The Conv1_2 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation, 1 layer of activation operation and 1 layer of Dropout operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and selects the nonlinear activation function LeakyReLU with a parameter of 0.01 as the activation function and the Dropout with a parameter of 0.5 for operation;

[0120] Step 1.2 Create and initialize subnetwork N HSPE , the sub-network N HSPE Depend on Figure 2 The shown one bidirectional spectrum Mamba module BSM is composed;

[0121] The BSM performs the following 7 steps:

[0122] (a) The original hyperspectral image is input into the global average pooling layer GAP to obtain the feature Among them, B is the batch size of the input network, C1 is the number of channels of the hyperspectral image;

[0123] (b) Use reshape operation to change feature F gap The shape of

[0124] (c) For feature F g ' ap Perform projection transformation operation, and use 2d convolution layer Conv2_0 to perform feature mapping to obtain feature F pro , where the size of the convolution kernel in the convolution layer is 1×1, and each convolution kernel performs convolution operation with a step size of 1 pixel;

[0125] (d) For feature F pro Perform position encoding and then flip to get feature F p ' ro ;

[0126] (e) The features F pro and F p ' ro The input parameter is shared in the spectrum Mamba module to obtain the feature F s and F s ′, where the spectral Mamba module consists of 1 Linear2_0 layer, 1 1d convolution layer Conv2_1, 1 activation function layer using nonlinear activation function SiLU, 1 S6_2_0 layer, 1 Linear2_1 layer, 1 activation function layer using nonlinear activation function SiLU, 1 element-wise multiplication operation and 1 Linear2_2 layer;

[0127] (f) The feature F s and F s ′ is added element by element, and the average operation is performed along the second dimension to obtain the feature F m ;

[0128] (g) The feature F m The input is processed in the 2d convolution layer Conv2_2 to obtain the output feature F of the BSM module bsm, where the size of the convolution kernel in the convolution layer is 1×1, and each convolution kernel performs convolution operation with a step size of 1 pixel;

[0129] Step 1.3 Create and initialize subnetwork N MSPA , the sub-network N MSPA It consists of 2 center-aware Mamba modules, 1 residual connection, 3 groups of 2d convolution modules and 1 cross-network residual connection. The 2 center-aware Mamba modules are CAM_1 and CAM_2, and the 3 groups of 2d convolution modules are Conv3_0, Conv3_1, and Conv3_2.

[0130] The CAM_1 is as Figure 3 Perform the following 5 steps as shown:

[0131] (a1) The original multispectral image MS is input into the projection transformation layer to obtain the feature F l , where the projection transformation layer contains one 2d convolution layer Conv3_3 and one normalization layer LayerNorm3_0. The convolution kernel size of Conv3_3 is 1×1, and each convolution kernel performs convolution operation with a step size of 1 pixel;

[0132] (b1) For feature F l Perform dual position encoding, then perform a Dropout operation with a parameter of 0.01, and then input it into the normalization layer LayerNorm3_1 to obtain the feature F n ;

[0133] (c1) The feature F n Input into the linear layer Linear3_0, and use the nonlinear activation function SiLU to obtain the gated branch feature F gate ;

[0134] (d1) The feature F n The input is sent to the linear layer Linear3_1, and then processed by the depthwise separable convolution layer DWConv3_0 and the nonlinear activation function SiLU to obtain the feature F dw , where the size of the DWConv3_0 convolution kernel is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the feature F dwIt is sent to the four-way central spiral scanning FDCSS_0 for feature extraction, where FDCSS_0 starts from the upper left, lower left, lower right and upper right pixels of the current patch, and scans the entire image patch from the outside to the inside in a clockwise spiral order, and forms four Token sequences S1, S2, S3 and S4 with different arrangement orders. Then the four different Token sequences are input into the S6_3_0, S6_3_1, S6_3_2 and S6_3_3 models for processing, and the sequences are reorganized into image features F according to the original arrangement order. s1 、F s2 、F s3 and F s4 According to formula (1), the output feature F of the FDCSS_0 module is calculated fdc ,in and is the parameter automatically learned by the network, and then the feature F fdc The input normalization layer LayerNorm3_2 is processed to obtain the main branch feature F main ;

[0135]

[0136] (e1) According to formula (2), the gated branch feature F is integrated gate and the main branch feature F main Get feature F fu , then the feature F fu The input is sent to the linear layer Linear3_2 and processed by a 2D convolution layer Conv3_4 to obtain the output feature F of the CAM_1 module. c1 , where the convolution kernel size of Conv3_4 is 1×1, and each convolution kernel performs convolution operation with a step size of 1 pixel;

[0137] F fu =F gate ×F main (2)

[0138] The CAM_2 is as Figure 4 Perform the following 5 steps as shown:

[0139] (a2) The output feature F of CAM_1 module c1 Input projection transformation layer processing to obtain feature F l ′, where the projection transformation layer contains one 2D convolution layer Conv3_5 and one normalization layer LayerNorm3_3. The convolution kernel size of Conv3_5 is 1×1, and each convolution kernel performs convolution operation with a step size of 1 pixel.

[0140] (b2) For feature F l′ performs dual position encoding, then performs a Dropout operation with a parameter of 0.01, and then inputs the normalization layer LayerNorm3_4 to obtain the feature F n ′;

[0141] (c2) The feature F n ′ is input into the linear layer Linear3_3, and processed by the nonlinear activation function SiLU to obtain the gated branch feature F g ' ate ;

[0142] (d2) The feature F n ′ is input into the linear layer Linear3_4, and then processed by the depthwise separable convolution layer DWConv3_1 and the nonlinear activation function SiLU to obtain the feature F d ' w , where the size of the DWConv3_1 convolution kernel is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the feature F d ' w It is sent to the four-way central spiral scanning FDCSS_1 for feature extraction, where FDCSS_1 starts from the upper left, lower left, lower right and upper right pixels of the current patch, scans the entire image patch from outside to inside in a clockwise spiral order, and forms four Token sequences S1′, S2′, S3′ and S4′ with different arrangement orders. Then the four different Token sequences are input into the S6_3_4, S6_3_5, S6_3_6 and S6_3_7 models for processing, and the sequences are reorganized into image features F according to the original arrangement order. s ′1, F s ′2, F s ′3 and F s ′4, according to formula (3), the output feature F of the FDCSS_1 module is calculated f ' dc ,in and is the parameter automatically learned by the network, and then the feature F f ' dc The input normalization layer LayerNorm3_5 is processed to obtain the main branch feature F m ' ain ;

[0143]

[0144] (e2) According to formula (4), the gated branch feature F is fused g ' ate and the main branch feature F m ' ain Get feature F f' u , then the feature F f ' u The input is sent to the linear layer Linear3_5 and processed by a 2D convolution layer Conv3_6 to obtain the output feature F of the CAM_2 module. c2 , where the convolution kernel size of Conv3_6 is 1×1, and each convolution kernel performs convolution operation with a step size of 1 pixel;

[0145] F f ' u =F g ' ate ×F m ' ain (4)

[0146] The residual connection is calculated according to formula (5), where θ1 and θ2 are parameters automatically learned by the network;

[0147] F cam =θ1×MS+θ2×F c2 (5)

[0148] The Conv3_0 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation;

[0149] The Conv3_1 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation;

[0150] The Conv3_2 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation, 1 layer of activation operation and 1 layer of Dropout operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and selects the nonlinear activation function LeakyReLU with a parameter of 0.01 as the activation function and the Dropout with a parameter of 0.5 for operation;

[0151] The cross-network residual connection is calculated according to formula (6), where θ3 and θ4 are parameters automatically learned by the network, and F conv3_2 It is the output feature of convolution module Conv3_2;

[0152] Fcres =θ3×F bsm +θ4×F conv3_2 (6)

[0153] Step 1.4 Create and initialize subnetwork N RSPA , the sub-network N RSPA It consists of three groups of 2d convolution modules, namely Conv4_0, Conv4_1, and Conv4_2;

[0154] The Conv4_0 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation;

[0155] The Conv4_1 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation;

[0156] The Conv4_2 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation, 1 layer of activation operation and 1 layer of Dropout operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and selects the nonlinear activation function LeakyReLU with a parameter of 0.01 as the activation function and the Dropout with a parameter of 0.5 for operation;

[0157] Step 1.5 Establish and initialize the sub-network N for feature fusion TF , the sub-network N TF like Figure 5 As shown, it contains a tri-modal fusion Mamba module TFM, which consists of a projection transformation layer, 3 visual Mamba modules, an element-by-element summation operation and a 2d convolution layer Conv5_0. The projection transformation layer contains a 2d convolution layer Conv5_1 and a normalization layer LayerNorm5_0. The convolution kernel sizes of Conv5_0 and Conv5_ are both 1×1, and each convolution kernel performs convolution operations with a step size of 1 pixel. The three visual Mamba modules are ViM_1, ViM_2 and ViM_3 respectively.

[0158] The sub-network N TF Follow the 3 steps below:

[0159] (a) The hyperspectral feature F hspa , multi-spectral feature F mspa and radar signature F rspa The input projection transformation layer processes the feature F h ' spa 、F m ' spa and F r ' spa ;

[0160] (b) The feature F h ' spa 、F m ' spa and F r ' spa The input is processed in three visual Mamba modules ViM_1, ViM_2 and ViM_3;

[0161] The ViM_1, ViM_2 and ViM_3 simultaneously perform the following 8 steps:

[0162] (b1) The feature F h ' spa 、F m ' spa and F r ' spa The normalization layers LayerNorm5_1, LayerNorm5_2, and LayerNorm5_3 are input for processing, and then the position encoding is added to obtain the feature F. h 、F m and F r ;

[0163] (b2) The feature F h 、F m and F r Input into normalization layers LayerNorm5_4, LayerNorm5_5 and LayerNorm5_6 respectively to obtain feature F hl 、F ml and F rl ;

[0164] (b3) The feature F hl Input into the linear layer Linear5_0, and use the nonlinear activation function SiLU to obtain the gated branch feature F of ViM_1 h_gate , the feature F ml Input into the linear layer Linear5_1, and use the nonlinear activation function SiLU to obtain the gated branch feature F of ViM_2 m_gate , the feature F rlInput into the linear layer Linear5_2, and use the nonlinear activation function SiLU to obtain the gated branch feature F of ViM_3 r_gate ;

[0165] (b4) The feature F hl The input is sent to the linear layer Linear5_3, and then processed by the depthwise separable convolution layer DWConv5_0 and the nonlinear activation function SiLU to obtain the feature F hdw , the feature F ml The input is sent to the linear layer Linear5_4, and then processed by the depthwise separable convolution layer DWConv5_1 and the nonlinear activation function SiLU to obtain the feature F mdw , the feature F rl The input is sent to the linear layer Linear5_5, and then processed by the depthwise separable convolution layer DWConv5_2 and the nonlinear activation function SiLU to obtain the feature F rdw , where the convolution kernel size of DWConv5_0, DWConv5_1 and DWConv5_2 is 3×3, and each convolution kernel performs convolution operation with a step size of 1 pixel;

[0166] (b5) The feature F hdw 、F mdw and F rdw The input is sent to the cross-modal selective scanning module CMSS for deep modeling. CMSS converts the feature F hdw 、F mdw and F rdw The three modal pixels at each position {(i,j)|i=(0,1,···,H-1),j=(0,1,···,W-1)} are regarded as a scan path. When the image size is H×W, there are H×W scan paths. According to formula (7), H×W token sequences are generated, and each sequence contains only three pixels. and They come from the same position of different pixels, where D is the feature dimension after projection transformation; then the H×W Token sequences are input into the corresponding S6 k=0,1,...,H×W The feature modeling is performed in the output sequence, and finally the image reorganization operation is performed on the output sequence, and all pixels in the same order in the output sequence are combined into an image to obtain the output feature F of CMSS. hc 、F mc and F rc ;

[0167]

[0168] (b6) The feature F hcInput normalization layer LayerNorm5_7 to obtain the main branch feature F of ViM_1 h_main , the feature F mc Input normalization layer LayerNorm5_8 to obtain the main branch feature F of ViM_2 m_main , the feature F rc Input normalization layer LayerNorm5_9 to obtain the main branch feature F of ViM_3 r_main ;

[0169] (b7) Calculate the fusion features F of ViM_1, ViM_2 and ViM_3 according to formulas (8), (9) and (10) respectively V1 、F V2 and F V3 ;

[0170] F V1 =F h_gate ×F h_main (8)

[0171] F V2 =F m_gate ×F m_main (9)

[0172] F V3 =F r_gate ×F r_main (10)

[0173] (b8) The feature F V1 、F V2 and F V3 Input into linear layers Linear5_6, Linear5_7 and Linear5_8 respectively to obtain feature F V ′1, F V ′2 and F V ′3, and finally the output features F of ViM_1, ViM_2 and ViM_3 are calculated according to formulas (11), (12), (13) V1_out 、F V2_out and F V3_out ;

[0174] F V1_out =F V ′1+F h (11)

[0175] F V2_out =F V ′2+F m (12)

[0176] F V3_out =F V ′3+Fr (13)

[0177] (c) According to formula (14), the fusion feature F V1_out 、F V2_out and F V3_out And get the sub-network N TF Output:

[0178]

[0179] Step 1.6 Create and initialize subnetwork N CLS , the sub-network N CLS It consists of two groups of convolution modules, namely Conv6_0 and Conv6_1;

[0180] The Conv6_0 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 1×1, each convolution kernel performs convolution operation with a step size of 1 pixel, and the nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation;

[0181] The Conv6_1 includes 1 AdaptiveAvgPool pooling operation, 1 layer of 2d convolution operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 1×1, each convolution kernel performs convolution operation with a step size of 1 pixel, and selects Softmax as the activation function for operation.

[0182] Step 2. Input the training set H of hyperspectral images, the training set M of multispectral images, the training set R of radar images, the manually annotated pixel point coordinate set and label set, and HMF Conduct training;

[0183] Step 2.1. Extract all pixel points with labels from the training set H of the hyperspectral image based on the manually labeled pixel point coordinates. Extract all pixel points with labels from the training set M of multispectral images Extract all pixel points with labels from the radar image training set R in, Represents X H The i1th pixel point in Represents X M The i1th pixel point in Represents X R The i1th pixel point in , N represents the total number of pixel points with labels;

[0184] Step 2.2. Use X HEach pixel point is taken as the center to divide H into a series of hyperspectral pixel blocks of size P′P X M Each pixel point is taken as the center to divide M into a series of multispectral pixel blocks of size P′P X R Each pixel point is the center and R is divided into a series of radar pixel blocks of size P′P.

[0185] Will and As the training set of the fusion classification neural network, the samples in the training set are integrated into tuples The form of network input is Represents the pixel pairs composed of hyperspectral images, multispectral images and radar images in the training set, which have the same spatial coordinates. express and Corresponding true category label, set the number of iterations iter←1, and execute steps 2.3 to 2.7;

[0186] Step 2.3. Use subnetwork N HSPA 、N HSPE 、N MSPA and N RSPA Extract features of the training set;

[0187] Step 2.3.1 Use subnetwork N HSPA Training set for hyperspectral images Perform feature extraction to obtain the spatial features F of the hyperspectral image hspa ;

[0188] Step 2.3.2 Use subnetwork N HSPE Training set for hyperspectral images Perform feature extraction to obtain the spectral feature F of the hyperspectral image hspe ;

[0189] Step 2.3.3 Use subnetwork N MSPA Training set for multispectral images Perform feature extraction to obtain the spatial features F of the multispectral image mspa ;

[0190] Step 2.3.4 Use subnetwork N RSPA Training set for radar images Perform feature extraction to obtain the spatial feature F of the radar image rspa ;

[0191] Step 2.4. Use subnetwork N TF Fusion feature Fhspa 、F mspa and F rspa Get joint features

[0192] Step 2.5. Use subnetwork N CLS Joint Features Perform classification and calculate the classification prediction result TR pred ;

[0193] Step 2.6. Calculate the cross entropy loss function according to the definition of formula (14);

[0194]

[0195] Among them, C represents the number of categories, p n,c is the predicted probability, y n,c is the probability vector corresponding to the label;

[0196] In step 2.7, if all pixel blocks in the training set have been processed, proceed to step 2.8; otherwise, take a group of unprocessed pixel blocks from the training set and return to step 2.3;

[0197] Step 2.8 Let iter←iter+1. If the number of iterations iter>Total_iter, the trained hybrid Mamba network N is obtained. HMF , go to step 3, otherwise, use the reverse error propagation algorithm based on stochastic gradient descent and the prediction loss L CE Update N HMF , go to step 2.3 to reprocess all pixel blocks in the training set, and Total_iter represents the preset number of iterations.

[0198] Step 3 Input the hyperspectral image H′, multispectral image M′ and radar image R¢ to be tested, and use the trained network N HMF Complete pixel classification.

[0199] Step 3.1. Extract all pixel points in H′ to form a set Extract all pixel points in M′ to form a set Extract all pixel points in R¢ to form a set in, Indicates T H The i2th pixel of Indicates T M The i2th pixel, t R,i2 Indicates T R The i2th pixel of , U represents the total number of all pixels;

[0200] Step 3.2. Take TH Each pixel point is taken as the center to divide H′ into a series of hyperspectral pixel blocks of size P′P, forming a hyperspectral image set to be tested. T M Each pixel point is taken as the center to divide M′ into a series of multispectral pixel blocks of size P′P, forming a multispectral image set to be tested T R Each pixel point of is taken as the center to divide R¢ into a series of radar pixel blocks of size P′P, forming the radar image set to be tested

[0201] Step 3.3. Use the fully trained sub-network N HSPA 、N HSPE 、N MSPA and N RSPA Extract the features of the image to be tested;

[0202] Step 3.3.1 Use the fully trained sub-network N HSPA The hyperspectral image collection to be tested Perform feature extraction to obtain the spatial features F of the hyperspectral image h ' spa ;

[0203] Step 3.3.2 Use the fully trained sub-network N HSPE The hyperspectral image collection to be tested Perform feature extraction to obtain the spectral feature F of the hyperspectral image h ' spe ;

[0204] Step 3.3.3 Use the fully trained sub-network N MSPA The multispectral image collection to be tested Perform feature extraction to obtain the spatial features F of the multispectral image m ' spa ;

[0205] Step 3.3.4 Use the fully trained sub-network N RSPA The radar image collection to be tested Perform feature extraction to obtain the spatial feature F of the radar image r ' spa ;

[0206] Step 3.4. Use the fully trained sub-network N TF Fusion feature F h ' spa 、F m ' spa and F r ' spa Get joint features

[0207] Step 3.5. Use the fully trained sub-network N CLS Joint Features Classify and calculate the classification prediction result TE pred .

[0208] To verify the effectiveness of the present invention, experiments were conducted using the public Houston 2013 dataset and Augsburg City dataset as examples. Overall Accuracy (OA), Average Accuracy (AA) and Kappa Coefficient (Kappa) were used as objective indicators to evaluate the classification results. The evaluation results of the present invention were compared with those of the AM3Net method, TBCNN method, MFT method, MDAS method, Fusion-CNN method and S2FL method.

[0209] Tables 1 and 2 give the OA, AA and Kappa evaluation indicators of different algorithms on two data sets. The method proposed in the present invention achieved the result closest to GroundTruth. As can be seen from Tables 1 and 2, the classification accuracy of the present invention achieved the highest result, and it is better than the comparison algorithm in all objective evaluation indicators on the two data sets. Among them, the OA index of the present invention is improved by 3% and 5.74% respectively compared with the suboptimal algorithm, the AA index is improved by 3.94% and 4.85% respectively compared with the suboptimal algorithm, and the Kappa index is improved by 3.28% and 6.18% respectively compared with the suboptimal algorithm. By deeply extracting and integrating the advantageous features of different data sources, and fully exploring the correlation and complementarity between multi-source heterogeneous features, the redundant information interference and feature loss problems in the feature fusion process are reduced, and the final land cover classification accuracy is significantly improved.

[0210] The classification results of the Houston2013 dataset by the embodiment of the present invention are compared with those of the AM3Net method, the TBCNN method, the MFT method, the MDAS method, the Fusion-CNN method, and the S2FL method. Figure 6 As shown in Table 1. Figure 6In the figure, (a) is a pseudo-color image of the hyperspectral image of the Houston2013 dataset; (b) is a pseudo-color image of the multispectral image of the Houston2013 dataset; (c) is a digital representation model of the LiDAR image of the Houston2013 dataset; (d) is the classification result of the AM3Net method on the Houston2013 dataset; (e) is the classification result of the TBCNN method on the Houston2013 dataset; (f) is the classification result of the MFT method on the Houston2013 dataset; (g) is the classification result of the MDAS method on the Houston2013 dataset; (h) is the classification result of the Fusion-CNN method on the Houston2013 dataset; (i) is the classification result of the S2FL method on the Houston2013 dataset; (j) is the classification result of the method of the present invention on the Houston2013 dataset; (k) is the real object distribution label of the Houston2013 dataset.

[0211] Table 1 Objective evaluation indicators of Houston2013 dataset under different algorithms

[0212]

[0213]

[0214] The classification results of the Augsburg City dataset are compared between the present invention and the AM3Net method, TBCNN method, MFT method, MDAS method, Fusion-CNN method and S2FL method. Figure 7 As shown in Table 2. Figure 7In the figure, (a) is a pseudo-color image of the hyperspectral image of the Augsburg City dataset; (b) is a true-color image of the multispectral image of the Augsburg City dataset; (c) is a pseudo-color image of the SAR image of the Augsburg City dataset; (d) is the classification result of the AM3Net method on the Augsburg City dataset; (e) is the classification result of the TBCNN method on the Augsburg City dataset; (f) is the classification result of the MFT method on the Augsburg City dataset; (g) is the classification result of the MDAS method on the Augsburg City dataset; (h) is the classification result of the Fusion-CNN method on the Augsburg City dataset; (i) is the classification result of the S2FL method on the Augsburg City dataset; (j) is the classification result of the method of the present invention on the Augsburg City dataset; (k) is the real object distribution label of the Augsburg City dataset.

[0215] from Figure 6 , Figure 7 It can be seen that the classification result map of the present invention has the best visual effect, and the boundaries are clearer and more coherent. Especially in the complex area in the lower half of the Houston2013 dataset, most of the existing methods generally divide the area into one or two dominant categories, almost covering the entire area, and ignoring the detailed differences inside the objects. On the contrary, the present invention has the most detailed classification results in this complex area, and the contour information of buildings and roads can be clearly seen, which also verifies the effectiveness of the method of the present invention. For the Augsburg City dataset, the classification result maps of different methods can be directly compared with Ground Truth. It can be seen that the method of the present invention is closest to Ground Truth and presents a significant advantage. The method of the present invention not only utilizes three-source remote sensing data, but also designs an efficient feature extraction scheme for different data sources, and fully exploits the advantageous features of hyperspectral, multispectral and radar data. In addition, the efficient fusion and collaborative optimization of the three-source heterogeneous features are also fully considered, and the interdependence between different data sources is deeply explored.

[0216] Table 2 Objective evaluation indicators of the AugsburgCity dataset under different algorithms

[0217]

[0218]

[0219] Comprehensive Table 1, Table 2, Figure 6 , Figure 7From the comparison results, it can be seen that the present invention constructs a three-source remote sensing image fusion classification model based on a hybrid Mamba network, effectively extracts the advantageous features of multi-source data, optimizes the feature fusion effect, reduces the redundant information interference in the fusion process, and achieves higher-precision land cover classification.

Claims

1. A three-source remote sensing image fusion classification method based on a hybrid Mamba network, characterized in that Proceed as follows: Step 1. Create and initialize the hybrid Mamba network N HMF , the N HMF Contains 4 sub-networks N for feature extraction HSPA 、N HSPE 、N MSPA and N RSPA , 1 sub-network N for feature fusion TF and 1 subnetwork N for classification CLS ; Step 2. Input the training set H of hyperspectral images, the training set M of multispectral images, the training set R of radar images, the manually annotated pixel point coordinate set and label set, and HMF Conduct training; Step 3. Input the hyperspectral image H′, multispectral image M′ and radar image R¢ to be tested, and use the trained network N HMF Complete pixel classification.

2. The three-source remote sensing image fusion classification method based on hybrid Mamba network according to claim 1 is characterized in that The step 1 is specifically as follows: Step 1.1 Create and initialize subnetwork N HSPA , the sub-network N HSPA It consists of three groups of 2d convolution modules, namely Conv1_0, Conv1_1, and Conv1_2; Conv1_0 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and a nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation; Conv1_1 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation; The Conv1_2 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation, 1 layer of activation operation and 1 layer of Dropout operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and selects the nonlinear activation function LeakyReLU with a parameter of 0.01 as the activation function and the Dropout with a parameter of 0.5 for operation; Step 1.2 Create and initialize subnetwork N HSPE , the sub-network N HSPE It consists of 1 bidirectional spectrum Mamba module BSM; The BSM performs the following 7 steps: (a) The original hyperspectral image is input into the global average pooling layer GAP to obtain the feature Among them, B is the batch size of the input network, C1 is the number of channels of the hyperspectral image; (b) Use reshape operation to change feature F gap The shape of (c) For feature F′ gap Perform projection transformation operation, and use 2d convolution layer Conv2_0 to perform feature mapping to obtain feature F pro , where the size of the convolution kernel in the convolution layer is 1×1, and each convolution kernel performs convolution operation with a step size of 1 pixel; (d) For feature F pro Perform position encoding and then flip to obtain feature F′ pro ; (e) The features F pro and F′ pro The input parameter is shared in the spectrum Mamba module to obtain the feature F s and F′ s , where the spectral Mamba module consists of 1 Linear2_0 layer, 1 1d convolution layer Conv2_1, 1 activation function layer using nonlinear activation function SiLU, 1 S6_2_0 layer, 1 Linear2_1 layer, 1 activation function layer using nonlinear activation function SiLU, 1 element-wise multiplication operation and 1 Linear2_2 layer; (f) The feature F s and F′ s Add element by element and average along the second dimension to get feature F m ; (g) The feature F m The input is processed in the 2d convolution layer Conv2_2 to obtain the output feature F of the BSM module bsm , where the size of the convolution kernel in the convolution layer is 1×1, and each convolution kernel performs convolution operation with a step size of 1 pixel; Step 1.3 Create and initialize subnetwork N MSPA , the sub-network N MSPA It consists of 2 center-aware Mamba modules, 1 residual connection, 3 groups of 2d convolution modules and 1 cross-network residual connection. The 2 center-aware Mamba modules are CAM_1 and CAM_2, and the 3 groups of 2d convolution modules are Conv3_0, Conv3_1, and Conv3_2. The CAM_1 performs the following five steps: (a1) The original multispectral image MS is input into the projection transformation layer to obtain the feature F l , where the projection transformation layer contains one 2d convolution layer Conv3_3 and one normalization layer LayerNorm3_0. The convolution kernel size of Conv3_3 is 1×1, and each convolution kernel performs convolution operation with a step size of 1 pixel; (b1) For feature F l Perform dual position encoding, then perform a Dropout operation with a parameter of 0.01, and then input it into the normalization layer LayerNorm3_1 to obtain the feature F n ; (c1) The feature F n Input into the linear layer Linear3_0, and use the nonlinear activation function SiLU to obtain the gated branch feature F gate ; (d1) The feature F n The input is sent to the linear layer Linear3_1, and then processed by the depthwise separable convolution layer DWConv3_0 and the nonlinear activation function SiLU to obtain the feature F dw , where the size of the DWConv3_0 convolution kernel is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the feature F dw It is sent to the four-way central spiral scanning FDCSS_0 for feature extraction, where FDCSS_0 starts from the upper left, lower left, lower right and upper right pixels of the current patch, and scans the entire image patch from the outside to the inside in a clockwise spiral order, and forms four Token sequences S1, S2, S3 and S4 with different arrangement orders. Then the four different Token sequences are input into the S6_3_0, S6_3_1, S6_3_2 and S6_3_3 models for processing, and the sequences are reorganized into image features F according to the original arrangement order. s1 、F s2 、F s3 and F s4 According to formula (1), the output feature F of the FDCSS_0 module is calculated fdc ,in and is the parameter automatically learned by the network, and then the feature F fdc The input normalization layer LayerNorm3_2 is processed to obtain the main branch feature F main ; (e1) According to formula (2), the gated branch feature F is integrated gate and the main branch feature F main Get feature F fu , then the feature F fu The input is sent to the linear layer Linear3_2 and processed by a 2D convolution layer Conv3_4 to obtain the output feature F of the CAM_1 module. c1 , where the convolution kernel size of Conv3_4 is 1×1, and each convolution kernel performs convolution operation with a step size of 1 pixel; F fu =F gate ×F main (2) The CAM_2 performs the following five steps: (a2) The output feature F of CAM_1 module c1 Input projection transformation layer processing to obtain feature F l ′, where the projection transformation layer contains one 2D convolution layer Conv3_5 and one normalization layer LayerNorm3_3. The convolution kernel size of Conv3_5 is 1×1, and each convolution kernel performs convolution operation with a step size of 1 pixel. (b2) For feature F l ′ performs dual position encoding, then performs a Dropout operation with a parameter of 0.01, and then inputs the normalization layer LayerNorm3_4 to obtain the feature F′ n ; (c2) The feature F′ n Input into the linear layer Linear3_3, and use the nonlinear activation function SiLU to obtain the gated branch feature F′ gate ; (d2) The feature F′ n The input is sent to the linear layer Linear3_4, and then processed by the depthwise separable convolution layer DWConv3_1 and the nonlinear activation function SiLU to obtain the feature F′ dw , where the size of the DWConv3_1 convolution kernel is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the feature F′ dw It is sent to the four-way central spiral scanning FDCSS_1 for feature extraction, where FDCSS_1 starts from the upper left, lower left, lower right and upper right pixels of the current Patch, and scans the entire image Patch from the outside to the inside in a clockwise spiral order, and forms four Token sequences S′1, S′2, S′3 and S′4 with different arrangement orders. Then the four different Token sequences are input into the S6_3_4, S6_3_5, S6_3_6 and S6_3_7 models for processing, and the sequences are reorganized into image features F′ according to the original arrangement order. s1 , F′ s2 , F′ s3 and F′ s4 , the output feature F′ of the FDCSS_1 module is calculated according to formula (3) fdc ,in and is the parameter automatically learned by the network, and then the feature F′ fdc The input normalization layer LayerNorm3_5 is processed to obtain the main branch feature F′ main ; (e2) According to formula (4), the gated branch feature F′ is integrated gate and the main branch feature F′ main Get feature F′ fu , then the feature F′ fu The input is sent to the linear layer Linear3_5 and processed by a 2D convolution layer Conv3_6 to obtain the output feature F of the CAM_2 module. c2 , where the convolution kernel size of Conv3_6 is 1×1, and each convolution kernel performs convolution operation with a step size of 1 pixel; F′ fu =F′ gate ×F′ main (4) The residual connection is calculated according to formula (5), where θ1 and θ2 are parameters automatically learned by the network; F cam =θ1×MS+θ2×F c2 (5) The Conv3_0 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation; The Conv3_1 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation; The Conv3_2 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation, 1 layer of activation operation and 1 layer of Dropout operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and selects the nonlinear activation function LeakyReLU with a parameter of 0.01 as the activation function and the Dropout with a parameter of 0.5 for operation; The cross-network residual connection is calculated according to formula (6), where θ3 and θ4 are parameters automatically learned by the network, and F conv3_2 It is the output feature of convolution module Conv3_2; F cres =θ3×F bsm +θ4×F conv3_2 (6) Step 1.4 Create and initialize subnetwork N RSPA , the sub-network N RSPA It consists of three groups of 2d convolution modules, namely Conv4_0, Conv4_1, and Conv4_2; The Conv4_0 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation; The Conv4_1 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and the nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation; The Conv4_2 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation, 1 layer of activation operation and 1 layer of Dropout operation, wherein the size of the convolution kernel in the convolution layer is 3×3, each convolution kernel performs convolution operation with a step size of 1 pixel, and selects the nonlinear activation function LeakyReLU with a parameter of 0.01 as the activation function and the Dropout with a parameter of 0.5 for operation; Step 1.5 Establish and initialize the sub-network N for feature fusion TF , the sub-network N TF It contains a tri-modal fusion Mamba module TFM, which consists of a projection transformation layer, 3 visual Mamba modules, an element-by-element summation operation and a 2d convolution layer Conv5_0. The projection transformation layer contains a 2d convolution layer Conv5_1 and a normalization layer LayerNorm5_0. The convolution kernel sizes of Conv5_0 and Conv5_ are both 1×1, and each convolution kernel performs convolution operations with a step size of 1 pixel. The three visual Mamba modules are ViM_1, ViM_2 and ViM_3. The sub-network N TF Follow the 3 steps below: (a) The hyperspectral feature F hspa , multi-spectral feature F mspa and radar signature F rspa The input projection transformation layer processes the feature F′ hspa , F′ mspa and F′ rspa ; (b) The feature F′ hspa , F′ mspa and F′ rspa The input is processed in three visual Mamba modules ViM_1, ViM_2 and ViM_3; The ViM_1, ViM_2 and ViM_3 simultaneously perform the following 8 steps: (b1) The feature F′ hspa , F′ mspa and F′ rspa The normalization layers LayerNorm5_1, LayerNorm5_2, and LayerNorm5_3 are input for processing, and then the position encoding is added to obtain the feature F. h 、F m and F r ; (b2) The feature F h 、F m and F r Input into normalization layers LayerNorm5_4, LayerNorm5_5 and LayerNorm5_6 respectively to obtain feature F hl 、F ml and F rl ; (b3) The feature F hl Input into the linear layer Linear5_0, and use the nonlinear activation function SiLU to obtain the gated branch feature F of ViM_1 h_gate , the feature F ml Input into the linear layer Linear5_1, and use the nonlinear activation function SiLU to obtain the gated branch feature F of ViM_2 m_gate , the feature F rl Input into the linear layer Linear5_2, and use the nonlinear activation function SiLU to obtain the gated branch feature F of ViM_3 r_gate ; (b4) The feature F hl The input is sent to the linear layer Linear5_3, and then processed by the depthwise separable convolution layer DWConv5_0 and the nonlinear activation function SiLU to obtain the feature F hdw , the feature F ml The input is sent to the linear layer Linear5_4, and then processed by the depthwise separable convolution layer DWConv5_1 and the nonlinear activation function SiLU to obtain the feature F mdw , the feature F rl The input is sent to the linear layer Linear5_5, and then processed by the depthwise separable convolution layer DWConv5_2 and the nonlinear activation function SiLU to obtain the feature F rdw , where the convolution kernel size of DWConv5_0, DWConv5_1 and DWConv5_2 is 3×3, and each convolution kernel performs convolution operation with a step size of 1 pixel; (b5) The feature F hdw 、F mdw and F rdw The input is sent to the cross-modal selective scanning module CMSS for deep modeling. CMSS converts the feature F hdw 、F mdw and F rdw The three modal pixels at each position {(i,j)|i=(0,1,···,H-1),j=(0,1,···,W-1)} are regarded as a scan path. When the image size is H×W, there are H×W scan paths. According to formula (7), H×W token sequences are generated, and each sequence contains only three pixels. and They come from the same position of different pixels, where D is the feature dimension after projection transformation; then the H×W Token sequences are input into the corresponding S6 k=0,1,...,H×W The feature modeling is performed in the output sequence, and finally the image reorganization operation is performed on the output sequence, and all pixels in the same order in the output sequence are combined into an image to obtain the output feature F of CMSS. hc 、F mc and F rc ; (b6) The feature F hc Input normalization layer LayerNorm5_7 to obtain the main branch feature F of ViM_1 h_main , the feature F mc Input normalization layer LayerNorm5_8 to obtain the main branch feature F of ViM_2 m_main , the feature F rc Input normalization layer LayerNorm5_9 to obtain the main branch feature F of ViM_3 r_main ; (b7) Calculate the fusion features F of ViM_1, ViM_2 and ViM_3 according to formulas (8), (9) and (10) respectively V1 、F V2 and F V3 ; F V1 =F h_gate ×F h_main (8) F V2 =F m_gate ×F m_main (9) F V3 =F r_gate ×F r_main (10) (b8) The feature F V1 、F V2 and F V3 Input into linear layers Linear5_6, Linear5_7 and Linear5_8 respectively to obtain feature F′ V1 , F′ V2 and F′ V3 Finally, the output features F of ViM_1, ViM_2 and ViM_3 are calculated according to formulas (11), (12), and (13) V1_out 、F V2_out and F V3_out ; F V1_out =F′ V1 +F h (11) F V2_out =F′ V2 +F m (12) F V3_out =F′ V3 +F r (13) (c) According to formula (14), the fusion feature F V1_out 、F V2_out and F V3_out And get the sub-network N TF Output: Step 1.6 Create and initialize subnetwork N CLS , the sub-network N CLS It consists of two groups of convolution modules, namely Conv6_0 and Conv6_1; The Conv6_0 includes 1 layer of 2d convolution operation, 1 layer of BatchNorm normalization operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 1×1, each convolution kernel performs convolution operation with a step size of 1 pixel, and the nonlinear activation function LeakyReLU with a parameter of 0.01 is selected as the activation function for operation; The Conv6_1 includes 1 AdaptiveAvgPool pooling operation, 1 layer of 2d convolution operation and 1 layer of activation operation, wherein the size of the convolution kernel in the convolution layer is 1×1, each convolution kernel performs convolution operation with a step size of 1 pixel, and selects Softmax as the activation function for operation.

3. The three-source remote sensing image fusion classification method based on hybrid Mamba network according to claim 2 is characterized in that The step 2 is specifically as follows: Step 2.

1. Extract all pixel points with labels from the training set H of the hyperspectral image based on the manually labeled pixel point coordinates. Extract all pixel points with labels from the training set M of multispectral images Extract all pixel points with labels from the radar image training set R in, Represents X H The i1th pixel point in Represents X M The i1th pixel point in Represents X R The i1th pixel point in , N represents the total number of pixel points with labels; Step 2.

2. Use X H Each pixel point is taken as the center to divide H into a series of hyperspectral pixel blocks of size P′P X M Each pixel point is taken as the center to divide M into a series of multispectral pixel blocks of size P′P X R Each pixel point is the center and R is divided into a series of radar pixel blocks of size P′P. Will and As the training set of the fusion classification neural network, the samples in the training set are integrated into tuples The form of network input is Represents the pixel pairs composed of hyperspectral images, multispectral images and radar images in the training set, which have the same spatial coordinates. express and Corresponding true category label, set the number of iterations iter←1, and execute steps 2.3 to 2.7; Step 2.

3. Use subnetwork N HSPA 、N HSPE 、N MSPA and N RSPA Extract features of the training set; Step 2.3.1 Use subnetwork N HSPA Training set for hyperspectral images Perform feature extraction to obtain the spatial features F of the hyperspectral image hspa ; Step 2.3.2 Use subnetwork N HSPE Training set for hyperspectral images Perform feature extraction to obtain the spectral feature F of the hyperspectral image hspe ; Step 2.3.3 Use subnetwork N MSPA Training set for multispectral images Perform feature extraction to obtain the spatial features F of the multispectral image mspa ; Step 2.3.4 Use subnetwork N RSPA Training set for radar images Perform feature extraction to obtain the spatial feature F of the radar image rspa ; Step 2.

4. Use subnetwork N TF Fusion feature F hspa 、F mspa and F rspa Get joint features Step 2.

5. Use subnetwork N CLS Joint Features Perform classification and calculate the classification prediction result TR pred ; Step 2.

6. Calculate the cross entropy loss function according to the definition of formula (14); Among them, C represents the number of categories, p n,c is the predicted probability, y n,c is the probability vector corresponding to the label; In step 2.7, if all pixel blocks in the training set have been processed, proceed to step 2.8; otherwise, take a group of unprocessed pixel blocks from the training set and return to step 2.3; Step 2.8 Let iter←iter+1. If the number of iterations iter>Total_iter, the trained hybrid Mamba network N is obtained. HMF , go to step 3, otherwise, use the reverse error propagation algorithm based on stochastic gradient descent and the prediction loss L CE Update N HMF , go to step 2.3 to reprocess all pixel blocks in the training set, and Total_iter represents the preset number of iterations.

4. The three-source remote sensing image fusion classification method based on hybrid Mamba network according to claim 3 is characterized in that The step 3 is as follows: Step 3.

1. Extract all pixel points in H′ to form a set Extract all pixel points in M′ to form a set Extract all pixel points in R¢ to form a set in, Indicates T H The i2th pixel of Indicates T M The i2th pixel of Indicates T R The i2th pixel of , U represents the total number of all pixels; Step 3.

2. Take T H Each pixel point is taken as the center to divide H′ into a series of hyperspectral pixel blocks of size P′P, forming a hyperspectral image set to be tested. T M Each pixel point is taken as the center to divide M′ into a series of multispectral pixel blocks of size P′P, forming a multispectral image set to be tested T R Each pixel point of is taken as the center to divide R¢ into a series of radar pixel blocks of size P′P, forming the radar image set to be tested Step 3.

3. Use the fully trained sub-network N HSPA 、N HSPE 、N MSPA and N RSPA Extract the features of the image to be tested; Step 3.3.1 Use the fully trained sub-network N HSPA The hyperspectral image collection to be tested Perform feature extraction to obtain the spatial feature F′ of the hyperspectral image hspa ; Step 3.3.2 Use the fully trained sub-network N HSPE The hyperspectral image collection to be tested Perform feature extraction to obtain the spectral feature F′ of the hyperspectral image hspe ; Step 3.3.3 Use the fully trained sub-network N MSPA The multispectral image collection to be tested Perform feature extraction to obtain the spatial feature F′ of the multispectral image mspa ; Step 3.3.4 Use the fully trained sub-network N RSPA The radar image collection to be tested Perform feature extraction to obtain the spatial feature F′ of the radar image rspa ; Step 3.

4. Use the fully trained sub-network N TF Fusion feature F′ hspa , F′ mspa and F′ rspa Get joint features Step 3.

5. Use the fully trained sub-network N CLS Joint Features Classify and calculate the classification prediction result TE pred .

Citation Information

Cited By

  • Hyperspectral multispectral image fusion method and system based on spectral state fusion tree Mamba and storage medium

    CN120387929A

  • A hyperspectral and multispectral image fusion method, system and storage medium based on spectral state fusion tree Mamba

    CN120387929B

  • Image fusion method based on spectral space fusion state attention Mama model

    CN120563988A

  • An image fusion method based on spectral-spatial fusion state attention Mamba model

    CN120563988B

  • Urban illegal building expansion remote sensing change monitoring method based on three-flow graph neural network

    CN120877104A