Multi-modal remote sensing image fusion method based on multi-scale Mamba architecture
Through the multimodal remote sensing image fusion method based on the multi-scale Mamba architecture, the data heterogeneity problem in the fusion of hyperspectral and lidar is solved, the key information of lidar is retained, the computing efficiency is optimized, and the accuracy and efficiency of image fusion are improved.
Patent Information
- Application Number
- CN202510558730.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-15
AI Technical Summary
The existing methods fail to fully consider data heterogeneity when fusion of hyperspectral and lidar data, resulting in the weakening of important information in lidar data, and the computational volume of Transformer increases with the expansion of model scale, limiting data processing efficiency.
A multimodal remote sensing image fusion method based on the multi-scale Mamba architecture is adopted, and the laser radar data is processed through gradient calculation and normalization, and sliding window segmentation and feature extraction modules of different sizes are designed, combined with the Mamba architecture for feature stitching and dimensionality reduction, and information fusion is used by the attention mechanism.
The key spatial structural characteristics in lidar data are effectively preserved, the modal gap of multimodal data is eliminated, the computing efficiency is optimized, the long-distance dependence of the data is captured, and the accuracy and efficiency of image fusion are improved.
Smart Images

Figure CN120495815A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing method, and in particular to a multimodal remote sensing image fusion method based on a multi-scale Mamba architecture. Background Art
[0002] Remote sensing technology is a technology that uses long-distance detection equipment to obtain information about ground objects. Its essence is to capture electromagnetic wave information reflected or radiated by the surface through airborne or satellite-borne sensors to invert the properties of ground objects.
[0003] To address the limitations of single-sensor data representation, it is necessary to fuse multiple remote sensing images (such as optical, radar, and lidar). The fused image overcomes the shortcomings of single-sensor images and enables remote sensing images to better represent ground objects. Because lidar and hyperspectral imaging naturally complement each other, joint hyperspectral and lidar classification has become a research hotspot in the field of multimodal remote sensing.
[0004] Two common fusion frameworks for hyperspectral and lidar are convolutional neural network (CNN) and Transformer-based fusion architectures. Convolutional neural network (CNN) uses a sliding window to extract features, which limits its ability to capture global information and leads to inaccurate image fusion. Transformer, by introducing a self-attention mechanism, can globally capture multimodal images. However, as the model scale expands and the sequences to be processed grow longer, the computational complexity of the Transformer increases quadratically with the length of the context, limiting data processing efficiency.
[0005] The Mamba architecture cleverly combines the advantages of state-space models and deep learning, efficiently capturing the global context of a target with linear complexity and low computational effort. However, its unique potential in the joint classification of hyperspectral and lidar data has yet to be fully explored. Hyperspectral data is optical information, while lidar data is elevation information. Most existing methods directly use Mamba to fuse hyperspectral and lidar data, without fully considering the heterogeneity of the two data types. For example, since hyperspectral data is much more abundant than lidar data, direct fusion can weaken important information in the lidar data. In summary, this proposal proposes a multimodal remote sensing image fusion method based on the multi-scale Mamba architecture. Summary of the Invention
[0006] In order to solve the above problems, the present invention proposes a multimodal remote sensing image fusion method based on the multi-scale Mamba architecture to solve the above cross-modal feature alignment problem and the problem of guided feature fusion.
[0007] To achieve the above object, a multi-modal remote sensing image fusion method based on a multi-scale Mamba architecture according to the present invention proposes the following technical solutions:
[0008] A multi-modal remote sensing image fusion method based on a multi-scale Mamba architecture includes the following steps:
[0009] S1: Obtain the hyperspectral and lidar original images in the remote sensing image, denoted as Ar and Br respectively;
[0010] S2: For the original lidar image Br, perform gradient joint calculation to obtain the derived image Dr;
[0011] S3: Normalize the hyperspectral original image Ar to obtain the normalized image
[0012] S4: Dimension-expand the original lidar image Br and the derived image Dr respectively by adding a single-channel dimension to obtain the lidar-expanded image and the derived-expanded image
[0013] S5: Use sliding windows of different sizes to and perform segmentation respectively; the window sizes are: s1×s1, s2×s2, s3×s3 (s1 < s2 < s3); use the s1×s1 window to and segment into A1, B1 and D1; use the s2×s2 window to and segment into A2, B2 and D2; use the s3×s3 window to and segment into A3, B3 and D3;
[0014] S6: Use different feature extraction modules to extract features from the images of different sizes in S5 to obtain the images A” k 、B” k and D” k ;
[0015] S7: Perform bilinear interpolation upsampling on the images after feature extraction in S6 to obtain the upsampled images;
[0016] S8: Stitch the images after bilinear interpolation upsampling in S7 with the images after feature extraction in S6 to obtain the corresponding stitched images; the stitched images include: the stitched normalized image, the stitched lidar-expanded image and the stitched derived-expanded image;
[0017] S9: perform dimensionality reduction on the features concatenated in S8 to obtain a dimensionality-reduced image;
[0018] S10: Based on the Mamba architecture, the image after dimensionality reduction in S9 is converted into a feature sequence, the feature sequence is linearized and then embedded into a trainable position information matrix, and local sequence modeling is performed to obtain a high-level linear feature sequence.
[0019] S11: Fuse the high-level linear feature sequence in S10 to obtain the elevation information feature sequence and the edge contour information feature sequence;
[0020] S12: Based on the Mamba architecture, obtain the local-global features of the spectral-spatial features, elevation information feature sequence, and edge contour information feature sequence in S11;
[0021] S13: Convert the local-global features in S12 into feature maps, fuse the feature maps, and output the final fused feature image.
[0022] Furthermore, step S6 includes the following steps:
[0023] S6.1: Normalized image after segmentation Perform feature extraction to obtain an image after feature extraction;
[0024] The formula is:
[0025]
[0026] Where A k Represents a normalized image Images of different sizes after segmentation; A1, A2, and A3 represent normalized images respectively Segmented small-size image, medium-size image and large-size image;
[0027] A' k Represents a normalized image The first output of the feature extraction submodule;
[0028] Represents a 2D convolutional layer with C input channels and 32 output channels; BN represents a batch normalization layer; ReLU represents an activation function layer;
[0029] A” k Represents a normalized image Image after feature extraction; A”1, A”2 and A”3 represent normalized images respectively Small-size image, medium-size image and large-size image after feature extraction;
[0030] conv 3dRepresents a 3D convolutional layer with 32 input channels and 64 output channels;
[0031] S6.2: Extending the Segmented Radar Image Perform feature extraction to obtain an image after feature extraction;
[0032]
[0033] Where B k Represents radar extended image Images of different sizes after segmentation; B1, B2, and B3 represent radar expanded images respectively Segmented small-size image, medium-size image and large-size image;
[0034] B' k Represents radar extended image The first output of the feature extraction submodule;
[0035] Represents a 2D convolutional layer with 1 input channel and 32 output channels; BN represents a batch normalization layer; ReLU represents an activation function layer;
[0036] B” k Represents radar extended image Image after feature extraction; B”1, B”2 and B”3 represent radar expanded images respectively Small-size image, medium-size image and large-size image after feature extraction;
[0037] Represents a 2D convolutional layer with 32 input channels and 64 output channels;
[0038] S6.3: Extending the Derived Image After Segmentation Perform feature extraction to obtain an image after feature extraction;
[0039] The formula is:
[0040]
[0041] Where D k Denotes a derived extended image Images of different sizes after segmentation; D1, D2 and D3 represent derived extended images respectively Segmented small-size image, medium-size image and large-size image;
[0042] D' k Denotes a derived extended image The first output of the feature extraction submodule;
[0043] Represents a 2D convolutional layer with 1 input channel and 32 output channels; BN represents a batch normalization layer; ReLU represents an activation function layer;
[0044] D” k Denotes a derived extended image Image after feature extraction; D”1, D”2 and D”3 represent the derived extended images respectively Small-size image, medium-size image and large-size image after feature extraction;
[0045] Represents a 2D convolutional layer with 32 input channels and 64 output channels.
[0046] Furthermore, in step S7, the upsampling process formula is:
[0047]
[0048] Where, X” s (x i ,y i , C) represents the pixel in the i-th row and j-th column in the C-th channel of the small-scale image or the medium-sized image after feature extraction; (x i ,y i ) represents the coordinates of the pixel at the i-th row and j-th column;
[0049] X” s_up (x,y,c) represents X” s (x i ,y i ,C) mapped pixels after upsampling; where W i,j Represents the weight coefficient, representing X" s (x i ,y i ,c)’s four neighboring pairs map to pixel X” s (x i ,y i ,c) the degree of contribution.
[0050] Furthermore, in step S8, the formulas for the spliced normalized image, the spliced radar extended image, and the spliced derived extended image are:
[0051] A cat =Concat(A” 1_up ,A” 2_up ,A”3)
[0052] B cat =Concat(B” 1_up ,B” 2_up ,B”3)
[0053] Dcat =Concat(D” 1_up ,D” 2_up ,D”3)
[0054] Where Concat() means splicing by channel; A cat represents the normalized image after splicing; B cat represents the radar expanded image after stitching; D cat Represents the derived extended image after splicing.
[0055] Furthermore, in step S9, the image formula after dimensionality reduction is:
[0056]
[0057] Where SE represents the attention mechanism module; Indicates channel-by-channel multiplication; X fuse Represents the image after dimensionality reduction; Wx represents the weight matrix used for dimensionality reduction.
[0058] Furthermore, step S10 includes the following steps:
[0059] S10.1: Convert the feature image after dimensionality reduction in S9 into a feature sequence;
[0060] The formula is:
[0061] X seq_origin =reshape(X fuse ,(b,c,s3 2 ))
[0062] reshape() represents the reshaping operation;
[0063] S10.2: Replace the linear feature X in S10.1 seq_origin Embed a trainable position information matrix to form a feature sequence of position information;
[0064] The formula is:
[0065] X seq_temp =X seq_origin +P,X∈{A,B,D}
[0066] Where P is a trainable parameter matrix; X seq_temp is the position information feature sequence;
[0067] S10.3: Then X seq_temp Input a linear transformation module, perform local sequence modeling on it, and obtain a high-level linear feature sequence
[0068]
[0069] Among them, conv is a 2D convolutional layer with 64 input channels and 64 output channels.
[0070] Furthermore, the formulas for the elevation information feature sequence and edge profile information feature sequence in step S11 are:
[0071]
[0072] Where ⊙ represents element-by-element multiplication, and They are elevation information feature sequence and edge contour information feature sequence respectively.
[0073] Furthermore, in step S12, the local-global features of the spectral-spatial features, the elevation information feature sequence, and the edge profile information feature sequence are:
[0074]
[0075] Where A seq_f Represents the local-global features extracted from high-level feature sequences containing spectral-spatial features based on the Mamba architecture;
[0076] B seq_f Represents the local-global features extracted from high-level feature sequences containing elevation information based on the Mamba architecture;
[0077] D seq_f Represents the local-global features extracted from high-level feature sequences containing edge contour information based on the Mamba architecture;
[0078] Mamba() means extracting local-global features under the Mamba architecture.
[0079] Furthermore, step S13 includes the following steps:
[0080] S13.1: A seq_f 、B seq_f and D seq_f Convert to feature map A f 、B f and D f ;
[0081] The formula is:
[0082] X f =reshape(X seq_f ))
[0083] S13.2: Feature map A f 、B f and D fSplice along the channel to obtain the fused feature map;
[0084] The formula is:
[0085] F cat =Concat(A f ,B f ,D f )
[0086] Where, F cat Represents the features after splicing;
[0087] S13.3: Apply the attention mechanism to the fused feature map in S13.2 to increase information interaction and obtain the final fused feature image;
[0088] The formula is:
[0089] F = ReLU(BN(SE(F cat )))
[0090] In the formula, SE represents the attention mechanism module; BN represents the batch normalization layer; ReLU represents the activation function layer.
[0091] The beneficial effects that can be achieved by adopting the above technical content are:
[0092] 1. By mining the elevation features of LiDAR data, a derived image containing enhanced edge contour information is obtained, which effectively retains key spatial structural features such as building edges and terrain undulations, and effectively avoids the weakening of LiDAR data during the multimodal fusion process.
[0093] 2. Based on the characteristics of data of three different modalities, different feature extraction modules are designed to eliminate the modal gap of multimodal data and effectively capture the effective features of data of different modalities.
[0094] 3. The latest Mamba architecture was introduced to replace the mainstream Transformer. Based on its unique linear computational complexity characteristics, it effectively captures the long-distance dependencies of data features while optimizing computational efficiency and reducing complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0095] Figure 1 is a flow chart of this method. DETAILED DESCRIPTION
[0096] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0097] A multimodal remote sensing image fusion method based on a multi-scale Mamba architecture includes the following steps:
[0098] S1: Obtain hyperspectral and lidar raw images from remote sensing images.
[0099] Assume that the height of the original hyperspectral image is H, the image width is W, and the number of spectral channels is C;
[0100] Then, the acquired hyperspectral original image Ar is expressed as:
[0101]
[0102] In formula (1), a ij Represents the pixel in the i-th row and j-th column of the original hyperspectral image;
[0103] Assume that the height of the original lidar image is H and the image width is W;
[0104] Then, the acquired laser radar original image Br is expressed as:
[0105]
[0106] In formula (2), b ij Represents the pixel in the i-th row and j-th column of the original lidar image;
[0107] S2: Gradient joint calculation is performed on the original lidar image Br to obtain the derived image Dr.
[0108] S2 specifically includes the following steps:
[0109] S2.1: Calculate the gradient information of all pixels in the horizontal direction x of the original lidar image Br to obtain the horizontal gradient matrix Dx.
[0110] The horizontal gradient matrix Dx is expressed as:
[0111]
[0112] in, Represents the gradient information of the original lidar image Br in the horizontal direction x; Δx represents the distance between two adjacent pixels in the horizontal direction (the distance is 1);
[0113] Expressed as:
[0114]
[0115] An example of calculation for formula (3):
[0116] In this example, it is assumed that the original lidar image Br is:
[0117]
[0118] After calculation using formula (3),
[0119]
[0120] S2.2: Calculate the gradient information of all pixels in the vertical direction y of the original lidar image Br to obtain the vertical gradient matrix Dy.
[0121] The vertical gradient matrix Dy is expressed as:
[0122]
[0123] in, Represents the gradient information of the original lidar image Br in the horizontal direction y; Δy represents the distance between two adjacent pixels in the vertical direction (the distance is 1);
[0124] Expressed as:
[0125]
[0126] Continuing with the example in S2.1, let’s calculate:
[0127] After calculation using formula (4),
[0128]
[0129] S2.3: Calculate the comprehensive change of the gradient in the x and y directions, that is, the gradient intensity d ij , construct the derived image Dr;
[0130] The derived image Dr is expressed as:
[0131]
[0132] d ij is the gradient strength, expressed as:
[0133]
[0134] Following the examples in S2.1 and S2.2, after calculation using formula (5):
[0135]
[0136] S3: Normalize the hyperspectral original image Ar to obtain the normalized image
[0137] The formula is:
[0138]
[0139] Where, represents the normalized hyperspectral image; Represents the pixel at the i-th row and j-th column in the normalized hyperspectral image.
[0140] in, The calculation formula is:
[0141]
[0142] Among them, A min Represents pixel a ij The minimum spectral value in the C channels, A max Represents pixel a ij The maximum spectral value among the C channels of ;
[0143] S4: Expand the dimensions of the original lidar image Br and the derived image Dr, add a single-channel dimension, and record them as lidar expanded images after expansion and derived extended images
[0144] Radar extended image The height is H, the image width is W, and the number of spectral channels is 1.
[0145] Derived Extended Image The height is H, the image width is W, and the number of spectral channels is 1.
[0146] S5: Using sliding windows of different sizes, the normalized hyperspectral image Radar extended image and derived extended images Separately split.
[0147] Specifically, three sizes of sliding windows are used to transform the above images and Divide them respectively, and the window sizes are: s1×s1, s2×s2, s3×s3 (s1 <s2<s3)。
[0148] When the sliding window of s1×s1 is and When segmenting, normalize the image Radar extended image and derived extended images The image segmented by s1×s1 window is recorded as a small-size image. Then, the normalized image Radar extended image and derived extended images After segmentation, the three corresponding small-size images are: A1, B1 and D1.
[0149] Similarly, the image segmented by the s2×s2 window is recorded as the medium-sized image. Then, the normalized image Radar extended image and derived extended images The corresponding medium-sized images are A2, B2 and D2.
[0150] Similarly, the image segmented by the s3×s3 window is recorded as the large-size image; then, the normalized image Radar extended image and derived extended images The corresponding large-size images are A3, B3 and D3.
[0151] S6: Using different feature extraction modules, perform feature extraction on the images of different sizes after segmentation in S5 to obtain feature-extracted images.
[0152] S6 specifically includes the following steps:
[0153] S6.1: Normalized image after segmentation Perform feature extraction to obtain an image after feature extraction.
[0154] The formula is:
[0155]
[0156] Where A k Represents a normalized image Images of different sizes after segmentation; A1, A2, and A3 represent normalized images respectively Segmented small-size image, medium-size image and large-size image;
[0157] A' k Represents a normalized image The first output of the feature extraction submodule;
[0158] Represents a 2D convolutional layer with C input channels and 32 output channels; BN represents a batch normalization layer; ReLU represents an activation function layer;
[0159] A” k Represents a normalized image Image after feature extraction; A”1, A”2 and A”3 represent normalized images respectively Small-size image, medium-size image and large-size image after feature extraction;
[0160] conv 3d Represents a 3D convolutional layer with 32 input channels and 64 output channels;
[0161] S6.2: Extending the Segmented Radar Image Perform feature extraction to obtain an image after feature extraction.
[0162] The formula is:
[0163]
[0164] Where B k Represents radar extended image Images of different sizes after segmentation; B1, B2, and B3 represent radar expanded images respectively Segmented small-size image, medium-size image and large-size image;
[0165] B' k Represents radar extended image The first output of the feature extraction submodule;
[0166] Represents a 2D convolutional layer with 1 input channel and 32 output channels; BN represents a batch normalization layer; ReLU represents an activation function layer;
[0167] B” k Represents radar extended image Image after feature extraction; B”1, B”2 and B”3 represent radar expanded images respectively Small-size image, medium-size image and large-size image after feature extraction;
[0168] Represents a 2D convolutional layer with 32 input channels and 64 output channels;
[0169] S6.3: Extending the Derived Image After Segmentation Perform feature extraction to obtain an image after feature extraction.
[0170] The formula is:
[0171]
[0172] Where D k Denotes a derived extended image Images of different sizes after segmentation; D1, D2 and D3 represent derived extended images respectively Segmented small-size image, medium-size image and large-size image;
[0173] D' k Denotes a derived extended image The first output of the feature extraction submodule;
[0174] Represents a 2D convolutional layer with 1 input channel and 32 output channels; BN represents a batch normalization layer; ReLU represents an activation function layer;
[0175] D” k Denotes a derived extended image Image after feature extraction; D”1, D”2 and D”3 represent the derived extended images respectively Small-size image, medium-size image and large-size image after feature extraction;
[0176] Represents a 2D convolutional layer with 32 input channels and 64 output channels;
[0177] S7: The image after feature extraction in S6 is subjected to bilinear interpolation upsampling processing to obtain an upsampled image.
[0178] Pixel X in the upsampled image" s_up (x,y,c) is expressed as:
[0179]
[0180] Where, X” s (x i ,y i , C) represents the pixel in the i-th row and j-th column in the C-th channel of the small-scale image or the medium-sized image after feature extraction. (x i ,y i ) represents the coordinates of the pixel in the i-th row and j-th column.
[0181] X” s_up (x,y,c) represents X” s (x i ,y i ,C) Mapping pixel after upsampling. Among them, W i,j Represents the weight coefficient, representing X" s (x i ,y i ,c)’s four neighboring pairs map to pixel X”s (x i ,y i ,c) the degree of contribution.
[0182] S8: stitching the image processed by bilinear interpolation upsampling in S7 with the image processed by feature extraction in S6 to obtain a corresponding stitched image; the stitched image includes: a stitched normalized image, a stitched radar extended image, and a stitched derived extended image.
[0183] The formula is:
[0184] A cat =Concat(A” 1_up ,A” 2_up ,A”3)
[0185] B cat =Concat(B” 1_up ,B” 2_up ,B”3)
[0186] D cat =Concat(D” 1_up ,D” 2_up ,D”3) (13)
[0187] In the formula, Concat() means splicing by channel. cat represents the normalized image after splicing; B cat represents the radar expanded image after stitching; D cat Represents the derived extended image after splicing.
[0188] S9: Perform dimensionality reduction on the image spliced in S8 to obtain a dimensionality-reduced image.
[0189] The formula is:
[0190]
[0191] Where SE represents the attention mechanism module; Indicates channel-by-channel multiplication; X fuse Represents the image after dimensionality reduction; Wx represents the weight matrix used for dimensionality reduction.
[0192] S10: Based on the Mamba architecture, the image after dimensionality reduction in S9 is converted into a feature sequence, the feature sequence is linearized and then embedded into a trainable position information matrix, and local sequence modeling is performed to obtain a high-level linear feature sequence.
[0193] The specific steps include:
[0194] S10.1: Since the Mamba architecture can process long sequence data, but cannot directly process image data, first fuse The image after dimensionality reduction is directly converted into the feature sequence X seq_origin ;
[0195] The formula is:
[0196] X seq_origin =reshape(X fuse )
[0197] Where, reshape() represents the reshaping operation;
[0198] S10.2: To avoid subsequent seq_origin The processing of the feature image causes the position information to be lost, so the linear feature X seq_origin Embed a trainable position information matrix to form a feature sequence of position information.
[0199] The formula is:
[0200] X seq_temp =X seq_origin +P,X∈{A,B,D} (15)
[0201] Where P is a trainable parameter matrix; X seq_temp is the position information feature sequence.
[0202] S10.3: Then X seq_temp Input a linear transformation module, perform local sequence modeling on it, and obtain a high-level linear feature sequence
[0203] The formula is:
[0204]
[0205] Among them, conv is a 2D convolutional layer with 64 input channels and 64 output channels.
[0206] S11: Advanced linear feature sequences in S10 The fusion is performed to obtain the elevation information feature sequence and the edge contour information feature sequence.
[0207] The formula is:
[0208]
[0209] Where ⊙ represents element-by-element multiplication, and They are respectively the high-level feature sequence of elevation information and the high-level feature sequence of edge contour information.
[0210] S12: Based on the Mamba architecture, obtain the local-global features of the spectral-spatial features, elevation information feature sequence, and edge contour information feature sequence in S11;
[0211] The formula is:
[0212]
[0213] Where A seq_f Represents the local-global features extracted from high-level feature sequences containing spectral-spatial features based on the Mamba architecture;
[0214] B seq_f Represents the local-global features extracted from high-level feature sequences containing elevation information based on the Mamba architecture;
[0215] D seq_f Represents the local-global features extracted from high-level feature sequences containing edge contour information based on the Mamba architecture;
[0216] Mamba() means extracting local-global features under the Mamba architecture.
[0217] S13: Convert the local-global features in S12 into feature maps, fuse the feature maps, and output the final fused feature image.
[0218] S13.1: Specifically include the following steps:
[0219] S13.1: A seq_f 、B seq_f and D seq_f Convert to feature map A f 、B f and D f ;
[0220] X f =reshape(X seq_f ))
[0221] S13.2: Feature map A f 、B f and D f Splice along the channel to obtain the fused feature map;
[0222] The formula is:
[0223] F cat =Concat(A f ,B f ,D f ) (18)
[0224] Where, F cat Represents the fused feature map;
[0225] S13.3: Apply the attention mechanism to the fused feature map in S13.2 to increase information interaction and obtain the final fused feature image;
[0226] The formula is:
[0227] F = ReLU(BN(SE(F cat ))) (19)
[0228] In the formula, SE represents the attention mechanism module; BN represents the batch normalization layer; ReLU represents the activation function layer;
[0229] The fused image combines the characteristics of different sensors to make up for the defects of a single sensor (such as optical images are affected by weather and SAR images have low resolution). The fused image can have high spatial resolution and rich spectral information at the same time, or combine multi-dimensional features such as texture, polarization, and elevation, thereby improving the recognition of ground details in remote sensing technology.
[0230] According to the final fused feature image, the ground object classification result can be obtained.
[0231] This operation is an existing technology, which is used to analyze the ground object category represented by each pixel by finally fusing the feature images.
[0232] The final fused feature image is a two-dimensional matrix (size is H×W×C), which is flattened to form a one-dimensional vector (length is H×W×C), which is convenient for processing by the fully connected layer; that is, the spatial structure information of the final fused feature image is converted into a feature set in vector form, which is convenient for processing by the fully connected layer;
[0233] This one-dimensional vector is input into the fully connected layer, which maps the final fused feature image to the category space to obtain H×W one-dimensional vectors. For a certain one-dimensional vector, calculate its classification score for each category (set as K categories), a total of K;
[0234] This one-dimensional vector is input to the Softmax layer, which converts the classification score into a probability distribution.
[0235] The argmax function is used to process the probability distribution output by the Softmax layer, and the category with the maximum probability of each one-dimensional vector is output, that is, the predicted category.
[0236] For example, for a one-dimensional vector, after softmax, the probability of each category is: the probability of category 1 is 0.1, the probability of category 2 is 0.05, and the probability of category 3 is 0.85. The softmax layer outputs that the one-dimensional vector is category 3. In other words, the pixel corresponding to the one-dimensional vector is category 3. The meaning of category is the classification of land objects (such as plants, buildings, etc.).
[0237] With the above-described preferred embodiments of the present invention as a guide, and with reference to the above description, relevant personnel are fully capable of making various changes and modifications without departing from the technical scope of this invention. The technical scope of this invention is not limited to the contents of the specification and must be determined according to the scope of the claims.
Claims
1. A multimodal remote sensing image fusion method based on a multi-scale Mamba architecture, characterized in that: The following steps are involved: S1: Obtain hyperspectral and lidar raw images from remote sensing images, denoted as Ar and Br respectively; S2: For the original image Br of the lidar, the gradient joint calculation is used to obtain the derived image Dr; S3: Normalize the original hyperspectral image Ar to obtain a normalized image S4: Expand the dimensions of the original lidar image Br and the derived image Dr, add a single-channel dimension, and obtain the lidar expanded image and derived extended images S5: Using sliding windows of different sizes, and are separately segmented; the window sizes are: s1×s1, s2×s2, s3×s3 (s1 < s2 < s3); using the s1×s1 window, and are segmented into A1, B1, and D1; using the s2×s2 window, and are segmented into A2, B2, and D2; using the s3×s3 window, and are segmented into A3, B3, and D3; S6: Use different feature extraction modules to extract features from images of different sizes in S5, and obtain the feature-extracted image A" k , B” k and D” k ; S7: The image after feature extraction in S6 is subjected to bilinear interpolation upsampling processing to obtain an upsampled image; S8: stitching the image processed by bilinear interpolation upsampling in S7 with the image processed by feature extraction in S6 to obtain a corresponding stitched image; the stitched image includes: a stitched normalized image, a stitched radar extended image, and a stitched derived extended image; S9: perform dimensionality reduction on the features concatenated in S8 to obtain a dimensionality-reduced image; S10: Based on the Mamba architecture, the image after dimensionality reduction in S9 is converted into a feature sequence, the feature sequence is linearized and then embedded into a trainable position information matrix, and local sequence modeling is performed to obtain a high-level linear feature sequence. S11: Fuse the high-level linear feature sequence in S10 to obtain the elevation information feature sequence and the edge contour information feature sequence; S12: Based on the Mamba architecture, obtain the local-global features of the spectral-spatial features, elevation information feature sequence, and edge contour information feature sequence in S11; S13: Convert the local-global features in S12 into feature maps, fuse the feature maps, and output the final fused feature image.
2. The multimodal remote sensing image fusion method based on the multi-scale Mamba architecture according to claim 1 is characterized in that: Step S6 includes the following steps: S6.1: Normalized image after segmentation Perform feature extraction to obtain an image after feature extraction; The formula is: "A" k =ReLU(BN(conv 3d (A' k ))),k∈{1,2,3} (9) Where A k Represents a normalized image Images of different sizes after segmentation; A1, A2, and A3 represent normalized images respectively Segmented small-size image, medium-size image and large-size image; A' k Represents a normalized image The first output of the feature extraction submodule; Represents a 2D convolutional layer with C input channels and 32 output channels; BN represents a batch normalization layer; ReLU represents an activation function layer; A” k Represents a normalized image Image after feature extraction; A”1, A”2 and A”3 represent normalized images respectively Small-size image, medium-size image and large-size image after feature extraction; conv 3d Represents a 3D convolutional layer with 32 input channels and 64 output channels; S6.2: Extending the Segmented Radar Image Perform feature extraction to obtain an image after feature extraction; Where B k Represents radar extended image Images of different sizes after segmentation; B1, B2, and B3 represent radar expanded images respectively Segmented small-size image, medium-size image and large-size image; B' k Represents radar extended image The first output of the feature extraction submodule; Represents a 2D convolutional layer with 1 input channel and 32 output channels; BN represents a batch normalization layer; ReLU represents an activation function layer; B” k Represents radar extended image Image after feature extraction; B”1, B”2 and B”3 represent radar expanded images respectively Small-size image, medium-size image and large-size image after feature extraction; Represents a 2D convolutional layer with 32 input channels and 64 output channels; S6.3: Extending the Derived Image After Segmentation Perform feature extraction to obtain an image after feature extraction; The formula is: Where D k Denotes a derived extended image Images of different sizes after segmentation; D1, D2 and D3 represent derived extended images respectively Segmented small-size image, medium-size image and large-size image; D' k Denotes a derived extended image The first output of the feature extraction submodule; Represents a 2D convolutional layer with 1 input channel and 32 output channels; BN represents a batch normalization layer; ReLU represents an activation function layer; D” k Denotes a derived extended image Image after feature extraction; D”1, D”2 and D”3 represent the derived extended images respectively Small-size image, medium-size image and large-size image after feature extraction; Represents a 2D convolutional layer with 32 input channels and 64 output channels.
3. The multimodal remote sensing image fusion method based on the multi-scale Mamba architecture according to claim 2 is characterized in that: In step S7, the upsampling process formula is: Where, X” s (x i ,y i , C) represents the pixel in the i-th row and j-th column in the C-th channel of the small-scale image or the medium-sized image after feature extraction; (x i ,y i ) represents the coordinates of the pixel at the i-th row and j-th column; X” s_up (x,y,c) represents X” s (x i ,y i ,C) Mapped pixels after upsampling; Among them, W i,j Represents the weight coefficient, representing X" s (x i ,y i ,c)’s four neighboring pairs map to pixel X” s (x i ,y i ,c) the degree of contribution.
4. The multimodal remote sensing image fusion method based on the multi-scale Mamba architecture according to claim 3 is characterized in that: In step S8, the formulas for the spliced normalized image, the spliced radar extended image, and the spliced derived extended image are: A cat =Concat(A” 1_up ,A” 2_up ,A”3) B cat =Concat(B” 1_up ,B” 2_up ,B”3) D cat =Concat(D” 1_up ,D” 2_up ,D”3) Where Concat() means splicing by channel; A cat represents the normalized image after splicing; B cat represents the radar expanded image after stitching; D cat Represents the derived extended image after splicing.
5. The multimodal remote sensing image fusion method based on the multi-scale Mamba architecture according to claim 4 is characterized in that: In step S9, the image formula after dimensionality reduction is: Where SE represents the attention mechanism module; Indicates channel-by-channel multiplication; X fuse Represents the image after dimensionality reduction; Wx represents the weight matrix used for dimensionality reduction.
6. The multimodal remote sensing image fusion method based on the multi-scale Mamba architecture according to claim 5 is characterized in that: Step S10 includes the following steps: S10.1: Convert the feature image after dimensionality reduction in S9 into a feature sequence; The formula is: X seq_origin =reshape(X fuse ,(b,c,s3 2 )) reshape() represents the reshaping operation; S10.2: Replace the linear feature X in S10.1 seq_origin Embed a trainable position information matrix to form a feature sequence of position information; The formula is: X seq_temp =X seq_origin +P,X∈{A,B,D} Where P is a trainable parameter matrix; X seq_temp is the position information feature sequence; S10.3: Then X seq_temp Input a linear transformation module, perform local sequence modeling on it, and obtain a high-level linear feature sequence Among them, conv is a 2D convolutional layer with 64 input channels and 64 output channels.
7. The multimodal remote sensing image fusion method based on the multi-scale Mamba architecture according to claim 6, characterized in that: The formulas for the elevation information feature sequence and edge profile information feature sequence in step S11 are: Where ⊙ represents element-by-element multiplication, and They are elevation information feature sequence and edge contour information feature sequence respectively.
8. The multimodal remote sensing image fusion method based on the multi-scale Mamba architecture according to claim 7 is characterized in that: In step S12, the local-global features of the spectral-spatial features, the elevation information feature sequence, and the edge profile information feature sequence are: Where A seq_f Represents the local-global features extracted from high-level feature sequences containing spectral-spatial features based on the Mamba architecture; B seq_f Represents the local-global features extracted from high-level feature sequences containing elevation information based on the Mamba architecture; D seq_f Represents the local-global features extracted from high-level feature sequences containing edge contour information based on the Mamba architecture; Mamba() means extracting local-global features under the Mamba architecture.
9. The multimodal remote sensing image fusion method based on the multi-scale Mamba architecture according to claim 8, characterized in that: Step S13 includes the following steps: S13.1: A seq_f 、B seq_f and D seq_f Convert to feature map A f 、B f and D f ; The formula is: X f =reshape(X seq_f )) S13.2: Feature map A f 、B f and D f Splice along the channel to obtain the fused feature map; The formula is: F cat =Concat(A f ,B f ,D f ) Where, F cat Represents the features after splicing; S13.3: Apply the attention mechanism to the fused feature map in S13.2 to increase information interaction and obtain the final fused feature image; The formula is: F=ReLU(BN(SE(F cat ))) In the formula, SE represents the attention mechanism module; BN represents the batch normalization layer; ReLU represents the activation function layer.
Citation Information
Cited By
Hyperspectral image and laser radar data classification method based on dynamic fusion network
CN121640285A
Hyperspectral image and lidar data classification method based on dynamic fusion network
CN121640285B