Hyperspectral and LiDAR data adaptive fusion collaborative classification method based on Mama structure

Through the adaptive fusion collaborative classification method of hyperspectral and LiDAR data based on Mamba structure, the problem of difficult to take into account global dependence modeling and computing efficiency in multimodal remote sensing image classification is solved, and higher classification accuracy and computing efficiency are achieved.

CN119992269AActive Publication Date: 2025-05-13FUZHOU UNIV

Patent Information

Application Number
CN202510067194.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

In the field of multimodal remote sensing image classification, existing methods are difficult to take into account both global dependency modeling and computing efficiency, and the potential of the Mamba model in the field of multimodal classification has not yet been fully explored.

Method used

Adaptive fusion collaborative classification method of hyperspectral and LiDAR data based on Mamba structure is adopted to achieve efficient feature extraction and fusion through data preprocessing, dual-branch depth feature extraction, spatial context marking and optimization, and multimodal Mamba fusion and classification.

Benefits of technology

It realizes that while maintaining the Transformer level modeling capability, it significantly improves computing efficiency and achieves higher classification accuracy in multimodal remote sensing image classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992269A_ABST
    Figure CN119992269A_ABST
Patent Text Reader

Abstract

The invention provides a hyperspectral and LiDAR data adaptive fusion collaborative classification method based on a Mama structure. The method comprises the following steps: firstly, extracting space-spectrum joint features of HSI data and elevation semantic information of LiDAR data by using a double-branch depth feature extraction architecture; then, feature aggregation is performed and spatial representation is optimized through a spatial context marker. In a feature fusion stage, a global dependency relationship is captured through a dual-channel collaborative attention module DCCAM based on a Mama structure, meanwhile, the consistency of heterogeneous features is ensured by utilizing parameter sharing, finally, multi-source features are effectively integrated through an adaptive fusion module AF, and joint representation of information is enhanced. Compared with an existing multi-mode remote sensing image classification algorithm, the method can achieve higher classification precision and remarkably improve the calculation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimodal remote sensing image classification, and in particular to a hyperspectral and LiDAR data adaptive fusion collaborative classification method based on a Mamba structure. Background Art

[0002] In recent years, due to the advancement of satellite technology, it has become possible to capture multi-source images from various satellite sensors, and the information fusion application of hyperspectral imaging (HSI) and laser radar (LiDAR) has attracted widespread attention from researchers. HSI obtains the spectral curve of ground objects and can distinguish them by the spectral differences of different ground objects. Therefore, it is widely used in environmental monitoring, national defense, precision agriculture, mineral exploration and other fields. However, HSI lacks height information, which makes it impossible to distinguish ground objects with similar spectra but different heights.

[0003] LiDAR can obtain terrain height information and construct digital surface model data DSM (Digital Surface Model), providing the missing high-dimensional and geometric information for HSI. As different heterogeneous remote sensing data, LiDAR data and hyperspectral data can make full use of their complementary spectral, spatial details and other information to effectively make up for the shortcomings of insufficient information in single-modality data, and help significantly improve classification results.

[0004] Although the traditional Transformer model has significant advantages in global dependency modeling and parallel computing, the self-attention mechanism also brings the problem of quadratic computational complexity. Especially when processing high-dimensional data such as hyperspectral images, this computational cost increases dramatically, posing a huge challenge to the practical application of the model. In recent years, with the rapid development of the state space model (SSM) in the field of deep learning, this parallel processing time series modeling method has gradually attracted attention. The computational complexity and memory usage of SSM are linearly related to the sequence length, which is crucial for achieving efficient training and reasoning.

[0005] Recently, a new framework based on SSM, the Mamba model, was proposed. Mamba has excellent performance in computational efficiency and feature extraction capabilities, surpassing the Transformer model of the same scale in multiple audio and text tasks. Its application in the field of remote sensing has gradually become a research topic. However, the research on the Mamba model mainly focuses on single-modal remote sensing data processing.

[0006] In summary, in the field of multimodal remote sensing image classification, existing methods find it difficult to take into account both global dependency modeling and computational efficiency at the same time, and the potential of the Mamba model in the field of multimodal classification has not yet been fully explored. Summary of the invention

[0007] In view of this, the purpose of the present invention is to provide a hyperspectral and LiDAR data adaptive fusion collaborative classification method based on the Mamba structure, which can significantly improve the computational efficiency while achieving the same modeling capability as Transformer.

[0008] To achieve the above object, the present invention adopts the following technical solution: a hyperspectral and LiDAR data adaptive fusion collaborative classification method based on Mamba structure, comprising the following steps:

[0009] Step S1: Data preprocessing stage: perform data preprocessing on the hyperspectral image HSI data and the laser radar LiDAR data to convert the original data into image space blocks that meet the model input requirements;

[0010] Step S2: Dual-branch deep feature extraction: A dual-branch deep feature extraction architecture is constructed to act on the preprocessed HSI data and LiDAR data respectively; the spatial-spectral joint features of the HSI data and the elevation semantic information of the LiDAR data are extracted to capture the height differences of the objects in the vertical direction and the terrain undulation characteristics;

[0011] Step S3: Spatial context labeling and optimization: The features extracted from the two branches are aggregated by the designed spatial context labeler, and the spatial representation of the features is optimized by comprehensively considering the context information of the features in the spatial neighborhood;

[0012] Step S4: Multimodal Mamba fusion and classification: The feature sequence processed by steps S1-S3 is input into the designed multimodal Mamba fusion encoder for encoding and fusion, and then the final classification result of the ground object is output through the classification layer.

[0013] Preferably, the data preprocessing process described in step S1 is specifically as follows:

[0014] Let HSI data be X H ∈R M×N×B ,LiDAR data is Where M and N represent the width and height of the image, respectively; B represents the number of spectral bands of HSI; and C represents the number of LiDAR data channels. i ∈{1,2,…,K}, where K is the number of true labels; a single pixel in the image is represented by x i,j ∈X H ,x i,j =[x i,j,1 ,…,x i,j,B ], i=1,...M, j=1,...,N;

[0015] Using the spatial-spectral joint strategy, the spatial block is used as the input of the model. In the preprocessing stage, the HSI data X H Perform the maximum and minimum normalization operation, and then extract the cube centered at pixel (i, j) In parallel, from Li DAR data X L Extract the corresponding spatial image centered at pixel (i, j) The size of the adjacent region is P×P.

[0016] Preferably, the dual-branch deep feature extraction architecture described in step S2 includes an HSI joint feature extraction module And LiDAR spatial feature extraction module

[0017] Modules It includes the sequentially arranged Conv3D layer, batch normalization BN layer, activation layer, HetConv structure, batch normalization BN layer, activation layer, Conv2D layer, batch normalization BN layer and activation layer;

[0018] Modules It includes Conv2D layer, batch normalization BN layer, activation layer, Conv2D layer, batch normalization BN layer and activation layer arranged in sequence.

[0019] Preferably, the module In the module, the first two activation layers use the ReLU function, and the last activation layer uses the GELU function; In , both activation layers use GELU function.

[0020] Preferably, the HetConv structure is specifically as follows: two parallel Conv2D layers are used, one of which performs group convolution and the other performs point-by-point convolution, and the outputs of the two convolution layers are then element-by-element added.

[0021] Preferably, the extraction of spatial-spectral joint features of HSI data and elevation semantic information of LiDAR data is specifically as follows:

[0022] The HSI cube of size P×P×B Reshape into 1×P×P×B and pass through the module The kernel size is 3×3×9 and the padding is 1×1×0. The feature representation X is generated by the batch normalization BN layer and the activation layer. in_1 ;

[0023] The feature X in_1 After being reshaped to P×P×32, they are passed through two parallel Conv2D layers of the HetConv structure, one of which is for the reshaped feature Xin_1 Perform a group convolution with a kernel size of 3×3, a group number of 8, and a padding of 1×1, and another convolution on the reshaped feature X in_1 Perform point-by-point convolution with a kernel size of 1×1, a group number of 1, and a padding of 0×0. Add the outputs of the two convolutional layers element-wise and obtain the feature representation X through a batch normalization BN layer and an activation layer. in_2 ;

[0024] X in_2 It is input into the Conv2D layer with a kernel size of 3×3, and the spatial-spectral joint feature sequence X of HSI data is generated through the batch normalization BN layer and the activation layer. out_H ;

[0025] The LiDAR spatial image of size P×P×C Input Module For feature extraction, the kernel size of the two Conv2D layers is 3×3, the padding is 1×1×0, and the module Generate the elevation feature sequence X of LiDAR data out_L .

[0026] Preferably, the spatial context marker in step S3 specifically performs the following operations: first, a spatial-spectral joint feature sequence X of size P×P×32 is out_H and the elevation feature sequence X out_L The cube is flattened to P 2 ×32 feature blocks, and then downsampled by average pooling to smooth local information while reducing noise; and then embed the learnable absolute position code into the feature block:

[0027]

[0028] in is a learnable absolute position encoding, Flatten(·) represents a flattening operation, AP(·) represents a pooling layer, and represents the processed feature sequence, Represents element-by-element addition.

[0029] Preferably, the multimodal Mamba fusion encoder in step S4 comprises a stackable dual-channel collaborative attention module DCCAM based on the Mamba structure and an adaptive fusion module AF; wherein the dual-channel collaborative attention module DCCAM is used to perform the multimodal Mamba fusion on the feature sequence. and While capturing long-distance dependencies, parameter sharing is used to ensure feature sequence and Consistency of information processing, output characteristics and As the input of the adaptive fusion module AF; the adaptive fusion module AF uses learnable weights to perform weighted fusion of the two input features and optimizes the fusion result through layer normalization.

[0030] Preferably, the dual-channel collaborative attention module DCCAM specifically performs the following operations:

[0031] For input features Perform layer normalization, then use the linear layer for processing, and then split the features to obtain the main features and auxiliary features for gated MLP

[0032]

[0033] Where i~(H,L) represents the modality type, Chunk(·) represents the block operation, Linear(·) represents the linear layer, and LN(·) represents the layer normalization;

[0034] The SSM module is then used to model the I / O relationship of the backbone features and then perform layer normalization. The auxiliary features are processed through a parameter-sharing CNN layer, Share = Conv2D(·), and the SiLU function is selected for activation. The activated features are element-wise multiplied with the backbone features after layer normalization to obtain the weighted backbone features. The specific process is as follows:

[0035]

[0036] Among them, Share(·) represents the parameter sharing layer, represents element-wise multiplication, Represents the backbone features after layer normalization. Represents the auxiliary features after SiLU function activation.

[0037] Preferably, the adaptive fusion module AF specifically performs the following operations:

[0038] The adaptive fusion module AF initializes a random vector with the same length as the feature sequence The random vector v is activated by the sigmoid function and the output of the dual-channel collaborative attention module DCCAM and The weighted features are weighted, and finally the weighted features are added to complete the feature fusion and the layer normalization operation is performed through the softmax function. The specific process is expressed as follows:

[0039]

[0040] In the formula, softmax(·) represents the softmax function; sigmoid(·) represents the sigmoid function;

[0041] The fused feature X after normalization operation fused is fed into the classification head to obtain the final classification result.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] In view of the fact that existing methods are difficult to take into account both global dependency modeling and computational efficiency at the same time, and the potential of the Mamba model in the field of multimodal classification has not been fully explored, the present invention constructs a hyperspectral and LiDAR data adaptive fusion collaborative classification network based on the Mamba structure for classification. Compared with the existing multimodal remote sensing image classification algorithm, the method of the present invention can achieve higher classification accuracy and significantly improve computational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A schematic diagram of a network architecture of a preferred embodiment of the present invention;

[0045] Figure 2 A schematic diagram of a feature extraction process in a preferred embodiment of the present invention;

[0046] Figure 3 Schematic diagram of the feature fusion process of a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0047] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings.

[0048] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.

[0049] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or their combinations.

[0050] like Figure 1-3 As shown, this embodiment provides a hyperspectral and LiDAR data adaptive fusion collaborative classification method based on Mamba structure, comprising the following steps:

[0051] Step S1: Data preprocessing stage: perform data preprocessing on the hyperspectral image HSI data and the laser radar LiDAR data to convert the original data into image space blocks that meet the model input requirements;

[0052] Step S2: Dual-branch deep feature extraction: A dual-branch deep feature extraction architecture is constructed, which is applied to the pre-processed HSI data and LiDAR data respectively. For HSI data, the focus is on mining its spatial-spectral joint features; for LiDAR data, the focus is on extracting its elevation semantic information, accurately capturing the height differences of objects in the vertical direction and the terrain undulation characteristics;

[0053] Step S3: Spatial context labeling and optimization: The features extracted from the dual branches are aggregated by the designed spatial context labeler, and the spatial representation of the features is further optimized by comprehensively considering the context information of the features in the spatial neighborhood;

[0054] Step S4: Multimodal Mamba fusion and classification: The feature sequence processed by the above steps is input into the designed multimodal Mamba fusion encoder for encoding and fusion, and then the final classification result of the ground object is output through the classification layer.

[0055] In this embodiment, the data preprocessing process described in step S1 is specifically as follows:

[0056] Let HSI data be X H ∈R M×N×B ,LiDAR data is Where M and N represent the width and height of the image, B represents the number of spectral bands of HSI, and C represents the number of LiDAR data channels. i ∈{1,2,…,K}, where K is the number of true labels. A single pixel in an image can be represented as x i,j ∈X H ,x i,j =[x i,j,1 ,…,x i,j,B ], B is the number of spectral bands, i=1,...M, j=1,...,N.

[0057] The method of the present invention uses the strategy of space-spectrum combination, takes the space block as the input of the model, and firstly processes the HSI data X H Perform the maximum and minimum normalization operation, and then extract the cube centered at pixel (i, j) In parallel, L Extract the corresponding spatial image centered at pixel (i, j) The size of the adjacent region is (P×P).

[0058] In this embodiment, the dual-branch deep feature extraction architecture described in step S2 includes an HSI joint feature extraction module And LiDAR spatial feature extraction module

[0059] Modules It includes the sequentially arranged Conv3D layer, batch normalization BN layer, ReLU function activation layer, HetConv structure, batch normalization BN layer, ReLU function activation layer, Conv2D layer, batch normalization BN layer and GELU activation layer;

[0060] The HetConv structure is as follows: two parallel Conv2D layers are used, one of which performs group convolution and the other performs point-by-point convolution, and the outputs of the two convolutional layers are then element-wise added ( ).

[0061] Modules It includes Conv2D layer, batch normalization BN layer, GELU activation layer, Conv2D layer, batch normalization BN layer and GELU activation layer arranged in sequence.

[0062] Extract the spatial-spectral joint features of HSI data and the elevation semantic information of LiDAR data, specifically:

[0063] First, the HSI cube of size (P×P×B) is Reshape into (1×P×P×B) and pass the module It is processed by a Conv3D layer with a kernel size of (3×3×9), and then the feature representation X is generated through a batch normalization BN layer and an activation layer. in_1 ,In order to keep the spatial height and width of the image unchanged, the padding used in the Conv3D layer is (1×1×0).

[0064] Next, feature X in_1 After being reshaped to P×P×32, they are passed through two parallel Conv2D layers of the HetConv structure, one of which is for the reshaped feature X in_1 Perform a group convolution with a kernel size of 3×3, a group number of 8, and a padding of 1×1, and another convolution on the reshaped feature X in_1 Perform a point-by-point convolution with a kernel size of 1×1, a group number of 1, and a padding of 0×0, and add the outputs of the two convolutional layers element-wise ( ), and obtain the feature representation X through batch normalization BN layer and activation layer in_2

[0065] Finally, X in_2It is input into a Conv2D layer with a kernel size of 3×3, and passes through a batch normalization BN layer and an activation layer to generate the spatial-spectral joint feature sequence X of the HSI data. out_H .

[0066] The LiDAR spatial image of size P×P×C Input Module For feature extraction, the kernel size of the two Conv2D layers is 3×3 and the padding is 1×1×0 to generate the elevation feature sequence X of the LiDAR data. out_L .

[0067] In this embodiment, the spatial context marker described in step S3 is processed as follows: first, the spatial-spectral joint feature sequence X of size (P×P×32) is out_H and the elevation feature sequence X out_L The cube is flattened to (P 2 ×32) feature blocks, which are then downsampled by average pooling to smooth local information while reducing noise. The learnable absolute position encoding is then embedded into the feature blocks, which enables the model to capture the position information of each element in the input data, thereby better understanding the order or spatial relationship of the data. The above process can be summarized by the following equation:

[0068]

[0069] in is a learnable absolute position encoding, Flatten(·) represents a flattening operation, AP(·) represents a pooling layer, and represents the processed feature sequence, Represents element-by-element addition.

[0070] In this embodiment, the multimodal Mamba fusion encoder described in step S4 includes a stackable dual channel collaborative attention module DCCAM (Dual Channel Collaborative Attention Module) based on the Mamba structure and an adaptive fusion module AF (Adaptive Fusion Block). DCCAM is used to analyze the HSI and LiDAR information features (feature sequence and ) captures long-range dependencies while using parameter sharing to ensure the consistency of HSI and LiDAR information processing, and maximizes the correlation and complementarity between HSI and LiDAR information. Output features and As the input of the adaptive fusion module AF; AF uses learnable weights to perform weighted fusion of the two input features and optimizes the fusion result through layer normalization.

[0071] In this embodiment, the specific processing process of DCCAM is as follows: first, the input feature Perform layer normalization, then use the linear layer for processing, and then split the features to obtain the main features and auxiliary features for gated MLP

[0072]

[0073] Where i~(H,L) represents the modality type, Chunk(·) represents the block operation, Linear(·) represents the linear layer, and LN(·) represents layer normalization.

[0074] After that, the SSM module is used to model the I / O (Input and Output) relationship of the backbone features and then perform layer normalization, while the auxiliary features are processed through a parameter-sharing CNN layer, Share = Conv2D (·). The SiLU function is selected for activation, and the activated features are element-wise multiplied with the backbone features after layer normalization to obtain the weighted backbone features. The specific process is as follows:

[0075]

[0076] Among them, Share(·) represents the parameter sharing layer, represents element-wise multiplication, Represents the backbone features after layer normalization. Represents the auxiliary features after SiLU function activation.

[0077] In this embodiment, the specific processing process of the AF module is as follows: the AF module initializes a random vector with the same length as the feature sequence The random vector v is activated by the sigmoid function and output by the DCCAM module and The weighted features are weighted, and finally the weighted features are added to complete the feature fusion and the layer normalization operation is performed through the softmax function. The process can be expressed as:

[0078]

[0079] In the formula, softmax(·) represents the softmax function; sigmoid(·) represents the sigmoid function;

[0080] Finally, the fused feature X after normalization operation fused is fed into the classification head to obtain the final classification result.

[0081] In summary, in view of the fact that existing methods are difficult to take into account both global dependency modeling and computational efficiency at the same time, and the potential of the Mamba model in the field of multimodal classification has not been fully explored, the hyperspectral and LiDAR data adaptive fusion collaborative classification network based on the Mamba structure can achieve higher classification accuracy and significantly improve computational efficiency compared with the existing multimodal remote sensing image classification algorithms.

[0082] The above description is only a preferred embodiment of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention should fall within the scope of the present invention.

Claims

1. A hyperspectral and LiDAR data adaptive fusion collaborative classification method based on Mamba structure, characterized by: The following steps are involved: Step S1: Data preprocessing stage: perform data preprocessing on the hyperspectral image HSI data and the laser radar LiDAR data to convert the original data into image space blocks that meet the model input requirements; Step S2: Dual-branch deep feature extraction: A dual-branch deep feature extraction architecture is constructed to act on the preprocessed HSI data and LiDAR data respectively; the spatial-spectral joint features of the HSI data and the elevation semantic information of the LiDAR data are extracted to capture the height differences of the objects in the vertical direction and the terrain undulation characteristics; Step S3: Spatial context labeling and optimization: The features extracted from the two branches are aggregated by the designed spatial context labeler, and the spatial representation of the features is optimized by comprehensively considering the context information of the features in the spatial neighborhood; Step S4: Multimodal Mamba fusion and classification: The feature sequence processed by steps S1-S3 is input into the designed multimodal Mamba fusion encoder for encoding and fusion, and then the final classification result of the ground object is output through the classification layer.

2. The adaptive fusion collaborative classification method of hyperspectral and LiDAR data based on Mamba structure according to claim 1 is characterized in that: The data preprocessing process described in step S1 is specifically as follows: Let HSI data be X H ∈R M×N×B , LiDAR data is Where M and N represent the width and height of the image, respectively; B represents the number of spectral bands of HSI; and C represents the number of LiDAR data channels. i ∈{1, 2, ..., K}, where K is the number of true labels; a single pixel in the image is represented by x i,j ∈X H , x i,j =[x i,j,1 , ..., x i,j,B ], i=1,...M, j=1,...,N; Using the spatial-spectral joint strategy, the spatial block is used as the input of the model. In the preprocessing stage, the HSI data X H Perform the maximum and minimum normalization operation, and then extract the cube centered at pixel (i, j) In parallel, L Extract the corresponding spatial image centered at pixel (i, j) The size of the adjacent region is P×P.

3. The adaptive fusion collaborative classification method of hyperspectral and LiDAR data based on Mamba structure according to claim 2 is characterized in that: The dual-branch deep feature extraction architecture described in step S2 includes the HSI joint feature extraction module And LiDAR spatial feature extraction module Modules It includes the sequentially arranged Conv3D layer, batch normalization BN layer, activation layer, HetConv structure, batch normalization BN layer, activation layer, Conv2D layer, batch normalization BN layer and activation layer; Modules It includes Conv2D layer, batch normalization BN layer, activation layer, Conv2D layer, batch normalization BN layer and activation layer arranged in sequence.

4. The method for adaptive fusion and collaborative classification of hyperspectral and LiDAR data based on Mamba structure according to claim 3 is characterized in that: The module In the module, the first two activation layers use the ReLU function, and the last activation layer uses the GELU function; In , both activation layers use GELU function.

5. The method for adaptive fusion and collaborative classification of hyperspectral and LiDAR data based on Mamba structure according to claim 4 is characterized in that: The HetConv structure is specifically as follows: two parallel Conv2D layers are used, one of which performs group convolution and the other performs point-by-point convolution, and the outputs of the two convolution layers are then element-by-element added.

6. The method for adaptive fusion and collaborative classification of hyperspectral and LiDAR data based on Mamba structure according to claim 5 is characterized in that: The extraction of the spatial-spectral joint features of the HSI data and the elevation semantic information of the LiDAR data is specifically as follows: The HSI cube of size P×P×B Reshape into 1×P×P×B and pass through the module The kernel size is 3×3×9 and the padding is 1×1×0. The feature representation X is generated by the batch normalization BN layer and the activation layer. in_1 ; The feature X in_1 After being reshaped to P×P×32, they are passed through two parallel Conv2D layers of the HetConv structure, one of which is for the reshaped feature X in_1 Perform a group convolution with a kernel size of 3×3, a group number of 8, and a padding of 1×1, and another convolution on the reshaped feature X in_1 Perform point-by-point convolution with a kernel size of 1×1, a group number of 1, and a padding of 0×0. Add the outputs of the two convolutional layers element-wise and obtain the feature representation X through a batch normalization BN layer and an activation layer. in_2 ; X in_2 It is input into the Conv2D layer with a kernel size of 3×3, and the spatial-spectral joint feature sequence X of HSI data is generated through the batch normalization BN layer and the activation layer. out_H ; The LiDAR spatial image of size P×P×C Input Module For feature extraction, the kernel size of the two Conv2D layers is 3×3, the padding is 1×1×0, and the module Generate the elevation feature sequence X of LiDAR data out_L .

7. The method for adaptive fusion and collaborative classification of hyperspectral and LiDAR data based on Mamba structure according to claim 1, characterized in that: The spatial context tagger in step S3 specifically performs the following operations: first, a spatial-spectral joint feature sequence X of size P×P×32 is out_H and the elevation feature sequence X out_L The cube is flattened to P 2 ×32 feature blocks, and then downsampled by average pooling to smooth local information while reducing noise; and then embed the learnable absolute position code into the feature block: in is a learnable absolute position encoding, Flatten(·) represents a flattening operation, AP(·) represents a pooling layer, and represents the processed feature sequence, Represents element-by-element addition.

8. The method for adaptive fusion and collaborative classification of hyperspectral and LiDAR data based on Mamba structure according to claim 1, characterized in that: The multimodal Mamba fusion encoder described in step S4 includes a stackable Mamba-based dual-channel collaborative attention module DCCAM and an adaptive fusion module AF; The dual-channel collaborative attention module DCCAM is used to analyze the feature sequence and While capturing long-distance dependencies, parameter sharing is used to ensure feature sequence and Consistency of information processing, output characteristics and As the input of adaptive fusion module AF; The adaptive fusion module AF uses learnable weights to perform weighted fusion of two input features and optimizes the fusion result through layer normalization.

9. The method for adaptive fusion and collaborative classification of hyperspectral and LiDAR data based on Mamba structure according to claim 8, characterized in that: The dual-channel collaborative attention module DCCAM specifically performs the following operations: For input features Perform layer normalization, then use the linear layer for processing, and then split the features to obtain the main features and auxiliary features for gated MLP Where i~(H, L) represents the modality type, Chunk(·) represents the chunking operation, Linear(·) represents the linear layer, and LN(·) represents the layer normalization; The SSM module is then used to model the I / O relationship of the backbone features and then perform layer normalization. The auxiliary features are processed through a parameter-sharing CNN layer, Share = Conv2D(·), and the SiLU function is selected for activation. The activated features are element-wise multiplied with the backbone features after layer normalization to obtain the weighted backbone features. The specific process is as follows: Among them, Share(·) represents the parameter sharing layer, represents element-wise multiplication, Represents the backbone features after layer normalization. Represents the auxiliary features after SiLU function activation.

10. The method for adaptive fusion and collaborative classification of hyperspectral and LiDAR data based on Mamba structure according to claim 9, characterized in that: The adaptive fusion module AF specifically performs the following operations: The adaptive fusion module AF initializes a random vector with the same length as the feature sequence The random vector v is activated by the sigmoid function and the output of the dual-channel collaborative attention module DCCAM and The weighted features are weighted, and finally the weighted features are added to complete the feature fusion and the layer normalization operation is performed through the softmax function. The specific process is expressed as follows: In the formula, softmax(·) represents the softmax function; sigmoid(·) represents the sigmoid function; After normalization, the fusion feature X fused is fed into the classification head to obtain the final classification result.

Citation Information

Patent Citations

  • Double-branch CNN-Transformer-based hyperspectral and LiDAR collaborative crop precise classification method

    CN117953259A

  • Hyperspectral and laser radar data fusion classification method based on interactive Transform-CNN (Convolutional Neural Network)

    CN117994616A

  • Unmanned vehicle robust position identification method based on look-around image

    CN119131740A

Cited By

  • Coronary artery affine registration method, equipment, medium and product

    CN121582307A

  • A coronary artery affine registration method, apparatus, medium, and product

    CN121582307B