Ground feature classification method based on hyperspectral image and LiDAR data fusion of improved Mama network
By improving the Mamba network, combining space-spectral adaptation and elevation enhancement modules, the fusion of hyperspectral images and LiDAR data is solved, and the problem of low span modal feature extraction and fusion efficiency in the existing technology is achieved, and higher geographic classification accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202510127008.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-28
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art has limitations in the fusion of hyperspectral images and LiDAR data, including the limited local perception of convolutional neural networks, the computational complexity of Transformers, and the efficiency of the Mamba architecture in cross-modal feature extraction and fusion.
The hyperspectral image and LiDAR data fusion method based on the improved Mamba network is used to extract the spatial-spectral features of the hyperspectral image through the spatial-spectral adaptive Mamba module, and the elevation enhancement Mamba module is used to extract the elevation characteristics of the LiDAR data, combined with the enhanced fusion module for feature fusion, and finally the classification results are generated by the multi-layer perceptron classifier.
Improve the accuracy and adaptability of land objects classification, overcome the limitations of a single data source, and demonstrate superior performance in complex terrain and different land cover scenarios.
Smart Images

Figure CN119942228A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of ground object classification methods, in particular to a ground object classification method based on the fusion of hyperspectral images and LiDAR data based on an improved Mamba network. Background Art
[0002] With the rapid development of remote sensing and imaging technology, it is now easier to capture a large amount of multi-band and multi-modal remote sensing data. These data not only help people observe the earth from different perspectives, but also bring challenges in processing and applying them to a wide range of remote sensing applications, such as ground material identification, mineral exploration and mapping, scene understanding, object detection, and change monitoring. In this paper, the present invention focuses on land cover classification, which is a key problem in ground material identification. It is becoming increasingly important to meet the growing needs of precision agriculture, forestry, urban planning, environmental monitoring, disaster response, and social governance.
[0003] Hyperspectral images are obtained through multi-band sensors that provide rich spectral information and can more accurately distinguish different types of materials. LiDAR data captures detailed spatial structural information of ground objects through laser scanning, accurately depicting features such as ground morphology and object height. Hyperspectral images and LiDAR data have highly complementary characteristics. Hyperspectral images provide detailed spectral information that helps to identify materials and components, while LiDAR data improves spatial accuracy and provides more detailed classification representations through precise spatial position and morphological features. Therefore, a joint classification method combining these two data sources can overcome the limitations of a single data source, improve classification accuracy, and show stronger adaptability and robustness in complex terrain or different land cover scenes. Combining hyperspectral images and LiDAR data to improve classification performance is becoming an increasingly popular research topic, while existing methods have the following problems: limited local perception of convolutional neural networks, quadratic computational complexity of Transformers, and how the Mamba architecture can effectively achieve cross-modal feature extraction and fusion. In order to solve the above problems, the present invention proposes a ground object classification method based on the fusion of hyperspectral images and LiDAR data based on an improved Mamba network. Summary of the invention
[0004] The purpose of the present invention is to provide a method for classifying objects by fusing hyperspectral images and LiDAR data based on an improved Mamba network, so as to solve the problems raised in the above-mentioned background technology.
[0005] To achieve the above-mentioned purpose, the present invention provides the following technical scheme: a method for classifying objects based on the fusion of hyperspectral images and LiDAR data using an improved Mamba network, comprising the following steps: step S1, extracting spatial-spectral features from the original hyperspectral image through a spatial-spectral adaptive Mamba module, and inputting the original LiDAR data into an elevation enhancement Mamba module to extract elevation features; step S2, using an enhanced fusion module to capture intra-modal and inter-modal relationships, and fusing the extracted features; step S3, inputting the fused feature set into a classifier to generate a final classification result.
[0006] Preferably: the feature extraction of the spatial-spectral adaptive Mamba module in step S1 uses pixel-level embedding to process the raw data, which includes a 2D convolution layer and a position encoding layer. The 2D convolution layer consists of a kernel size of 1×1 and an output channel of C. h The 2D convolution is composed of a layer of ReLU activation function and a batch normalization layer, and a set of learnable parameters are introduced as position encoding (PE). The initial value of the parameter is set to 1 and updated during the training process. The above process is expressed as follows:
[0007] X HSI =BN(ReLU(Conv2D(X hsi )))+PE
[0008] Where X hsi is the raw data of the hyperspectral image, and the above results are expressed as X HSI , X HSI The shape is (C h , P,P).
[0009] Preferably: the step S1 adopts four different methods to scan the patch block, wherein route 1 scans the patch block row by row from left to right and from top to bottom, the scanning direction of route 2 is opposite to the scanning direction of route 1, route 3 scans the patch block column by column from top to bottom and from left to right, and the scanning method of route 4 is opposite to the scanning method of route 3, data is read through the above routes and four corresponding different 1D sequences are obtained, and then the sequences are input into the spatial route Mamba for feature extraction.
[0010] Preferably: the step S1 applies a reshape operation to transpose the input data, and then scans the spectral data of the patch block from the front and back into 1D sequences, respectively referred to as path 1 and path 2, and then these sequences are input into the bidirectional route Mamba model for feature extraction. The operations performed on these sequences are as follows: first, a linear projection is applied to convert each sequence into a state space for selection, and then a 1D convolution layer is used to extract local features between spectral bands, mainly including a 1D convolution with a kernel size of 3 and a Sigmoid activation function. Then, the extracted features are further extracted through the S6 block to generate a Y for each sequence. route_i The resulting feature map is , where i represents the route order. The whole process is mathematically represented as follows:
[0011] Y route_i =S6(Sigmod(Conv1D(Linear(route i ))))
[0012] Among them, route i Indicates different routes. On the other hand, the route1 sequence also enters another branch, performs linear projection and SiLU activation function operations, and the output is marked as Y route_down , the above operation is expressed by the following formula:
[0013] Y route_down =SiLU(Linear(route 1 ))
[0014] Then, Y route_down Multiply by Y route_i Get the output, marked as Y route_out_i. . We integrate these feature maps by adding them together to obtain the result Y spe_out Finally, linear projection is applied to map these features back to the original space to obtain the spectral feature map Y out_spe , then Y out_spa and Y out_spe Multiply them to get Y out_hsi Then a residual connection is used to obtain the final hyperspectral image feature map Y HSI_out , including adding Y out_spa , Y out_spe and Y out_hsi , the above operation formula is expressed as:
[0015] Y out_hsi =Y out_spa ×Y out_spe
[0016] Y HSI_out =Y out_hsi +Y out_spa+Y out_spe .
[0017] Preferably: the elevation enhancement Mamba module in step S1 first divides the data into fixed-size patches using a center pixel filling method, and then uses pixel-level embedding to enhance the representation of elevation information, including a 2D convolutional layer and position encoding. The 2D convolutional layer is first used to change the elevation dimension of the data, including a kernel size of 1×1 and an output channel of C. l A 2D convolution, a ReLU activation function, and a batch normalization layer are then used to add positional encoding to the data. The result is denoted as X LiDAR , its shape is (p×p, C 1 ), the vector is scanned back and forth into two 1D sequences, marked as route 1 and route 2. The above operation formula is expressed as:
[0018] X lidar_route1 =S6(Couv1D(Linear(route1)))
[0019] X lidar_route2 =S6(Couv1D(Linear(route2)))
[0020] Where X lidar_route1 and X lidar_route2 They represent the outputs of the two routes respectively. On the other hand, route 1 performs the operation of linear projection and SiLU activation function, and the output is represented as X lidar_down Then, X lidar_down Multiply by X lidar_route1 and X lidar_route2 , respectively get X lidar_out1 and X lidar_out2 , the preliminary feature map is obtained by the following operations:
[0021] X lidar_fout =Linear(X lidar_out1 +X lidar_out2 )
[0022] Among them, Linear(·) is a linear projection that maps the feature map to its original space, X lidar_fout is the preliminary feature map. Finally, the residual connection is used to convert X lidar_fout and X LiDAR Add together to get denoted as Y LiDAR_out The final elevation feature map.
[0023] Preferably: In step S2, the hyperspectral image and LiDAR data feature maps are first enhanced using the Mamba architecture respectively, and the process is as follows: Linear projection maps the feature map to the specified space, and then these features pass through two branches. In the first branch, the Conv1D function with a convolution kernel size of 3 is used to further extract local features, and the Sigmoid activation function is applied to introduce nonlinearity. These features are then passed to the S6 block, and the result is represented as Y en_fe_up In the second branch, the SiLU activation function is used to introduce nonlinearity to enhance the features, and the result is expressed as Y en_fe_down , then multiply the outputs of the two branches to get the final output Y en_fe Finally, the enhanced features are mapped back to the original space using linear projection. The feature map is enhanced through the above process. The enhanced hyperspectral image and LiDAR data features are represented as Y en_HSI and Y en_LiDAR , the above operation can be expressed as:
[0024] Y en_fe =Linear( Yen_fe_up ×Y en_fe_down )
[0025] Y out =Y en_HSI +Y en_LiDAR +Y LiDAR_out +Y HSI_out
[0026] where Y en_fe is a pseudo symbol representing the enhanced hyperspectral image or LiDAR data feature map, Y out Represents the final extracted feature map.
[0027] Preferably: the classifier of step S3 is designed based on a simple multilayer perceptron framework. Specifically, the features are first flattened into (B, P×P×S), where B represents the batch size, P represents the patch size, and S is set to 64. The multilayer perceptron classification has two fully connected layers. The first layer performs a linear transformation to map the input feature vector to a hidden dimension, denoted as hid, and the dimension is regarded as a hyperparameter. A ReLU activation function is applied after the first layer to introduce nonlinearity, and a dropout operation is added between layers. The second fully connected layer acts as a classification head to project the output of the hidden layer to the final number of categories. Finally, a cross entropy loss function is used to calculate the difference between the probability distribution predicted by the model and the true label.
[0028] Compared with the prior art, the present invention has the following beneficial effects:
[0029] The present invention discloses a joint classification method for hyperspectral images and LiDAR data based on Mamba architecture, which is called multimodal fusion Mamba network. In the feature extraction stage, the present invention introduces spatial-spectral adaptive Mamba, which aims to capture rich spatial and spectral information from hyperspectral images. It combines the scanning mechanism with the enhanced Mamba block to effectively extract features from spatial and spectral dimensions. For LiDAR data, elevation enhanced Mamba is used to extract detailed elevation information. Next, the enhanced fusion module integrates intra-modal and inter-modal relationships at the pixel level. The classifier is constructed with a simple linear structure. The present invention verifies the network proposed in the present invention on MUUFL, Houston2013 and Trento datasets, demonstrates its superiority, and highlights the effectiveness of edge contour information in multimodal joint classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a flow chart of the method of the present invention;
[0031] Figure 2 It is the detailed architecture of the spatial-spectral adaptive Mamba of the present invention;
[0032] Figure 3 It is the Mamba structure of different routes of the present invention;
[0033] Figure 4 It is the structure of the enhanced fusion module of the present invention;
[0034] Figure 5 The MUUFL dataset of an embodiment of the present invention is as follows: (a) pseudo-color composite image (30, 15 and 7 bands); (b) LiDAR-derived DSM image; (c) ground truth map;
[0035] Figure 6 The Houston 2013 dataset of the embodiment of the present invention: (a) pseudo-color composite image (64, 30 and 20 bands); (b) LiDAR-derived DSM image; (c) ground truth map;
[0036] Figure 7 The Trento dataset of an embodiment of the present invention: (a) pseudo-color composite image (64, 30 and 20 bands); (b) LiDAR-derived DSM image; (c) ground truth map;
[0037] Figure 8 It is a full pixel classification map of the MUUFL dataset of the embodiment of the present invention. (a) CrossHL (b) Am3 (c) ExViT (d) HCT (e) CMFAEN (f) CoupledCNN (g) S2Enet (h) Ours;
[0038] Fig. 9 It is a full pixel classification map of the Houston2013 dataset in an embodiment of the present invention. (a) CrossHL (b) Am3 (c) ExViT (d) HCT (e) CMFAEN (f) CoupledCNN (g) S2Enet (h) Ours;
[0039] Fig.10 It is a full pixel classification map of the Trento dataset according to an embodiment of the present invention. (a) CrossHL (b) Am3 (c) ExViT (d) HCT (e) CMFAEN (f) CoupledCNN (g) S2Enet (h) Ours. DETAILED DESCRIPTION
[0040] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0041] Example
[0042] See also Figure 1 The method for classifying objects based on the fusion of hyperspectral image and LiDAR data based on the improved Mamba network shown in the figure includes the following steps: Step S1, the original hyperspectral image is subjected to the spatial-spectral adaptive Mamba module to extract spatial-spectral features, and the original LiDAR data is input into the elevation enhancement Mamba module to extract elevation features; Step S2, the enhanced fusion module is used to capture the intra-modal and inter-modal relationships, and the extracted features are fused; Step S3, the fused feature set is input into the classifier to generate the final classification result.
[0043] In this embodiment, a spatial-spectral adaptive Mamba tailored for hyperspectral images is proposed. Figure 2 The detailed architecture of the spatial-spectral adaptive Mamba is shown in Figure 2. Assuming that the hyperspectral image is represented by X hsi and Where H and W represent the height and width of the image, respectively, and L represents the spectral dimension of the hyperspectral image. In order to apply the Mamba architecture to the joint classification task of hyperspectral images and LiDAR data, the image is divided into multiple patches of equal size based on ViT. Specifically, a filling method based on the center pixel is adopted to ensure that the edge information of the image is preserved. Since hyperspectral images contain rich spectral bands, this brings challenges in terms of computational and memory requirements. To alleviate this problem, pixel-level embedding is used to process the raw data, which includes 2D convolutional layers and position encoding. The 2D convolutional layer consists of a kernel size of 1×1 and an output channel of C. h The 2D convolution is composed of a layer of ReLU activation function and a batch normalization layer, and a set of learnable parameters are introduced as position encoding (PE). The initial value of these parameters is set to 1 and updated during the training process. The above operation is expressed as follows:
[0044] X HSI =BN(ReLU(Conv2D(X hsi )))+PE
[0045] Where X hsi is the raw data of the hyperspectral image, and the result is expressed as X HSI , X HSI The shape is (C h , P, P) Subsequently, features are extracted from the spatial and spectral levels, respectively.
[0046] In order to effectively extract spatial information from hyperspectral images, a new scanning method is first designed. This method scans the patch blocks in four different ways. Route 1 scans the patch blocks row by row from left to right and from top to bottom. The scanning direction of route 2 is opposite to that of route 1. Route 3 scans the patch blocks column by column from top to bottom and from left to right, and the scanning method of route 4 is opposite to that of route 3. By using this method to read data in the horizontal and vertical directions, four different 1D sequences are obtained. These sequences can more comprehensively represent the spatial information and achieve a balance between the global and local perspectives. The sequences are then input into the spatial route Mamba for feature extraction, as shown in Figure 2. Figure 3 As shown, each route performs similar operations in parallel: linear projection maps the sequence to the state selection space denoted as mid_dim, 1D convolution with a kernel size of 3 is used to integrate local spatial features, Sigmoid activation function is used to introduce nonlinear factors, S6 block is used to extract the final features from the sequence, and the feature label extracted from each sequence is denoted as Y route_i , where i represents the route order. In addition, route 1 is sent to another branch, which includes a linear projection layer and a SiLU activation function. The above operation is expressed by the following formula:
[0047] Y route_i =S6(Sigmod(Conv1D(Linear(route i ))))
[0048] Y route_down =SiLU(Linear(route 1 ))
[0049] where Y route_down Marked as the obtained output, next, Y route_i Multiply by Y route_down , we get Y route_out_i , these feature maps are integrated through the following operations:
[0050]
[0051] where Y out_spa Represents the final spatial feature map.
[0052] Furthermore, in order to make full use of the rich spectral information of hyperspectral imaging, a reshaping operation is first applied to transpose the input data, and then a new bidirectional scanning method is proposed to enhance the features of the spectral information. This method involves scanning the spectral data of the patch block from the front and back sides into 1D sequences, called path 1 and path 2, respectively. These sequences are then input into the bidirectional route Mamba model for feature extraction, as shown in Figure 3 As shown in Figure 1, the operations performed on these sequences are as follows: First, a linear projection is applied to transform each sequence into a state space for selection. Then, a 1D convolution layer is used to extract local features between spectral bands, mainly including a 1D convolution with a kernel size of 3 and a Sigmoid activation function. The extracted features are then further extracted through the S6 block to generate a Y representation for each sequence. route_i The resulting feature map is , where i represents the route order. The whole process is mathematically represented as follows:
[0053] Y route_i =S6(Sigmod(Conv1D(Linear(route i ))))
[0054] Among them, route i Indicates different routes. On the other hand, the route1 sequence also enters another branch, performs linear projection and SiLU activation function operations, and the output is marked as Y route_down , the above operation is expressed by the following formula:
[0055] Y route_down =SiLU(Linear(route 1 ))
[0056] Then, Y route_down Multiply by Y route_i Get the output, marked as Y route_out_i ., integrate these feature maps by adding them together to obtain the result Y spe_out Finally, linear projection is applied to map these features back to the original space to obtain the spectral feature map Y out_spe .
[0057] In order to effectively integrate these rich spatial and spectral features, we first transform Y out_spa and Y out_spe Multiply them to get Y out_hsi , and then use residual connection to obtain the final hyperspectral image feature map Y HSI_out , including adding Y out_spa , Y out_spe and Y out_hsi , through these operations, the spatial and spectral features can be effectively integrated. The above operation formula is expressed as:
[0058] Y out_hsi =Y out_spa ×Y out_spe
[0059] Y HSI_out =Y out_hsi +Y out_spa +Y out_spe .
[0060] In this embodiment, LiDAR data provides rich elevation information, which can be used as a supplement to hyperspectral images to improve classification performance. By using the powerful reasoning ability of the state-space model, an elevation enhancement Mamba method based on the Mamba architecture is proposed to effectively enhance and extract elevation information from LiDAR data. The overall structure of the elevation enhancement Mamba is as follows: Figure 2 shown.
[0061] Assume that the LiDAR data is represented by X lidar and Where H and W represent the height and width of the image, respectively. Similar to the processing method of hyperspectral images, the data is first divided into fixed-size patches using the center pixel filling method, and then pixel-level embedding is used to enhance the representation of elevation information, which includes a 2D convolutional layer and position encoding. The 2D convolutional layer is first used to change the elevation dimension of the data, including a kernel size of 1×1 and an output channel of C. 12D convolution, a ReLU activation function and a batch normalization layer. Subsequently, the position encoding is added to the data, where it shares the position encoding of the hyperspectral image. The shared position encoding can provide more accurate position information, enhance the understanding ability of the model, and help to fuse multimodal features. The result is represented by X LiDAR , its shape is (p×p, C 1 ). The vector is scanned forward and backward into two 1D sequences, labeled as route 1 and route 2. The bidirectional route Mamba method is used to extract features from these sequences. Similar to spectral feature extraction, on the one hand, the route 1 and route 2 sequences perform the following operations: linear projection, conv1D layer and S6 block, respectively. In particular, the conv1D layer includes a 1D convolution with a kernel size of 5, Sigmod activation function and batch normalization. The above operation formula is expressed as:
[0062] X lidar_route1 =S6(Couv1D(Linear(route1)))
[0063] X lidar_route2 =S6(Couv1D(Linear(route2)))
[0064] Where X lidar_route1 and X lidar_route2 They represent the outputs of the two routes respectively. On the other hand, route 1 performs the operation of linear projection and SiLU activation function, and the output is represented as X lidar_down , then, X lidar_down Multiply by X lidar_route1 and X lidar_route2 , respectively get X lidar_out1 and X lidar_out2 The preliminary feature map is obtained by the following operations:
[0065] X lidar_fout =Linear(X lidar_out1 +X lidar_out2 )
[0066] Among them, Linear(·) is a linear projection that maps the feature map to its original space, X lidar_fout is a preliminary feature map. Finally, residual connections are used to alleviate gradient vanishing and gradient problems. lidar_fout and X LiDAR Add together to get denoted as Y LiDAR_out The final elevation feature map.
[0067] Furthermore, in order to integrate the intra-modal and inter-modal relationships and achieve pixel-level fusion, an enhanced fusion module is adopted, which uses a series of simplified Mamba blocks and residual connections to effectively fuse these heterogeneous features. Figure 4 A detailed description of the module structure is provided. The hyperspectral image and LiDAR data feature maps are first enhanced using the Mamba architecture respectively. The Mamba architecture works as follows: Linear projection maps the feature map to the specified space. Then, these features pass through two branches. In the first branch, the Conv1D function with a convolution kernel size of 3 is used to further extract local features. In order to enhance the expressiveness and generalization ability of the model, the Sigmoid activation function is applied to introduce nonlinearity. Then, these features are further enhanced by passing them to the S6 block. The result is represented as Y en_fe_up In the second branch, the SiLU activation function is used to introduce nonlinearity to enhance the features, and the result is expressed as Y en_fe_down , then multiply the outputs of the two branches to get the final output Y en_fe Finally, the enhanced features are mapped back to the original space using linear projection. Through this process, the feature map is enhanced. The enhanced hyperspectral image and LiDAR data features are represented as Y en_HSI and Y en_LiDAR , these multi-source feature maps are fused using addition operations, and then residual connections are introduced to improve the generalization ability of the model and accelerate convergence. The above operation can be expressed as:
[0068] Y en_fe =Linear(Y en_fe_up ×Y en_fe_down )
[0069] Y out =Y en_HSI +Y en_LiDAR +Y LiDAR_out +Y HSL_out
[0070] where Y en_fe is a pseudo symbol representing the enhanced hyperspectral image or LiDAR data feature map, Y out Represents the final extracted feature map.
[0071] In this embodiment, in order to improve computational efficiency and avoid overfitting, a classifier is designed using a simple multi-layer perceptron framework. Specifically, the features are first flattened into (B, P×P×S), where B represents the batch size, P represents the patch size, and S is set to 64. The multi-layer perceptron classification has two fully connected layers. The first layer performs a linear transformation to map the input feature vector to a hidden dimension, denoted as hid. The dimension is regarded as a hyperparameter and can be adjusted according to the experiment. The ReLU activation function is applied after the first layer to introduce nonlinearity so that the model can learn more complex patterns in the data. In order to further reduce the possibility of overfitting, a dropout operation is added between layers to effectively regularize the model by randomly deactivating a portion of neurons during the training process. The second fully connected layer acts as a classification head, projecting the output of the hidden layer to the final number of categories, thereby generating category probabilities. Finally, the cross entropy loss function is used to optimize the performance of the model. The function calculates the difference between the predicted class probability and the true label to guide the parameter update of the model during the training process.
[0072] This paper uses three public hyperspectral images and LiDAR datasets with different land cover to verify the effectiveness of the method proposed in this paper. These three datasets are MUUFL dataset, Houston2013 dataset and Trento dataset. The following is a brief introduction to these datasets:
[0073] (1) MUUFL: The MUUFL dataset consists of hyperspectral images and LiDAR data collected by the ITERS CASI-1500 sensor at the campus of the University of Southern Mississippi Gulf Park in Long Beach, Mississippi in November 2010. The spatial size of the dataset is 325 × 220 pixels with a spatial resolution of 0.54 × 1.0 m. The hyperspectral image data includes 64 spectral bands ranging from 375 to 1050 nm. It contains 11 different land cover classes, including trees, mostly grass, mixed ground, dirt and sand, roads, water, building shadows, buildings, sidewalks, yellow curbs, and cloth boards. Pseudo-color images of the hyperspectral images, ground truth maps, LiDAR-derived DSM images, and corresponding class colors are shown in Figure 2. Figure 5 The number of samples in each category is shown in Table 1.
[0074] (2) Houston2013: The Houston2013 dataset was collected by the ITERS CASI-1500 sensor in 2012 on the campus of the University of Houston and the surrounding urban area in Houston, Texas, USA. It consists of hyperspectral images and LiDAR data. The spatial size of the dataset is 349×1905 pixels, and the spatial resolution is about 2.5 meters. The hyperspectral image data includes 144 spectral bands spanning the range of 380 to 1050 nm. The LiDAR data provides information on the height of ground objects. The land cover is divided into 15 categories, including healthy grass, stressed grass, synthetic grass, trees, soil, water, residential, commercial, road, highway, railway, parking lot 1, parking lot 2, tennis court, and running track. The pseudo-color images of the hyperspectral image data, the ground truth map, the DSM image derived from the LiDAR, and the corresponding class color of the Houston2013 dataset are shown in Figure 2. Figure 6 The number of samples in each category is shown in Table 1.
[0075] (3) Trento: The Trento dataset was captured using the AISA Eagle sensor for hyperspectral imaging and the Optech ALTM 3100EA sensor for LiDAR-derived digital surface model data. It covers a rural area south of the city of Trento, Italy. The spatial extent of the dataset is 166 × 600 pixels with a spatial resolution of approximately 1 meter. The hyperspectral image data consists of 63 spectral bands with a wavelength range of 420 to 990 nm. The land cover types in the dataset are divided into six categories: apple trees, buildings, ground, forests, vineyards, and roads. The pseudo-color images of the hyperspectral image data, the corresponding ground truth maps, the LiDAR-derived DSM images, and the category color maps of the Trento dataset are shown in Figure 2. Figure 7 The number of samples in each category is shown in Table 2.
[0076] Table 1. Training samples, validation samples, and test samples of the Houston2013 and MUUFL datasets.
[0077]
[0078] Table 2. Training samples, validation samples, and test samples of the Trento dataset.
[0079]
[0080]
[0081] 1% of the samples are randomly selected for training and validation sets respectively, and the remaining 98% are used as the test set. Table 1 provides more details about the experimental data. To ensure a fair comparison, the mini-batch size is set to 64, the number of training cycles is set to 200, the learning rate is set to 0.0001, the patch size is set to 5, the mid_dim is set to 96, and the C H With C L The experiments involving the proposed method and other models were conducted within the PyTorch framework. To optimize the network, the Adam optimizer was used. The computations were performed on a laptop equipped with an AMD Ryzen 7 6800H CPU and Radeon Graphics (3.20GHz), 24GB RAM, and an NVIDIA GeForce RTX 4060 laptop GPU (6GB VRAM).
[0082] In order to evaluate the classification performance of the proposed framework against existing models, several metrics were calculated and analyzed, including overall accuracy (OA), average accuracy (AA), Kappa coefficient, and per-class accuracy. The OA metric is calculated by dividing the number of correctly predicted samples by the total number of samples, while the AA metric calculates the average accuracy of all classes. The Kappa coefficient is a statistical metric used to measure the consistency of classification. The following formulas show the calculation details of OA, AA, and Kappa, respectively.
[0083]
[0084] Among them C ii represents the number of correct classifications of category i, C ij represents the number of misclassifications of actual category i as category j, and N represents the number of categories. Therefore, OA provides a more comprehensive measure of the overall performance of the model, which is why it is the main metric emphasized in this invention.
[0085] In order to verify the effectiveness of the proposed method on multi-source remote sensing datasets, seven state-of-the-art methods are compared, including CrossHL, Am3, ExViT, HCT, CMFAEN, CoupleCNN and S2E. The brief descriptions of these algorithms are as follows:
[0086] (1) CrossHL: This method is a new multimodal deep learning framework that effectively integrates RS data and uses the CrossHL self-attention module to fuse their patch projections.
[0087] (2) AM3: This method improves the feature expression ability of multimodal data through fusion through adaptive mutual learning.
[0088] (3) ExViT: This method uses the self-attention mechanism to capture multi-dimensional long-distance dependencies and adopts an adaptive learning strategy to dynamically adjust the feature fusion method according to the characteristics of different data sources.
[0089] (4) HCT: This method first uses a convolutional neural network (CNN) to extract the spatial features of hyperspectral data, and then uses a transformer to further process the fused spatial and spectral features. Through the collaboration of convolutional neural networks and transformers, effective cross-modal feature fusion is achieved.
[0090] (5) CMFAEN: This method effectively combines the spectral features of hyperspectral data with the spatial geometric features of LiDAR data by designing a cross-modal feature aggregation mechanism, fully exploiting the complementary advantages of cross-modal feature aggregation and enhancement.
[0091] (6) CoupleCNN: This method extracts features through a coupled convolutional neural network architecture and fully combines the complementary information between the two data through a sharing or fusion mechanism.
[0092] (7) S2E: This method proposes a spatial enhancement module that enhances the spatial representation of hyperspectral data by LiDAR features and a spatial enhancement module that enhances the spectral representation of LiDAR data by hyperspectral features.
[0093] In these comparison methods, if the prior art provides model parameters and training process, the optimal parameters specified in the prior art will be used. If the prior art does not provide model parameters or training method, experiments will be conducted to obtain the optimal parameters of the method and use the same training conditions as the present invention.
[0094] Tables 3-5 show the OA, AA, Kappa and category accuracy values of the proposed method and all the compared methods on the three datasets. The best results are highlighted in bold. The results show that the proposed method significantly outperforms the other methods in terms of classification accuracy. Specifically, on the MUFFL dataset, the method of the present invention shows better OA, AA and Kappa values than the other methods. It is worth noting that the method of the present invention achieves the highest classification accuracy in 8 (buildings), 9 (sidewalks) and 10 (yellow curbs) categories, which are 94.52%, 76.49% and 50.28%, respectively. The performance of the tenth category is particularly impressive, because the present invention randomly selects 1% of the samples as training samples and validation samples, respectively. This proves the effectiveness of the proposed method in processing small sample data. This success can be attributed to the unique scanning strategy in the method, which enhances the balance between local and global features. The proposed method also performs well on the Houston2013 dataset, where the network of the proposed method achieves the best results in Class 6 (Water), Class 10 (Highway) and Class 12 (Parking Lot 1), slightly higher than other algorithms, which highlights the importance of LiDAR data and edge contour information in distinguishing between building and non-building categories. Although the results are promising, there are still some limitations. After the introduction of edge contour and height information, the category features of different heights are enhanced. However, this leads to the network of the proposed method failing in some categories (such as Class 1, Class 4 and Class 11), the performance is slightly worse than that of CrossHL (a), Am3 (b) and S2E (g) algorithms. Nevertheless, the method of the present invention still shows significant advantages in OA, AA and Kappa. On the Trento dataset, the method of the present invention achieves the best results in OA, AA and Kappa. However, slight performance changes are observed in specific categories compared with other methods, which is mainly due to the limited pixel-level fusion in the adopted fusion strategy.
[0095] Table 3. Classification accuracy obtained by different methods on the MUFFL dataset (the best results are shown in bold)
[0096]
[0097]
[0098] Table 4. Classification accuracy obtained by different methods on the Houston2013 dataset (the best results are shown in bold)
[0099]
[0100] Table 5. Classification accuracy obtained by different methods on the Trento dataset (the best results are shown in bold)
[0101]
[0102] Figure 8-10 The classification results of all methods on the three datasets are shown. In order to clearly demonstrate the classification effect, the present invention amplifies and displays the areas with significant differences in the classification map. Specifically, in the left area of the classification map of the MUUFL dataset, the Am3 (b), S2E (g) and the present invention (h) methods have high classification performance, CrossHL (a) and CMFAEN (e) show considerable salt and pepper noise because these methods do not fully consider the fusion of heterogeneous features. In the correct area, all methods deviate from the real terrain, but the method of the present invention is closest to the actual terrain map, which is consistent with CrossHL (a) and CMFAEN (e), HCT (d), The significant noises in CoupleCNN (g) and Ours (h) are different. The method of the present invention still performs well in the left area of the Houston 2013 dataset, showing smoother edges and more accurate classification results. In the right area, the method of the present invention shows the clearest classification results. Due to the characteristics of the Trento dataset, all methods perform well in the classification results. The method of the present invention has a more refined classification effect in the right area. In general, the method of the present invention achieves good classification performance on three different datasets, which is mainly attributed to the powerful reasoning ability of the feature extraction and fusion strategy Mamba architecture designed by the present invention.
[0103] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0104] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for classifying objects based on the fusion of hyperspectral images and LiDAR data based on an improved Mamba network, characterized in that: The steps include: Step S1, the original hyperspectral image is extracted with the spatial-spectral adaptive Mamba module to extract the spatial-spectral features, and the original LiDAR data is input into the elevation enhancement Mamba module to extract the elevation features; Step S2, using an enhanced fusion module to capture intra-modal and inter-modal relationships and fuse the extracted features; Step S3, input the fused feature set into the classifier to generate the final classification result.
2. The method for classifying objects based on the fusion of hyperspectral image and LiDAR data based on the improved Mamba network according to claim 1, characterized in that: The feature extraction of the spatial-spectral adaptive Mamba module in step S1 uses pixel-level embedding to process the raw data, which includes a 2D convolution layer and a position encoding layer. The 2D convolution layer consists of a kernel size of 1×1 and an output channel of C. h The 2D convolution is composed of a layer of ReLU activation function and a batch normalization layer, and a set of learnable parameters are introduced as position encoding (PE). The initial value of the parameter is set to 1 and updated during the training process. The above process is expressed as follows: X HSI =BN(ReLU(Conv2D(X hsi )))+PE Where X hsi is the raw data of the hyperspectral image, and the above results are expressed as X HSI , X HSI The shape is (C h , P, P).
3. The method for classifying objects based on the fusion of hyperspectral image and LiDAR data based on the improved Mamba network according to claim 2 is characterized by: The step S1 uses four different methods to scan the patch block, where route 1 scans the patch block row by row from left to right and from top to bottom, the scanning direction of route 2 is opposite to the scanning direction of route 1, route 3 scans the patch block column by column from top to bottom and from left to right, and the scanning method of route 4 is opposite to the scanning method of route 3. Data is read through the above routes and four corresponding different 1D sequences are obtained, and then the sequences are input into the spatial route Mamba for feature extraction.
4. The method for classifying objects based on the fusion of hyperspectral image and LiDAR data based on the improved Mamba network according to claim 3 is characterized by: The step S1 applies a reshape operation to transpose the input data, and then scans the spectral data of the patch block from the front and back into 1D sequences, respectively called path 1 and path 2. These sequences are then input into the bidirectional route Mamba model for feature extraction. The operations performed on these sequences are as follows: First, a linear projection is applied to transform each sequence into a state space for selection. Then, a 1D convolution layer is used to extract local features between spectral bands, mainly including a 1D convolution with a kernel size of 3 and a Sigmoid activation function. Then, the extracted features are further extracted through the S6 block to generate a Y for each sequence. route_i The resulting feature map is , where i represents the route order. The whole process is mathematically represented as follows: Y route_i =S6(Sigmod(Conv1D(Linear(route i )))) Among them, route i Indicates different routes. On the other hand, the route1 sequence also enters another branch, performs linear projection and SiLU activation function operations, and the output is marked as Y route_down , the above operation is expressed by the following formula: Y route_down =SiLU(Linear(route1)) Then, Y route_down Multiply by Y route_i Get the output, marked as Y route_out_i ., integrate these feature maps by adding them together to obtain the result Y spe_out Finally, linear projection is applied to map these features back to the original space to obtain the spectral feature map Y out_spe , then Y out_spa and Y out_spe Multiply them to get Y out_hsi Then a residual connection is used to obtain the final hyperspectral image feature map Y HSI_out , including adding Y our_spa , Y out_spe and Y out_hsi , the above operation formula is expressed as: AND out_hsi =And out_spa ×Y out_spe AND HSI_out =And out_hsi +Y out_spa +Y out_spe 。 5. The method for classifying objects based on the fusion of hyperspectral image and LiDAR data based on the improved Mamba network according to claim 4 is characterized in that: The elevation enhancement Mamba module in step S1 first divides the data into fixed-size patches using the center pixel filling method, and then uses pixel-level embedding to enhance the representation of elevation information, including a 2D convolutional layer and position encoding. The 2D convolutional layer is first used to change the elevation dimension of the data, including a kernel size of 1×1 and an output channel of C. l A 2D convolution, a ReLU activation function, and a batch normalization layer are then used to add positional encoding to the data. The result is denoted as X LiDAR , its shape is (p×p, C l ), the vector is scanned back and forth into two 1D sequences, marked as route 1 and route 2. The above operation formula is expressed as: X lidar_route1 =S6(Couv1D(Linear(route1))) X lidar_route2 =S6(Couv1D(Linear(route2))) Where X lidar_route1 and X lidar_route2 They represent the outputs of the two routes respectively. On the other hand, route 1 performs the operation of linear projection and SiLU activation function, and the output is represented as X lidar_down Then, X lidar_down Multiply by X lidar_route1 and X lidar_route2 , respectively get X lidar_out1 and X lidar_out2 , the preliminary feature map is obtained by the following operations: X lidar_fout =Linear(X lidar_out1 +X lidar_out2 ) Among them, Linear(·) is a linear projection that maps the feature map to its original space. X lidar_fout is the preliminary feature map, and finally, the residual is used to transform X lidar_fout and X LiDAR Add together to get denoted as Y LiDAR_out The final elevation feature map.
6. The method for classifying objects based on the fusion of hyperspectral images and LiDAR data based on the improved Mamba network according to claim 5, characterized in that: In step S2, the hyperspectral image and LiDAR data feature maps are first enhanced using the Mamba architecture respectively. The process is as follows: Linear projection maps the feature map to the specified space. Then, these features pass through two branches. In the first branch, the Conv1D function with a convolution kernel size of 3 is used to further extract local features. The Sigmoid activation function is applied to introduce nonlinearity. Then, these features are passed to the S6 block. The result is represented as Y en_fe_up In the second branch, the SiLU activation function is used to introduce nonlinearity to enhance the features, and the result is expressed as Y en_fe_down , then multiply the outputs of the two branches to get the final output Y en_fe Finally, the enhanced features are mapped back to the original space using linear projection. The feature map is enhanced through the above process. The enhanced hyperspectral image and LiDAR data features are represented as Y en_HSI and Y en_LiDAR , the above operation can be expressed as: AND en_fe =Linear(Y en_fe_up ×Y en_fe_down ) AND out =And en_HSI +Y en_LiDAR +Y LiDAR_out +Y HSI_out where Y en_fe is a pseudo symbol representing the enhanced hyperspectral image or LiDAR data feature map, Y out Represents the final extracted feature map.
7. The method for classifying objects based on the fusion of hyperspectral image and LiDAR data based on the improved Mamba network according to claim 6, characterized in that: The classifier of step S3 is designed based on a simple multi-layer perceptron framework. Specifically, the features are first flattened into (B, P×P×S), where B represents the batch size, P represents the patch size, and S is set to 64. The multi-layer perceptron classification has two fully connected layers. The first layer performs a linear transformation to map the input feature vector to a hidden dimension, denoted as hid. The dimension is regarded as a hyperparameter. The ReLU activation function is applied after the first layer to introduce nonlinearity, and dropout operations are added between layers. The second fully connected layer acts as a classification head to project the output of the hidden layer to the final number of categories. Finally, the cross entropy loss function is used to optimize the model performance.