Hyperspectral and laser radar data collaborative classification method based on multistage feature fusion

Through the multi-level feature fusion method, the differences in features, scale, resolution and heterogeneity between hyperspectral images and lidar data are solved, efficient data classification is achieved, and classification accuracy and reliability are improved.

CN120147748APending Publication Date: 2025-06-13HARBIN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510312769.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing classification methods that combine hyperspectral images and lidar data have limited classification performance due to data characteristics, scales, resolution differences and data heterogeneity.

Method used

The hyperspectral and lidar data collaborative classification method based on multi-stage feature fusion is adopted, and the dimensionality reduction is performed through principal component analysis, slice processing and a classification network is constructed. The training set is used to train the classification network to realize the joint processing and classification of hyperspectral images and LiDAR-DSM data.

Benefits of technology

Effectively capture and learn hyperspectral null spectrum joint features, extract accurate elevation features, and enhance classification performance and improve classification accuracy and reliability through adaptive asymmetric gating mechanism and null spectrum linear attention feature fusion module.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147748A_ABST
    Figure CN120147748A_ABST
Patent Text Reader

Abstract

The invention discloses a hyperspectral and laser radar data collaborative classification method based on multistage feature fusion, and belongs to the field of remote sensing image classification. According to the invention, the problem of poor classification performance of the existing classification method fusing the hyperspectral image and the laser radar data is solved. According to the method, the hyperspectral spatial-spectral joint features can be effectively captured and learned from the hyperspectral image data, and meanwhile, the accurate elevation features are extracted from the LiDAR-DSM data. In the feature extraction process, a self-adaptive asymmetric gating mechanism encoder deeply analyzes and comprehensively extracts local and global features in data, so that the comprehensiveness and accuracy of information are ensured; the spatial-spectral linear attention feature fusion module can enhance the classification performance by cooperatively fusing spectral features and spatial features in the hyperspectral data and elevation features in the LiDAR-DSM data. The method can be applied to remote sensing image classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of remote sensing image classification, and in particular relates to a hyperspectral and laser radar data collaborative classification method based on multi-level feature fusion. Background Art

[0002] With the rapid development of remote sensing technology, the traditional single remote sensing data model has gradually exposed its limitations in complex surface environments. Especially in the face of complex surface object types and diverse environmental factors, a single data source is often difficult to fully and accurately extract ground object information, thus affecting the accuracy and reliability of remote sensing information extraction. To solve this problem, multimodal remote sensing data fusion technology has emerged and has become one of the key technologies to improve the utilization efficiency and classification accuracy of remote sensing data. Multimodal fusion can give full play to the advantages of various types of data by combining heterogeneous data from different sensors, thereby effectively making up for the shortcomings of a single data source and providing a more comprehensive and accurate description of the environment.

[0003] Among the many research fields of remote sensing data fusion, the fusion classification task of hyperspectral image (HSI) and laser radar digital surface model (LiDAR-DSM) has become a typical research object. The combination of the two can significantly improve the classification accuracy. However, due to their large differences in data characteristics, scale, resolution, etc., how to coordinate the spatial spectrum characteristics in hyperspectral data with the elevation characteristics in LiDAR-DSM data to form an information-complementary and seamlessly integrated whole has become one of the core issues to be solved in this field. In addition, how to overcome the challenges brought by data heterogeneity while maintaining information integrity is still an important research direction in this field.

[0004] In summary, due to the large differences in data characteristics, scale, and resolution between hyperspectral images and lidar data, as well as the heterogeneity of the data, the performance of the existing classification method that integrates hyperspectral images and lidar data is still relatively limited. In order to improve the classification performance, it is necessary to propose a new classification method. Summary of the invention

[0005] The purpose of the present invention is to solve the problem of poor classification performance of the existing classification method of fusing hyperspectral images and laser radar data, and propose a collaborative classification method of hyperspectral and laser radar data based on multi-level feature fusion.

[0006] The technical solution adopted by the present invention to solve the above technical problems is: a hyperspectral and laser radar data collaborative classification method based on multi-level feature fusion, the method specifically comprising the following steps:

[0007] Step 1: Obtain LiDAR-DSM image data and hyperspectral image data from the dataset;

[0008] Step 2: Use the principal component analysis method to perform dimensionality reduction on the acquired hyperspectral image data to obtain the hyperspectral image data after dimensionality reduction processing;

[0009] Step 3: Perform slicing processing on the LiDAR-DSM image data and the hyperspectral image data after dimensionality reduction processing respectively, and then divide the sliced LiDAR-DSM image data and hyperspectral image data into two parts: a training set and a test set;

[0010] Step 4: Construct a classification network based on multi-level feature fusion, and use the training set to train the constructed classification network until the set maximum number of training times is reached or the classification accuracy of the classification network on the test set no longer improves, and then stop training to obtain a trained classification network;

[0011] Step 5: Use the trained classification network to jointly process the hyperspectral image and LiDAR-DSM image of the area to be classified to obtain a classification result.

[0012] The beneficial effects of the present invention are as follows:

[0013] The present invention can effectively capture and learn the hyperspectral spatial-spectral joint features from hyperspectral image data, and at the same time extract accurate elevation features from LiDAR-DSM data. In the feature extraction process, the adaptive asymmetric gating mechanism encoder deeply analyzes and comprehensively extracts the local and global features in the data, ensuring the comprehensiveness and accuracy of the information; the spatial-spectral linear attention feature fusion module can enhance the classification performance by synergistically fusing the spectral features, spatial features in hyperspectral data, and elevation features in LiDAR-DSM data. Description of the Drawings

[0014] Figure 1 is the structural diagram of the classification network based on multi-level feature fusion;

[0015] Figure 2 is the flow chart of a method for collaborative classification of hyperspectral and lidar data based on multi-level feature fusion according to the present invention;

[0016] Figure 3 is the structural diagram of the AGMF encoder;

[0017] Figure 4 is the structural diagram of the spatial-spectral linear attention feature fusion module;

[0018] Figure 5 is the spatial distribution and ground object category information of the Trento dataset;

[0019] Figure 6 is the spatial distribution and ground object category information of the MUUFL dataset;

[0020] Figure 7 It is the spatial distribution and feature class information of the Augsburg dataset;

[0021] Figure 8 It is the Trento dataset;

[0022] In the figure, (a) is the hyperspectral false color image, (b) is the DSM grayscale image, and (c) is the true value image;

[0023] Figure 9 It is the MUUFL dataset;

[0024] In the figure, (a) is the hyperspectral false color image, (b) is the DSM grayscale image, and (c) is the true value image;

[0025] Figure 10 It is the Augsburg dataset;

[0026] In the figure, (a) is the hyperspectral false color image, (b) is the DSM grayscale image, and (c) is the true value image;

[0027] Figure 11 It is the subjective classification result of different data by the classification network based on multi-level feature fusion;

[0028] In the figure, (a) is Trento, (b) is MUUFL, and (c) is Augsburg. Detailed implementation manners

[0029] Detailed implementation manner 1: In combination with Figure 2 This implementation manner is described. A hyperspectral and lidar data collaborative classification method based on multi-level feature fusion described in this implementation manner specifically includes the following steps:

[0030] Step 1: Obtain LiDAR-DSM image data and hyperspectral image data from the dataset;

[0031] For each hyperspectral image obtained from the dataset, LiDAR-DSM image data corresponding to the area of the hyperspectral image is simultaneously obtained;

[0032] Step 2: Use the principal component analysis method to perform dimensionality reduction processing on the obtained hyperspectral image data to obtain the hyperspectral image data after dimensionality reduction processing;

[0033] Step 3: Perform slicing processing on the LiDAR-DSM image data and the hyperspectral image data after dimensionality reduction processing respectively, and then divide the sliced LiDAR-DSM image data and hyperspectral image data into two parts: a training set and a test set;

[0034] When dividing, the LiDAR-DSM image data corresponding to the same area and the hyperspectral image data after dimensionality reduction processing should be divided into the training set or the test set at the same time, so that when inputting into the model later, the LiDAR-DSM image data corresponding to the same area and the hyperspectral image data after dimensionality reduction processing can be input into the model at the same time;

[0035] Step 4: Construct a classification network based on multi-level feature fusion, and use the training set to train the constructed classification network until the maximum number of training times set is reached or the classification accuracy of the classification network on the test set no longer improves, and then stop training to obtain a trained classification network;

[0036] Step 5: Use the trained classification network to jointly process the hyperspectral image and the LiDAR-DSM image of the area to be classified to obtain a classification result.

[0037] Before inputting into the classification network, the hyperspectral image and the LiDAR-DSM image of the area to be classified both need to go through preprocessing, that is, the hyperspectral image needs to go through dimensionality reduction and slicing processing, and the LiDAR-DSM image needs to go through slicing processing.

[0038] Specific Embodiment 2: The difference between this embodiment and Specific Embodiment 1 is that the classification network based on multi-level feature fusion specifically includes a LiDAR-DSM image data processing branch, a spatial processing branch of hyperspectral image data, and a spectral processing branch of hyperspectral image data.

[0039] Other steps and parameters are the same as those in Specific Embodiment 1.

[0040] Specific Embodiment 3: The difference between this embodiment and Specific Embodiment 1 or 2 is that the working process of the classification network based on multi-level feature fusion is as follows:

[0041] Take the LiDAR-DSM image data as the input of the LiDAR-DSM image data processing branch. In the LiDAR-DSM image data processing branch, the LiDAR-DSM image data first passes through the first two-dimensional convolutional layer with a kernel size of 1×1 (in the present invention, convolutional layers not specifically emphasized are all conventional convolutional layers);

[0042] Take the output of the first two-dimensional convolutional layer as the input of the first two-dimensional depth convolutional layer with a kernel size of 3×1, and then take the output of the first two-dimensional depth convolutional layer as the input of the second two-dimensional depth convolutional layer with a kernel size of 1×3;

[0043] Take the output of the second two-dimensional depth convolutional layer as the input of the first batch normalization layer, and then take the output of the first batch normalization layer as the input of the first ReLU activation function layer;

[0044] Flatten the output of the first ReLU activation function layer into a one-dimensional vector sequence, embed additional learned encodings at the head of the one-dimensional vector sequence to obtain an overall sequence, then embed position encodings for each element in the overall sequence respectively, and then pass the vector sequence after embedding the position encodings through the first Adaptive Asymmetric Gating Mechanism Encoder (AGMF);

[0045] Use the hyperspectral image data after dimensionality reduction processing as the input of the spectral processing branch of the hyperspectral image data. Within the spectral processing branch of the hyperspectral image data, the hyperspectral image data after dimensionality reduction processing first passes through the first three-dimensional convolutional layer with a convolutional kernel size of 1×1×3;

[0046] Use the output of the first three-dimensional convolutional layer as the input of the second batch normalization layer, and then use the output of the second batch normalization layer as the input of the second ReLU activation function layer;

[0047] Reorganize the three-dimensional data output by the second ReLU activation function layer into two-dimensional data, and use the reorganized two-dimensional data as the input of the second two-dimensional convolutional layer with a convolutional kernel size of 1×1;

[0048] Use the output of the second two-dimensional convolutional layer as the input of the third batch normalization layer, and then use the output of the third batch normalization layer as the input of the third ReLU activation function layer;

[0049] Flatten the output of the third ReLU activation function layer into a one-dimensional vector sequence, embed additional learned encodings at the head of the one-dimensional vector sequence to obtain an overall sequence, then embed position encodings for each element in the overall sequence respectively, and pass the vector sequence after embedding the position encodings through the second Adaptive Asymmetric Gating Mechanism Encoder (AGMF);

[0050] Use the hyperspectral image data after dimensionality reduction processing as the input of the spatial processing branch of the hyperspectral image data. Within the spatial processing branch of the hyperspectral image data, first pass the hyperspectral image data after dimensionality reduction processing through the second three-dimensional convolutional layer with a convolutional kernel size of 1×1×1;

[0051] Use the output of the second three-dimensional convolutional layer as the input of the first three-dimensional depth convolutional layer with a convolutional kernel size of 3×1×1;

[0052] Use the output of the first three-dimensional depth convolutional layer as the input of the second three-dimensional depth convolutional layer with a convolutional kernel size of 1×3×1;

[0053] Use the output of the second three-dimensional depth convolutional layer as the input of the fourth batch normalization layer, and then use the output of the fourth batch normalization layer as the input of the fourth ReLU activation function layer;

[0054] Recombine the three-dimensional data output by the fourth ReLU activation function layer into two-dimensional data, and then use the recombined two-dimensional data as the input of the third two-dimensional convolutional layer with a convolution kernel size of 1×1;

[0055] Use the output of the third two-dimensional convolutional layer as the input of the third two-dimensional depth convolutional layer with a convolution kernel size of 3×1;

[0056] Use the output of the third two-dimensional depth convolutional layer as the input of the fourth two-dimensional depth convolutional layer with a convolution kernel size of 1×3;

[0057] Use the output of the fourth two-dimensional depth convolutional layer as the input of the fifth batch normalization layer, and then use the output of the fifth batch normalization layer as the input of the fifth ReLU activation function layer;

[0058] Flatten the output of the fifth ReLU activation function layer into a one-dimensional vector sequence, embed additional learning codes at the head of the one-dimensional vector sequence to obtain the overall sequence, then embed position codes for each element in the overall sequence respectively, and pass the vector sequence after embedding the position codes through the third adaptive asymmetric gating mechanism encoder;

[0059] Send the output of the second adaptive asymmetric gating mechanism encoder and the output of the third adaptive asymmetric gating mechanism encoder into the first cross-attention module to obtain the outputs A and B of the first cross-attention module;

[0060] Send the output A and the output of the first adaptive asymmetric gating mechanism encoder into the second cross-attention module to obtain the outputs C and D of the second cross-attention module;

[0061] Send the output B and the output of the first adaptive asymmetric gating mechanism encoder into the third cross-attention module to obtain the outputs E and F of the third cross-attention module;

[0062] Stack the output C, the output E and the input of the first adaptive asymmetric gating mechanism encoder, and denote the stacked result as G;

[0063] Stack the output D and the input of the second adaptive asymmetric gating mechanism encoder, and denote the stacked result as H,

[0064] Stack the output F and the input of the third adaptive asymmetric gating mechanism encoder, and denote the stacked result as I;

[0065] Send the stacked result G, the stacked result H and the stacked result I into the empty-spectrum linear attention feature fusion module, and pass the output of the empty-spectrum linear attention feature fusion module through the first linear layer to obtain the final classification result.

[0066] Other steps and parameters are the same as those in the first or second specific implementation manner.

[0067] Specific implementation manner four: Figure 3 This specific implementation manner will be described in conjunction with

[0068] The input of the first adaptive asymmetric gating mechanism encoder is used as the input of the first encoding module, and then the output of the (i - 1)-th encoding module is used as the input of the i-th encoding module, and the output of the N-th encoding module is used as the output of the first adaptive asymmetric gating mechanism encoder.

[0069] Other steps and parameters are the same as those in one of the first to third specific implementation manners.

[0070] The structures of the second and third adaptive asymmetric gating mechanism encoders are the same as that of the first adaptive asymmetric gating mechanism encoder.

[0071] Specific implementation manner five: The difference between this specific implementation manner and one of the first to fourth specific implementation manners is that the working process of the first encoding module is as follows:

[0072] Step one: Map the input of the first encoding module into a query vector Q, a key vector K, and a value vector V.

[0073] Step two: Perform multi-head self-attention calculation on the query vector Q, the key vector K, and the value vector V to obtain the multi-head self-attention calculation result.

[0074] Step three: Perform layer normalization on the multi-head self-attention calculation result in step two, and perform residual connection between the result of layer normalization and the input of the first encoding module to obtain a residual connection result a.

[0075] Step four: Send the residual connection result a into the asymmetric gating feed-forward unit (AGF), then perform layer normalization on the output of the asymmetric gating feed-forward unit, and perform residual connection between the result of layer normalization and the residual connection result a to obtain a residual connection result b, and use the residual connection result b as the output of the first encoding module.

[0076] Other steps and parameters are the same as those in one of the first to fourth specific implementation manners.

[0077] The working processes of the second to N-th encoding modules are the same as the structure of the first encoding module.

[0078] Specific implementation manner six: The difference between this specific implementation manner and one of the first to fifth specific implementation manners is that the working process of the asymmetric gating feed-forward unit is as follows:

[0079] Inside the asymmetric gated feed-forward unit, first, the input of the asymmetric gated feed-forward unit is used as the input of the second linear layer;

[0080] Then, the output of the second linear layer is used as the input of the GELU activation function layer, and the output of the GELU activation function layer is recombined into two-dimensional data;

[0081] The two-dimensional data obtained by recombination is evenly divided into two parts along the channel dimension; the first half passes through the fourth two-dimensional convolutional layer with a kernel size of 1×1; the second half passes through the fifth two-dimensional convolutional layer with a kernel size of 1×1, and then the output of the fifth two-dimensional convolutional layer is used as the input of the fifth two-dimensional depth convolutional layer with a kernel size of 3×1, and the output of the fifth two-dimensional depth convolutional layer is used as the input of the sixth two-dimensional depth convolutional layer with a kernel size of 1×3;

[0082] The output of the fourth two-dimensional convolutional layer is multiplied by the output of the sixth two-dimensional depth convolutional layer, and the multiplication result is used as the input of the sixth two-dimensional convolutional layer with a kernel size of 1×1;

[0083] Random inactivation is performed on the two-dimensional data obtained by recombination, the random inactivation result is superimposed on the output of the sixth two-dimensional convolutional layer, and then the superimposed result is recombined into one-dimensional data. The recombined one-dimensional data is used as the input of the third linear layer (the linear layer reorganizes the one-dimensional data again to restore the channel position and obtain the data features after restoring the channel position), and the output of the third linear layer is used as the output of the asymmetric gated feed-forward unit.

[0084] Other steps and parameters are the same as those in any one of the first to fifth specific embodiments.

[0085] Specific embodiment seven: Combine Figure 4 This embodiment is described. The difference between this embodiment and any one of the first to sixth specific embodiments is that the empty-spectrum linear attention feature fusion module includes M parallel feature fusion units, and the outputs of the M feature fusion units are concatenated, and the concatenated result is used as the output of the empty-spectrum linear attention feature fusion module.

[0086] Other steps and parameters are the same as those in any one of the first to sixth specific embodiments.

[0087] Specific embodiment eight: Combine Figure 4 This embodiment is described. The difference between this embodiment and any one of the first to seventh specific embodiments is that the working process of each of the feature fusion units is as follows:

[0088] Pass the superimposed result H (hyperspectral spectral features) through the fourth linear layer, pass the superimposed result G (LiDAR-DSM features) through the fifth linear layer, multiply the output result of the fourth linear layer by the output result of the fifth linear layer to obtain the multiplication result c, and then pass the multiplication result c through the sixth linear layer, the seventh linear layer, and the Sigmod function layer in sequence;

[0089] And pass the output of the sixth linear layer through the Softmax function layer;

[0090] Pass the superimposed result G through the eighth linear layer, pass the superimposed result I (hyperspectral spatial features) through the ninth linear layer, then multiply the output result of the eighth linear layer by the output result of the ninth linear layer to obtain the multiplication result d, and then multiply the multiplication result d by the output of the Softmax function layer to obtain the multiplication result e;

[0091] Then multiply the multiplication result e by the output of the Sigmod function layer, and use the multiplication result as the output of the feature fusion unit.

[0092] Other steps and parameters are the same as those in any one of the first to seventh specific embodiments.

[0093] Embodiment

[0094] The present invention proposes a collaborative classification method for hyperspectral and lidar data based on multi-level feature fusion. The implementation process of the method of the present invention is shown in Table 1;

[0095] Table 1 Classification algorithm process based on multi-level feature fusion

[0096]

[0097] The classification method of the present invention is described in detail as follows:

[0098] Step 1: Obtain hyperspectral image data and LiDAR-DSM data (using publicly available data).

[0099] Step 2: Use the PCA principal component analysis method to perform dimensionality reduction processing on the obtained hyperspectral image data, which can reduce the number of bands in the third dimension of the three-dimensional original data, reduce data redundancy, and speed up the running time.

[0100] Step 3: Preprocess the LiDAR-DSM data and the dimensionality-reduced hyperspectral data. According to the number of pixel samples marked in the hyperspectral image and the LiDAR-DSM image, specify the number of training sets and test sets, and store them as three copies of data in the same format as the LiDAR-DSM data and the dimensionality-reduced hyperspectral data. Then, perform slicing processing on the three copies of data respectively.

[0101] Step 4: Training and Classification of the Classification Network Based on Multi-Level Feature Fusion

[0102] Step 4.1: Construct a classification network based on multi-level feature fusion. The training set and test set images used in the present invention both adopt a resolution of 11×11 pixels. As Figure 1 shown, H×W represents the spatial dimension, and B represents the third-dimensional band. After PCA processing and data preprocessing, a series of hyperspectral images of size 11×11×C (C represents the third-dimensional band after dimensionality reduction) and LiDAR-DSM data of size 11×11 are obtained.

[0103] The shallow feature extraction structure of the classification network based on multi-level feature fusion is mainly divided into feature extraction in the upper part and feature fusion in the lower part. The feature extraction part includes a LiDAR-DSM data processing branch, a spatial processing branch of hyperspectral image data, and a spectral processing branch of hyperspectral image data, which are respectively used to extract LiDAR-DSM data, spectral features of hyperspectral image data, and spatial features of hyperspectral image data. Each processing branch contains a convolution operation and an adaptive asymmetric gating mechanism encoder. The convolution kernel sizes of the convolution layers mainly include 1×1×1, 1×1×3, 3×1×1, 1×3×1, 1×1, 3×1, and 1×3.

[0104] The working process of the classification network based on multi-level feature fusion is as follows:

[0105] In the LiDAR-DSM data processing branch, the spatial processing branch of hyperspectral image data, and the spectral processing branch of hyperspectral image data, the inputs of the adaptive asymmetric gating mechanism encoders are linearly flattened into one-dimensional vector sequences, and additional learning codes are embedded at the head of the vector sequences to obtain the overall sequences, and then positional encoding is embedded in the overall sequences;

[0106] Cross-attention operations are performed on the outputs of the adaptive asymmetric gating mechanism encoders in the spatial processing branch of hyperspectral image data and the spectral processing branch of hyperspectral image data to obtain the spatial features and spectral features after fusing the spatial and spectral features;

[0107] The spatial features and spectral features after fusing the spatial and spectral features are respectively subjected to cross-attention operations with the outputs of the adaptive asymmetric gating mechanism encoders in the LiDAR-DSM data processing branch to obtain the spatial features after fusing LiDAR-DSM features, the spectral features after fusing LiDAR-DSM features, the LiDAR-DSM features of the fused spatial features, and the LiDAR-DSM features of the fused spectral features;

[0108] Overlay the spatial features after fusing LiDAR-DSM features with the input of the adaptive asymmetric gating mechanism encoder in the spatial processing branch of the hyperspectral image data; overlay the spectral features after fusing LiDAR-DSM features with the input of the adaptive asymmetric gating mechanism encoder in the spectral processing branch of the hyperspectral image data; overlay the LiDAR-DSM features with fused spatial features, the LiDAR-DSM features with fused spectral features with the input of the adaptive asymmetric gating mechanism encoder in the LiDAR-DSM data processing branch;

[0109] Input the overlay result into the spatial-spectral linear attention feature fusion module, perform a linear operation on the output of the spatial-spectral linear attention feature fusion module to obtain the classification result;

[0110] Step 5: Use the LiDAR-DSM data and hyperspectral image data of the area to be classified as inputs, and output the classification result through the trained classification network.

[0111] Experimental part

[0112] The network model training and classification result verification experiments of the present invention are both completed on the following platform:

[0113] The hardware configuration includes an Intel Core i9-12900K processor, 32GB of DDR5 5200MHz memory, an NVIDIA GeForce RTX 3090Ti graphics card. The storage system consists of a 1TB solid-state drive (SSD) and a 4TB hard disk drive (HDD), and the operating system is Windows 11 Professional Edition. The experiment uses three well-known multimodal remote sensing datasets, specifically including the Trento dataset, the Mississippi University and University of Florida Golf Course dataset (MUUFL), and the Augsburg dataset.

[0114] The spatial distribution and ground object category information of each dataset are respectively as Figure 5 , Figure 6 and Figure 7 shown, where different colors represent different ground object categories. The detailed statistical information and feature descriptions of the datasets are respectively as Figure 8 , Figure 9 and Figure 10 shown.

[0115] 1. The capture location of the Trento dataset is in the rural area around the city of Trento, Italy. The dataset contains hyperspectral images and LiDAR images. The hyperspectral images are sized 600×166 pixels, with a total of 63 bands, covering a spectral band range from 420.89 to 989.09 nanometers, a spectral resolution of 9.2 nanometers, and a spatial resolution of 1 meter. The LiDAR image is a single-channel image containing the elevation of the corresponding ground location, with the same image size as the hyperspectral image. There are a total of six ground object categories in the annotation information of the dataset.

[0116] 2. The MUUFL dataset is a registered aerial hyperspectral-LiDAR dataset. The two modal images of this dataset are acquired simultaneously during a single aerial flight in November 2010, located in Mississippi, USA. The image size is 325×220 pixels. The hyperspectral image contains 64 spectral bands. The LiDAR image is a single-channel image containing the elevation of the corresponding ground location, with the same image size as the hyperspectral image. The data annotation information contains 11 categories.

[0117] 3. The Augsburg dataset was captured over the city of Augsburg, Germany. The HSI data was obtained by the DAS-EOC HySpex sensor, and the LiDAR-DSM data was collected by the DLR-3K system. The spatial resolution of both images was downsampled to 30 meters to fully manage multimodal fusion. In this dataset, the HSI data consists of 180 bands, covering a band range from 0.4 to 2.5 micrometers, while the LiDAR-DSM data has only one raster. The size of the dataset is 332×485 pixels, and the dataset depicts seven different land cover categories.

[0118] For the evaluation metrics of the model, three of the most widely used objective evaluation metrics in the industry are selected: Overall Accuracy (OA), Average Accuracy (AA), and Kappa Coefficient (K). These three evaluation metrics are calculated based on the Confusion Matrix. The following will introduce these evaluation metrics separately:

[0119] (1) Confusion Matrix: The confusion matrix is an error matrix with n rows and n columns, which represents the standard form for calculating accuracy. The detailed content that makes up this matrix is obtained from Table 2. The confusion matrix is a specific matrix used to visualize the performance of an algorithm. Each column of the confusion matrix represents the predicted value, and each row represents the true category. The name of this matrix is because its composition form can clearly show whether there is confusion between multiple categories, that is, whether one category is predicted as another category.

[0120] Composition form of the confusion matrix in Table 2

[0121]

[0122] (2) Overall classification accuracy (OA): OA is a basic evaluation index, which is expressed as the ratio of the number of correctly classified samples to the total number of samples, that is, the overall classification accuracy. The overall classification accuracy represents the classification accuracy of the algorithm as a whole, which is the ratio of the number of correctly classified class pixels to the total number of classes:

[0123]

[0124] (3) Average classification accuracy (AA): AA represents the evaluation index of classification accuracy in the classification algorithm. First, sum the classification accuracies of each type of ground object, and then calculate the average value as the classification accuracy. The average classification accuracy is different from the overall classification accuracy, and it focuses on evaluating the classification accuracy of the algorithm for each type of ground object:

[0125]

[0126] (4) Kappa coefficient (K): The Kappa coefficient represents the proportion of reduction in the completely random classification error. It is also an evaluation index of classification accuracy. As a supplement to the fact that the classification accuracy cannot be clearly seen from the error matrix, it can clearly reflect the effect of classification accuracy:

[0127]

[0128] Among them, n represents the number of rows and columns of the classification matrix;

[0129] m ij represents the value in the i-th row and j-th column of the confusion matrix;

[0130] m i+ is the row sum of the confusion matrix;

[0131] m +i is the column sum of the confusion matrix;

[0132] N represents all elements included in the confusion matrix.

[0133] The objective classification results of the network model used in the present invention on three data sets are shown in Table 3. From left to right, the first column is the data set; the second column is the OA data of the classification result; the third column is the AA data of the classification result; the fourth column is the K×100 data of the classification result; the fifth column is the training duration of the classification; the sixth column is the test duration of the classification. The subjective classification results are as Figure 11As shown, (a) represents the classification effect of the Trento dataset, (b) represents the classification effect of the MUUFL dataset, and (c) represents the classification effect of the Augsburg dataset; the higher the classification accuracy, the less salt-and-pepper noise in the image.

[0134] Table 3

[0135]

[0136] The above examples of the present invention are only for explaining in detail the calculation model and calculation process of the present invention, and are not intended to limit the implementation manner of the present invention. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is impossible to list all the implementation manners here. Any obvious changes or variations derived from the technical solution of the present invention still fall within the protection scope of the present invention.

Claims

1. A collaborative classification method of hyperspectral and lidar data based on multi-level feature fusion, characterized in that: The method specifically comprises the following steps: Step 1: Obtain LiDAR-DSM image data and hyperspectral image data from the dataset; Step 2: Using principal component analysis to perform dimensionality reduction processing on the acquired hyperspectral image data to obtain hyperspectral image data after dimensionality reduction processing; Step 3, slicing the LiDAR-DSM image data and the hyperspectral image data after dimensionality reduction processing, and then dividing the sliced ​​LiDAR-DSM image data and the hyperspectral image data into a training set and a test set; Step 4: Construct a classification network based on multi-level feature fusion, and use the training set to train the constructed classification network until the maximum number of training times is reached or the classification accuracy of the classification network on the test set is no longer improved, and then stop training to obtain a trained classification network; Step 5: Use the trained classification network to jointly process the hyperspectral image and LiDAR-DSM image of the classification area to obtain the classification result.

2. The method for collaborative classification of hyperspectral and laser radar data based on multi-level feature fusion according to claim 1 is characterized in that: The classification network based on multi-level feature fusion specifically includes a LiDAR-DSM image data processing branch, a hyperspectral image data spatial processing branch, and a hyperspectral image data spectral processing branch.

3. The method for collaborative classification of hyperspectral and laser radar data based on multi-level feature fusion according to claim 2 is characterized in that: The working process of the classification network based on multi-level feature fusion is as follows: The LiDAR-DSM image data is used as the input of the LiDAR-DSM image data processing branch. In the LiDAR-DSM image data processing branch, the LiDAR-DSM image data first passes through the first two-dimensional convolution layer with a convolution kernel size of 1×1; The output of the first two-dimensional convolutional layer is used as the input of the first two-dimensional deep convolutional layer with a convolution kernel size of 3×1, and then the output of the first two-dimensional deep convolutional layer is used as the input of the second two-dimensional deep convolutional layer with a convolution kernel size of 1×3; The output of the second 2D depth convolutional layer is used as the input of the first batch normalization layer, and the output of the first batch normalization layer is used as the input of the first ReLU activation function layer; The output of the first ReLU activation function layer is linearly flattened into a one-dimensional vector sequence, and an additional learning code is embedded in the head of the one-dimensional vector sequence to obtain the overall sequence, and then each element in the overall sequence is embedded with a position code, and then the vector sequence after the embedded position code is passed through the first adaptive asymmetric gating mechanism encoder; The hyperspectral image data after dimensionality reduction processing is used as the input of the spectrum processing branch of the hyperspectral image data. In the spectrum processing branch of the hyperspectral image data, the hyperspectral image data after dimensionality reduction processing first passes through the first three-dimensional convolution layer with a convolution kernel size of 1×1×3; The output of the first 3D convolutional layer is used as the input of the second batch normalization layer, and the output of the second batch normalization layer is used as the input of the second ReLU activation function layer; Reorganize the three-dimensional data output by the second ReLU activation function layer into two-dimensional data, and use the reorganized two-dimensional data as the input of the second two-dimensional convolution layer with a convolution kernel size of 1×1; The output of the second 2D convolutional layer is used as the input of the third batch normalization layer, and the output of the third batch normalization layer is used as the input of the third ReLU activation function layer; The output of the third ReLU activation function layer is linearly flattened into a one-dimensional vector sequence, and an additional learning code is embedded in the head of the one-dimensional vector sequence to obtain the overall sequence. Then, each element in the overall sequence is embedded with a position code, and the vector sequence after the embedded position code is passed through the second adaptive asymmetric gating mechanism encoder; The hyperspectral image data after dimensionality reduction processing is used as the input of the spatial processing branch of the hyperspectral image data. In the spatial processing branch of the hyperspectral image data, the hyperspectral image data after dimensionality reduction processing is firstly passed through a second three-dimensional convolution layer with a convolution kernel size of 1×1×1; The output of the second 3D convolutional layer is used as the input of the first 3D depth convolutional layer with a convolution kernel size of 3×1×1; The output of the first 3D deep convolutional layer is used as the input of the second 3D deep convolutional layer with a convolution kernel size of 1×3×1; The output of the second three-dimensional depth convolution layer is used as the input of the fourth batch normalization layer, and the output of the fourth batch normalization layer is used as the input of the fourth ReLU activation function layer; The three-dimensional data output by the fourth ReLU activation function layer is reorganized into two-dimensional data, and then the reorganized two-dimensional data is used as the input of the third two-dimensional convolution layer with a convolution kernel size of 1×1; The output of the third 2D convolutional layer is used as the input of the third 2D depth convolutional layer with a convolution kernel size of 3×1; The output of the third 2D depth convolutional layer is used as the input of the fourth 2D depth convolutional layer with a convolution kernel size of 1×3; The output of the fourth two-dimensional depth convolution layer is used as the input of the fifth batch normalization layer, and the output of the fifth batch normalization layer is used as the input of the fifth ReLU activation function layer; The output of the fifth ReLU activation function layer is linearly flattened into a one-dimensional vector sequence, and an additional learning code is embedded in the head of the one-dimensional vector sequence to obtain the overall sequence, and then each element in the overall sequence is embedded with a position code, and the vector sequence after the embedded position code is passed through the third adaptive asymmetric gating mechanism encoder; Sending the output of the second adaptive asymmetric gating mechanism encoder and the output of the third adaptive asymmetric gating mechanism encoder to the first cross attention module to obtain outputs A and B of the first cross attention module; Send the output A and the output of the first adaptive asymmetric gating mechanism encoder to the second cross attention module to obtain the outputs C and D of the second cross attention module; Send the output B and the output of the first adaptive asymmetric gating mechanism encoder to the third cross attention module to obtain the outputs E and F of the third cross attention module; Superimpose the output C, the output E and the input of the first adaptive asymmetric gating mechanism encoder, and record the superposition result as G; The output D is superimposed with the input of the second adaptive asymmetric gating mechanism encoder, and the superposition result is recorded as H. The output F is superimposed with the input of the encoder of the third adaptive asymmetric gating mechanism, and the superposition result is recorded as I; The superposition results G, H and I are sent to the spatial-spectral linear attention feature fusion module, and the output of the spatial-spectral linear attention feature fusion module is passed through the first linear layer to obtain the final classification result.

4. The method for collaborative classification of hyperspectral and laser radar data based on multi-level feature fusion according to claim 3 is characterized in that: The first adaptive asymmetric gating mechanism encoder includes N encoding modules connected in series, wherein: The input of the first adaptive asymmetric gating mechanism encoder is used as the input of the first encoding module, the output of the i-1th encoding module is used as the input of the ith encoding module, and the output of the Nth encoding module is used as the output of the first adaptive asymmetric gating mechanism encoder.

5. The method for collaborative classification of hyperspectral and laser radar data based on multi-level feature fusion according to claim 4 is characterized in that: The working process of the first encoding module is: Step 1: Map the input of the first encoding module into a query vector Q, a key vector K, and a value vector V; Step 2: Perform multi-head self-attention calculation on the query vector Q, key vector K and value vector V to obtain the multi-head self-attention calculation result; Step 3: perform layer normalization on the multi-head self-attention calculation results in step 2, and perform residual connection on the layer normalization result and the input of the first encoding module to obtain the residual connection result a; Step 4: Send the residual connection result a to the asymmetric gated feedforward unit, then perform layer normalization on the output of the asymmetric gated feedforward unit, perform residual connection on the layer normalization result and the residual connection result a to obtain the residual connection result b, and use the residual connection result b as the output of the first encoding module.

6. The method for collaborative classification of hyperspectral and laser radar data based on multi-level feature fusion according to claim 5 is characterized in that: The working process of the asymmetric gated feedforward unit is: In the asymmetric gated feedforward unit, the input of the asymmetric gated feedforward unit is first used as the input of the second linear layer; The output of the second linear layer is used as the input of the GELU activation function layer, and the output of the GELU activation function layer is reorganized into two-dimensional data; The reorganized two-dimensional data is divided into two parts according to the channel dimension; the first half passes through the fourth two-dimensional convolution layer with a convolution kernel size of 1×1; the second half passes through the fifth two-dimensional convolution layer with a convolution kernel size of 1×1, and then the output of the fifth two-dimensional convolution layer is used as the input of the fifth two-dimensional deep convolution layer with a convolution kernel size of 3×1, and the output of the fifth two-dimensional deep convolution layer is used as the input of the sixth two-dimensional deep convolution layer with a convolution kernel size of 1×3; Multiply the output of the fourth two-dimensional convolutional layer with the output of the sixth two-dimensional depth convolutional layer, and then use the multiplication result as the input of the sixth two-dimensional convolutional layer with a convolution kernel size of 1×1; The reorganized two-dimensional data is randomly inactivated, and the random inactivation result is superimposed with the output of the sixth two-dimensional convolutional layer. The superimposed result is reorganized into one-dimensional data, and the reorganized one-dimensional data is used as the input of the third linear layer, and the output of the third linear layer is used as the output of the asymmetric gated feedforward unit.

7. The method for collaborative classification of hyperspectral and laser radar data based on multi-level feature fusion according to claim 6 is characterized in that: The spatial-spectral linear attention feature fusion module includes M parallel feature fusion units, and the outputs of the M feature fusion units are spliced, and the splicing result is used as the output of the spatial-spectral linear attention feature fusion module.

8. The method for collaborative classification of hyperspectral and laser radar data based on multi-level feature fusion according to claim 7 is characterized in that: The working process of the feature fusion unit is: The superposition result H passes through the fourth linear layer, the superposition result G passes through the fifth linear layer, the output result of the fourth linear layer is multiplied by the output result of the fifth linear layer to obtain the multiplication result c, and then the multiplication result c passes through the sixth linear layer, the seventh linear layer and the Sigmod function layer in sequence; And pass the output of the sixth linear layer through the Softmax function layer; The superposition result G is passed through the eighth linear layer, the superposition result I is passed through the ninth linear layer, and then the output result of the eighth linear layer is multiplied by the output result of the ninth linear layer to obtain the multiplication result d, and then the multiplication result d is multiplied by the output of the Softmax function layer to obtain the multiplication result e; Then multiply the multiplication result e with the output of the Sigmod function layer, and use the multiplication result as the output of the feature fusion unit.