Land cover classification method based on hyperspectral-lidar images

Through the multi-scale cross-correlation attention convolutional fusion network, the problem of insufficient spectral and spatial information utilization in HSI-LiDAR data fusion is solved, and a higher-precision land cover classification is achieved, especially in complex scenarios, which significantly improves the classification effect.

CN117975140BActive Publication Date: 2025-08-29QIQIHAR UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410152845.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-03
Publication Date
2025-08-29
Estimated Expiration
2044-02-03

AI Technical Summary

Technical Problem

The existing hyperspectral-lidar image fusion method fails to fully utilize the spectral information of HSI and the spatial information of LiDAR, resulting in insufficient accuracy and credibility of land cover classification, especially in complex scenarios, which is difficult to fully and accurately capture the characteristics of ground cover.

Method used

A multi-scale cross-correlation attention convolution fusion network is used to extract spectral, spatial and elevation features from HSI and LiDAR data through multi-scale residual convolution blocks, and feature fusion is used for cross-correlation attention blocks to generate joint semantic features, and finally land cover classification is performed.

Benefits of technology

It improves the accuracy and credibility of land cover classification, especially in complex scenarios, which can capture the characteristics of ground cover more comprehensively, and improves classification accuracy and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117975140B_ABST
    Figure CN117975140B_ABST
Patent Text Reader

Abstract

The land cover classification method for hyperspectral-lidar images solves the problem that existing land cover classification methods are difficult to effectively extract discriminant features, and belongs to the technical field of land cover classification. The present invention includes: inputting HSI and LiDAR data of land cover into a multi-scale cross-correlation attention convolutional fusion network to obtain joint semantic features of HSI and LiDAR data, inputting the joint semantic features into a classification module, and the classification module generates the final classification results; the multi-scale cross-correlation attention convolutional fusion network of the present invention uses multi-scale residual convolution blocks to extract spectral, spatial and elevation features from HSI and LiDAR data, and uses cross-correlation attention blocks to fuse the extracted spectral features with the spatial features of HSI and the elevation features of LiDAR, and also uses cross-correlation attention blocks to fuse the spatial features of HSI and LiDAR. The present invention can fully utilize the spectral information of HSI and the elevation information of LiDAR to improve the accuracy of land cover classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a land cover classification method for a hyperspectral-lidar image, and belongs to the technical field of land cover classification. Background Art

[0002] In recent years, remote sensing technology has played a vital role in Earth observation missions. With the advancement of sensor technology, the sources of remote sensing imaging have become increasingly diverse. While this implementation method utilizes a large amount of data from a variety of sensor sources, remote sensing data from each source can only capture one or a few specific attributes and cannot fully and comprehensively describe the observed scene information. This significantly limits subsequent applications. Multi-source remote sensing data fusion technology is a viable solution to this problem. By integrating remote sensing images from multiple data sources and achieving information complementarity, more complete scene information can be obtained, enabling more reliable and accurate execution of the target mission.

[0003] Depending on the source of remote sensing data, multi-source fusion techniques can be categorized as homogeneous fusion and heterogeneous fusion. Based on the fusion level, fusion techniques can be further divided into pixel-level fusion, feature-level fusion, and decision-level fusion. Compared to homogeneous remote sensing data, heterogeneous remote sensing data offers greater diversity and complementarity, and thus has great potential for practical applications. Due to the significant differences in observed target features, heterogeneous remote sensing data is typically processed using feature-level fusion or decision-level fusion methods. Taking hyperspectral-LiDAR imagery (HSI) as an example, hyperspectral imagery (HSI) possesses rich spectral information, enabling accurate classification and identification of land cover. However, hyperspectral imagery suffers from the problem of "same spectrum, different objects, same object, different spectra," which hinders its effective use in practical applications. Digital surface model (DSM) data provided by LiDAR imagery, based on laser detection and ranging technology, can accurately depict the elevation and three-dimensional spatial geometry of land cover, but lacks the ability to obtain detailed spectral information of land objects. The combined use of HSI and LiDAR data can significantly improve the accuracy and reliability of land cover classification, introducing a new approach for precisely identifying ground cover. Traditional HSI-LiDAR data fusion methods use manually designed feature extraction techniques to fuse HSI and LiDAR data. However, these methods only utilize the individual features of the HSI and LiDAR data and cannot achieve true feature-level fusion. To achieve feature-level fusion, an improved method utilizes feature dimensionality reduction and correlation subspace techniques to uniformly map HSI and LiDAR data into a shared feature space. The joint features of HSI and LiDAR can be well represented in the shared space. However, feature mapping is a challenging problem, and it is difficult to find an effective quantitative method to map land cover features that contain significant differences into an appropriate shared feature space. In addition to the aforementioned feature-level fusion methods, some researchers have proposed decision-level fusion methods, such as maximum voting fusion and additional decision fusion based on weighted majority voting. However, the practicality of these methods is limited by the quality of features extracted by the manually designed feature extraction methods, which restricts their application in complex scenes.

[0004] In recent years, deep learning-based methods have become a popular choice for HSI-LiDAR data fusion due to their excellent feature extraction capabilities. Several improved deep learning fusion methods have been proposed, effectively improving data fusion performance and enhancing the accuracy and precision of subsequent tasks. However, multi-source remote sensing data contains a rich variety of surface cover types and complex spatial and spectral structures, making it difficult for a single-scale network to fully and accurately capture the characteristics of ground cover. Furthermore, existing methods fail to fully utilize the spectral information of HSI and the complementary spatial information of HSI and LiDAR. This results in insufficient extraction of HSI spectral information and poor complementary fusion of HSI-LiDAR spatial information, resulting in low utilization of spectral and spatial information. Summary of the Invention

[0005] Aiming at the problem that the diversity of land cover objects makes the spatial and spectral structure of remote sensing data complicated and the existing land cover classification is difficult to effectively extract discriminant features, the present invention provides a land cover classification method for hyperspectral-lidar images.

[0006] A land cover classification method for hyperspectral-lidar images of the present invention comprises:

[0007] The hyperspectral image and the lidar image of the land cover are input into a multi-scale cross-correlation attention convolutional fusion network to obtain the joint semantic features of the hyperspectral image and the lidar image, and the joint semantic features are input into a classification module, which generates the final classification results;

[0008] The multi-scale cross-correlation attention convolution fusion network includes a multi-scale residual spatial-spectral convolution feature extraction module, a multi-scale residual No. 1 spatial convolution feature extraction module and a cross-correlation attention fusion module;

[0009] The spatial-spectral convolution feature extraction module includes a spectral convolution feature extraction module and a multi-scale residual No. 2 spatial convolution feature extraction module;

[0010] The mutual-correlation attention fusion module includes mutual-correlation attention block 1, mutual-correlation attention block 2, and mutual-correlation attention block 3;

[0011] The lidar image is input into the No. 1 spatial convolution feature extraction module, which outputs the spatial features of the lidar image.

[0012] The hyperspectral image is input to the spectral convolution feature extraction module and the No. 2 spatial convolution feature extraction module at the same time. The spectral convolution feature extraction module outputs the spectral features in the hyperspectral image; the No. 2 spatial convolution feature extraction module outputs the spatial features in the hyperspectral image;

[0013] Cross-correlation attention block 1 fuses the spectral features and spatial features in the hyperspectral image to obtain the spatial-spectral features of the hyperspectral image;

[0014] Cross-correlation attention block 2 fuses the spatial features of the lidar image with the spectral features of the hyperspectral image to obtain the spatial-spectral features of the lidar image. Cross-correlation attention block 3 fuses the spatial-spectral features of the hyperspectral image with the spatial-spectral features of the lidar image to obtain the joint spatial features of the hyperspectral image and the lidar image.

[0015] The joint spatial features of the hyperspectral image and the lidar image are fused with the spatial-spectral features of the hyperspectral image to obtain the joint semantic features;

[0016] The joint semantic features are input into the classification module to obtain the final classification results.

[0017] Preferably, the data cube of the lidar image 1 is the number of data cubes, h×w is the spatial scale of the data cube, and c is the number of spectral channels;

[0018] The spatial convolution feature extraction module No. 1 uses a 1×1×1 convolution layer to extract the data cube The operation is performed to increase the number of data cubes to f, and then three spatial multi-scale residual convolution blocks are connected in series to extract spatial features from the increased number of data cubes. The spatial features extracted by the three spatial multi-scale residual convolution blocks are merged using one-time aggregation, and a 1×1×1 convolution layer is used to convert the scale of the merged spatial features from 3f×h×w×1 to f×h×w×1. The residual method is used to retain shallow features to obtain the spatial features of the lidar image.

[0019] Preferably, a data cube of a hyperspectral image 1 is the number of data cubes, h×w is the spatial scale of the data cube, and c is the number of spectral channels;

[0020] The spectral convolution feature extraction module uses a 1×1×1 convolution layer to extract the data cube The operation is performed to reduce the number of spectral channels of the hyperspectral image to 1 and increase the number of data cubes to f. Then, three spectral multi-scale residual convolution blocks in series are used to extract spectral features from the increased data cubes. The features extracted by the three spectral multi-scale residual convolution blocks are merged using one-time aggregation. The merged spectral features are converted from 3f×h×w×c to f×h×w×c using a 1×1×1 convolutional layer, and the residual method is used to retain shallow features to obtain the spectral features of the hyperspectral image.

[0021] As a preference, the No. 2 spatial convolution feature extraction module uses a 1×1×c convolution layer to extract the data cube. Perform operations to reduce the number of spectral channels of the hyperspectral image to 1 and increase the data cube The number of spatial features is increased to f, and then three spatial multi-scale residual convolution blocks are connected in series to extract spatial features from the increased data cube. The spatial features extracted by the three spatial multi-scale residual convolution blocks are merged using one-time aggregation. A 1×1×1 convolutional layer is used to convert the scale of the merged spatial features from 3f×h×w×1 to f×h×w×1, and the residual method is used to retain shallow features to obtain the spatial features of the lidar image.

[0022] As a preference, the extraction process of the spatial multi-scale residual convolution block is:

[0023] F1=Concat(Conv1(FM in ),Conv2(FM in ),Conv3(FM in ))

[0024] FM out1 =Mish(BN(Conv4(Mish(BN(F1)))))+FM in

[0025] Among them, the input features Where f is the number of data cubes, h×w is the spatial scale, d is the number of spectral channels, Conv1 is a 3×3×1 convolutional layer, Conv2 is a 5×5×1 convolutional layer, Conv3 is a 7×7×1 convolutional layer, Conv4 is a 1×1×1 convolutional layer, Concat(·) represents the concatenation operator, Mish(·) represents the Mish activation function, BN(·) represents the batch normalization layer, and FM out1 Represents the output of the spatial multi-scale residual convolution block.

[0026] As a preference, the extraction process of the spectral multi-scale residual convolution block is:

[0027] F2=Concat(Conv5(FM in ),Conv6(FM in ),Conv7(FM in ))

[0028] FM out2 =Mish(BN(Conv8(Mish(BN(F2)))))+FM in

[0029] Among them, the input features Where f is the number of data cubes, h×w is the spatial scale, d is the number of spectral channels, Conv5 is a 1×1×3 convolutional layer, Conv6 is a 1×1×5 convolutional layer, Conv7 is a 1×1×7 convolutional layer, Conv8 is a 1×1×1 convolutional layer, Concat(·) represents the concatenation operator, Mish(·) represents the Mish activation function, BN(·) represents the batch normalization layer, and FM out2 Represents the output of the spectral multi-scale residual convolution block.

[0030] As a preference, the process of fusing the cross-correlated attention blocks is as follows:

[0031]

[0032] Q′=Reshape(Conv9(Q)),K′=Reshape(Conv10(K))

[0033]

[0034] F v =A×Reshape(Conv11(V))

[0035] FM out3 =Reshape(LN(Reshape(Q+Conv12(Reshape(F v )))))

[0036] in, is the input data, Conv9(·) and Conv10(·) are both 3×3 convolution layers, Reshape(·) is the reshape layer, softmax represents the softmax regression function, Conv11 and Conv12 are both 1×1 convolution layers, and LN(·) represents layer normalization.

[0037] Preferably, the classification module includes an average pooling layer, a batch normalization layer, a Mish activation function, a reshape layer, a dropout layer and a linear layer connected in sequence.

[0038] As a preference, a multi-scale cross-correlation attention convolutional fusion network and a classification module form a land cover classification network. During training, a cross entropy loss function is used to calculate the loss of the land cover classification network output.

[0039] Beneficial effects of the present invention: The present invention proposes a multi-scale cross-correlation attention convolutional fusion network for HSI-LiDAR land cover classification, which uses a multi-scale residual convolution block to extract spectral, spatial and elevation features from HSI and LiDAR data to ensure the comprehensiveness and accuracy of feature information. Furthermore, the application proposes a spectral multi-scale residual convolution block to extract the spectral features of HSI, and uses a cross-correlation attention block to fuse the extracted spectral features with the spatial features of HSI and the elevation features of LiDAR to make full use of the spectral data. In addition, the cross-correlation attention block is also used to fuse the spatial features of HSI and LiDAR. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 This is the overall structure diagram of the multi-scale cross-correlation attention convolution fusion network;

[0041] Figure 2 Detailed structure diagram of the multi-scale cross-correlation attention convolution fusion network;

[0042] Figure 3 Figure 2 is the structure diagram of the spatial-spectral multi-scale residual convolution block, where (a) is the spatial multi-scale residual convolution block and (b) is the spectral multi-scale residual convolution block.

[0043] Figure 4 This is the structure diagram of the cross-correlation attention block;

[0044] Figure 5 Trento dataset, where (a) is a pseudo-color image (31, 14, 2 bands); (b) is a LiDAR DSM image; (c) is a ground-marked map;

[0045] Figure 6 MUUFL dataset, where (a) is a pseudo-color image (31, 16, 6 bands); (b) is a LiDAR DSM image; (c) is a ground-marked map;

[0046] Figure 7 The Houston 2013 dataset includes (a) a pseudo-color image (69, 36, 12 bands), (b) a LiDARDSM image, and (c) a ground-marked map.

[0047] Figure 8 The full pixel classification map of the Trento dataset, (a) is SVM; (b) is HYSN; (c) is FusAtNet; (d) is CCNN; (e) is AM 3 net; (f) is HCTnet; (g) is Sal 2 RN; (h) the method of the present invention; (i) is a ground marking image; (j) is a pseudo-color image;

[0048] Figure 9 The full pixel classification map of the MUUFL dataset, (a) is SVM; (b) is HYSN; (c) is FusAtNet; (d) is CCNN; (e) is AM 3 net; (f) is HCTnet; (g) is Sal 2 RN; (h) is the method of the present invention; (i) is the ground marking map; (j) is the pseudo-color map;

[0049] Figure 10 The full pixel classification map of the Houston2013 dataset, (a) is SVM; (b) is HYSN; (c) is FusAtNet; (d) is CCNN; (e) is AM 3 net; (f) is HCTnet; (g) is Sal 2 RN; (h) is the method of the present invention; (i) is the ground marking map; (j) is the pseudo-color map;

[0050] Figure 11 The effects of hyperparameters on the proposed method on the Trento, MUUFL, and Houston2013 datasets, where (a) is the spatial scale of the data cube (patch); (b) is the number of median nodes in the data cube (f); (c) is the channel compression ratio (r); and (d) is the number of multi-head attention heads (head).

[0051] Figure 12 The effect of different numbers of training samples on the method of the present invention, (a) is the Trento dataset; (b) is the MUUFL dataset; (c) is the Houston2013 dataset;

[0052] Figure 13 Table 1. DETAILED DESCRIPTION

[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0054] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features therein may be combined with each other.

[0055] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but they are not intended to limit the present invention.

[0056] The land cover classification method of hyperspectral-lidar images in this embodiment creates a multi-scale cross-correlation attention convolution fusion network, such as Figure 1 As shown in Figure 2, the multi-scale cross-correlation attention convolution fusion network includes a multi-scale residual spatial-spectral convolution feature extraction module, a spatial convolution feature extraction module, and a cross-correlation attention fusion module.

[0057] The hyperspectral image and the lidar image of the land cover are input into a multi-scale cross-correlation attention convolutional fusion network to obtain the joint semantic features of the hyperspectral image and the lidar image, and the joint semantic features are input into a classification module, which generates the final classification results;

[0058] like Figure 2 As shown, the spatial-spectral convolution feature extraction module includes a spectral convolution feature extraction module and a multi-scale residual spatial convolution feature extraction module;

[0059] The mutual-correlation attention fusion module includes mutual-correlation attention block 1, mutual-correlation attention block 2 and mutual-correlation attention block 3;

[0060] The lidar image is input into the No. 1 spatial convolution feature extraction module, which outputs the spatial features of the lidar image.

[0061] The hyperspectral image is input to the spectral convolution feature extraction module and the No. 2 spatial convolution feature extraction module at the same time. The spectral convolution feature extraction module outputs the spectral features in the hyperspectral image; the No. 2 spatial convolution feature extraction module outputs the spatial features in the hyperspectral image;

[0062] Cross-correlation attention block 1 fuses the spectral features and spatial features in the hyperspectral image to obtain the spatial-spectral features of the hyperspectral image;

[0063] Cross-correlation attention block 2 fuses the spatial features of the lidar image with the spectral features of the hyperspectral image to obtain the spatial-spectral features of the lidar image. Cross-correlation attention block 3 fuses the spatial-spectral features of the hyperspectral image with the spatial-spectral features of the lidar image to obtain the joint spatial features of the hyperspectral image and the lidar image.

[0064] The joint spatial features of the hyperspectral image and the lidar image are fused with the spatial-spectral features of the hyperspectral image to obtain the joint semantic features;

[0065] The joint semantic features are input into the classification module to obtain the final classification results.

[0066] The multi-scale cross-correlation attention convolutional fusion network and classification module are trained and then applied to land cover classification.

[0067] The input of the multi-scale cross-correlation attention convolutional fusion network in this embodiment is a remote sensing dataset from two different sensor sources, namely the HSI dataset and the LiDAR dataset, and the output is the classification label of each pixel in the remote sensing image. Specifically, given an HSI dataset and its corresponding LiDAR dataset Where H and W are the height and width of the two datasets, and c is the number of spectral channels (the number of spectral dimensions) of the HSI. The input HSI cube can be described as in represents the HSI data of the i-th pixel, 1 is the number of data cubes, h×w is the spatial scale of the data cube, and c is the number of spectral channels.

[0068] The input LiDAR cube can be represented as That is, the LiDAR data of the i-th pixel. The i-th input data of the proposed network can be expressed as The true feature label of the i-th input data is expressed as Where C is the number of categories. The output of the network is expressed as

[0069] In this embodiment, the spatial convolution feature extraction module is used to extract spatial features from the LiDAR data set. Since the number of spectral channels of the LiDAR data is 1, the spatial convolution feature extraction module No. 1 uses a 1×1×1 convolution layer to extract spatial features from the data cube. The operation is performed to increase the number of data cubes to f, and then three spatial multi-scale residual convolution blocks are connected in series to extract spatial features from the increased number of data cubes. The spatial features extracted by the three spatial multi-scale residual convolution blocks are merged using one-time aggregation, and a 1×1×1 convolution layer is used to convert the scale of the merged spatial features from 3f×h×w×1 to f×h×w×1. The residual method is used to retain shallow features, and the batch normalization method and the Mish activation function are used after the convolution layer to avoid network overfitting and provide the network with nonlinear capabilities to obtain the spatial features of the lidar image.

[0070] In this embodiment, the spectral convolution feature extraction module uses a 1×1×1 convolution layer to extract the data cube. The operation is performed to reduce the number of spectral channels of the hyperspectral image to 1 and increase the number of data cubes to f. Then, three spectral multi-scale residual convolution blocks in series are used to extract spectral features from the increased number of data cubes. The features extracted by the three spectral multi-scale residual convolution blocks are merged using one-time aggregation. The merged spectral features are converted from 3f×h×w×c to f×h×w×c using a 1×1×1 convolutional layer. The residual method is used to retain shallow features to avoid network degradation. The batch normalization method and the Mish activation function are used after the convolutional layer to avoid network overfitting and provide the network with nonlinear capabilities to obtain the spectral features of the hyperspectral image.

[0071] In this embodiment, the spatial convolution feature extraction module No. 2 uses a 1×1×c convolution layer to extract the data cube Perform operations to reduce the number of spectral channels of the hyperspectral image to 1 and increase the data cube The number of spatial features is increased to f, and then three spatial multi-scale residual convolution blocks are connected in series to extract spatial features from the increased data cube. The spatial features extracted by the three spatial multi-scale residual convolution blocks are merged using one-time aggregation. A 1×1×1 convolutional layer is used to convert the scale of the merged spatial features from 3f×h×w×1 to f×h×w×1, and the residual method is used to retain shallow features to avoid network degradation. The batch normalization method and the Mish activation function are used after the convolution layer to avoid network overfitting and provide the network with nonlinear capabilities to obtain the spatial features of the lidar image.

[0072] In the proposed network, this embodiment uses convolution to extract spatial and spectral features from HSI and LiDAR datasets. In order to obtain data features with more expressive and discriminative capabilities, this embodiment uses two improved multi-scale residual convolution blocks to extract multi-scale features from HSI and LiDAR datasets. These two residual convolution blocks are named spatial multi-scale residual convolution block and spectral multi-scale residual convolution block respectively. Unlike the traditional multi-scale convolution kernel whose spatial and spectral scales vary in coordination, the convolution block designed in this embodiment clearly divides the convolution kernel into spatial convolution kernel and spectral convolution kernel. The scale of the convolution kernel only changes in the spatial or spectral dimension, which can reduce the computational complexity and make the network more focused on spatial or spectral features. In addition, the proposed convolution block also uses residual aggregation technology to enhance the stability and anti-degradation ability of the network.

[0073] This embodiment uses the spatial-spectral convolutional feature extraction module and the spatial convolutional feature extraction module to obtain three features: the spatial features of HSI, the spectral features of HSI, and the spatial features of LiDAR. These three features are fused using the cross-correlation attention fusion module. To fully utilize the spectral features of HSI, the proposed network uses two cross-correlation attention blocks to fuse the spectral features of HSI with the spatial features of HSI and LiDAR, respectively. The outputs of these two cross-correlation attention blocks are the HSI spatial-spectral features and the LiDAR spatial-spectral features, both of which are derived from the dependencies between spatial and spectral information obtained through the attention mechanism. This embodiment then fuses these two features using another cross-correlation attention block to obtain a joint spatial feature from HSI and LiDAR, named the HSI-LiDAR spatial feature map. The proposed network then uses a join operator to fuse the HSI spatial-spectral feature map and the HSI-LiDAR spatial feature map to avoid losing the rich and highly discriminative HSI data, generating a joint semantic feature representation after feature fusion.

[0074] In the network proposed in this embodiment, the structures of the spatial multi-scale residual convolution block and the spectral multi-scale residual convolution block are given, as shown in FIG. Figure 3 As shown, assuming the input features are Where f is the number of data cubes, h×w is the spatial scale, and d is the number of spectral channels. In the spatial multi-scale residual convolution block, three convolution layers (3×3×1), (5×5×1) and (7×7×1) are used to extract the spatial features of the remote sensing image. The number of data cubes is reduced to half of the original to reduce the complexity of the model. Then, this embodiment uses the connection operator to aggregate the obtained multi-scale features into one data feature, uses the (1×1×1) convolution layer to reduce the number of data cubes, and performs residual aggregation with the input features to maintain shallow features, making it easier for the network to converge. Batch Norm and Mish are used after the convolution layer to provide stability and nonlinear characteristics to the network. The extraction process of the spatial multi-scale residual convolution block is expressed by the following formula:

[0075] F1=Concat(Conv1(FM in ),Conv2(FM in ),Conv3(FM in ))

[0076] FM out1 =Mish(BN(Conv4(Mish(BN(F1)))))+FM in

[0077] Among them, the input features Where f is the number of data cubes, h×w is the spatial scale, d is the number of spectral channels, Conv1 is a 3×3×1 convolutional layer, Conv2 is a 5×5×1 convolutional layer, Conv3 is a 7×7×1 convolutional layer, Conv4 is a 1×1×1 convolutional layer, Concat(·) represents the concatenation operator, Mish(·) represents the Mish activation function, BN(·) represents the batch normalization layer, and FM out1 Represents the output of the spatial multi-scale residual convolution block.

[0078] The overall structure of the spectral multi-scale residual convolution block is similar to that of the spatial multi-scale residual convolution block. The difference is that the convolution kernel scales of the former are different in the spectral dimension, namely (1×1×3), (1×1×5), and (1×1×7). The extraction process of the spectral multi-scale residual convolution block is:

[0079] F2=Concat(Conv5(FM in ),Conv6(FM in ),Conv7(FM in ))

[0080] FM out2 =Mish(BN(Conv8(Mish(BN(F2)))))+FM in

[0081] Among them, the input features Where f is the number of data cubes, h×w is the spatial scale, d is the number of spectral channels, Conv5 is a 1×1×3 convolutional layer, Conv6 is a 1×1×5 convolutional layer, Conv7 is a 1×1×7 convolutional layer, Conv8 is a 1×1×1 convolutional layer, Concat(·) represents the concatenation operator, Mish(·) represents the Mish activation function, BN(·) represents the batch normalization layer, and FM out2 Represents the output of the spectral multi-scale residual convolution block.

[0082] In the network proposed in this embodiment, a cross-correlation attention module is designed to fuse HSI and LiDAR data features. Before considering global dependencies, the cross-correlation attention module needs to collect local features. Therefore, this embodiment replaces the multilayer perceptron widely used in standard attention modules with convolutional layers. Furthermore, this module employs channel compression and multi-head attention techniques to reduce network complexity and increase flexibility.

[0083] Specifically, the structure of the cross-correlation attention module is as follows Figure 4 As shown. Given input data and Where h×w and h′×w′ are spatial scales, d1 and d2 are spectral scales. The cross-correlation attention module first generates input data, namely query, key and value. The query (Q) is given by Generated, key (K) and value (V) by Generation. To collect local features between Q and K, the proposed module uses 2 (3×3) convolutional layers to transform the features. Then, scaled dot products are applied to provide multi-head attention weights A∈[0,1] hw×h′w′ The process of fusing the cross-correlated attention blocks:

[0084]

[0085] Q′=Reshape(Conv9(Q)),K′=Reshape(Conv10(K))

[0086]

[0087] The weighted V can be obtained, expressed as F v . F v is fed into a single feed-forward network (FFN) represented by a (1×1) convolutional layer to obtain further results. Finally, this embodiment uses residual aggregation technology and layer normalization (LayerNorm) method to obtain the final result (FM out ). In addition, the cross-correlation attention block contains two hyperparameters, namely the channel compression rate (r) and the number of heads of multi-head attention (head). The process can be expressed by the following formula:

[0088] F v =A×Reshape(Conv11(V))

[0089] FM out3 =Reshape(LN(Reshape(Q+Conv12(Reshape(F v )))))

[0090] in, is the input data, Conv9(·) and Conv10(·) are both 3×3 convolution layers, Reshape(·) is the reshape layer, softmax represents the softmax regression function, Conv11 and Conv12 are both 1×1 convolution layers, and LN(·) represents layer normalization.

[0091] In this embodiment, the classification module includes an average pooling layer, a batch normalization layer, a Mish activation function, a reshape layer, a dropout layer, and a linear layer connected in sequence.

[0092] The average pooling layer is used to concentrate the information of joint semantic features. Batch Norm is used to standardize the data, and the formula can be expressed as

[0093]

[0094] Where B={x 1…m} is the value of x in a mini-batch, μ B is the mean of the mini-batch, is the variance of the mini-batch, is x i Normalization, BN(x i ) is x i The batch norm transformation, γ and β are the parameters to be learned. The Mish activation function is used to provide the network with nonlinear mapping capabilities and can be expressed as

[0095] softplus(x)=ln(1+e x )

[0096] Mish(x)=x×tanh(softplus(x))

[0097] The reshape layer is used to eliminate the dimensions of the data features. The dropout layer is used to improve the generalization ability of the network. The linear layer is used to generate the output vector.

[0098] In this embodiment, a multi-scale cross-correlation attention convolution fusion network and a classification module constitute a land cover classification network. During training, a cross entropy loss function is used to calculate the loss of the land cover classification network output.

[0099] The cross entropy loss function is applied to calculate the loss of the network output and can be expressed as

[0100]

[0101] Among them, L i is the cross entropy loss of the i-th pixel. In addition, this embodiment also adopts early stopping and dynamic learning rate techniques to shorten the training time of the network and provide better network convergence.

[0102] Experimental verification:

[0103] 1. Data Description

[0104] In the experiments, this implementation uses three HSI-LiDAR datasets with different land cover types to evaluate the effectiveness of the proposed network: the Trento dataset, the MUUFL dataset, and the Houston 2013 dataset. A brief introduction to these datasets is as follows:

[0105] 1. Trento dataset: The Trento dataset is an HSI-LiDAR dataset, in which the HSI data is collected by the AISAEagle sensor and the LiDAR DSM data is collected by the Optech ALTM 3100EA sensor. The dataset was collected in a rural area south of the city of Trento, Italy. The spatial scale of the Trento dataset is 166×600, and the spatial resolution is about 1 meter. The HSI data contains 63 bands with spectral wavelengths from 420 nanometers to 990 nanometers. LiDAR DSM data can reflect the height of ground cover. Land cover is divided into 6 categories, including Apple trees, Buildings, Ground, Woods, Vineyard and Roads. In the experiment, this embodiment randomly selects 1% of the labeled samples as training samples and verification samples, and the remaining labeled samples are used as test samples. The pseudo-color images, LiDAR DSM images and ground labeled maps of the Trento dataset are as follows: Figure 5 Table 1 lists the dataset categories, colors and sample numbers in detail.

[0106] 2. MUUFL dataset: The MUUFL dataset was collected by the ITERS CASI-1500 sensor at the Gulf Park campus of the University of Southern Mississippi in Long Beach, Mississippi in November 2010, and includes an HSI dataset and a LiDAR dataset. The spatial scale of HSI and LiDAR is 325×220, and the spatial resolution is 0.54×1.0 meters. HSI contains 64 available bands with a wavelength range of 375 to 1050 nanometers. LiDAR data can reflect the height of ground cover. Surface cover is divided into 11 categories, including Trees, Mostly grass, Mixed ground surface, Dirt and sand, Road, Water, Building Shadow, Building, Sidewalk, Yellow curb, and Cloth panels. This implementation method randomly extracts 1% of the labeled samples as training samples and verification samples, respectively, and the remaining labeled samples are test samples. The pseudo-color image of HSI, LiDARDSM image, and ground labeled map are as follows: Figure 6 Table 1 shows the categories, colors and number of samples of the dataset.

[0107] 3. Houston2013 dataset: The Houston2013 dataset was acquired by the ITERS CASI-1500 sensor over the University of Houston campus and surrounding urban areas in Houston, Texas, USA in 2012. This dataset consists of HSI data and LiDAR DSM data. The HSI data contains 144 spectral channels from 380 to 1050 nanometers. LiDAR data can reflect the height of ground cover. Land cover is divided into 15 categories, including Healthy grass, Stressed grass, Syntheticgrass, Trees, Soil, Water, Residential, Commercial, Road, Highway, Railway, Parking Lot 1, Parking Lot 2, Tennis Court, and Running track. In this embodiment, 1% of the labeled samples are randomly selected as training samples and verification samples, and the remaining labeled samples are used as test samples. The pseudo-color images of HSI data, LiDAR DSM images, and ground labeled maps of the Houston2013 dataset are shown in Figure 2. Figure 7 Table 1 gives the categories, colors and number of samples.

[0108] 2. Experimental Setup and Evaluation Metrics

[0109] This implementation collects seven representative methods, including SVM, HYSN, FusAtNet, CCNN, AM3net, HCTnet, and Sal2RN, for comparison. SVM and HYSN represent traditional classification methods and convolutional networks, respectively. FusAtNet, CCNN, AM3net, HCTnet, and Sal2RN represent the most advanced multi-source remote sensing dataset fusion networks. In the experiment, SVM with RBF kernel was used, where the penalty parameter C and RBF kernel width σ were selected by Grid SearchCV, both in the range of (10 -2 ,10 2 ). All deep learning-based methods use the default parameter settings specified in their original papers. In this implementation, the data enhancement function of FusAtNet is removed in the experiment. The HSI and LiDAR datasets are concatenated and used as input data for SVM and HYSN. For the method of the present invention, the spatial scale of the data cube is set to 7×7, the batch size is set to 32, the number of data cubes (f) is set to 24, the channel compression rate (r) is set to 2, the number of heads of multi-head attention is set to 2, and the dropout rate is set to 0.5. Other hyperparameters are set as follows: epoch is set to 200, the initial learning rate is set to 5.0×10 -4, the decay rate of Adam optimizer is set to (0.9, 0.999), and the fuzziness factor is set to 10 -8 , the epoch of the cosine annealing technique is set to 200, and the epoch of the early stopping technique is set to 50.

[0110] In this implementation, the overall classification accuracy (OA), average classification accuracy (AA) and Kappa are used to quantitatively measure the performance of the comparison method. All experimental results are repeated ten times independently, and the average value is taken as the final result. The experimental hardware environment is a computer equipped with an Intel Xeon E5-2680v4 processor 2.4GHz and an NVIDIA TM A workstation with a GeForce RTX2080Ti GPU. The software environment is CUDAv11.2, PyTorch 1.10, and Python 3.8.

[0111] 3. Experimental Results

[0112] This implementation first compares the performance of various methods on the Trento dataset. The classification results are shown in Table 2, and the full pixel classification diagram is shown in Figure 8 . This embodiment marks the best classification accuracy of each category as well as the overall classification accuracy, average classification accuracy, Kappa and training time in bold in the table. By observing the classification accuracy of each method, this embodiment can find that the Trento dataset is relatively easy to classify. The overall classification accuracy of all fusion networks tends to saturation (greater than 98%). The overall classification accuracy of SVM is the lowest (91.97%), indicating that the pixel classification method using only spectral features is not as effective as the method using spatial-spectral features. Specifically, SVM has difficulty classifying C1 and C3 (75.39% and 87.95%). This shows that these two types of ground cover are difficult to distinguish by spectral features alone. The overall classification accuracy of HYSN (95.78%) is higher than that of SVM. Among them, the classification accuracy of C1 and C3 is 21.80% and 6.29% higher than that of SVM, respectively. This shows that the introduction of spatial information helps to improve the discrimination ability of ground cover. However, the classification accuracy of HYSN for C2 is 14.82% lower than that of SVM. Observation Figure 9The classification plot in (b) shows that some buildings are classified as ground cover types such as roads and vineyards. This indicates that the introduction of spatial information can affect the ground cover classification results by surrounding objects, thereby reducing classification accuracy. This problem can be alleviated by introducing additional discriminative information to help the deep network identify ground cover, such as LiDAR images with elevation information. Fusion networks such as FusAtNet, CCNN, AM3net, HCTnet, Sal2RN, and the proposed method are used to fuse information from HSI and LiDAR datasets and provide ground cover classification results. Experimental results show that the fusion network achieves higher classification accuracy than SVM and HYSN. Comparing various fusion networks, the proposed method achieves the highest overall classification accuracy (98.83%) among the compared methods. The proposed method achieves classification accuracies of 97.67%, 96.02%, and 99.90% for C1, C2, and C3, respectively, demonstrating that the inclusion of LiDAR data can effectively distinguish these three ground cover types. The SVM has the best training time (4.63 seconds) among the compared methods. Comparing deep neural networks, HCTnet has the shortest training time (10.07 seconds). In contrast, FusAtNet has the longest training time (68.03 seconds). The training time of the proposed method is 30.98 seconds.

[0113] Table 2 Classification results and training schedule of the Trento dataset

[0114]

[0115] To further test the performance of the proposed method, this implementation was experimented on the MUUFL dataset, which contains complex terrain and an unbalanced number of labeled samples. The classification results and full pixel classification maps are shown in Tables 3 and Figure 9. Unlike the experimental results on the Trento dataset, the overall classification accuracy of SVM is 86.36%, which is 2.68%, 2.63%, 0.86 and 2.26% higher than HSYN, FusAtNet, HCTnet and Sal2RN, respectively. From the classification accuracy of each category, C3, C7, C9 and C10 are difficult to classify in all methods. After introducing spatial information (cube-based pixel classification) and elevation information (LiDAR dataset), the classification accuracy of C3, C7 and C9 is improved by 3.51% (AM3net), 14.47% (CCNN) and 16.48% (CCNN) compared with SVM, respectively. However, the accuracy of C10 decreases from 67.29% (SVM) to 5.98% (HCTnet). Looking closely at C10, it can be seen that the number of labeled samples in this embodiment is 183, and 2 pixels are used as training samples when training the network. By observing the ground marking map, this embodiment can see that C10 represents yellow curbstone, which is the ground cover next to the road and has a smaller scale. Taking into account the number of labeled samples and the spatial distribution of C10, this embodiment summarizes the reasons for the decline in accuracy as insufficient network training and the difficulty of the cube-based classification method in extracting fine spatial features. Comparing the classification accuracy of deep neural networks for C10, the accuracy of the method of the present invention (23.68%) is relatively higher than that of other methods. Comparing the overall classification accuracy of various methods, the overall classification accuracy of the method of the present invention (89.24%) is the highest among all compared methods. SVM has the shortest training time (11.75 seconds). FusAtNet has the longest training time (225.94 seconds).

[0116] Table 3 Classification results and training schedule of MUUFL dataset

[0117]

[0118] In order to further test the performance of the proposed network under small sample conditions, this implementation was experimented on the Houston2013 dataset, where the number of training samples varied from 6 to 12. The classification results and full pixel classification maps are shown in Tables 4 and Figure 10 . It can be seen that the overall classification accuracy values ​​of SVM, HYSN and FusAtNet (79.89%, 79.63% and 77.30%) are lower than those of other fusion networks. This shows that these three classification methods perform poorly under conditions with fewer training samples. AM3net, HCTnet and Sal2RN achieved similar overall classification accuracies (86.55%, 89.29% and 88.84%). CCNN and the method of the present invention achieved the highest overall classification accuracy values ​​(91.73% and 92.32%). This shows that CCNN and the method of the present invention can be trained under small sample conditions to obtain classification models with strong generalization capabilities.

[0119] Table 4 Classification results and training schedule of Houston2013 dataset

[0120]

[0121] Influence of hyperparameters: To examine the impact of hyperparameters on the proposed network, this implementation tested four hyperparameters in the experiment, including the spatial scale of the data cube (patch), the number of data cubes (f), the channel compression rate (r), and the number of heads in the attention block (head). The experimental results are shown in Figure 2. Figure 11 Specifically, Figure 12 (a) shows the overall classification accuracy of the method of the present invention for the Trento, MUUFL, and Houston2013 datasets under data cubes of different spatial scales. This embodiment clearly shows that as the spatial scale of the data cube increases, the overall classification accuracy of the method of the present invention on these three datasets decreases. This phenomenon is easy to understand, because the larger the spatial scale of the data cube, the more spatial information is introduced, which affects the judgment of the center pixel, resulting in a decrease in overall classification accuracy. In the experiment, the default value of the spatial scale is set to patch = 7. Figure 12 (b) shows the overall classification accuracy of the method of the present invention on three HSI-LiDAR datasets with different numbers of data cubes appearing in the multi-scale residual convolution block. The more data cubes there are, the more learnable parameters there are and the more complex the network structure is. From the experimental results, it can be seen that when the number of data cubes is 32 (f=32), the overall classification accuracy of the method of the present invention on the three datasets is the highest. However, as the number of data cubes continues to increase (f=48), the overall classification accuracy of the method of the present invention decreases. This phenomenon shows that complex networks cannot provide better classification results. This can be explained by the fact that complex networks are difficult to be fully trained in limited labeled samples, resulting in a decrease in the generalization ability of the network. In combination with the complexity and accuracy of the network, this embodiment selects f=24 as the default hyperparameter setting in actual experiments. Figure 12 (c) shows the overall classification accuracy of the method of the present invention under different channel compression rates in the attention block. The greater the compression rate, the smaller the channel width in the attention block. From the experimental results, this embodiment can happily find that on the Trento and Houston2013 datasets, the higher the compression ratio (r=4), the higher the overall classification accuracy, while on the MUUFL dataset, the overall classification accuracy is slightly reduced (0.17%) compared with r=2. The experimental results provide confidence for this embodiment to improve the channel compression rate in the attention block. In the experiment, this embodiment sets the default value of the channel compression rate to r=2. Finally, this embodiment compares the performance of the method of the present invention under different numbers of heads in the attention block. The experimental results are as follows: Figure 12 As shown in (d). This embodiment shows that the number of heads in the multi-head attention has little impact on the Trento (98.77%, 98.83%, 98.95%) and MUUFL (89.15%, 89.24%, 88.96%) datasets, but has a greater impact on the Houston2013 dataset (90.15%, 92.32%, 92.83%). Theoretically, the more heads there are, the more flexible the attention mechanism is. In the experiment, this embodiment sets the default number of heads to head = 2.

[0122] The impact of the training sample ratio on the method:

[0123] In this section, this embodiment will examine the performance of the comparison method under different training sample ratios on three HSI-LiDAR datasets. In order to comprehensively analyze the performance of the comparison method under different training sample numbers, this embodiment randomly selects 0.5%, 1%, 2%, 4%, and 5% of labeled samples in each class to form the training sample set. The classification results are shown in Figure 2. Figure 12 As shown in Figure 2 . Generally speaking, a large number of training samples allows supervised learning methods to be fully trained and provides sufficient discriminative information, thereby helping the methods improve their discriminative capabilities. Experimental results confirm this: when the training sample ratio is small, the overall classification accuracy of the comparison methods is low, but when the number of training samples increases, the overall classification accuracy increases. When the training sample ratio is small (0.5%), the overall classification accuracy of CCNN, AM3net, and the proposed method decreases less than that of other methods. This indicates that under small sample conditions, these methods are better able to extract effective data features. As the training sample ratio increases, the classification accuracy of the proposed method remains at the forefront of the comparison methods. When the training sample ratio reaches 5%, the classification accuracy of all methods reaches its maximum. Among them, the proposed method achieves the highest classification accuracy in all three datasets, reaching saturation accuracy in the Trento (99.7%) and Houston2013 (98.6%) datasets, and an accuracy of 94.01% in the MUUFL dataset. The experimental results once again demonstrate the superiority of convolutional networks combined with attention mechanisms in remote sensing image pixel classification tasks, and provide ideas for the design of multi-source dataset fusion networks.

[0124] This embodiment proposes a multi-scale cross-correlation attention convolutional fusion network for pixel-level HSI-LiDAR classification, which is designated as the inventive method. The proposed method consists of four modules: a spatial-spectral convolutional feature extraction module, a spatial convolutional feature extraction module, a cross-correlation attention fusion module, and a classification module. The spatial-spectral convolutional feature extraction module extracts spatial and spectral features from the HSI dataset. The spatial convolutional feature extraction module extracts spatial features from the LiDAR dataset. The cross-correlation attention fusion module fuses HSI and LiDAR data features using cross-modal representation learning to generate a joint semantic information representation. The classification module maps semantic features into classification results. To improve the network's stability, convergence, and generalization, this embodiment employs techniques such as multi-scale residual convolution, cross-correlation attention, one-shot aggregation, residual aggregation, batch norm, and Mish activation function in the proposed network. To validate the effectiveness of the proposed method, three HSI-LiDAR datasets with varying land cover and spectral spatial resolutions were used. Seven related methods were also introduced for comparison. Experimental results demonstrate that the proposed method has superior discriminative capabilities compared to comparable methods. Furthermore, this implementation discusses the impact of hyperparameters and training sample ratios on the proposed method. In the future, this implementation will consider employing networks with different architectures to address multi-source remote sensing dataset fusion tasks.

[0125] Although the present invention is described herein with reference to specific embodiments, it should be understood that these embodiments are merely illustrative of the principles and applications of the invention. It should be understood that many modifications may be made to the illustrative embodiments, and that other arrangements may be devised, without departing from the spirit and scope of the invention as defined by the appended claims. It should be understood that the various dependent claims and features described herein may be combined in ways other than those described in the original claims. It should also be understood that features described in conjunction with individual embodiments may be employed in conjunction with other described embodiments.

Claims

1. A land cover classification method based on hyperspectral-lidar images, characterized in that: The method comprises: The hyperspectral image and the lidar image of the land cover are input into a multi-scale cross-correlation attention convolutional fusion network to obtain the joint semantic features of the hyperspectral image and the lidar image, and the joint semantic features are input into a classification module, which generates the final classification results; The multi-scale cross-correlation attention convolution fusion network includes a multi-scale residual spatial-spectral convolution feature extraction module, a multi-scale residual No. 1 spatial convolution feature extraction module and a cross-correlation attention fusion module; The spatial-spectral convolution feature extraction module of multi-scale residuals includes a spectral convolution feature extraction module and a No. 2 spatial convolution feature extraction module of multi-scale residuals; The mutual-correlation attention fusion module includes mutual-correlation attention block 1, mutual-correlation attention block 2, and mutual-correlation attention block 3; The lidar image is input into the No. 1 spatial convolution feature extraction module of the multi-scale residual, and the No. 1 spatial convolution feature extraction module of the multi-scale residual outputs the spatial features of the lidar image; The hyperspectral image is simultaneously input into the spectral convolution feature extraction module and the multi-scale residual No. 2 spatial convolution feature extraction module. The spectral convolution feature extraction module outputs the spectral features in the hyperspectral image; the multi-scale residual No. 2 spatial convolution feature extraction module outputs the spatial features in the hyperspectral image. Cross-correlation attention block 1 fuses the spectral features and spatial features in the hyperspectral image to obtain the spatial-spectral features of the hyperspectral image; Cross-correlation attention block 2 fuses the spatial features of the lidar image with the spectral features of the hyperspectral image to obtain the spatial-spectral features of the lidar image. Cross-correlation attention block 3 fuses the spatial-spectral features of the hyperspectral image with the spatial-spectral features of the lidar image to obtain the joint spatial features of the hyperspectral image and the lidar image. The joint spatial features of the hyperspectral image and the lidar image are fused with the spatial-spectral features of the hyperspectral image to obtain the joint semantic features; The joint semantic features are input into the classification module to obtain the final classification results.

2. The land cover classification method of hyperspectral-lidar image according to claim 1, characterized in that: Data cube of lidar image 1 is the number of data cubes, h×w is the spatial scale of the data cube, c is the number of spectral channels, and the number of spectral channels c is 1; the No. 1 spatial convolution feature extraction module of the multi-scale residual uses a 1×1×1 convolution layer to extract the data cube. The operation is performed to increase the number of data cubes to f, and then three spatial multi-scale residual convolution blocks are connected in series to extract spatial features from the increased number of data cubes. The spatial features extracted by the three spatial multi-scale residual convolution blocks are merged using one-time aggregation, and a 1×1×1 convolution layer is used to convert the scale of the merged spatial features from 3f×h×w×1 to f×h×w×1. The residual method is used to retain shallow features to obtain the spatial features of the lidar image.

3. The land cover classification method of hyperspectral-lidar image according to claim 1, characterized in that: Data cube of hyperspectral image 1 is the number of data cubes, h×w is the spatial scale of the data cube, and c is the number of spectral channels; The spectral convolution feature extraction module uses a 1×1×1 convolution layer to extract the data cube The operation is performed to reduce the number of spectral channels of the hyperspectral image to 1 and increase the number of data cubes to f. Then, three spectral multi-scale residual convolution blocks in series are used to extract spectral features from the increased data cubes. The spectral features extracted by the three spectral multi-scale residual convolution blocks are merged using one-time aggregation. The merged spectral features are converted from 3f×h×w×c to f×h×w×c using a 1×1×1 convolutional layer, and the residual method is used to retain shallow features to obtain the spectral features of the hyperspectral image.

4. The land cover classification method of hyperspectral-lidar image according to claim 3, characterized in that: The multi-scale residual No. 2 spatial convolution feature extraction module uses a 1×1×c convolution layer to extract the data cube Perform operations to reduce the number of spectral channels of the hyperspectral image to 1 and increase the data cube The number of spatial features is increased to f, and then three spatial multi-scale residual convolution blocks are connected in series to extract spatial features from the increased data cube. The spatial features extracted by the three spatial multi-scale residual convolution blocks are merged using one-time aggregation. A 1×1×1 convolutional layer is used to convert the scale of the merged spatial features from 3f×h×w×1 to f×h×w×1, and the residual method is used to retain shallow features to obtain the spatial features of the lidar image.

5. The land cover classification method based on hyperspectral-lidar images according to claim 2 or 4, characterized in that: The extraction process of the spatial multi-scale residual convolution block is: F1=Concat(Conv1(FM in ),Conv2(FM in ),Conv3(FM in )) FM out1 =Mish(BN(Conv4(Mish(BN(F1)))))+FM in Among them, the input features Where f is the number of data cubes, h×w is the spatial scale, d is the number of spectral channels, Conv1 is a 3×3×1 convolutional layer, Conv2 is a 5×5×1 convolutional layer, Conv3 is a 7×7×1 convolutional layer, Conv4 is a 1×1×1 convolutional layer, Concat(·) represents the concatenation operator, Mish(·) represents the Mish activation function, BN(·) represents the batch normalization layer, and FM out1 Represents the output of the spatial multi-scale residual convolution block.

6. The land cover classification method of hyperspectral-lidar image according to claim 3, characterized in that: The extraction process of the spectral multi-scale residual convolution block is: F2=Concat(Conv5(FM in ),Conv6(FM in ),Conv7(FM in )) <h2 style=";text-align:left;direction:ltr">FM<h2 style=";text-align:left;direction:ltr"> out2 <h2 style=";text-align:left;direction:ltr"> =Mish(BN(Conv8(Mish(BN(F2)))))+FM<h2 style=";text-align:left;direction:ltr"> in Among them, the input features Where f is the number of data cubes, h×w is the spatial scale, d is the number of spectral channels, Conv5 is a 1×1×3 convolutional layer, Conv6 is a 1×1×5 convolutional layer, Conv7 is a 1×1×7 convolutional layer, Conv8 is a 1×1×1 convolutional layer, concat(·) represents the concatenation operator, Mish(·) represents the Mish activation function, BN(·) represents the batch normalization layer, and FM out2 Represents the output of the spectral multi-scale residual convolution block.

7. The land cover classification method of hyperspectral-lidar image according to claim 1, characterized in that: The process of fusing the No. 1 mutual-correlation attention block, the No. 2 mutual-correlation attention block, or the No. 3 mutual-correlation attention block is as follows: Q′=Reshape(Conv9(Q)),K′=Reshape(Conv10(K)) F v =A×Reshape(Conv11(V)) FM out3 =Reshape(LN(Reshape(Q+Conv12(Reshape(F v ))))) in, is the input data, Conv9(·) and Conv10(·) are both 3×3 convolution layers, Reshape(·) is the reshape layer, softmax represents the softmax regression function, Conv11 and Conv12 are both 1×1 convolution layers, LN(·) represents layer normalization, and d1 is spectral scale.

8. The land cover classification method of hyperspectral-lidar image according to claim 1, characterized in that: The classification module includes an average pooling layer, a batch normalization layer, a Mish activation function, a reshape layer, a dropout layer, and a linear layer connected in sequence.

9. The land cover classification method of hyperspectral-lidar image according to claim 1, characterized in that: The multi-scale cross-correlation attention convolution fusion network and the classification module constitute the land cover classification network. During training, the cross entropy loss function is used to calculate the loss of the land cover classification network output.

10. A land cover classification device for hyperspectral-lidar images, comprising a storage device, a processor, and a computer program stored in the storage device and executable on the processor, characterized in that: The processor executes the computer program to implement the land cover classification method for hyperspectral-lidar images according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Hyperspectral and LiADR data collaborative classification method based on double branches

    CN114429564A

  • Hyperspectral and laser radar multimode image spatial-spectral fusion ground object identification method and device

    CN117036879A