Hyperspectral and laser radar image ground object coverage classification method based on multi-scale split reconstruction collaborative fusion network

Through the method of reconstructing a collaborative fusion network based on multi-scale splitting, the problem of insufficient multi-scale information extraction in the data fusion of hyperspectral images and lidar images is solved, and higher classification accuracy and more comprehensive feature representation are achieved.

CN120070966APending Publication Date: 2025-05-30QIQIHAR UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510129016.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient multi-scale information extraction, limitations in shallow feature extraction and insufficient data integration technology in the multi-source remote sensing data fusion of hyperspectral images and lidar images, resulting in low classification accuracy.

Method used

The method of reconstructing a collaborative fusion network based on multi-scale split reconstructing is adopted, and multi-scale features are extracted and fused through multi-scale hierarchical inverted pyramid module, cross-modal collaborative fusion module and classification module, and global relationship modeling and feature fusion are carried out through self-attention and cross-attention mechanisms.

Benefits of technology

It effectively improves the geographic coverage classification accuracy of hyperspectral and lidar images, and enhances the characterization ability of complex structures and semantic information of multi-source remote sensing data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070966A_ABST
    Figure CN120070966A_ABST
Patent Text Reader

Abstract

The invention relates to the field of ground feature coverage classification, and discloses a hyperspectral and laser radar image ground feature coverage classification method based on a multi-scale split reconstruction collaborative fusion network, and the model adopted by the method comprises a plurality of parallel multi-scale layered inverted pyramid modules MHIP, a cross-modal collaborative fusion module CCFM and a classification module. The MHIP module internally comprises a plurality of convolution splitting reconstruction modules CSRB of different scales and is used for extracting and fusing differentiated multi-scale features, the CCFM uses a self-attention mechanism to carry out global relation modeling on data of different modals by utilizing the consistency of spatial scales, and the CSRB is used for extracting and fusing the multi-scale features of different modalities. And then complementarity information of the heterogeneous data is fully fused by using a cross attention mechanism, a classification module maps fused features to a final classification result, and the model of the method has better generalization and effectiveness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of ground object coverage classification, and specifically to a method for classifying ground object coverage of hyperspectral and lidar images based on a multi-scale splitting and reconstruction collaborative fusion network. Background Art

[0002] In recent years, with the continuous progress of remote sensing technology, earth observation through remote sensing images has become a research hotspot. Using remote sensing data from different sensors to obtain remote sensing data from different sources for land cover identification is an important way to solve complex surface classification problems. Rich multi-source remote sensing data not only plays an important role in earth observation, but also provides new methods and challenges for a wide range of applications in the field of remote sensing. As one of the research hotspots in the field of remote sensing, hyperspectral images (HSIs) have received increasing attention because they contain a large number of continuous spectral bands and can provide rich spectral and spatial information. Due to its unique properties, HSIs are widely used in many fields such as environmental detection, mineral analysis, and precision agriculture. However, as a passive remote sensing technology, HSIs are easily affected by the atmosphere and light, resulting in various changes in spectral bands in real scenes. This characteristic makes HSIs face many challenges when applied to complex scenes. In addition, single remote sensing data can only capture its specific attributes and cannot effectively describe the overall information of the scene.

[0003] To solve the above problems, in the past decade, researchers have fully explored the fusion of multi-source remote sensing data. Lidar, as an active sensing technology, can accurately obtain the spatial structure and elevation information of ground objects. In addition, lidar is not easily affected by factors such as lighting conditions, weather changes, and surface material reflectivity, and has strong environmental adaptability. Due to different sensor types, multi-source remote sensing data can be divided into homogeneous data and heterogeneous data. HSIs and lidar data are two types of heterogeneous data that contain spatial-spectral information and spatial-elevation information, respectively. Research shows that for ground cover objects with similar spectral curves but different heights, the elevation information of lidar data can effectively alleviate the occurrence of the phenomenon of different objects with the same spectrum and the same object with different spectra in HSIs. Therefore, the spectral information of HSIs and the elevation information of lidar data can complement each other in land cover classification, effectively improving the classification accuracy of land use and land cover (LULC).

[0004] Although CNN- and Transformer-based methods have made significant progress in HSI-LiDAR joint classification, there are still some key challenges. First, the insufficient extraction of multi-scale information leads to the inability of existing networks to fully exploit the local detail information and global context information in multi-source remote sensing data, resulting in the correlation between the local and the whole being often overlooked. Second, existing methods are mostly limited to the extraction and fusion of shallow features and fail to fully utilize deep features to comprehensively represent the complex structure and semantic information of multi-source data. Finally, there is also a lack of sufficient data integration techniques to fully fuse HSI and LiDAR data. To solve the above problems, the present invention proposes a method for classifying ground object coverage of hyperspectral and lidar images based on a multi-scale splitting and reconstruction collaborative fusion network. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for classifying ground object coverage of hyperspectral and lidar images based on a multi-scale splitting and reconstruction collaborative fusion network to solve the problems raised in the above background art.

[0006] To achieve the above object, the present invention provides the following technical solution: A method for classifying ground object coverage of hyperspectral and lidar images based on a multi-scale splitting and reconstruction collaborative fusion network, the method is carried out by using a multi-scale splitting and reconstruction collaborative fusion network, the multi-scale splitting and reconstruction collaborative fusion network includes a plurality of parallel multi-scale hierarchical inverted pyramid modules, a cross-modal collaborative fusion module and a classification module, each multi-scale hierarchical inverted pyramid module contains a plurality of convolutional splitting and reconstruction modules with different scales, and the convolutional splitting and reconstruction modules are used to extract and fuse discriminative multi-scale features, the cross-modal collaborative fusion module, based on the consistency of the spatial scale, uses the self-attention mechanism to model the global relationship of data in different modalities, and then uses the cross-attention mechanism to fully fuse the complementary information of heterogeneous data, and the classification module maps the fused features to the final classification result.

[0007] Preferably: The data extraction process of the multi-scale hierarchical inverted pyramid module, i.e., the MHIP module, is as follows. The data is first input into a convolutional layer for downsampling operation to obtain a feature map F H , and the obtained feature map is input into the first convolutional splitting and reconstruction module, i.e., CSRB, of each layer of the MHIP module. After being processed by the first CSRB of the first layer of the MHIP module, the output feature map will be passed to the next CSRB of this layer for processing. In addition, the feature map output by the first CSRB of the first layer and the feature map output by the first CSRB of the second layer are cascaded, and then the fused feature map The second CSRB is passed to the second layer. The third layer follows this pattern and operates in the same way as the second layer. The above process can be expressed by the following formula:

[0008]

[0009] where F H represents the HSI cube input into the MHIP module, and C 7 , C 5 , C 3 represent the CSRBs of convolutional kernels of different scales respectively. represent the feature maps output by the first CSRB of each layer respectively.

[0010]

[0011] where represents the concatenation operation of the feature map output by the first CSRB of the first layer and the feature map output by the first CSRB of the second layer. represents the concatenation operation of the feature map output by the first CSRB of the second layer and the feature map output by the first CSRB of the third layer.

[0012]

[0013]

[0014] where represent the feature maps output by the second CSRB of each layer respectively, and F spe represents the final output result after concatenating the above three feature maps.

[0015] Preferably, a 3D convolution is used to increase the number of channels of the hyperspectral image, i.e., the HSI cube, before the data enters the CSRB. The output feature map is represented by . FI is divided into P groups along the channel dimension, and each group has K = c / P channels, where i ∈ {1, 2, 3,..., P}. The specific process is as follows: where ConvP3D(·) represents pseudo-3D convolution, Split(·) represents the splitting operation along the channel dimension. Among the P groups of split feature maps,

[0016]

[0017] represents the feature map of the i = 1-st group, and FI is directly connected to the next layer without any further processing. The feature maps of other groups are first fed into a pseudo-3D convolution to extract features, and represents the feature map of the i = 1-st group, and denoted, where j ∈ {1, 2,..., P - 1}. From The extracted feature map is denoted as fj. Subsequently, it is divided into two subgroups along the channel dimension, denoted as fj_1 and fj_2 respectively. Subsequently, fj_2 is concatenated with Fi+1 of the next group and then fed into , and finally, F1 and each fj_1 are concatenated along the channel dimension to form the final output feature map Fout. The above process is expressed by the formula:

[0018]

[0019]

[0020] Preferably: The hyperspectral image, i.e., HSI, and the lidar image, i.e., LiDAR, are subjected to a convolutional channel compression module, i.e., C 3 B for convolutional operation. C 3 B contains a standard convolutional module and a compression module. After being processed by C 3 B, the spectral features, spatial features, and elevation features of LiDAR data are respectively flattened into one-dimensional vectors. Each one-dimensional vector is tokenized to obtain the spectral feature token spatial feature token and LiDAR feature token where N represents the number of tokens,

[0021]

[0022] An additional learnable cls-token needs to be connected before each group of feature tokens Finally, through the method of positional embedding (P), the sequence information of the spectral, spatial, and LiDAR embedding tokens is retained, and the position of each token is encoded to capture the position information in the sequence.

[0023]

[0024] Preferably: The cross-modal collaborative fusion module, i.e., CCFM, is divided into two stages. The first stage is three encoders based on multi-head self-attention, i.e., MHSA, which respectively perform global relationship modeling on the spectral-spatial features of HSI and the elevation features of LiDAR data. The three feature embeddings They are respectively fed into an encoder composed of MHSA, layer normalization (LN) and multi-layer perceptron (MLP), and three learnable weights, namely Wq, Wk and Wv, are predefined. The feature tokens are respectively multiplied by the three learnable weights and then linearly encapsulated into three different weight matrices, namely query vector Q, key vector K and value vector V.

[0025]

[0026] Q and K are used to calculate attention scores, and softmax is used to convert the attention scores into weight probabilities. MHSA is unified as the following formula:

[0027]

[0028] MHSA(Q, K, V) = Cat(SA1, SA2,..., SAh)Wo(19)

[0029] Among them, SA represents the self-attention mechanism, dk represents the dimension of K, h is the number of attention heads, and Wo represents a parameter matrix. The encoder is used to encode the feature sequences of different modalities and is expressed by the formula:

[0030]

[0031] In addition, and are respectively the outputs of the multi-head self-attention encoder. Among them, and respectively represent the new cls-tokens containing land cover classes.

[0032] Preferably, in the second stage of the cross-modal collaborative fusion module, two Transformer encoders based on multi-head cross-attention (MHCA) are used to perform deep feature fusion on the inherent features between HSI and LiDAR data. Specifically, the two outputs in the first stage and are fed into one of the MHCA encoders. Before the feature tokens are input into the MHCA encoder, three weight matrices QSpa, KSpe and VSpe are generated, and the three generated weight matrices are input into the MHCA for cross-modal feature fusion. The process is as follows:

[0033]

[0034] MHCA(QSpa, KSpe, VSpe) = Cat(CA1, CA2,..., CAh)Wz(24)

[0035]

[0036] Among them, CA represents the cross-attention mechanism, represents the output of the MHCA encoder that fuses spectral and spatial features. In addition, the two outputs in the first stage and are fed into another MHCA encoder. Similarly, before inputting the feature tokens into the MHCA encoder, three weight matrices QSpe, KL, and VL are generated. These three weight matrices are input into the MHCA for cross-modal feature fusion again. This process is expressed as:

[0037]

[0038] MHCA(QSpe, KL, VL) = Cat(CA1, CA2,..., CAh)Wu(27)

[0039]

[0040] Among them, represents the output of the MHCA encoder for fusing spectral and elevation features. Finally, the main element addition is used to fuse the outputs of the two MHCA encoders. After reshaping the fused data, it is fed into a 3D convolution to form the final output result.

[0041]

[0042] Preferably: After the data is processed by the MHIP module and CCFM, and the spatial-spectral features and elevation features are cascaded and then input into an adaptive average pooling layer with batch normalization BatchNorm and Mish activation function. Finally, a linear layer is used to form the final classification result, and cross-entropy loss is used as the loss function during network training.

[0043] Compared with the prior art, the beneficial effects of the present invention are:

[0044] A multi-scale splitting and reconstruction collaborative fusion network (MSRCFNet) based on CNN and attention mechanism proposed by the present invention is used for the land cover classification of HSI and LiDAR. This network consists of three parallel multi-scale hierarchical inverted pyramid (MHIP) modules, a cross-modal collaborative fusion module (CCFM), and a classification module. The MHIP module is used to extract multi-scale spatial, spectral, and elevation features of HSI and LiDAR data. The CCFM adaptively fuses the differential features of multi-source remote sensing data by using different attention mechanisms respectively. The classification module maps the semantic features to the final classification results. Technologies such as multi-scale convolution, splitting-reconstruction fusion strategy, attention mechanism, layer normalization, and Mish activation function are used to effectively improve the classification performance of the network. Description of the Drawings

[0045] Figure 1 It is the overall structure diagram of the model of the present invention;

[0046] Figure 2 It is the architecture diagram of the MHIP module and CSRB of the present invention;

[0047] Figure 3 It is the C 3 structure of B;

[0048] Figure 4 It is the structure of the CCFM of the present invention;

[0049] Figure 5 It is the Trento dataset of the embodiment of the present invention: (a) Pseudo-color composite image (bands 31, 14, and 2); (b) LiDAR image; (c) Ground truth map;

[0050] Figure 6 It is the Houston2013 dataset of the embodiment of the present invention: (a) Pseudo-color composite image (bands 64, 30, and 20); (b) LiDAR image; (c) Ground truth map;

[0051] Figure 7 It is the MUUFL dataset of the embodiment of the present invention: (a) Pseudo-color composite image (bands 31, 16, and 6); (b) LiDAR image; (c) Ground truth map;

[0052] Figure 8 It is the full-pixel classification map of the Trento dataset of the embodiment of the present invention. (a) HDDA; (b) MCFN; (c) CoupledCNN; (d) S 2 ENet; (e) HCT; (f) Sal 2 RNet; (g) MS2CANet; (h) Cross-HL; (i) MSRCFNet;

[0053] Figure 9 is the full-pixel classification map of the Houston2013 dataset in the embodiments of the present invention. (a) HDDA; (b) MCFN; (c) CoupledCNN; (d) S 2 ENet; (e) HCT; (f) Sal 2 RNet; (g) MS2CANet; (h) Cross-HL; (i) MSRCFNet;

[0054] Figure 10 is the full-pixel classification map of the MUUFL dataset in the embodiments of the present invention. (a) HDDA; (b) MCFN; (c) CoupledCNN; (d) S 2 ENet; (e) HCT; (f) Sal 2 RNet; (g) MS2CANet; (h) Cross-HL; (i) MSRCFNet;

[0055] Figure 11 are the classification results of different classification methods for each dataset under different sample sizes in the embodiments of the present invention. (a) Trento dataset; (b) Houston2013 dataset; (c) MUUFL dataset. Detailed implementation manners

[0056] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0057] Embodiment

[0058] Please refer to Figure 1 , the method for classifying ground object coverage of hyperspectral and lidar images based on a multi-scale splitting reconstruction collaborative fusion network in the figure. The method is carried out by using a multi-scale splitting reconstruction collaborative fusion network. The multi-scale splitting reconstruction collaborative fusion network includes a plurality of parallel multi-scale hierarchical inverted pyramid modules, a cross-modal collaborative fusion module, and a classification module. Each multi-scale hierarchical inverted pyramid module contains a plurality of convolution splitting reconstruction modules with different scales, and the convolution splitting reconstruction modules are used to extract and fuse discriminative multi-scale features. The cross-modal collaborative fusion module, based on the consistency of the spatial scale, uses the self-attention mechanism to perform global relationship modeling on data of different modalities, and then uses the cross-attention mechanism to fully fuse the complementary information of heterogeneous data. The classification module maps the fused features to the final classification result.

[0059] In this embodiment, the above network includes three MHIP modules for extracting multi-scale features of multi-modal data, a CCFM for fusing heterogeneous data, and a final classification module. MSRCFNet performs multi-scale feature extraction and interactive feature fusion of heterogeneous data on HSI and LiDAR data in a parallel manner.

[0060] Given an HSI, denoted as and a corresponding LiDAR image, denoted as where C represents the number of channels, H and W represent the width and height of the image respectively, and D represents the number of spectral bands in the HSI. To reduce the redundant spectral information in the HSI, PCA is used to perform dimensionality reduction on the original HSI, and the dimensionality-reduced HSI can be denoted as To make full use of the spatial information of the image, the HSI is classified using an adjacent cube of size s×s centered on each pixel, and the Lidar data is classified using an adjacent patch of size s×s centered on each pixel. Each pixel of the HSI is extracted through a three-dimensional module to obtain a small cube Each pixel of the Lidar image is extracted through a two-dimensional module and finally obtains a small patch The index of the central pixel is used to label each patch. For edge pixels, padding with a width of (s - 1) / 2 is used. Suppose there are T labeled training samples, denoted as The corresponding label is where yi ∈ {1, 2,..., C}, and C represents the number of land cover types in the remote sensing data summary. After removing the pixel blocks with label zero, the remaining sample blocks are divided into a training set, a validation set, and a test set.

[0061] Furthermore, in the field of neuroscience, the receptive field refers to the specific area of input data that a neuron can process or respond to. In other words, a smaller receptive field can capture lower-level local texture features, and a larger receptive field can capture more global context information. Based on this, a new MHIP module is designed to capture multi-scale features in HSI and LiDAR data, as Figure 2 shown, the MHIP module internally contains multiple CSRBs based on pseudo-3D convolution of different scales. The CSRBs of different scales can fully extract local texture features and global context features. In addition, compared with traditional 3D convolution, pseudo-3D convolution can extract the spatial and spectral features of the HSI respectively. This method can effectively reduce the computational complexity of the model and improve the training efficiency of the model.

[0062] In the MHIP module for extracting spectral multi-scale features, the two CSRBs in the first layer use a convolutional kernel of size (1×1×7), the two CSRBs in the second layer use a convolutional kernel of size (1×1×5), and the two CSRBs in the third layer use a convolutional kernel of size (1×1×3). The MHIP for extracting spatial multi-scale features and the MHIP for extracting multi-scale features of LiDAR data have the same structure. The sizes of the convolutional kernels used by each layer of CSRB are (7×7×1), (5×5×1), and (3×3×1) respectively. By adjusting the size of the convolutional kernel, each layer of CSRB can adapt to the feature extraction requirements of different scales, thereby enhancing the network's processing ability for HSI and LiDAR data. This refined structure configuration makes the MHIP module more flexible and effective in processing complex data. The data processing processes of the three MHIP modules are the same. Here, the implementation details of the MHIP module for extracting HSI spectral features are specifically introduced. The data is first input into a convolutional layer for downsampling operation to obtain a feature map FH. The obtained feature map is input into the first CSRB of each layer of the MHIP module. The feature map output after being processed by the first CSRB in the first layer of the MHIP module will be passed to the next CSRB in this layer for processing. In addition, the feature map output by the first CSRB in the first layer is concatenated with the feature map

[0063]

[0064] output by the first CSRB in the second layer, and then the fused feature map is passed to the second CSRB in the second layer. The third layer follows this rule and operates in the same way as the second layer. Such a design aims to fully extract and fuse features of different scales, thereby enhancing the complementarity between features of different scales. The formula summary of the above process is:

[0065]

[0066] where, FH represents the HSI cube input into the MHIP module, and C7, C5, C3 represent the CSRBs with convolutional kernels of different scales respectively, represent the feature maps output by the first CSRB in each layer respectively, represents the concatenation operation of the feature map output by the first CSRB in the first layer and the feature map output by the first CSRB in the second layer,

[0067]

[0068] Among them, respectively represent the feature maps output from the second CSRB in each layer, and F spe represents the final output result after cascading the above three feature maps.

[0069] Among them, the details of the CSRB are as Figure 2 shown. Before entering the CSRB, a 3D convolution is used to increase the number of channels of the HSI cube The output feature map is represented by Dividing F I along the channel dimension into P groups, each group has K = c / P channels, where i ∈ {1, 2, 3,..., P}. Equation (5) summarizes the above process:

[0070]

[0071] Among them, Conv P3D (·) represents a pseudo-3D convolution, and Split(·) represents a splitting operation along the channel dimension.

[0072] Among the P groups of split feature maps, represents the feature map of the i = 1-st group, and no further processing is performed on F 1 which is directly connected to the next layer. The feature maps of other groups are first fed into a pseudo-3D convolution to extract features, represented by where j ∈ {1, 2,..., P - 1}, and the feature map extracted by is represented as f j Subsequently, it is divided into two subgroups along the channel dimension, represented by f j_1 and f j_2 respectively. Subsequently, f j_2 is cascaded with the next group of F i+1 and then fed into Finally, F 1 and each f j_1 are cascaded along the channel dimension to form the final output feature map F out The formula for the above process is summarized as:

[0073]

[0074]

[0075] In this embodiment, the core of the CSRB is two splitting and reconstruction operations. In the first splitting operation, the feature map F ISplit into multiple sub - feature maps F along the channel dimension i , the purpose of which is to refine the band information, so as to capture more fine - grained feature information. The second split further refines and decomposes the feature map into f j_1 and f j_2 , so that the network can capture higher - resolution detailed features while retaining the original features. Using the concatenation operation to reconstruct the split feature maps can well express the feature representation ability between different features. The purpose of the first reconstruction is to fuse the feature maps between different groups, so as to improve the network's expression ability for diverse features. The second reconstruction operation fuses all the sub - feature maps obtained through the split operation, integrates the detailed information of all sub - feature maps, and thus restores the global feature representation. For the spatial features of HSI and the elevation features of LiDAR data, the CSRB module focuses on feature extraction and fusion in the spatial dimension. On the basis of retaining the original spatial features, through segmentation and recombination operations, it further deeply excavates the spatial features of HSI and LiDAR data, thereby significantly enhancing the model's expression ability for complex ground object features.

[0076] In this embodiment, HSI contains rich high - dimensional information, and each pixel contains many spectral bands. This high - dimensional characteristic not only increases the complexity of data processing, but also limits the application of many deep learning methods. Convolutional operations can not only provide better visual representations for Transformers, but also enable HSI and LiDAR data to better adapt to subsequent processing. Therefore, a convolutional channel compression module (C 3 B) is used to further enhance the network's representation ability. The detailed structure of C 3 B is as shown in Figure 3 . C 3 B contains a standard convolutional module and a compression module. During the process of tokenization, a large number of channels are usually required for token embedding. If a constant - size convolution is used, it will inevitably generate a large amount of computational cost. Therefore, a compression factor θ is set in the two convolutional layers of the compression module to control the degree of channel compression, thereby reducing the computational cost.

[0077] Before embedding the features output by C 3 B into the subsequent fusion module, a tokenization operation is required, as shown in Figure 1 . After being processed by C 3 , the spectral features, spatial features, and elevation features of LiDAR data are respectively flattened into one - dimensional vectors. After performing the tokenization operation on each one - dimensional vector, spectral feature tokens spatial feature tokens and LiDAR feature tokens Among them, N represents the number of tokens,

[0078]

[0079] An additional learnable cls-token needs to be connected before each group of feature tokens So as to aggregate all feature information through the attention mechanism. Finally, through the method of positional embedding (P), the sequence information of spectral, spatial, and LiDAR embedding tokens is retained, and the position of each token is encoded, thereby effectively capturing the position information in the sequence.

[0080]

[0081] The above input method effectively retains spatial information during the learning process of the network and combines local spectral context information, so as to facilitate the transfer of features to the subsequent fusion module for processing.

[0082] Furthermore, due to the large amount of complementary information among multi-modal remote sensing data, feature fusion of heterogeneous data can effectively utilize the advantages of different modal data, thereby enhancing the network's discriminative ability for ground object features. Therefore, a cross-modal collaborative fusion module (CCFM) is designed to fuse HSI and LiDAR data, as Figure 4 shown, CCFM is divided into two stages. The first stage is three encoders based on multi-head self-attention (MHSA), which respectively perform global relationship modeling on the spectral-spatial features of HSI and the elevation features of LiDAR data. In this stage, the self-attention mechanism (SA) is used to capture the long-range dependence relationships within different modal data, generating richer spectral, spatial, and elevation feature representations, providing more global information for the deep feature fusion in the second stage.

[0083] Three feature embeddings are respectively fed into the encoder composed of MHSA, layer normalization (LN), and multi-layer perceptron (MLP). In order to learn the relationships between feature tokens, three learnable weights are predefined, namely W q , W k and W v . The feature tokens are respectively multiplied by the three learnable weights and then linearly encapsulated into three different weight matrices, namely query vector (Q), key vector (K), and value vector (V).

[0084]

[0085] Q and K are used to calculate attention scores, and softmax is used to convert the attention scores into weight probabilities. MHSA is unified as the following formula:

[0086]

[0087] MHSA(Q, K, V) = Cat(SA 1 , SA 2 ,..., SA h )W o (19)

[0088] where SA represents the self-attention mechanism, d k represents the dimension of K, h is the number of attention heads, and W o represents a parameter matrix. The encoder is used to encode the feature sequences of different modalities and is represented by the formula:

[0089]

[0090] Specifically, and are the outputs of the multi-head self-attention encoder respectively, where and represent the new cls-tokens containing land cover categories respectively.

[0091] In the first stage, three MHSA-based Transformer encoders are designed to capture the long-range dependence relationships inherent in heterogeneous data, thereby enhancing the discriminative ability of the model. However, since HSI and LiDAR data are captured at exactly the same spatial resolution in the same area, the feature fusion of the two types of data can give full play to their complementarity, thereby enhancing the information expression ability of the model. In the second stage, two multi-head cross-attention (MHCA)-based Transformer encoders are used to perform deep feature fusion on the inherent features between HSI and LiDAR data. Specifically, in order to learn the relationship between feature tokens of different modalities, the two outputs and in the first stage are fed into one of the MHCA encoders. Three weight matrices Q Spa , K Spe and V Spe are generated before inputting the feature tokens into the MHCA encoder. The three generated weight matrices are input into the MHCA for cross-modal feature fusion. The following formula summarizes this process:

[0092]

[0093] MHCA(Q Spa , K Spe , V Spe ) = Cat(CA 1 , CA2 ,..., CA h )W z (24)

[0094]

[0095] Among them, CA represents the cross-attention mechanism, represents the output of the MHCA encoder that fuses spectral and spatial features. In addition, the two outputs in the first stage and are fed into another MHCA encoder. Similarly, three weight matrices Q Spe , K L and V L are generated before inputting the feature tokens into the MHCA encoder. These three weight matrices are input into the MHCA for cross-modal feature fusion again. The formula for this process is summarized as:

[0096]

[0097] MHCA(Q Spe , K L , V L ) = Cat(CA 1 , CA 2 ,..., CA h )W u (27)

[0098]

[0099] Among them, represents the output of the MHCA encoder used to fuse spectral and elevation features. Finally, the outputs of the two MHCA encoders are fused using element-wise addition . After reshaping the fused data, it is fed into a 3D convolution to form the final output result.

[0100]

[0101] Among them, the spatial-spectral features and elevation features after being processed by the MHIP module and CCFM are concatenated and then input into an adaptive average pooling layer with batch normalization (BatchNorm) and Mish activation function. Finally, a linear layer is used to form the final classification result. It should be noted that the cross-entropy loss is used as the loss function during network training, and it implicitly contains the probability distribution of the labels, so there is no need to use softmax to obtain the final classification result.

[0102] In this embodiment, the Trento dataset, the Houston2013 dataset, and the MUUFL dataset are used to conduct extensive experiments to evaluate the effectiveness of MSRCFNet. The detailed descriptions of the three datasets are as follows:

[0103] (1) Trento dataset: The Trento dataset is one of the commonly used benchmark datasets for HSI and LiDAR data fusion classification. The HSI of this dataset is acquired by the AISAEagle sensor, and the LiDAR DSM data is acquired by the Optech ALTM3100EA sensor. This dataset is acquired in the rural area south of Trento, Italy. The spatial size of this data is 600×166, and the spatial resolution is 1.0 m. The HSI contains 63 bands, and the spectral range is from 400 to 990 nm. At the same time, the LiDAR data provides high-precision terrain and elevation information, with 6 ground cover classes. In the experiment, 1% of the labeled samples are randomly selected as training samples and validation samples respectively, and the remaining labeled samples are used as test samples. The detailed information is shown in Table I and Figure 5 ;

[0104] (2) Houston2013 dataset: The Houston2013 dataset is captured by the ITRES CASI-1500 imaging sensor on the campus of the University of Houston and its surrounding rural areas in Houston, Texas, USA. This data is mainly composed of two data sources and is publicly available in IEEE GRSS DFC 2013. The HSI and LiDAR data cover an area of 349×1905 pixels, with a spatial resolution of 2.5 m, including 15 categories such as healthy grass, stressed grass, synthetic grass, trees, soil, water, residential, commercial, road, highway, railway, parking lot 1, parking lot 2, tennis court, and runway. The HSI data includes 144 spectral bands, and the spectral resolution is from 0.38 to 1.05 μm. The LiDAR data consists of a single band. In the experiment, 1% of the labeled samples are randomly selected as training samples and validation samples respectively, and the remaining labeled samples are used as test samples. The detailed information is shown in Table I and Figure 6 ;

[0105] (3) MUUFL Dataset: The MUUFL dataset was collected by the ITERS CASI-1500 sensor at the University of Southern Mississippi Gulf Park in Long Beach, Mississippi in November 2010. It contains an HSI dataset and a LiDAR dataset. The spatial dimensions of this dataset are 325×220 pixels, and the spatial resolution is 0.54×1.0 m. The HSI includes 64 available spectral bands in the range of 375 to 1050 nm and contains 11 land cover classes. In the experiment, 1% of the labeled samples were randomly selected as training samples and validation samples respectively, and the remaining labeled samples were used as test samples. The detailed information is shown in Table Ⅰ. Figure 7 。

[0106] Table Ⅰ

[0107] Categories and Number of Samples in Trento, Housotn2013 and MUUFL Datasets

[0108]

[0109]

[0110] Experiments were conducted on eight comparison methods and MSRCFNet respectively on three datasets. To ensure the fairness of the experiment, all methods were experimented with the same parameters on the same dataset. The training samples were randomly selected, and different data were used for the training set, validation set, and test set without overlap. To avoid contingency, the average value was taken after each method was run 5 times. The classification results of the three datasets are shown in Tables Ⅱ - Ⅳ and Figures 8 - 10 , in Tables Ⅱ - Ⅳ, the best classification results are indicated in bold. It can be seen from the classification results that the performance of the method proposed in this paper is the best.

[0111] Table Ⅱ

[0112] Classification Results, Standard Deviation, OA (%), AA (%) and Kappa of Trento Dataset (Best Results are Highlighted in Bold)

[0113]

[0114] Table Ⅲ

[0115] Classification Results, Standard Deviation, OA (%), AA (%) and Kappa of Houston2013 Dataset (Best Results are Highlighted in Bold)

[0116]

[0117]

[0118] Table Ⅳ

[0119] Classification results, standard deviation, OA (%), AA (%), and Kappa of the MUUFL dataset (the best results are highlighted in bold)

[0120]

[0121] The experimental results and analysis are as follows:

[0122] (1) Experimental results and analysis of the Trento dataset: The Trento dataset has sufficient training samples and a small number of land cover classes, so all algorithms achieved high classification accuracies. The specific classification results and full-pixel classification maps are shown in Table II and Figure 8 respectively. The best classification accuracy for each class, as well as OA, AA, and Kappa, are highlighted in bold in the table. Through multiple experiments, we set the hyperparameters on this dataset as follows: PCA was set to 30, the learning rate was set to 9E-4, the patch size was set to 7×7, and the encoder depth was set to 6. These settings enabled the network to exhibit the best classification performance on this dataset. Observing the classification accuracies of each method, for HDDA and MCFN that only used HSI as input, OA reached 96.39% and 97.86% respectively. In contrast, the OA of most multi-source fusion methods that used both HSI and LiDAR data basically reached over 98%. The OA of the proposed MSRCFNet in this paper reached 99.21%, which was the highest among all methods. This indicates that introducing LiDAR data can effectively compensate for the deficiencies of HSI, enabling the network to capture complex land cover features more accurately, thereby reducing the occurrence of misclassifications. However, compared with HDDA and MCFN, fusion classification algorithms such as CoupledCNN, HCT, Sal 2 RNet, MSCANet, and Cross-HL had a decrease in the accuracies of C1 and C3 after introducing LiDAR elevation information. This indicates that the introduction of LiDAR data may affect the network's judgment of the central pixel class. However, the proposed MSRCFNet showed good classification performance in these two classes, indicating that the proposed method can effectively capture and fuse the correlation and complementary information of the two types of data. By observing Figure 8 it can be found that compared with other classification methods, the full-pixel classification map ( Figure 8 (i)) generated by the proposed MSRCFNet can express the detailed information of land cover more accurately, further demonstrating the superiority of the proposed method;

[0123] (2) Experimental results and analysis of the Houston2013 dataset: To further test the classification performance of the proposed method under small samples, experiments were conducted on the Houston2013 dataset using 1% of the training samples. The classification results and the full-pixel classification map are shown in Table III and Figure 9 , through multiple experiments, we set the hyperparameters on this dataset as follows: PCA was set to 30, the learning rate was set to 3E-4, the patch size was set to 7×7, and the encoder depth was set to 6. These settings enabled the network to exhibit the best classification performance on this dataset. Since the Houston2013 dataset has a relatively small sample size and a relatively large number of classes, the classification accuracy of most algorithms is relatively low. Nevertheless, the proposed MSRCFNet still achieved satisfactory classification results, with OA, AA, and Kappa reaching 90.49%, 92.05%, and 89.72% respectively. The OA of HDDA and MCFN using HSI as input reached 87.12% and 87.37% respectively. Compared with these two methods, the OA of the proposed MSRCFNet increased by 3.37% and 3.12% respectively. This indicates that the proposed MSRCFNet can effectively fuse HSI and LiDAR data, enhancing the network's ability to identify detailed features and complex terrains. Among the fusion classification networks using HSI and LiDAR data as input, S 2 ENet and HCT achieved relatively good classification results, with OA reaching 89.15% and 89.13% respectively. This may be because S 2 ENet and HCT can effectively fuse the key information of the two types of data after extracting the features of HSI and LiDAR data. However, for C5, C9, and C15, S 2 the classification accuracy of ENet and HCT is lower than that of the proposed MSRCFNet. This may be because the proposed MSRCFNet uses a multi-scale method to extract features of different scales, thus improving the model's classification ability for complex land cover classes. From Figure 6 (a), it can be found that due to the influence of clouds in some areas, this will cause quite a large interference to the classification accuracy. Nevertheless, by observing Figure 9 , compared with other methods, in the full-pixel classification map generated by the proposed MSRCFNet, the boundaries between different classes are clearer and smoother. This indicates that the proposed MSRCFNet can still better capture the information between different classes in complex scenarios, further verifying the effectiveness and generalization of the method;

[0124] (3) Experimental results and analysis of the MUUFL dataset: Table IV and Figure 10The classification results and full-pixel classification maps of different classification methods on the MUUFL dataset are shown respectively. Through multiple experiments, we set the hyperparameters on this dataset as follows: PCA is set to 30, the learning rate is set to 3E-4, the patch size is set to 7×7, and the encoder depth is set to 8. These settings enable the network to exhibit the best classification performance on this dataset. Due to the complex terrain and uneven distribution of labeled samples in the MUUFL dataset, this poses a severe challenge to classification methods. Nevertheless, the proposed method still obtains the best classification results among all classification methods, with OA, AA, and Kappa reaching 89.76%, 79.88%, and 86.41% respectively. By observing the classification accuracy of each method, the OA of some multi-source fusion methods is lower than that of the method using only HSI as input. This may be because the scattered distribution of ground cover and the similarity of elevation information result in LiDAR data being unable to provide additional discriminative information. However, MSRCFNet alleviates this impact through an effective feature fusion strategy, enabling LiDAR data to be fully complementary to HSI, thus exceeding other classification methods in terms of overall classification performance. In addition, due to the insufficient number of labeled samples for C9 and C10 and their adjacency in spatial location, most methods have difficulty classifying C9 and C10. However, the proposed MSRCFNet can still exhibit good classification performance for C9 and C10. This may be because multi-scale feature extraction and cross-modal feature fusion effectively integrate global and local features, achieving relatively good classification results even when the number of labeled samples is small. By observing the ground truth map and the full-pixel classification maps of all methods, the proposed MSRCFNet generates a clearer land cover classification map. The consistency between the classification results and the full-pixel classification maps further demonstrates the superiority of the proposed MSRCFNet in multi-source data fusion tasks.

[0125] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.

[0126] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A hyperspectral and lidar image feature coverage classification method based on a multi-scale split-reconstruction collaborative fusion network, wherein the method is performed using a multi-scale split-reconstruction collaborative fusion network and is characterized by: The multi-scale split-reconstruction collaborative fusion network includes multiple parallel multi-scale hierarchical inverted pyramid modules, a cross-modal collaborative fusion module and a classification module. Each of the multi-scale hierarchical inverted pyramid modules contains multiple convolutional split reconstruction modules of different scales, and the convolutional split reconstruction module is used to extract and fuse discriminative multi-scale features. The cross-modal collaborative fusion module uses a self-attention mechanism to model the global relationship of data of different modalities based on the consistency of spatial scale, and then uses a cross-attention mechanism to fully fuse the complementary information of heterogeneous data. The classification module maps the fused features to the final classification result.

2. The hyperspectral and laser radar image coverage classification method based on a multi-scale split-reconstruction collaborative fusion network according to claim 1 is characterized by: The data extraction process of the multi-scale hierarchical inverted pyramid module, i.e., the MHIP module, is as follows: First, it is input into a convolutional layer for downsampling to obtain the feature map F H The obtained feature map is input into the first convolution split reconstruction module of each layer of the MHIP module, namely CSRB. The feature map output after being processed by the first CSRB of the first layer of the MHIP module is It will be passed to the next CSRB in this layer for processing. In addition, the feature map output by the first CSRB in the first layer The feature map output by the first CSRB in the second layer Perform a cascade operation and then merge the feature maps The second CSRB passed to the second layer, the third layer follows this rule, consistent with the second layer operation, the above process formula is expressed as: Among them, F H represents the HSI cube input to the MHIP module, C7, C5, and C3 represent the CSRB of convolution kernels of different scales, Respectively represent the feature maps output by the first CSRB of each layer in, It means that the feature map output by the first CSRB in the first layer is cascaded with the feature map output by the first CSRB in the second layer. It means that the feature map output by the first CSRB in the second layer is cascaded with the feature map output by the first CSRB in the third layer. in, Respectively represent the feature maps output by the second CSRB of each layer, F spe It represents the final output result after cascading the above three feature maps.

3. The hyperspectral and laser radar image coverage classification method based on a multi-scale split-reconstruction collaborative fusion network according to claim 2 is characterized by: The data is augmented with a 3D convolution to create a hyperspectral image, or HSI cube, before entering the CSRB. The number of channels, the output feature map is Indicates that F I Divided into P groups along the channel dimension, each group With K = c / P channels, where i∈(1, 2, 3, ..., P), the process is as follows: Among them, Conv P3D (·) represents pseudo 3D convolution, Split(·) represents splitting operation along the channel dimension, and in the split P group feature map, represents the feature map of the i=1-st group, and no longer I Any processing is directly connected to the next layer. The feature maps of other groups are first fed into a pseudo-3D convolution to extract features. Represented by , where j∈{1, 2, ..., P-1}. The extracted feature map is represented as f j , then, it is divided into two subgroups along the channel dimension, respectively represented by f j_1 and f j_2 Subsequently, f j_2 With the next group of F i+1 Cascade operation and then feed to Finally, F1 and each f j_1 Cascade operations are performed along the channel dimension to form the final output feature map F out , the above process formula is expressed as:

4. The hyperspectral and laser radar image coverage classification method based on a multi-scale split-reconstruction collaborative fusion network according to claim 3 is characterized by: The hyperspectral image (HSI) and the laser radar image (LiDAR) are compressed using a convolutional channel compression module (C 3 B performs convolution operation, C 3 B contains a standard convolution module and a compression module, and C 3 The spectral features, spatial features and elevation features of LiDAR data after processing by B are flattened into one-dimensional vectors respectively, and each one-dimensional vector is tokenized to obtain the spectral feature tokens respectively. Spatial feature token and LiDAR feature tokens Where N represents the number of tokens. An additional learnable cls-token needs to be connected before each set of feature tokens Finally, through the position embedding (P) method, the sequence information of the spectral, spatial and LiDAR embedded markers is retained, and the position of each marker is encoded to capture the position information in the sequence.

5. The hyperspectral and laser radar image coverage classification method based on a multi-scale split-reconstruction collaborative fusion network according to claim 4 is characterized by: The cross-modal collaborative fusion module (CCFM) is divided into two stages. The first stage is based on three encoders of multi-head self-attention (MHSA), which respectively model the global relationship between the spectral-spatial features of HSI and the elevation features of LiDAR data. They are fed into the encoder composed of MHSA, layer normalization (LN) and multi-layer perceptron (MLP), and three learnable weights are pre-defined, namely W q , W k and W v , the feature labels are multiplied by three learnable weights respectively, and then linearly encapsulated into three different weight matrices, namely the query vector Q, the key vector K and the value vector V, Q and K are used to calculate the attention score, and softmax is used to convert the attention score into a weighted probability. MHSA is unified into the following formula: Among them, SA represents the self-attention mechanism, d k represents the dimension of K, h is the number of attention heads, and W o Represents a parameter matrix. The encoder is used to encode the features of different modal sequences, which is expressed by the formula: also, and are the outputs of the multi-head self-attention encoder, where and Respectively represent new cls-tokens containing land cover categories.

6. The hyperspectral and laser radar image coverage classification method based on a multi-scale split-reconstruction collaborative fusion network according to claim 5 is characterized by: In the second stage of the cross-modal collaborative fusion module, two Transformer encoders based on multi-head cross attention (MHCA) are used to perform deep feature fusion on the inherent features between HSI and LiDAR data. Specifically, the two outputs in the first stage are and Feed it into one of the MHCA encoders, and generate three weight matrices Q before inputting the feature token into the MHCA encoder Spa , K Spe and V Spe , the three generated weight matrices are input into MHCA for cross-modal feature fusion. The specific process is as follows: MHCA(Q Spa ,K Spe ,V Spe )=Cat(CA1,CA2,...,CA h )W z (24) Among them, CA represents the cross attention mechanism, represents the output of the MHCA encoder that fuses spectral and spatial features. In addition, the two outputs in the first stage are and Feed it into another MHCA encoder. Similarly, three weight matrices Q are generated before the feature token is input into the MHCA encoder. Spe , K L and V L , these three weight matrices are input into MHCA for cross-modal feature fusion again. The process is expressed as: MHCA(Q Spe ,K L ,V L )=Cat(CA1,CA2,...,CA h )W u (27) in, Represents the output of the MHCA encoder used to fuse spectral and elevation features. Finally, principal element addition is used The outputs of the two MHCA encoders are fused, and the fused data is reshaped and fed into a 3D convolution to form the final output result.

7. The hyperspectral and laser radar image coverage classification method based on a multi-scale split-reconstruction collaborative fusion network according to claim 6 is characterized by: After the data is processed by the MHIP module and CCFM, the spatial-spectral features and elevation features are cascaded and input into the adaptive flat pooling layer with batch normalization BatchNorm and Mish activation function. Finally, a linear layer is used to form the final classification result, and the cross entropy loss is used as the loss function during network training.

Citation Information

Cited By

  • Cross-modal guided hyperspectral image classification framework and method under multiple degradation conditions

    CN120953709A

  • Joint classification method for hyperspectral image and laser radar data

    CN121305244A

  • A method for joint classification of hyperspectral images and lidar data

    CN121305244B