Hyperspectral and laser radar image classification method based on pyramid progressive cross fusion network
By using a pyramid-gradual cross-fusion network in multi-source remote sensing image classification, multi-scale features are extracted and fused, the problems of feature loss and insufficient fusion in the prior art are solved, and image classification effect with high precision and strong generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510060655.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
The existing multi-source remote sensing image classification method based on deep learning is difficult to effectively utilize the multi-scale information in the image, resulting in feature loss and network convergence difficulties. At the same time, a single simple fusion method cannot fully integrate multi-source features.
Using a method based on a pyramid progressive cross-fusion network, multi-scale features are extracted from hyperspectral and lidar data through a pyramid expansion compression module, and a full fusion of multi-source features is achieved through an asymmetry cross-fusion module, and finally image classification is used using an attention enhancement classifier.
Effectively utilizing multi-scale features and differentiated features improves the accuracy and generalization ability of image classification, avoids the shortcomings of a single fusion method, and significantly improves the performance of multi-source remote sensing image classification.
Smart Images

Figure CN119992360A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image classification, and in particular to a hyperspectral and laser radar image classification method based on a pyramid progressive cross-fusion network. Background Art
[0002] With the increasing availability of satellite sensors and the advancement of imaging technology, it has become easier to obtain multi-source remote sensing images. Among all types of remote sensing images, hyperspectral (HS) data has been widely used in many fields, such as precision agriculture, mineral exploration, anomaly detection, etc., thanks to its rich spectral information. Image classification has become a key step in these applications because it can assign unique class labels to different land cover types. However, HS data often face the challenges of different objects with the same spectrum and different spectra for the same object. Light Detection and Ranging (LiDAR) data provides rich elevation and three-dimensional spatial information, which can effectively supplement the limitations of HS data in terms of spectral similarity and variability. However, LiDAR data has limitations in capturing spectral information of different land covers. Therefore, by integrating the rich spectral information of HS data and the detailed spatial and elevation information of LiDAR, the advantages of HS and LiDAR data can be complemented, which is a promising method to improve the accuracy of land cover classification.
[0003] Taking advantage of the powerful representation learning ability of deep neural networks (DNNs), many DNN-based methods have been proposed for the joint classification of HS and LiDAR data. Based on the network architecture, these methods can be roughly divided into three categories: convolutional neural network (CNN)-based, Transformer-based, and hybrid CNN-Transformer-based methods. At the same time, due to the scale differences of object types in remote sensing images, single-scale feature extraction is often insufficient to fully capture the spatial and spectral information of various targets. In order to address this limitation, a series of multi-scale feature extraction methods have been proposed to enhance the generalization ability of the model and thus improve the classification accuracy. In addition to efficient feature extraction, the effective use of the complementary and differential features of HS and LiDAR data to achieve efficient fusion is also the key to accurate classification of multi-source remote sensing images. At present, deep learning-based fusion methods include pixel-level methods, feature-level methods, and decision-level methods. Among them, feature-level fusion methods are often used due to their excellent fusion capabilities.
[0004] However, deep learning-based classification methods still have many limitations and challenges in multi-source remote sensing classification. On the one hand, the complexity and diversity of classification objects place demands on the network's multi-scale feature extraction capabilities. Existing methods usually enhance the network's feature extraction capabilities by deepening or widening the network. However, they still cannot effectively utilize the multi-scale information in the image, resulting in feature loss and challenges to network convergence. On the other hand, effective fusion methods are crucial for the final joint classification. Existing methods usually use a single simple fusion method that cannot effectively integrate multi-source features. There are also some cross-attention-based fusion methods that only consider the correlation between features when calculating the attention score, but ignore the use of difference features. Therefore, this application proposes a hyperspectral and lidar image classification method based on a pyramid progressive cross-fusion network. Summary of the invention
[0005] The purpose of the present invention is to provide a hyperspectral and lidar image classification method based on a pyramid progressive cross-fusion network to solve the problems raised in the above background technology.
[0006] To achieve the above object, the present invention provides the following technical solution: a hyperspectral and laser radar image classification method based on a pyramid progressive cross fusion network, the method is implemented using a pyramid progressive cross fusion network, and comprises the following steps:
[0007] Step S1, using a pyramid expansion and compression module of a pyramid progressive cross-fusion network to extract spatial spectral features from HS data and elevation features from LiDAR data, and using an expansion and compression method to remove redundant features and compress multi-scale feature representations;
[0008] Step S2, using a progressive cross fusion module of a pyramid progressive cross fusion network to fuse multi-scale features;
[0009] Step S3, using the attention enhancement classifier of the pyramid progressive cross fusion network to classify the image.
[0010] Preferably: the pyramid expansion and compression module of step S1 has a dual-branch structure, and the pyramid expansion and compression module includes a pyramid spectral feature extraction block composed of a multi-scale 3D convolution block, a 3D batch normalization layer and a ReLU activation layer, and a pyramid spatial feature extraction block composed of a multi-scale 2D convolution block, a 2D batch normalization layer and a ReLU activation layer.
[0011] Preferably: the pyramid spectrum feature extraction block and the pyramid space feature extraction block are respectively expressed as:
[0012] F pse (X) = R(BN(cat(Conv3D 3×1×1(X), Conv3D 5×1×1 (X), Conv3D 7×1×1 (X))))
[0013] F psa (X) = R(BN(cat(Conv2D 3×3 (X), Conv2D 5×5 (X), Conv2D 7×7 (X))))
[0014] Where Conv3D() and Conv2D() represent 3D convolutional layers and 2D convolutional layers with different kernel sizes, cat() represents a cascade operation, BN() represents batch normalization, and R() represents ReLU.
[0015] F pse () and F psa () respectively represent the operation process of pyramid spectral feature extraction block and pyramid spatial feature extraction block.
[0016] Preferably: the progressive cross fusion module of step S2 includes a difference feature fusion block and a pair of common feature fusion blocks.
[0017] Preferably: the specific process of the difference feature fusion block is to calculate the similarity matrix between Q and K, and then multiply it by V to infer the common feature CFM1 between Q and V, which can be expressed as:
[0018]
[0019] Where Dropout() represents the Dropout layer, SoftmQx() is the SoftMax activation function, and d k represents the scaling factor, Q1, K1, and V1 represent the query, key, and value obtained by linear transformation of different modes, K1 T Represents the transposed matrix of K1, V1 is subtracted from CFM1, and the difference feature DFM1 between Q and V is obtained. This process can be expressed as:
[0020] DFM1=V1-CFM1
[0021] In order to, the difference features are injected into Q to obtain complementary information from HS data, which is expressed as:
[0022]
[0023] Where Linear(·) represents a linear mapping operation. The feature map passes through the normalization layer and MLP in sequence, and finally performs a residual connection to obtain the output F1 of the DFF block.
[0024] Preferably: the common feature fusion block is composed of an interactive attention layer, an LN and an MLP, and the specific steps are: first, injecting the common features of the lidar data to enrich the fusion features, specifically, using the fusion feature F1 containing the difference between the HS and LiDAR data as Q2, and using the LiDAR features as K2 and V2, and further fusing the LiDAR features as:
[0025]
[0026] F2=MLP(LN(CFM2))+CFM2(19)
[0027] Where F2 represents the output of the first CFF block, and then F2 is further fused with the HS feature. The process is the same as the above equations (18) and (19).
[0028] Preferably: the specific process of the attention enhancement classifier in step S3 is that before the fused features are input into the MLP head composed of the LN layer and the FC layer for final classification, the attention weight A is calculated by introducing the attention mechanism to compress the feature map, and then multiplied by the fused feature F as an incentive. The process can be expressed as
[0029] A=Softmax(MLP(F))
[0030] F′=A⊙F
[0031] Finally, the following multi-source training samples are used to optimize the network parameters based on multi-class cross entropy loss.
[0032]
[0033] where |Ω train | represents the cardinality of the training index set, represents the output unnormalized score for the i-th pixel of each class.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] The present invention proposes a PyionNet network model for joint classification of HS and LiDAR. Firstly, the PEC module is introduced to extract multi-scale features from HS and LiDAR data, and the extension compression method is used to aggregate and extract features of various scales, eliminate redundant information, and enhance the network feature representation capability. Then, the PCF module is used to achieve full fusion of multi-source features. In this module, the multi-source features are input into a DFF block and a pair of CFF blocks to ensure that the multi-source features are effectively fused, which not only avoids the insufficient fusion of a single simple fusion, but also improves the utilization rate of complementary and differential features between different modal data, showing excellent performance and strong generalization ability compared with the existing methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 This is the overall architecture diagram of the network model of the present invention;
[0037] Figure 2 It is a structural diagram of the DFF block and the CFF block of the present invention;
[0038] Figure 3 is a schematic diagram of the classifier of the present invention;
[0039] Figure 4 It is a classification diagram of different classification methods in the embodiments of the present invention on the Trento data set;
[0040] Figure 5 It is a classification diagram of different classification methods of the embodiments of the present invention on the Houston2013 data set;
[0041] Figure 6 It is a classification diagram of different classification methods of the embodiments of the present invention on the MUUFL dataset. DETAILED DESCRIPTION
[0042] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0043] Example
[0044] See also Figure 1 , a hyperspectral and lidar image classification method based on a pyramid progressive cross fusion network is shown in the figure. The method is implemented using a pyramid progressive cross fusion network and includes the following steps:
[0045] Step S1, using a pyramid expansion and compression module of a pyramid progressive cross-fusion network to extract spatial spectral features from HS data and elevation features from LiDAR data, and using an expansion and compression method to remove redundant features and compress multi-scale feature representations;
[0046] Step S2, using a progressive cross fusion module of a pyramid progressive cross fusion network to fuse multi-scale features;
[0047] Step S3, using the attention enhancement classifier of the pyramid progressive cross fusion network to classify the image.
[0048] Furthermore, in order to better extract multi-scale features and achieve comprehensive feature representation, a pyramid expansion and compression module (PEC) module is proposed. Specifically, 3D convolution blocks and 2D convolution blocks are used to extract spatial and spectral features respectively, and then convolution layers with different convolution kernel sizes are used to process the input data in parallel for multi-scale feature extraction. The expansion-compression method is used to eliminate redundant features and enhance feature representation capabilities.
[0049] The PEC module has a dual-branch structure, which extracts features from HS and LiDAR data by performing multiple expansion and compression operations on feature maps. The pyramid spectral feature extraction (PSeFE) block and the pyramid spatial feature extraction (PSaFE) block are used to increase the feature richness when performing spectral expansion operations. Specifically, the PSeFE block consists of a multi-scale 3D convolution block, a 3D batch normalization (BN) layer, and a ReLU activation layer. In the multi-scale 3D convolution block, the size of the convolution kernel in the spatial dimension is set to 1×1, and the size of the spectral dimension is set to b∈{3,5,7} in turn. After parallel multi-scale feature extraction, three feature maps with the same number of channels are obtained, and then the spectral dimension is compressed. The cascade operation is performed to achieve channel expansion, and the expansion rate is set to λ1, which is a hyperparameter. Finally, the obtained feature map passes through the BN layer and the activation layer to improve the training speed, generalization ability and model expression ability of the network. Similarly, the PSaFE block includes a multi-scale 2D convolution block, a 2DBN and a ReLU activation layer. The sizes of the three convolution kernels of the multi-scale 2D convolution block are set to (3×3, 5×5, 7×7) according to experience. The expansion rate of the PSeFE block is set to λ2. The PSeFE block and the PSeFE block can be expressed as, respectively.
[0050] F pse (X) = R(BN(cat(Conv3D 3×1×1 (X), Conv3D 5×1×1 (X), Conv3D 7×1×1 (X))))
[0051] F psa (X) = R(BN(cat(Conv2D 3×3 (X), Conv2D 5×5 (X), Conv2D 7×7 (X))))
[0052] Among them, Conv3D() and Conv2D() represent 3D convolutional layers and 2D convolutional layers with different kernel sizes, cat() represents a cascade operation, BN() represents batch normalization, R() represents ReLU, and F pse () and F psa () represent the operation process of PSeFE block and PSeFE block respectively.
[0053] Furthermore, given a set of X H ∈R 1×B×H×W HS data represented by H L ∈R H×W Represents the corresponding area LiDAR data, H and W represent the height and width of the original image, B represents the number of channels of HS data, and each pixel in the original image is used as the center pixel to generate a data block of size N×N. Then, the HS data block and LiDAR data blocks At the same time, the PEC module is input to extract multi-scale features. For the HS data block First, a 3D convolution layer with a convolution kernel of 16@11×3×3 is used, where the stride is set to (3, 1, 1), and in order to keep the data aligned, the padding is set to (5, 1, 1). The size of the output data is 16×(B / 3)×N×N.
[0054] Next, the spectral feature map is obtained through the PSeFE block The expansion rate λ1 = 1.5. Then, a 3D convolution layer with a convolution kernel of 16@3×3×3 is used to compress the feature map. The output data size is 16×(B / 3)×N×N. In order to facilitate the subsequent 2D convolution operation, the feature map is obtained by reshape operation. Then, As the input of the PSaFE block, spatial feature extraction and channel expansion are performed. The output data size of the PSaFE block is λ2C×N×N, where λ2 represents the expansion rate and C represents the number of channels in the final feature map. In the present invention, C is set to 64 and λ2 is set to 1.5. Finally, the features are compressed through a 2D convolution layer with a convolution kernel of C@1×1, and the final HS spatial spectrum feature X′ is obtained through a maximum pooling layer. H ∈R C×P×P , where P = (N+1) / 2. The maximum pooling layer is used to compress the spatial information, reduce the computational complexity of subsequent operations, and enhance important spatial information. All convolutional layers are followed by BN layers and ReLU activation layers to improve the network's training speed, generalization ability, and model expression ability. The process can be described by the following formula:
[0055]
[0056] where F M () indicates the maximum pooling layer.
[0057] Similarly, LiDAR data blocks The spatial and elevation features are extracted by two PSaFE blocks in sequence. The two PSaFE blocks can also be used to gradually expand the feature map, where the output feature size of the first FSaEF block is set to (λ2C / 2)×N×N. Next, in order to further extract the elevation features and align the feature maps, the elevation features are As the input data of the convolution layer, the convolution kernel is C@1×1, and then the final elevation feature X′ is obtained through the maximum pooling layer L ∈R C×P×P , the process can be described as follows:
[0058]
[0059] Among them, considering the correlation and complementarity between HS and LiDAR data, a feature-level fusion method PCF module is adopted after extracting features respectively. The PCF module consists of a difference feature fusion (DFF) block and a pair of common feature fusion (CFF) blocks. The PCF module can give full play to the complementary advantages of the two, realize feature enhancement, and obtain comprehensive features that are more conducive to accurate classification.
[0060] Traditional single fusion may not be able to fully realize the complementarity between multimodal features. A progressive fusion transformer strategy is proposed to ensure a more complete fusion of different modal features through multiple iterations. Traditional Transformer cross-fusion methods usually only consider the correlation between features, but fail to effectively utilize the complementarity and difference information between features, which is not conducive to the completion of the final classification task. In order to solve this problem, a new block DFF block is proposed, which considers the common information and difference information of the fused classification.
[0061] Similar to the feature block embedding operation in ViT, the two feature maps are first converted to P with length C. 2 vectors, each vector represents a pixel in the feature map. Next, the vectors are segmented according to the traditional ViT rules. The process is shown in the following formula:
[0062]
[0063] where i∈{H,L},X i Denotes HS feature X′ H or LiDAR feature X′ L , [;] represents the cascade operation along the first dimension, represents a learnable position embedding, which adds position information to help the model better understand and process the positional relationship between elements in the data. DP(·) represents the Dropout layer, which prevents the model from overfitting and improves the generalization ability of the model. The Z obtained after tokenization is H and Z lAs the input data of the DFF block, cross-attention fusion is performed. By cross-fusing the common features and difference features between different modalities multiple times, the complementary features between the modalities are fully utilized to achieve full fusion, and finally the global dependency feature F3 is obtained. They can be expressed as:
[0064] F1=DFFB(E Q (Z L ), E K (Z H ), E V (Z H ))
[0065] F2=CFFB(E Q (F1), E K (Z L ), E V (Z L ))
[0066] F3=CFFB(E Q (Z H ), E K (F2), E V (F2)
[0067] Where E i (·) indicates that the input data is linearly transformed according to different weight matrices to obtain the three key elements of Transformer: query Q, key K and value V. DFFB(·) indicates a series of DFF block operations, and CFFB(·) indicates a series of CFF block operations. Before classification, linear operations are first performed on the spatial spectral features Z_H of the original HS data and the fusion features F3 extracted by three cross-fusions. Next, the final fusion is completed by element-by-element addition operations. Finally, the final fusion feature F is obtained, and the final classification is performed based on the fusion feature. The process can be expressed by the following formula:
[0068] F=Linear(Z H )+Linear(F3).
[0069] Furthermore, in order to capture and use the difference information in the fusion process and extract the common information at the same time, a new structure DFF block is proposed, such as Figure 2 As shown, in order to explore the common features of HS and LiDAR data, the similarity matrix between Q and K is calculated and then multiplied by V to infer the common feature CFM1 between Q and V, which can be expressed as:
[0070]
[0071] Where Dropout() represents the Dropout layer, Softmax() is the SoftMax activation function, and d k represents the scaling factor, Q1, K1, and V1 represent the query, key, and value obtained by linear transformation of different modes, K1 T Represents the transposed matrix of K1. Then, V1 is subtracted from CFM1 to obtain the difference feature DFM1 between Q and V. This process is expressed as:
[0072] DFM1=V1-CFM1
[0073] In order to obtain complementary information from HS data, differential features are injected into Q, and they are expressed as:
[0074]
[0075] Linear(·) represents a linear mapping operation. Similar to the traditional Transformer, the feature map passes through the normalization layer and MLP in sequence, and finally performs a residual connection to obtain the output F1 of the DFF block.
[0076] After the DFF block, data F1 is further extracted from the HS and LiDAR data through the CFF block. Then, the common features are gradually fused to enrich the fused features. The structure of the CFF block is as follows: Figure 2 As shown in Figure 1, it consists of an interactive attention layer, LN, and MLP. First, the common features of the LiDAR data are injected to enrich the fusion features. Specifically, the fusion feature F1 containing the difference between HS and LiDAR data is used as Q2, and the LiDAR features are used as K2 and V2. The process of further fusing the LiDAR features is expressed as:
[0077]
[0078] F2=MLP(LN(CFM2))+CFM2(19)
[0079] Where F2 represents the output of the first CFF block, which is further combined with the LiDAR features. Next, F2 is further fused with the HS features to enrich the fused features. The process is the same as Equations 18 and 19.
[0080] Attention-based classifiers are used at the end of the network, which aims to improve classification accuracy by selectively focusing on the most relevant parts of the input data, such as Figure 3 As shown, before the fused features are input into the MLP head composed of LN layer and FC layer for final classification, the attention weight A is calculated by introducing the attention mechanism to compress the feature map, and then multiplied with the fused feature F as an incentive. The process can be as follows:
[0081] A=Softmax(MLP(F))
[0082] F′=A⊙F
[0083] Finally, the parameters of the network of the present invention are optimized using the following multi-source training samples based on multi-class cross entropy loss:
[0084]
[0085] Where |Ω train | represents the cardinality of the training index set, represents the output unnormalized score for the i-th pixel of each class.
[0086] In order to verify the superiority of the network model proposed in this paper, the following datasets are used for experiments: (1) Trento dataset: The Trento dataset consists of HS and LiDAR data collected from southern Trento, Italy. The HS data is acquired using the Eagle sensor, an airborne hyperspectral imaging system, which captures 63 spectral bands from 0.42 to 0.99 μm. The LiDAR data is collected using the Optech Airborne Laser Topography (ALTM) 3100EA sensor. The scene consists of 166 × 600 pixels with a spatial resolution of 1 meter. The dataset includes six Land cover type, including a total of 30,214 real samples. In the experiment, training samples and validation samples were randomly selected from the labeled samples at a ratio of 1%, and the rest were used as test samples; (2) Houton2013 dataset: The Houston2013 dataset includes HS and LiDAR data. This dataset is provided by the IEEE GRSS data fusion competition. The dataset was collected by the National Airborne Laser Mapping Center in June 2012 using the ITRESCASI-1500 imaging sensor over the campus of the University of Houston. The HS data consists of 144 spectral bands, covering 0.38 The HS and LiDAR data are provided in the wavelength range of 349 to 1.05 μm, while the LiDAR data are provided as a single band. The size of both HS and LiDAR data is 349 × 1905 pixels, with a spatial resolution of 2.5 m. The dataset includes 15 categories and a total of 15,029 real samples. The training samples and validation samples are randomly selected from the labeled samples at a ratio of 1%, and the rest are used as test samples. (3) MUUFL dataset: The MUUFL dataset includes HS and LiDAR data. These data were collected in November 2010 at the Gulf Park campus area of the University of Southern Mississippi in Long Beach, Mississippi, USA. The S data were collected using the ITRES Research Limited (ITRES) compact airborne spectral imager (CASI-1500) sensor, which includes 64 available bands in the range of 375 to 1050 nm. The spatial size of the MUUFL dataset is 325×220 pixels and the spatial resolution is 0.54×1.0 meters. The dataset includes 11 categories and a total of 53,687 real samples. The training samples and validation samples are randomly selected from the labeled samples at a ratio of 1%, and the rest are used as test samples. The details of the sample size of each category are shown in Table 1.
[0087] Table 1: Number, name and quantity of samples on three datasets
[0088]
[0089]
[0090] To verify the effectiveness of the proposed PyionNet, it is compared with eight representative methods on three datasets, including SVM, HResNet, DBCTNet, CCNN, ExViT, HCTNet, MS2CANet and MHST. SVM, HResNet and DBCTNet are representatives of classic machine learning methods, traditional deep CNN methods and methods combining CNN and Transformer for HS classification tasks, respectively. Among the remaining methods, CCNN and MS2CANet are based on deep CNN architectures, while ExViT, HCTNet and MHST use CNN-Transformer architectures. To ensure the fairness of the experiments, the hyperparameters in all methods are consistent with the original parameters in the paper.
[0091] In the experiment, the AdamW optimizer with a decay rate of (0.9, 0.999) is used to optimize the network, the epoch is set to 200, and the batchsize is configured to 128. Considering the different data scales and spatial resolutions of the three datasets, the initial learning rate and patch size are adjusted separately for each dataset. For the Trento dataset, the initial learning rate and patch size are set to 5e-4 and 7x7, respectively. For the Houston2013 dataset, they are set to 1e-3 and 11x11, respectively. For the MUUFL dataset, they are set to 5e-4 and 11x11, respectively.
[0092] Overall accuracy (OA), average accuracy (AA) and Kappa coefficient (Kappa) are used as quantitative measures of the performance of each method. Considering that the random instability of each experiment will undermine the fairness of the experiment, all experiments are repeated five times independently, and the final result is the average of the five experimental results. All experiments are performed using Intel Core i7-7700HQ CPU and NVIDIA GeForce RTX 2080Ti GPU. The software environment is CUDA version 11.6, PyTorch 2.1.1 and python 3.11.5.
[0093] Table 2: Classification accuracy of each classification method on the Trento dataset OA, AA, Kappa (k)
[0094]
[0095] The classification performance of each method is evaluated on the Trento, Houston2013 and MUUFL datasets, and the results are shown in Tables 2, 3 and 4. The highest OA, AA, KAPPA and the best classification results of each category are highlighted in bold. The experimental results show that the proposed PyionNet outperforms other methods in OA, AA and KAPPA. On the Trento dataset, the methods that only rely on HS classification tend to have lower classification results compared with the methods that combine HS and LiDAR data classification. Specifically, the OA of SVM, HResNet and DBCTNet designed for HS classification are 93.67%, 98.67% and 97.60%, respectively, which are lower than the OA of the proposed PyionNet and other networks for HS and LiDAR classification. This not only demonstrates the beneficial effect of integrating LiDAR data to supplement elevation information, but also emphasizes the necessity of extracting dedicated features from LiDAR data. On the Houston2013 dataset, the OA of SVM is 93.67%, which is 5.39% lower than that of the method in this paper. This is because SVM is a pixel classification method that only uses spectral information for classification. On the ston2013 dataset, SVM has the lowest classification accuracy for "C13", which shows that "C13" relies heavily on spatial information for accurate classification. Thanks to the powerful feature representation ability of CNN, the classification method based on CNN can improve the classification accuracy of "C13" to more than 85%. Among them, multi-scale feature extraction technology is usually better than single-scale feature extraction method. For example, MS2CANet using multi-scale pyramid convolution has the highest OA of 97.18% for this category. The proposed PyionNet also uses multi-scale feature extraction and achieved 95.2% for this category. 3% better classification effect, this is because the multi-scale feature extraction method provides a more comprehensive feature representation, which can improve the model's ability to identify objects of different sizes, and also better promote the exploration and utilization of spatial information. For the MUUFL dataset, AA is usually much lower than OA. This is due to the small number of "C9", "C10" and "C11" in the dataset, resulting in serious sample imbalance. For example, the classification accuracy of MS2CANet for these three categories is only 46.73%, 0.00% and 16.22%, resulting in AA of only 61.91% at a medium level of OA.Therefore, alleviating sample imbalance and improving AA is a severe challenge for the MUUFL dataset. The method based on hybrid convolution and Transformer can improve AA. Specifically, on the MUUFL dataset, compared with the CNN-based methods (including HResNet, CCNN and MS2CANet), the CNN and Transformer-based methods (including DBCTNet, ExViT, HCTNet, MHST and PyionNet) can obtain higher AA. The proposed method has the highest AA of 78.31%, which can effectively alleviate the negative impact of sample distribution imbalance on classification results. Combining the above advantages, PyionNet first uses multi-scale convolution blocks to extract shallow features, and then uses the Transformer-based fusion method to extract global features and perform feature fusion. Finally, the fused features are used for classification, thereby obtaining excellent classification accuracy.
[0096] From visual comparison and analysis: Figure 4 , 5 6 shows the classification maps generated by various methods on three datasets, where the local expansion operation is used to more clearly show the differences in the classification results of each method. The results show that the classification maps generated by the proposed PyionNet are more consistent with the ground truth maps compared with other methods, thereby improving the classification results. In addition, the proposed method generates smoother classification maps, while SVM, ExViT and other methods tend to have more discrete data points. Specifically, for Figure 5 For the Houston2013 dataset shown in Figure 1, the classification graph of the proposed PyionNet has clearer boundaries. For example, for the boundaries between “C3”, “C4” and “C10”, the classification results of the proposed PyionNet are more accurate, while the boundaries generated by other methods are not only more blurred but also difficult to distinguish. Almost all methods will misclassify the “C9” part in the MUUFL dataset, such as Figure 6 As shown, in the partially enlarged image, it can be seen that SVM almost completely misclassifies "C9" as "C5", while other methods partially misclassify "C9" as "C3". Compared with other methods, the classification results of PyionNet are closer to the real objects, which is also consistent with the results shown in Table 3. Overall, the proposed PyionNet shows excellent performance and good classification ability in both quantitative and visual analysis.
[0097] Table 3: Classification accuracy OA, AA, Kappa(k) on Houston2013 dataset
[0098]
[0099] Table 4: Classification accuracy of each classification method on the Houston2013 dataset: OA, AA, Kappa(k)
[0100]
[0101]
[0102] Table 5: Ablation experiment results.
[0103]
[0104] The proposed PyionNet is mainly composed of PEC module and PCF module. The PCF module mainly includes DFF block and CFF. It adopts progressive cross fusion method to fully fuse multi-scale features. In order to evaluate the effectiveness of the above components, ablation experiments are carried out on three datasets. The results of these experiments are shown in Table 5. Groups 1-3 of experiments demonstrate the role of the main parts of the PCF module. In the first group of experiments, CFF block is used for single cross fusion, in the second group of experiments, DFF block is used for single cross fusion, and in the third group of experiments, three CFF blocks are used for progressive cross fusion. It can be seen from the results that the performance of DFF block and CFF block on different datasets is different. For example, for the Trento dataset, the classification accuracy of a single cross fusion using the DFF block is higher than that of a single cross fusion using the CFF block. However, for the Houston2013 and MUUFL datasets, a single cross fusion using the CFF block achieves better classification results. The classification performance of the progressive cross fusion method is better than that of the single cross fusion method. Specifically, on the three datasets, the classification accuracy of the sixth group of experiments using the progressive cross fusion method is higher than that of the fourth group of experiments using the single cross fusion method. The results of these two groups of experiments illustrate the effectiveness of the progressive cross fusion strategy. The fourth group of experiments uses the complete network and shows the best classification performance.
[0105] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0106] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A hyperspectral and laser radar image classification method based on a pyramid progressive cross fusion network, the method is implemented using a pyramid progressive cross fusion network, characterized in that: The steps include: Step S1, using a pyramid expansion and compression module of a pyramid progressive cross-fusion network to extract spatial spectral features from HS data and elevation features from LiDAR data, and using an expansion and compression method to remove redundant features and compress multi-scale feature representations; Step S2, using a progressive cross fusion module of a pyramid progressive cross fusion network to fuse multi-scale features; Step S3, using the attention enhancement classifier of the pyramid progressive cross fusion network to classify the image.
2. The hyperspectral and laser radar image classification method based on a pyramid progressive cross fusion network according to claim 1 is characterized in that: The pyramid expansion and compression module of step S1 has a dual-branch structure, and the pyramid expansion and compression module includes a pyramid spectral feature extraction block composed of a multi-scale 3D convolution block, a 3D batch normalization layer and a ReLU activation layer, and a pyramid spatial feature extraction block composed of a multi-scale 2D convolution block, a 2D batch normalization layer and a ReLU activation layer.
3. The hyperspectral and laser radar image classification method based on a pyramid progressive cross fusion network according to claim 2 is characterized in that: The pyramid spectrum feature extraction block and the pyramid space feature extraction block are respectively expressed as: F pse (X)=R(BN(cat(Conv3D 3×1×1 (X),Conv3D 5×1×1 (X),Conv3D 7×1×1 (X)))) F psa (X)=R(BN(cat(Conv2D 3×3 (X),Conv2D 5×5 (X),Conv2D 7×7 (X)))) Among them, Conv3D() and Conv2D() represent 3D convolutional layers and 2D convolutional layers with different convolution kernel sizes, cat() represents a cascade operation, BN() represents a batch normalization layer, R() represents a ReLU activation function, and F pse () and F psa () respectively represent the operation process of pyramid spectral feature extraction block and pyramid spatial feature extraction block.
4. The hyperspectral and laser radar image classification method based on a pyramid progressive cross fusion network according to claim 3 is characterized in that: The progressive cross fusion module of step S2 includes a difference feature fusion block and a pair of common feature fusion blocks.
5. The hyperspectral and laser radar image classification method based on a pyramid progressive cross fusion network according to claim 4 is characterized in that: The specific process of the difference feature fusion block is to calculate the similarity matrix between Q and K, and then multiply it by V to infer the common feature CFM1 between Q and V, which can be expressed as: Where Dropout() represents the Dropout layer, Softmax() is the SoftMax activation function, and d k represents the scaling factor, Q1, K1, and V1 represent the query, key, and value obtained by linear transformation of different modes, K1 T Represents the transposed matrix of K1, V1 is subtracted from CFM1, and the difference feature DFM1 between Q and V is obtained. This process can be expressed as: DFM1=V1-CFM1 The difference features are injected into Q to obtain complementary information from HS data, expressed as: Where Linear(·) represents a linear mapping operation. The feature map passes through the normalization layer and MLP in sequence, and finally performs a residual connection to obtain the output F1 of the DFF block.
6. The hyperspectral and laser radar image classification method based on a pyramid progressive cross fusion network according to claim 5 is characterized in that: The common feature fusion block consists of an interactive attention layer, LN and MLP. The specific steps are as follows: first, the common features of HS and LiDAR data are injected to enrich the fused features. Specifically, the fused feature F1 containing the difference between HS and LiDAR data is used as Q2, and the LiDAR features are used as K2 and V2. The process of further fusing LiDAR features is expressed as: F2=MLP(LN(CFM2))+CFM2 (19) Where F2 represents the output of the first common feature fusion block, and then F2 is further fused with the HS feature. The process is the same as the above equations (18) and (19).
7. The hyperspectral and laser radar image classification method based on a pyramid progressive cross fusion network according to claim 6 is characterized in that: The specific process of the attention enhancement classifier in step S3 is that before the fused features are input into the MLP head composed of the LN layer and the FC layer for final classification, the attention weight A is calculated by introducing the attention mechanism to compress the feature map, and then multiplied by the fused feature F as an incentive. The process can be expressed as A=Softmax(MLP(F)) F′=A⊙F Finally, the following multi-source training samples are used to optimize the network parameters based on multi-class cross entropy loss. Where |Ω train | represents the cardinality of the training index set, represents the output unnormalized score for the i-th pixel of each class.
Citation Information
Cited By
Neural network construction method for hyperspectrum and laser radar fusion classification
CN120259794A
Breast cancer new auxiliary information analysis system based on multi-modal fusion
CN120452755A