Axial Transform multi-source remote sensing image classification method based on hierarchical multi-modal feature aggregation
By introducing hierarchical multimodal feature aggregation and multi-head axial attention Transformer in multi-source remote sensing image classification, the problem of feature-level processing limitations in the prior art is solved, and more efficient multimodal feature fusion and classification performance improvement is achieved.
Patent Information
- Application Number
- CN202510085241.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
Existing multi-source remote sensing image classification methods are usually limited to single-level feature-level processing, ignoring the deep fusion between multimodal features at different levels, making it difficult to continue to improve classification performance.
A multi-head axial attention Transformer (HMAT) method based on hierarchical multi-modal feature aggregation is proposed. Cross-level feature interaction and multi-scale feature fusion are realized through hierarchical multi-modal feature aggregation module, pyramid-inverse pyramid convolution module and multi-head axial attention component.
By deeply fusion of multi-level features, the performance of multi-source remote sensing image classification is significantly improved. The experimental results show that HMAT has better classification performance than current advanced methods on three publicly available data sets.
Smart Images

Figure CN120014453A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a multi-source remote sensing image classification method. Background Art
[0002] In hyperspectral images (HSIs), each sample corresponds to a complete spectral curve, and this spectral information can be used to identify the type or state of land cover. Therefore, HSI is widely used in land monitoring, urban planning, medical diagnosis, precision agriculture and other fields. In the hyperspectral image classification (HSIC) task, the corresponding category label is assigned to each pixel in the image by analyzing the data features in HSI. However, HSI data is easily disturbed by factors such as atmospheric conditions and lighting conditions during the acquisition process, which leads to incorrect classification. Light detection and ranging (LiDAR) data can record the elevation information of the measured object and is less disturbed by atmospheric conditions. Therefore, it can be used in combination with HSI data to alleviate the limitations of using the two data separately. With the continuous breakthroughs of multimodal AI in the field of artificial intelligence, studying the joint image classification method of HSI and LiDAR data has also become a hot topic in the field of remote sensing image classification.
[0003] Early studies on the classification of HSI and LiDAR data were mostly based on hand-crafted features. A generalized feature fusion method based on graphs was proposed, which combines the original spectral information with the morphological profile features in multimodal fusion data for classification tasks. A method based on extended morphological profiles was proposed to classify HSI and LiDAR data by superimposing features extracted from different modal data. A method based on support vector machines was proposed to classify complex forest areas. Although these methods are easy to implement, they rely too much on prior knowledge and cannot adaptively characterize the intrinsic characteristics of different modal data.
[0004] In recent years, deep learning technology has made continuous breakthroughs in the field of multimodal AI, especially in the sub-field of multi-source remote sensing image classification, where deep learning models have achieved many remarkable results. Convolutional neural networks (CNNs) have been favored by researchers for their powerful feature extraction capabilities and parameter sharing characteristics. A method based on coupled convolutional neural network (CoupledCNN) was proposed, which fuses HSI and LiDAR data at the feature level through a weight sharing strategy and explores multiple decision-level fusion methods. A method called interleaving perception convolutional neural network (IP-CNN) was proposed, which combines traditional CNN with information fusion technology to solve the joint classification problem of multi-source heterogeneous data. A convolutional neural network with cross-channel reconstruction module (CCR-Net) based on cross-channel reconstruction was proposed, which learns reconstruction strategies across different remote sensing data sources to more effectively exchange information with each other. A method called spatial-spectral cross-modal enhancement network (S2ENet) was proposed, which promotes information interaction between the two modalities by enhancing the spatial features in HSI and the spectral features in LiDAR respectively. A method called cross-modal semantic enhancement network (CMSE) was proposed, which realizes the mining and fusion of similar high-order semantic information and complementary discriminant information in multimodal data.
[0005] Generative adversarial network (GAN) effectively improves the robustness and generalization of the learning model through adversarial training between the generator and the discriminator. A coupled adversarial learning based classification (CALC) method is proposed, which effectively integrates the high-level semantic information and complementary information in HSI and LiDAR data through coupled adversarial learning and multi-level feature processing.
[0006] Graph attention network (GAT) is a variant of graph neural network (GNN), which emphasizes the importance of using attention mechanism when propagating and updating information between nodes in the graph. A graph-attention based multimodal fusion network (GAMF) based on GAT is proposed, which uses graph attention mechanism to construct undirected graph to alleviate the long-distance dependency problem in HSI and LiDAR data.
[0007] Transformer is a deep learning model for processing sequence data, which has had a revolutionary impact in the field of natural language processing (NLP). After its core idea was adjusted and improved, it has also been widely used in the field of computer vision (CV) and achieved remarkable success. A method called global-local transformer network (GLT-Net) was proposed, which combines the advantages of CNN in extracting local spatial features and the ability of transformer in learning long-distance dependencies. A method called multimodal fusion transformer (MFT) was proposed. Different from traditional feature fusion techniques, MFT improves the generalization ability of the model by introducing multimodal data as external classification tokens into the transformer encoder. A method called hierarchical convolutional neural network and transformer (HCT) was proposed, which fuses the features of two modal data through specially designed cross-token attention. A method called spectral-spatial-elevation fusion transformer (S2EFT) was proposed, which alleviated the shortcomings of transformer in processing local information. A method called cross hyperspectral and lidar attention transformer (Cross-HL) was proposed, which promoted the accurate exchange of information between different modalities by extending the self-attention mechanism. A method called height information guided hierarchical fusion-and-separation network (HFSNet) was proposed, which designed a dual-structure encoder to capture the spectral sequence information in HSI and the spatial information in LiDAR respectively, and used the elevation information to guide the mutual learning between modalities.
[0008] Although the above methods have achieved good classification performance, most joint classifications are often limited to a single-level feature-level processing, while ignoring the deep fusion of multimodal features at different levels. This limitation makes it difficult to further improve the classification performance of these methods. Summary of the invention
[0009] The purpose of this invention is to solve the problem that although existing methods have achieved good classification performance, most joint classifications are often limited to a single level of feature level processing, while ignoring the deep fusion between multimodal features at different levels. This limitation makes it difficult to further improve the classification performance of these methods. A multi-source remote sensing image classification method based on axial Transformer of hierarchical multimodal feature aggregation is proposed.
[0010] The specific process of the multi-source remote sensing image classification method based on axial Transformer of hierarchical multimodal feature aggregation is as follows:
[0011] Step 1: Obtain hyperspectral image data HSI. Hyperspectral image data HSI is expressed as
[0012] Get the light detection and ranging data LiDAR, the light detection and ranging data LiDAR is expressed as
[0013]
[0014] Where H and W represent the height and width of the input data, respectively, and B′ and B″ represent the number of bands of HSI input and LiDAR input, respectively; represents a real number;
[0015] Step 2: Construct HMAT network model;
[0016] The HMAT network model represents a multi-head axial attention Transformer model based on hierarchical multimodal feature aggregation;
[0017] The HMAT network model includes a channel modulation module, a hierarchical multimodal feature aggregation module, a weighted local feature extraction module, a global feature extraction module, and a local-global feature fusion classification module;
[0018] The specific working process of the HMAT network model is as follows:
[0019] The hyperspectral image data HSI is input into the channel modulation module, and the channel modulation module outputs the pre-processed HSI data;
[0020] The light detection and ranging data LiDAR is input into the channel modulation module, and the channel modulation module outputs the pre-processed LiDAR data;
[0021] The preprocessed HSI data and the preprocessed LiDAR data are input into the hierarchical multimodal feature aggregation module, and the hierarchical multimodal feature aggregation module outputs data X*;
[0022] The data X* is input into the weighted local feature extraction module, and the output data of the weighted local feature extraction module is respectively input into the global feature extraction module and the local-global feature fusion classification module;
[0023] The output data of the global feature extraction module is input into the local-global feature fusion classification module;
[0024] The output data of the local-global feature fusion classification module is used as the classification result output by the HMAT network model;
[0025] Step 3: Based on the constructed HMAT network model, the hyperspectral image data and light detection and ranging data pairs are predicted to obtain the classification results.
[0026] The beneficial effects of the present invention are:
[0027] To alleviate the above problems, a hierarchical multi-modal feature aggregation based multi-head axial attention transformer (HMAT) is proposed for the joint classification task of HSI and LiDAR data. Firstly, a hierarchical multi-modal feature aggregation module (HMFA) is proposed to extract and fuse feature information from different levels, realize cross-level feature interaction and layer-by-layer progressive feature fusion. Secondly, a pyramid-inverted pyramid convolution module (PIP) is designed to further enhance the model's ability to extract local proximity and detail features of elements in the image. Finally, a multi-head axial attention component (MHAA) is built, which can capture and fuse features at multiple scales, thereby more effectively improving the performance of the model in classification tasks. The proposed HMAT is fully experimentally verified on three publicly available datasets. Experimental results show that the proposed method has better classification performance than some current advanced methods.
[0028] In order to more effectively process cross-level feature information, a hierarchical multimodal feature aggregation module is proposed. This module builds multi-level feature interactions between multimodal data and realizes feature fusion from shallow to deep and layer by layer.
[0029] In order to enhance the model's ability to extract local proximity and detail features of elements, a pyramid-reverse pyramid convolution module is designed. This module not only enriches the level and details of local feature representation through two complementary structures, but also realizes mutual enhancement and complementation between features.
[0030] In order to capture the global context information in the fused features at different scales, a multi-head axial attention component is built, which makes full use of features from different scales and significantly enhances the performance of the model in classification tasks.
[0031] A multimodal feature processing framework based on hierarchical multimodal feature aggregation is proposed, and extensive experiments are carried out on three datasets. The experimental results demonstrate that the proposed method can provide excellent classification performance for the joint classification of hyperspectral and radar data.
[0032] This paper proposes a multi-head axial attention Transformer based on hierarchical multimodal feature aggregation for the joint classification task of HSI and LiDAR data. First, the multimodal data is unified and standardized to eliminate the scale differences between different data sources. After that, the modulated multimodal data is subjected to feature extraction and feature fusion at different levels using the hierarchical multimodal feature aggregation module. Next, the model is enhanced to extract local proximity and detail features of elements through the pyramid-inverse pyramid convolution module. Then, the Transformer encoder based on multi-head axial attention is used to process multimodal fusion features at different scales to extract more effective global features. Finally, the local-global features are fed into the classification head to obtain the final result through an adaptive fusion strategy. Experimental results on three datasets show that, with limited training samples, the proposed HMAT exhibits superior classification performance compared with the current state-of-the-art methods in the field of joint classification of HSI and LiDAR data. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 The overall structure diagram of the multi-head axial attention HMAT based on hierarchical multimodal feature aggregation proposed in the present invention;
[0034] Figure 2 This is the structural diagram of the pyramid-inverse pyramid convolution module PIP;
[0035] Figure 3 This is the spatial attention structure diagram;
[0036] Figure 4 The detailed structure diagram of the multi-head axial attention component MHAA, T0cls represents the classification mark, T1 represents the first element in the feature sequence, T2 represents the second element in the feature sequence, and Td +1 represents the (d+1)th element in the feature sequence, Matmul represents matrix multiplication, Scale represents scaling, s represents the sequence length in the row direction or column direction, n represents the length of the spatial sequence, c represents the length of the channel sequence; d represents the length of the channel sequence;
[0037] Figure 5 The classification results of all methods on the Houston2013 dataset are as follows. (a)-(n) are: pseudo-color image of HSI, DSM image of LiDAR, true label image, CoupledCNN, CALC, S2ENet, GLT, MFT, HCT, S2EFT, GAMF, CrossHL, HFSNet, Proposed; Healthygrass represents healthy lawn, Stressedgrass represents stressed lawn, Syntheticgrass represents artificial lawn, Trees represents trees, Soil represents soil, Water represents water, Residential represents residence, Commercial represents business, Road represents road, Highway represents highway, Railway represents railway, ParkingLot1 represents parking lot 1, ParkingLot2 represents parking lot 2, Tennis Court represents tennis court, RunningTrack represents running track;
[0038] Figure 6 The classification results of all methods on the Trento dataset. (a)-(n) are: pseudo-color image of HSI, DSM image of LiDAR, true label image, CoupledCNN, CALC, S2ENet, GLT, MFT, HCT, S2EFT, GAMF, CrossHL, HFSNet, Proposed; Apple Trees represents apple trees, Buildings represents buildings, Ground represents ground, Woods represents trees, Vineyard represents vineyard, and Roads represents roads;
[0039] Figure 7The classification results of all methods on the MUUFL dataset, (a)-(n) are: pseudo color image of HSI, DSM image of LiDAR, true label image, CoupledCNN, CALC, S2ENet, GLT, MFT, HCT, S2EFT, GAMF, CrossHL, HFSNet, Proposed; Trees represents trees, Mostly Trees represents grass, Mixed groundsurface represents mixed ground, Dirt and sand represents dirt and sand, Road represents road, Water represents water, Building shadow represents building shadow, Building represents building, Sidewalk represents sidewalk, YellowCurb represents yellow curb, Clothpanels represents cloth panels;
[0040] Figure 8 Comparison between MHAA and MHSA on three datasets, (a)-(c) are: Houston 2013, Trento, MUUFL, OA represents the overall accuracy, AA represents the average accuracy, K represents the confusion matrix, MHAA represents multi-head axial attention, MHSA represents multi-head self-attention, and OverallAccuracy represents the overall accuracy;
[0041] Fig. 9 Feature visualization on the Houston 2013 dataset. (a)-(f) are GLT, MFT, HCT, CrossHL, HFSNet, and Proposed.
[0042] Fig.10 Feature visualization on the Trento dataset. (a)-(f) are: GLT, MFT, HCT, CrossHL, HFSNet, Proposed;
[0043] Fig.11 Feature visualization on the MUUFL dataset. (a)-(f) are: GLT, MFT, HCT, CrossHL, HFSNet, Proposed;
[0044] Fig.12 Confusion matrix on Houston 2013 dataset. (a)-(f) are GLT, MFT, HCT, CrossHL, HFSNet, Proposed. True Classes indicates the real category, and predicted Classes indicates the predicted category.
[0045] Fig.13Confusion matrix on the Trento dataset, (a)-(f) are: GLT, MFT, HCT, CrossHL, HFSNet, Proposed;
[0046] Fig.14 Confusion matrix on the MUUFL dataset, (a)-(f) are: GLT, MFT, HCT, CrossHL, HFSNet, Proposed. DETAILED DESCRIPTION
[0047] Specific implementation method 1: The specific process of the multi-source remote sensing image classification method based on the axial Transformer of hierarchical multimodal feature aggregation in this implementation method is as follows:
[0048] With the significant breakthrough of multimodality in the field of artificial intelligence, multi-source remote sensing image classification has also become a research hotspot. However, most joint classification methods for multi-source data are limited to single-level feature-level processing, ignoring the deep fusion of cross-level feature information. This limitation seriously restricts the exchange and utilization of effective information between different modal data. To alleviate this problem, this paper proposes a hierarchical multi-modal feature aggregation based multi-headaxial attention transformer (HMAT) for the joint classification task of HSI and LiDAR data. First, a hierarchical multi-modal feature aggregation module (HMFA) is proposed to process and fuse feature information across levels. Secondly, a pyramid-inverted pyramid convolution module (PIP) is designed to enhance the model's extraction of local proximity and detail features of elements. Finally, a multi-head axial attention component (MHAA) is built to capture fusion features at different scales, thereby more effectively improving the classification ability. The proposed HMAT is fully experimented on three public datasets. The experimental results show that the proposed method has better classification performance than some current advanced methods.
[0049] Step 1: Obtain hyperspectral image data HSI. Hyperspectral image data HSI is expressed as
[0050] Get the light detection and ranging data LiDAR, the light detection and ranging data LiDAR is expressed as
[0051]
[0052] Where H and W represent the height and width of the input data, respectively, and B′ and B″ represent the number of bands of HSI input and LiDAR input, respectively; represents a real number;
[0053] Step 2: Construct HMAT network model;
[0054] The HMAT network model represents the hierarchical multi-modal feature aggregation based multi-head axial attention transformer (HMAT).
[0055] The HMAT network model includes a channel modulation module, a hierarchical multimodal feature aggregation module, a weighted local feature extraction module, a global feature extraction module, and a local-global feature fusion classification module;
[0056] The overall structure of the proposed HMAT is as follows Figure 1 As shown in the figure, its main process includes five parts: channel modulation, hierarchical multimodal feature aggregation, weighted local feature extraction, global feature extraction, and local-global feature fusion classification.
[0057] The specific working process of the HMAT network model is as follows:
[0058] The hyperspectral image data HSI is input into the channel modulation module, and the channel modulation module outputs the pre-processed HSI data;
[0059] The light detection and ranging data LiDAR is input into the channel modulation module, and the channel modulation module outputs the pre-processed LiDAR data;
[0060] The preprocessed HSI data and the preprocessed LiDAR data are input into the hierarchical multimodal feature aggregation module, and the hierarchical multimodal feature aggregation module outputs data X*;
[0061] The data X* is input into the weighted local feature extraction module, and the output data of the weighted local feature extraction module is respectively input into the global feature extraction module and the local-global feature fusion classification module;
[0062] The output data of the global feature extraction module is input into the local-global feature fusion classification module;
[0063] The output data of the local-global feature fusion classification module is used as the classification result output by the HMAT network model;
[0064] Step 3: Based on the constructed HMAT network model, the hyperspectral image data and light detection and ranging data pairs are predicted to obtain the classification results.
[0065] Since HSI data and LiDAR data have different feature representation spaces, especially obvious differences in data dimensions, directly inputting these raw data into the network will make it difficult for the network to capture effective common features, which will have an adverse impact on subsequent processing.
[0066] Therefore, in the channel modulation part, the number of channels of different modal data is unified using 3×3 convolution kernels. In addition, these data are standardized through batch normalization operations, which effectively eliminates the scale and distribution differences caused by different data sources.
[0067] Afterwards, in the hierarchical multimodal feature aggregation part, the proposed HMFA module is used to extract and preliminarily fuse the multimodal features at different levels; then, through the feature aggregation strategy, these fused features are effectively aggregated into low-dimensional feature representations containing information at different levels. Next, in the weighted local feature extraction part, the proposed PIP module is used to perform multi-scale feature extraction and spatial dimension attention weighting on the low-dimensional output of the previous stage. Subsequently, the weighted output is fed into two branches, the first branch is the subsequent global feature extraction part to capture global contextual information; the other branch uses inverted pyramid convolution to extract features that complement the input. Then, in the global feature extraction part, through the proposed MHAA component, Transformer can extract information on different scales of the input data, and then output more representative global features.
[0068] Finally, in the local-global feature fusion classification stage, local features and global features are processed into feature vectors of the same shape. Subsequently, an adaptive fusion strategy is used to set dynamically adjusted weighting coefficients for global features and local features respectively to effectively fuse the two features, thereby obtaining a more accurate classification result.
[0069] Specific implementation method 2: This implementation method is different from the specific implementation method 1 in that the specific working process of the channel modulation module is as follows:
[0070] The HSI data and LiDAR data are preprocessed respectively to obtain the preprocessed HSI data and LiDAR data; they are expressed as:
[0071]
[0072] in, represents a 2D convolutional layer of size 3, F BN (·) represents the batch normalization layer, δ represents the ReLU activation function, * represents the convolution operator, W represents the weight of the convolution kernel, b represents the bias, Xin represents the HSI data or LiDAR data, and Xi′n represents the preprocessed HSI data or LiDAR data;
[0073] The other steps and parameters are the same as those in the first embodiment.
[0074] Specific implementation method three: The difference between this implementation method and specific implementation method one or two is that: the hierarchical multimodal feature aggregation module includes a first 3×3 convolutional layer, a first BN layer, a first ReLU activation function layer, a second 3×3 convolutional layer, a second BN layer, a second ReLU activation function layer, a third 3×3 convolutional layer, a third BN layer, a third ReLU activation function layer, a fourth 3×3 convolutional layer, a fourth BN layer, a fourth ReLU activation function layer, a fifth 3×3 convolutional layer, a fifth BN layer, a fifth ReLU activation function layer, a sixth 3×3 convolutional layer, a sixth BN layer, a sixth ReLU activation function layer, a seventh 3×3 convolutional layer, The seventh BN layer, the seventh ReLU activation function layer, the eighth 3×3 convolution layer, the eighth BN layer, the eighth ReLU activation function layer, the ninth convolution layer, the ninth BN layer, the ninth Sigmoid activation function layer, the first 2×2 average pooling layer, the second 2×2 average pooling layer, the third 2×2 average pooling layer, the fourth 2×2 average pooling layer, the fifth 2×2 average pooling layer, the sixth 2×2 average pooling layer, the first connection layer (Concatenate), the second connection layer (Concatenate), the third connection layer (Concatenate), and the fourth connection layer (Concatenate).
[0075] The other steps and parameters are the same as those in the first or second embodiment.
[0076] Specific implementation method 4: This implementation method is different from any one of the specific implementation methods 1 to 3 in that: the specific working process of the hierarchical multimodal feature aggregation module HMFA is as follows:
[0077] 1) HSI data is sequentially input into the first 3×3 convolutional layer, the first BN layer, the first ReLU activation function layer, the second 3×3 convolutional layer, the second BN layer, the second ReLU activation function layer, and the first 2×2 average pooling layer. The first 2×2 average pooling layer outputs feature A. Feature A is sequentially input into the third 3×3 convolutional layer, the third BN layer, the third ReLU activation function layer, and the second 2×2 average pooling layer. The second 2×2 average pooling layer outputs feature B. Feature B is sequentially input into the fourth 3×3 convolutional layer, the fourth BN layer, the fourth ReLU activation function layer, and the third 2×2 average pooling layer. The third 2×2 average pooling layer outputs feature C.
[0078] 2) LiDAR data is sequentially input into the fifth 3×3 convolutional layer, the fifth BN layer, the fifth ReLU activation function layer, the sixth 3×3 convolutional layer, the sixth BN layer, the sixth ReLU activation function layer, and the fourth 2×2 average pooling layer. The fourth 2×2 average pooling layer outputs feature A′. Feature A′ is sequentially input into the seventh 3×3 convolutional layer, the seventh BN layer, the seventh ReLU activation function layer, and the fifth 2×2 average pooling layer. The fifth 2×2 average pooling layer outputs feature B′. Feature B′ is sequentially input into the eighth 3×3 convolutional layer, the eighth BN layer, the eighth ReLU activation function layer, and the sixth 2×2 average pooling layer. The sixth 2×2 average pooling layer outputs feature C′.
[0079] 3) Feature A and feature A′ are input into the first connection layer (Concatenate), and the first connection layer (Concatenate) outputs feature xs;
[0080] Feature B and feature B′ are input into the second connection layer (Concatenate), and the second connection layer (Concatenate) outputs feature xm;
[0081] Features C and C′ are input into the third connection layer (Concatenate), and the third connection layer (Concatenate) outputs feature xd;
[0082] 4) Flatten and linearly map the feature xs in turn to obtain the feature
[0083] Flatten and linearly map the feature xm in turn to obtain the feature
[0084] Flatten and linearly map the feature xd in turn to obtain the feature
[0085] 5) The features feature feature Input the fourth connection layer (Concatenate), the fourth connection layer (Concatenate) outputs the stacked data
[0086] It is expressed as:
[0087]
[0088] Among them, Fconcat(·) represents cascade (the fourth connection layer (Concatenate));
[0089] 6) The stacked data Input the ninth convolutional layer, the ninth BN layer, and the ninth Sigmoid activation function layer in sequence;
[0090] This setting ensures that each convolution operation can cover the feature information from the shallow, middle and deep layers, so as to effectively integrate more hierarchical details in the process of low-dimensional data fusion;
[0091] The output features of the ninth Sigmoid activation function layer are deformed and re-transformed into a matrix format to obtain data X* for the convenience of processing in subsequent modules;
[0092] It is expressed as:
[0093]
[0094] in, represents a 2D convolutional layer with a shape of 1×3, F BN (·) represents the batch normalization layer, σ represents the Sigmoid activation function, and F ε (·) represents a deformation operation.
[0095] Hierarchicalmulti-modalfeature aggregation module
[0096] At present, the fusion methods of multi-source data are generally limited to single-level feature-level processing, ignoring the deep fusion of cross-level feature information. This limitation greatly hinders the establishment of effective information interaction between different modal data. In order to alleviate this problem, a hierarchical multi-modal feature aggregation module (HMFA) is designed. Its detailed structure is shown in Figure 1The light red background part in the middle. In the previous steps, the feature data from different modalities have been normalized. On this basis, a symmetrically designed dual-branch structure is adopted for subsequent feature extraction. Specifically, the shallow features of different modal data are first extracted using a 3×3 convolution block, and then average pooling is introduced to effectively reduce the spatial dimension of the feature (the original window size is h×w) to h / 2×w / 2, so as to achieve spatial compression of the feature and preliminary aggregation of information. Finally, the features of different modalities are stacked together using a cascade operation. This process can be expressed as:
[0097]
[0098] Among them, Favg(·) represents the average pooling layer;
[0099] This process is then repeated at a deeper level, but the spatial dimensions of the features are adjusted after each processing to adapt to the feature representation requirements at different levels. In detail, in the mid-level feature extraction stage, the window size is reduced to h / 4×w / 4 after the pooling operation. At this point, the feature map is further compressed. In the deep feature extraction stage, the window size is adjusted to h / 8×w / 8, which is designed to extract the most abstract and representative features.
[0100] Through the above hierarchical feature extraction process, three different levels of multimodal fusion features are captured, namely: shallow features Mid-level features Deep Features Next, we will perform a deeper feature fusion process on these data. The detailed process is as follows: Figure 2 shown.
[0101] Specifically, these multi-dimensional data are first flattened and converted into a unified one-dimensional vector form. Then, the lengths of the three different feature sequences are adjusted and aligned using linear mapping. At this point, the data from different modalities will also be initially fused. This process can be expressed as:
[0102]
[0103] in, represents the flattening function, F L inear(·) represents the linear layer. Then, a new dimension is expanded for each of the three initially fused features, namely Then, on this newly added dimension, these three features are stacked in order to obtain Then, the stacked data is processed using a banded convolution kernel. Specifically, the size of the convolution kernel is set to 1×3, which is exactly the same as The stacked feature dimensions match each other. This setting ensures that each convolution operation can cover the feature information from the shallow, middle and deep layers, so that more hierarchical details can be effectively integrated into the low-dimensional data fusion process. Finally, the data after the hierarchical feature fusion is Reshape it into matrix format, and we get To facilitate the processing of subsequent modules.
[0104] The other steps and parameters are the same as those in Specific Embodiments 1 to 3.
[0105] Specific implementation method 5: This implementation method is different from any one of the specific implementation methods 1 to 4 in that: the data X* is input into the weighted local feature extraction module, and the output data of the weighted local feature extraction module is respectively input into the global feature extraction module and the local-global feature fusion classification module; the specific process is:
[0106] The weighted local feature extraction module is a pyramid-inverted pyramid module;
[0107] The pyramid-inverted pyramid module (PIP) includes a pyramid convolution block, a spatial attention block, and an inverse pyramid convolution block;
[0108] The pyramid convolution block includes a tenth 1×1 convolution layer, an eleventh 3×3 convolution layer, a twelfth 5×5 convolution layer, and a thirteenth 7×7 convolution layer;
[0109] The tenth 1×1 convolution layer is divided into 1 group, and the size of the convolution kernel of each group is 1×1 (the number of channels becomes: the number of channels divided by the number of groups);
[0110] The eleventh 3×3 convolutional layer is divided into 2 groups, and the size of the convolution kernel of each group is 3×3 (the number of channels becomes: the number of channels divided by the number of groups);
[0111] The twelfth 5×5 convolution layer is divided into 3 groups, and the size of the convolution kernel of each group is 5×5 (the number of channels becomes: the number of channels divided by the number of groups);
[0112] The thirteenth 7×7 convolutional layer is divided into 4 groups, and the size of the convolution kernel of each group is 7×7 (the number of channels becomes: the number of channels divided by the number of groups);
[0113] The inverse pyramid convolution block includes a fourteenth 7×7 convolution layer, a fifteenth 5×5 convolution layer, a sixteenth 3×3 convolution layer, and a seventeenth 1×1 convolution layer;
[0114] The fourteenth 7×7 convolutional layer is divided into 4 groups, and the size of the convolution kernel of each group is 7×7 (the number of channels becomes: the number of channels divided by the number of groups);
[0115] The fifteenth 5×5 convolutional layer is divided into 3 groups, and the size of the convolution kernel of each group is 5×5 (the number of channels becomes: the number of channels divided by the number of groups);
[0116] The sixteenth 3×3 convolution layer is divided into 2 groups, and the size of the convolution kernel of each group is 3×3 (the number of channels becomes: the number of channels divided by the number of groups);
[0117] The seventeenth 1×1 convolution layer is divided into 1 group, and the size of the convolution kernel of each group is 1×1 (the number of channels becomes: the number of channels divided by the number of groups);
[0118] The working process of the pyramid-inverted pyramid convolution module PIP (PIP) is as follows:
[0119] Data X* is input into the pyramid convolution block, and the pyramid convolution block outputs feature X P conv;
[0120] Feature X P conv inputs spatial attention, and spatial attention outputs data X′j;
[0121] Data X′j is added element by element to data X* to obtain feature X′ P ′conv;
[0122] The feature X′ P ′conv is reshaped into Xvit (dimensional transformation, the second dimension and the third dimension are exchanged) for subsequent global feature processing;
[0123] The feature X′ P ′conv is fed into the inverted pyramid convolution block, and the inverted pyramid convolution block outputs the feature Xcnn;
[0124] Feature Xvit input global feature extraction module;
[0125] The inverted pyramid convolution block outputs the feature Xcnn which is input into the local-global feature fusion classification module.
[0126] The other steps and parameters are the same as those in Specific Implementation 1 to 4-1.
[0127] Specific implementation method 6: This implementation method is different from any one of the specific implementation methods 1 to 5 in that: the data X* is input into the pyramid convolution block, and the pyramid convolution block outputs the feature X P conv; the specific process is:
[0128] The data X* is input into the tenth 1×1 convolution layer which is divided into a group, and the tenth 1×1 convolution layer outputs the feature D;
[0129] The data X* is input into the eleventh 3×3 convolutional layer divided into two groups, and the eleventh 3×3 convolutional layer outputs the feature E;
[0130] The data X* is input into the twelfth 5×5 convolutional layer divided into three groups, and the twelfth 5×5 convolutional layer outputs the feature F;
[0131] The data X* is input into the thirteenth 7×7 convolutional layer divided into four groups, and the thirteenth 7×7 convolutional layer outputs the feature G;
[0132] Cascade feature D, feature E, feature F, and feature G to get feature X P conv;
[0133] It is expressed as:
[0134] X P conv=Fconcat(F1* ×1 (X*),F3* ×3 (X*),F5* ×5 (X*),F7* ×7 (X*))
[0135] Among them, F5* ×5 (·) represents a 2D convolution kernel of size 5, F7* ×7 (·) represents a 2D convolution kernel of size 7, F3* ×3 (·) represents a 2D convolution kernel of size 3, F1* ×1 (·) represents a 2D convolution kernel of size 1, X P conv represents the output of pyramid convolution.
[0136] The other steps and parameters are the same as those in Specific Implementation Methods 1 to 5-1.
[0137] Specific implementation method 7: This implementation method is different from any one of the specific implementation methods 1 to 6 in that: the feature X′ P ′conv is sent to the inverted pyramid convolution block, and the inverted pyramid convolution block outputs the feature Xcnn; the specific process is:
[0138] Data X′ P ′conv input is divided into four groups of four 7×7 convolutional layers, and the fourteenth 7×7 convolutional layer outputs features D′;
[0139] Data X′ P 'conv input is divided into three groups of fifteenth 5×5 convolutional layer, and the fifteenth 5×5 convolutional layer outputs the feature E';
[0140] Data X′ P ′conv input is divided into two groups of sixteenth 3×3 convolutional layers, and the sixteenth 3×3 convolutional layer outputs features F′;
[0141] Data X′ P ′conv input is divided into a group of seventeenth 1×1 convolutional layers, and the seventeenth 1×1 convolutional layer outputs the feature Xcnn;
[0142] It is expressed as:
[0143] Xcnn=Fconcat(F7* ×7 (X′ P ′conv),F5* ×5 (X′ P ′conv),F3* ×3 (X′ P ′conv),F1* ×1 (X′ P ′conv)).
[0144] The other steps and parameters are the same as those in Specific Embodiments 1 to 6.
[0145] Pyramid-invertedpyramidconvolutionmodule
[0146] The Transformer model uses a self-attention mechanism to allow each element in the sequence to interact with all other elements. Although this global view is powerful, it also reduces the direct attention to the local proximity and detailed features of the elements, limiting the ability to extract local features. To alleviate this problem, a pyramid-inverted pyramid convolution module (PIP) was proposed. Its detailed structure is as follows Figure 3 shown.
[0147] First, the pyramid convolution block is used to transform the low-dimensional output X of the previous stage * Perform multi-scale feature extraction. The number of pyramid convolution layers is set to 4, and each layer of convolution adopts a grouping strategy, with the number of groups in each layer being (1, 2, 4, 8). The larger the size of the convolution kernel, the more groups are selected. This allows for the extraction of as much feature information of different scales as possible while maintaining computational efficiency. The number of output channels after each branch is processed is set to 1 / 4 of the number of input channels, and the feature maps processed by different branches are concatenated to form the final output feature map.
[0148] Afterwards, the position attention module is used to weight the local features in the feature map to distinguish the importance of different positions. Its detailed structure is as follows Figure 4 Specifically, we first use convolution to transform the cascade output X of the previous process. Pconv Perform nonlinear transformation to obtain two outputs Q and K, where Afterwards, reshape these two outputs into Where n is the number of elements in the spatial dimension. Next, Q and K T Multiply and calculate the position attention map The score range is (0,1). The calculation process of the attention map is
[0149]
[0150] Among them, Mij represents the influence of the i-th position on the j-th position. Then, the attention map is multiplied with the original input and reshaped into In order to keep some key original information from being lost, a skip connection is introduced, which fuses the original information with the weighted features to generate the final weighted output X′ P conv. This process can be expressed as
[0151]
[0152] Wherein, α represents the dynamically adjusted weighting coefficient.
[0153] Next, the weighted feature is fused with the original input X* and the result is divided into two branches. The first branch reshapes the feature into It is used for subsequent global feature processing. The other branch sends the feature to the inverted pyramid convolution block. The structure of this module is symmetric with the pyramid convolution block mentioned above. This symmetric structure aims to use larger convolution kernels to integrate the fine-grained information previously captured by smaller kernels, while using smaller convolution kernels to further refine the coarse-grained information previously obtained by the larger kernels. The features extracted at the two scales complement each other, enriching the level and details of feature representation. Finally, these features are also cascaded in sequence to output a result containing local features. This process can be expressed as
[0154] X′ P ′conv=F ε (X′ P conv)+X*
[0155]
[0156] Among them, X′ P ′conv represents the weighted features of the position attention and the original input X * fusion result.
[0157] Specific implementation eight: This implementation differs from any one of specific implementations one to seven in that: the feature Xvit is input into the global feature extraction module; the specific process is:
[0158] The global feature extraction module is a multi-head axial attention mechanism module MHAA (multi-head axial attention, MHAA);
[0159] Output of PIP As the input of MHAA; n represents the length of the spatial sequence, c represents the length of the channel sequence;
[0160] Add learnable tokens to the input sequence Xvit and integrate positional encoding information to obtain information
[0161]
[0162] Subsequently, a weight sharing strategy is adopted to map the information X′vit through two parallel linear layers to map the query vector Xq and the key vector Xk, that is, n′ represents the length of the spatial sequence after adding the learnable token;
[0163] This shared weight setting is designed to promote the correlation between the query and the key vector. At the same time, in order to retain more original information, a separate linear layer is used to map the information X′vit to map the value vector Xv, that is,
[0164] This process can be expressed as
[0165] X′vit=Fconcat(Xcls,Xvit)+X PE
[0166]
[0167] Among them, X PE represents positional encoding, Xcls represents learnable tokens, Fl′inear(·) represents a linear layer with weight sharing; Fconcat represents concatenation; Flinear(·) represents a linear layer;
[0168] Next, before calculating the attention map, the query vector Xq and the key vector Xk are reshaped into matrix shapes, i.e. h represents the length of the channel sequence in the row direction, and w represents the length of the channel sequence in the column direction;
[0169] The value vector Xv is further divided into data of multiple attention heads to form To process information of different subspaces in parallel; d represents the length of the channel sequence;
[0170] Then, the column pooling operation is implemented on the query vector Xq, aiming to capture the key information of the vector in the vertical direction;
[0171] Accordingly, the row pooling operation is implemented on the key vector Xk to extract important features in the horizontal direction;
[0172] Next, the pooled results are used to construct an attention map, and the column pooling and row pooling results are converted into normalized attention scores through scaling operations and the Softmax function.
[0173] Expressed as
[0174]
[0175] in,
[0176] Fa′vg(·) represents the average pooling in the column direction, F ε (·) represents the deformation operation (reshaping operation), dq represents the dimension of Xq, Mcol represents the attention map in the column direction, and Fsoftmax(·) represents the Softmax function;
[0177] Fa″vg(·) represents the average pooling in the row direction, dk represents the dimension of Xk, and Mrow represents the attention map in the row direction;
[0178] After that, the value vector Xv is matrix multiplied with the attention maps of two different scales, Mcol and Mrow. This step combines the information in the value vector with the attention weights to generate a weighted feature representation. Next, the outputs of different attention heads are integrated, and the weighted results of the two branches are reshaped into a sequence. The two sequences are merged through matrix addition, and the fused feature representation is remapped to the original shape of the input data. Subsequently, layer normalization is used to stabilize the distribution of the data and fuse the shallow features from the skip connection to obtain the feature This process is represented by Xmhaa=F LN (Flinear(F ε (Xv×Mrow)+F ε (Xv×Mcol)))+Xv′it
[0179] Among them, F LN (·) represents layer normalization;
[0180] The feature Xmhaa is sent to the feedforward perceptron to perform nonlinear transformation on the vector of each position; the layer normalization is used to stabilize the distribution of the data and fuse the shallow features from the jump connection to obtain the feature; this encoding process will be repeated δ times (5 times) to extract the deepest and most abstract global features This process is represented by
[0181] X′v′it=F LN (F MLP(Xmhaa)+Xmhaa
[0182] Among them, F MLP (·) represents a feedforward perceptron.
[0183] When processing input sequences, the self-attention mechanism often operates on each token at the same scale, and each element is assigned a fixed and identical receptive field. This processing method limits the model's potential to capture features of different scales to a certain extent. Especially in sequences that integrate spectral and elevation data, a single-scale feature extraction mechanism is difficult to fully utilize information. To alleviate this problem, a multi-head axial attention mechanism (MHAA) was designed to replace the multi-head self-attention in Transformer. Its detailed structure is shown in the figure. Figure 5 shown.
[0184] The other steps and parameters are the same as those in Specific Embodiments 1 to 7.
[0185] Specific implementation method 9: This implementation method is different from any one of the specific implementation methods 1 to 8 in that the working process of the local-global feature fusion classification module is:
[0186] X G =F LN (F MLP (X′v′it))
[0187]
[0188] Xout=η·X G +(1-η)·X L
[0189] Among them, F MLP (·) represents a multi-layer perceptron, F LN (·) represents layer normalization; represents a 2D convolutional layer of size 1, F BN (·) represents the batch normalization layer, δ represents the ReLU activation function, Fa″′vg(·) represents the adaptive average pooling layer; η represents the dynamically adjusted weighting coefficient, · represents the Hadamard product; X G represents the feature after global feature processing, X L represents the features after feature Xcnn processing, and Xout represents the output features of the local-global feature fusion classification module.
[0190] The other steps and parameters are the same as those in Specific Embodiments 1 to 8.
[0191] Specific embodiment 10: This embodiment differs from any one of specific embodiments 1 to 9 in that: in step 3, based on the constructed HMAT network model; the hyperspectral image data and the light detection and ranging data to be measured are predicted to obtain the classification result; the specific process is:
[0192] Based on the hyperspectral image data and light detection and ranging data obtained in step 1, the constructed HMAT network model is trained to obtain a trained HMAT network model;
[0193] Based on the trained HMAT network model, the hyperspectral image data and light detection and ranging data pairs to be tested are predicted to obtain the classification results.
[0194] The other steps and parameters are the same as those in Specific Embodiments 1 to 9.
[0195] Embodiment 1:
[0196] Dataset Description:
[0197] Houston2013 dataset: This dataset originated from a mapping project conducted by the National Airborne Laser Mapping Center in June 2012. The project used a compact airborne spectral imager sensor to collect data on the University of Houston campus and its surrounding urban areas. Among them, HSI contains 144 bands with a wavelength range from 0.38 to 1.05 microns, a spatial resolution of 2.5 meters, and a size of 349×1905 pixels. LiDAR data has the same size and resolution as HSI data, but only contains 1 band, providing height information of ground structures. The entire dataset contains 15,029 field-verified samples, which are classified into 15 different categories, aiming to promote the recognition and classification of objects in complex urban environments. Figure 5 (a)-(n) show the pseudo-color image of HSI, the DSM image of LiDAR, and the true label image on this dataset, respectively.
[0198] Trento Dataset: This dataset originates from a rural area south of Trento, Italy. HIS data is collected by the Airborne Hyperspectral Imaging System (AISA) Eagle sensor with 63 spectral bands ranging from 0.42 to 0.99 microns. Correspondingly, LiDAR data is acquired by the Optech Airborne Laser Topography (ALTM) 3100EA sensor. The size of the dataset is 166×600 pixels with a spatial resolution of 1 meter. The entire dataset contains 30,214 samples, which are classified into 6 categories. Figure 6 (a)-(n) show the pseudo-color image of HSI, the DSM image of LiDAR, and the true label image on this dataset, respectively.
[0199] MUUFL dataset: This dataset was collected by CASI-1500 sensor on the campus of University of Southern Mississippi Gulf Park in November 2010. The HSI data covers 72 spectral bands from 0.38 to 1.05 microns, but 8 bands were removed due to noise. The LiDAR data contains two bands, one is the digital elevation model (DEM) and the other is the laser pulse intensity image. The size of the dataset is 325×220 pixels, and it contains 53,687 samples in total, covering 11 different land cover types. Figure 7 (a)-(n) show the pseudo-color image of HSI, the DSM image based on LiDAR, and the true label image on this dataset, respectively.
[0200] The land cover category names, the number of training set samples, and the number of test set samples of the above three datasets are listed in detail in Table 1.
[0201] Table 1 Number of training and testing samples for Houston2013, Trento and MUUFL
[0202]
[0203]
[0204] Experimental setup
[0205] Hardware configuration: The proposed HMAT model is implemented in the PyCharm development environment, which is configured with PyTorch version 1.10.1 and Python version 3.7.0. In addition, in order to ensure the stability and efficiency of the experiment, the CPU uses an Intel Core i9-9900K processor with 128GB of memory and an NVIDIA GeForce RTX 3090 graphics card with 24GB of large-capacity video memory.
[0206] Evaluation indicators: In order to evaluate the classification performance of each method more comprehensively and objectively, multiple evaluation indicators were selected, including: overall accuracy (OA), average accuracy (AA), Kappa coefficient (κ×100), parameter count (Params) and floating-point operations (FLOPs).
[0207] Parameter settings: In order to fairly evaluate the performance of HMAT, all the compared methods use the hyperparameter settings recommended in their respective papers. For the proposed method, we use the Adam optimizer to train the network with a learning rate of 5e-3, and the decay parameter of the learning rate is set to 0.9. In addition, the number of batches, window size, and training rounds are set to 64, 16, and 300, respectively. The optimal size settings of the convolution kernels of each layer in the PIP module on the three datasets are shown in Table 2.
[0208] Table 2 Comparison of convolution kernels of different sizes in PIP
[0209]
[0210] Quantitative experiments: In order to verify the effectiveness of the proposed HMAT, a variety of currently state-of-the-art joint classification methods were selected for comparison with the proposed method, including: CoupledCNN, CALC, S2ENet, GLT, MFT, HCT, S2EFT, GAMF, CrossHL and HFSNet. Among them, CoupledCNN and S2ENet are CNN-based methods. CALC is a method that applies coupled adversarial learning. GAMF uses a graph attention network framework. Other comparison methods and the proposed HMAT combine CNN and Transformer models. In addition, CALC, GLT, HFSNet and the proposed method all adopt a multi-level feature fusion strategy. Tables 3-4 show the classification results of all methods in quantitative experiments on three datasets. Figure 5-Figure 7 Classification plots of all methods on three datasets are shown.
[0211] Although CoupledCNN integrates multimodal data at both the feature level and the decision level, its architecture is very basic and consists only of convolutional layers and fully connected layers, which limits its ability to extract and utilize global feature dependencies, thus affecting the overall classification performance. S2EF directly integrates features of different modal data at the pixel level, but lacks processing of differences between modalities, resulting in poor classification performance and obvious errors in the classification graph. The situation of MFT is similar to that of S2EFT, but it effectively alleviates the problem of data imbalance by differentially processing different modal data, thereby achieving improvement in classification performance. Based on MFT, CrossHL pre-processes HSI and LiDAR data in a targeted manner before fusing them and then sending them to Transformer for feature extraction, which further improves its classification performance. Although S2ENet enhances the interaction between multimodal data with the help of different modal processing modules, its classification performance is also limited by the relatively basic network architecture. This suggests that in complex tasks, model design with a certain depth is crucial to improving performance. HCT transforms multimodal data into unified low-dimensional data through a feature labeling module, thereby achieving relatively good classification performance. GAMF encodes multimodal data into graph data and uses a graph attention mechanism to complete the classification of samples.
[0212] The remaining methods all use multi-level feature processing. Among them, the CALC model introduces coupled adversarial learning to effectively extract high-level semantic information from multimodal data. Its classification performance is better than most comparative methods, highlighting the effectiveness of multi-level feature processing. GLT improves the prediction ability of the model through multi-scale spatial feature processing and unsupervised loss reconstruction. HFSNet uses multi-level feature processing and edge prediction to obtain relatively good classification results. The proposed HMAT, based on multi-level feature processing, further strengthens detail capture through PIP and enhances global context understanding through MHAA, thus achieving the best level on all three datasets.
[0213] Specifically, on the Houston 2013 dataset, the proposed method shows a clear advantage, with a performance improvement of about 0.5% in the three indicators of OA, AA and Kappa compared to the suboptimal method. In addition, among all 15 categories, 7 categories achieved the best performance. Figure 5 In the classification diagram shown in Figure 2, the proposed method also has the best classification results. In particular, in the local detail enlargement diagram, the proposed HMAT retains complete detail information on the "railway" category. Furthermore, the proposed HMAT also performs well on the Trento dataset, as shown in Figure 2. Figure 6 As shown in Figure 2. Not only is it the best overall, but also 4 of the 6 categories have reached the best level. On the MUUFL dataset, HMAT still achieved the best classification results, especially in the AA indicator, which is about 1.5% higher than the second-best method. Figure 7 As shown, in the local detail image of the “building” category, HMAT successfully retains the most complete edge information.
[0214] In contrast, other comparison methods have different degrees of category confusion in the classification diagrams of the three datasets. In particular, methods with simpler structures such as CoupledCNN and S2EFT have obvious shortcomings under the same test conditions. These methods cannot achieve the same level of classification effect when dealing with complex textures and subtle differences, resulting in more serious misclassification in the classification diagrams. The proposed method, with its excellent feature extraction and classification capabilities, has demonstrated excellent classification performance on all three datasets.
[0215] Table 3 Classification performance of different methods on the Houston2013 dataset (the best results are shown in bold)
[0216]
[0217]
[0218] Table 4 Classification performance of different methods on the Trento dataset (the best results are shown in bold)
[0219]
[0220] Table 5 Classification results of different methods on MUUFL dataset (the best results are shown in bold)
[0221]
[0222] Comparison of computational costs: In order to evaluate the proposed method more objectively, in this section, the computational costs of different models are compared using two indicators, Params and FLOPs. The comparison results are shown in Table 6. Although CoupledCNN shows the lowest computational cost on the three datasets, the network structure is too simple, which makes the model unable to understand complex features, and ultimately affects the improvement of classification performance. This situation is also reflected in S2ENet and S2EFT. On the other hand, CALC introduces coupled adversarial learning, which improves performance, but also bears additional unsupervised learning costs, resulting in increased Params and FLOPs. HFSNet significantly improves performance by adopting a multi-level fusion-separation structure, but this design is also accompanied by an increase in computational costs, resulting in higher Params. It is worth noting that GAMF has the highest Params and FLOPs on the three datasets, especially on the Houston 2013 dataset, where its Params are as high as 7.31M. This phenomenon is mainly attributed to its integration of graph neural networks under the batch training framework. This process inevitably requires the construction of a large number of adjacency matrices, which greatly increases the computational burden. Finally, compared with the suboptimal methods, the proposed method shows excellent classification performance while keeping the computational cost at a low level. Compared with the methods with simpler network structures, the proposed HMAT ensures that the network has a good understanding of complex features without making its computational cost too high.
[0223] Table 6 Comparison of parameters and FLOPs of different methods
[0224]
[0225] Ablation experiment: In order to verify the effectiveness of the three modules proposed in HMAT, this part carried out 8 groups of ablation experiments with OA as the main evaluation index. The experimental results are shown in Table 7. First, in the first group of basic control experiments, the network only contains the baseline model structure, that is, only the channel modulation stage and the classification head are integrated. The classification performance at this time is at the lowest level in all test scenarios. Then, from the second to the fourth group of experiments, the three modules proposed in HMAT are introduced one by one. The experimental results show that no matter which module is added alone, the classification performance of the model can be improved to varying degrees. Compared with the baseline model, the performance improvement is obvious, which preliminarily verifies the effectiveness of the independent role of each module. Further, in the fifth to seventh groups of experiments, the combination effect between modules is verified. By combining these modules in pairs, the classification performance of the model is further improved compared with the case of only one module, which shows that different modules can complement each other. Finally, in the eighth group of experiments, HMAT contains all three modules. The experimental data show that the classification performance of the network reaches the highest point at this time, which proves that the three modules proposed can effectively cooperate with each other.
[0226] Table 7 Ablation experiment
[0227]
[0228]
[0229] In order to verify the advantages of the proposed MHAA over MHSA in extracting global features when processing multimodal fusion data, three sets of comparative experiments were designed. The experimental results are shown in the following table. Figure 8 As shown in the figure, it can be observed that regardless of the data set, MHAA is better than MHSA in terms of OA, AA and κ×100. This result proves that the MHAA component has higher adaptability than MHSA for the processing of fused multimodal data.
[0230] t-SNE visualization: This experiment uses the t-SNE algorithm to visualize the features of the proposed method and five comparative methods with better performance on three datasets. The experimental results are shown in the figure below. Figures 9 to 11As shown in the figure. Specifically, on the Houston 2013 dataset, compared with other methods, the proposed HMAT shows obvious advantages, with smaller intra-class distance, larger inter-class distance, and the lowest inter-class confusion. Especially in the 8th "commercial advertisement" category, HMAT shows the slightest dispersion of the same category. Secondly, on the Trento dataset, compared with other methods, HMAT still achieves the best clustering effect. Especially in the fifth "vineyard" category, HMAT's clustering effect is the most complete, accurately clustering the same samples tightly, and effectively distinguishing the heterogeneous samples. In contrast, other comparison methods have different degrees of dispersion of the same samples or confusion of heterogeneous samples. Finally, for the MUUFL dataset with greater classification difficulty, all methods have produced a certain degree of category confusion, but HMAT still maintains the lowest degree of confusion, which further proves the effectiveness of the proposed method.
[0231] Confusion matrix comparison: This experiment compares the confusion matrices of the proposed method and five comparative methods with better performance on three datasets. The experimental results are as follows: Figure 12 to Figure 14 Specifically, on the Houston 2013 dataset, all five comparison methods show obvious category confusion, especially in the 9th category. Fig.12 As shown in (a), even the relatively excellent GLT method, although it achieved a high recognition rate of 97% on this category, also mistakenly classified 5 samples from other categories into it, reflecting the difficulty of identifying this category. In contrast, the proposed HMAT method significantly reduced the confusion level on this category, showing a stronger ability to distinguish. On the Trento dataset, although the confusion matrices of all methods showed good performance, the proposed method still had an advantage. For example, on the second category, except for HMAT and HFSNet, the other methods all had more than 5% category confusion, while HMAT maintained a lower confusion rate. As for the MUUFL dataset, which is more difficult to classify, the confusion matrices of all methods are relatively general. However, compared with other methods, the proposed HMAT still maintains a relatively low degree of confusion. Overall, whether in easy-to-distinguish or extremely challenging datasets, the HMAT method shows better performance than other comparison methods, fully verifying its effectiveness in the HSIC task.
[0232] The present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.
Claims
1. A multi-source remote sensing image classification method based on axial Transformer with hierarchical multimodal feature aggregation, characterized by: The specific process of the method is: Step 1: Obtain hyperspectral image data HSI. Hyperspectral image data HSI is expressed as Get the light detection and ranging data LiDAR, the light detection and ranging data LiDAR is expressed as Where H and W represent the height and width of the input data, respectively, and B′ and B″ represent the number of bands of HSI input and LiDAR input, respectively; represents a real number; Step 2: Construct HMAT network model; The HMAT network model represents a multi-head axial attention Transformer model based on hierarchical multimodal feature aggregation; The HMAT network model includes a channel modulation module, a hierarchical multimodal feature aggregation module, a weighted local feature extraction module, a global feature extraction module, and a local-global feature fusion classification module; The specific working process of the HMAT network model is as follows: The hyperspectral image data HSI is input into the channel modulation module, and the channel modulation module outputs the pre-processed HSI data; The light detection and ranging data LiDAR is input into the channel modulation module, and the channel modulation module outputs the pre-processed LiDAR data; The preprocessed HSI data and the preprocessed LiDAR data are input into the hierarchical multimodal feature aggregation module, and the hierarchical multimodal feature aggregation module outputs data X*; The data X* is input into the weighted local feature extraction module, and the output data of the weighted local feature extraction module is respectively input into the global feature extraction module and the local-global feature fusion classification module; The output data of the global feature extraction module is input into the local-global feature fusion classification module; The output data of the local-global feature fusion classification module is used as the classification result output by the HMAT network model; Step 3: Based on the constructed HMAT network model, the hyperspectral image data and light detection and ranging data pairs are predicted to obtain the classification results.
2. The multi-source remote sensing image classification method based on axial Transformer of hierarchical multimodal feature aggregation according to claim 1 is characterized in that: The specific working process of the channel modulation module is as follows: The HSI data and LiDAR data are preprocessed respectively to obtain the preprocessed HSI data and LiDAR data; they are expressed as: in, represents a 2D convolutional layer of size 3, F BN (·) represents the batch normalization layer, δ represents the ReLU activation function, * represents the convolution operator, W represents the weight of the convolution kernel, b represents the bias, Xin represents the HSI data or LiDAR data, and Xi′n represents the preprocessed HSI data or LiDAR data.
3. The multi-source remote sensing image classification method based on axial Transformer of hierarchical multimodal feature aggregation according to claim 2 is characterized by: The hierarchical multimodal feature aggregation module includes a first 3×3 convolutional layer, a first BN layer, a first ReLU activation function layer, a second 3×3 convolutional layer, a second BN layer, a second ReLU activation function layer, a third 3×3 convolutional layer, a third BN layer, a third ReLU activation function layer, a fourth 3×3 convolutional layer, a fourth BN layer, a fourth ReLU activation function layer, a fifth 3×3 convolutional layer, a fifth BN layer, a fifth ReLU activation function layer, a sixth 3×3 convolutional layer, a sixth BN layer, a sixth ReLU activation function layer. The seventh layer, the seventh 3×3 convolution layer, the seventh BN layer, the seventh ReLU activation function layer, the eighth 3×3 convolution layer, the eighth BN layer, the eighth ReLU activation function layer, the ninth convolution layer, the ninth BN layer, the ninth Sigmoid activation function layer, the first 2×2 average pooling layer, the second 2×2 average pooling layer, the third 2×2 average pooling layer, the fourth 2×2 average pooling layer, the fifth 2×2 average pooling layer, the sixth 2×2 average pooling layer, the first connection layer, the second connection layer, the third connection layer, and the fourth connection layer.
4. The multi-source remote sensing image classification method based on axial Transformer of hierarchical multimodal feature aggregation according to claim 3 is characterized by: The specific working process of the hierarchical multimodal feature aggregation module HMFA is as follows: 1) HSI data is sequentially input into the first 3×3 convolutional layer, the first BN layer, the first ReLU activation function layer, the second 3×3 convolutional layer, the second BN layer, the second ReLU activation function layer, and the first 2×2 average pooling layer. The first 2×2 average pooling layer outputs feature A. Feature A is sequentially input into the third 3×3 convolutional layer, the third BN layer, the third ReLU activation function layer, and the second 2×2 average pooling layer. The second 2×2 average pooling layer outputs feature B. Feature B is sequentially input into the fourth 3×3 convolutional layer, the fourth BN layer, the fourth ReLU activation function layer, and the third 2×2 average pooling layer. The third 2×2 average pooling layer outputs feature C. 2) LiDAR data is sequentially input into the fifth 3×3 convolutional layer, the fifth BN layer, the fifth ReLU activation function layer, the sixth 3×3 convolutional layer, the sixth BN layer, the sixth ReLU activation function layer, and the fourth 2×2 average pooling layer. The fourth 2×2 average pooling layer outputs feature A′. Feature A′ is sequentially input into the seventh 3×3 convolutional layer, the seventh BN layer, the seventh ReLU activation function layer, and the fifth 2×2 average pooling layer. The fifth 2×2 average pooling layer outputs feature B′. Feature B′ is sequentially input into the eighth 3×3 convolutional layer, the eighth BN layer, the eighth ReLU activation function layer, and the sixth 2×2 average pooling layer. The sixth 2×2 average pooling layer outputs feature C′. 3) Feature A and feature A′ are input into the first connection layer, and the first connection layer outputs feature xs; Features B and B′ are input into the second connection layer, and the second connection layer outputs feature xm; Features C and C′ are input into the third connection layer, and the third connection layer outputs feature xd; 4) Flatten and linearly map the feature xs in turn to obtain the feature Flatten and linearly map the feature xm in turn to obtain the feature Flatten and linearly map the feature xd in turn to obtain the feature 5) The features feature feature Input the fourth connection layer, and the fourth connection layer outputs the stacked data It is expressed as: Among them, Fconcat(·) represents cascade; 6) The stacked data Input the ninth convolutional layer, the ninth BN layer, and the ninth Sigmoid activation function layer in sequence; The output features of the ninth Sigmoid activation function layer are deformed and re-transformed into a matrix format to obtain data X*; It is expressed as: in, represents a 2D convolutional layer with a shape of 1×3, F BN (·) represents the batch normalization layer, σ represents the Sigmoid activation function, and F ε (·) represents a deformation operation.
5. The multi-source remote sensing image classification method based on axial Transformer of hierarchical multimodal feature aggregation according to claim 4 is characterized in that: The data X* is input into the weighted local feature extraction module, and the output data of the weighted local feature extraction module is respectively input into the global feature extraction module and the local-global feature fusion classification module; the specific process is: The weighted local feature extraction module is a pyramid-inverted pyramid module; The pyramid-inverted pyramid module includes a pyramid convolution block, a spatial attention block, and an inverse pyramid convolution block; The pyramid convolution block includes a tenth 1×1 convolution layer, an eleventh 3×3 convolution layer, a twelfth 5×5 convolution layer, and a thirteenth 7×7 convolution layer; The tenth 1×1 convolution layer is divided into 1 group, and the size of the convolution kernel of each group is 1×1; The eleventh 3×3 convolutional layer is divided into 2 groups, and the size of the convolution kernel of each group is 3×3; The twelfth 5×5 convolutional layer is divided into 3 groups, and the size of the convolution kernel of each group is 5×5; The thirteenth 7×7 convolutional layer is divided into 4 groups, and the size of the convolution kernel of each group is 7×7; The inverse pyramid convolution block includes a fourteenth 7×7 convolution layer, a fifteenth 5×5 convolution layer, a sixteenth 3×3 convolution layer, and a seventeenth 1×1 convolution layer; The fourteenth 7×7 convolutional layer is divided into 4 groups, and the size of the convolution kernel of each group is 7×7; The fifteenth 5×5 convolutional layer is divided into 3 groups, and the size of the convolution kernel of each group is 5×5; The sixteenth 3×3 convolutional layer is divided into 2 groups, and the size of the convolution kernel of each group is 3×3; The seventeenth 1×1 convolutional layer is divided into 1 group, and the size of the convolution kernel of each group is 1×1; The working process of the pyramid-inverted pyramid module is as follows: Data X* is input into the pyramid convolution block, and the pyramid convolution block outputs feature X P conv; Feature X P conv inputs spatial attention, and spatial attention outputs data X′j; Data X′j is added element by element to data X* to obtain feature X′ P ′conv; The feature X′ P ′conv is reshaped into Xvit; The feature X′ P ′conv is fed into the inverted pyramid convolution block, and the inverted pyramid convolution block outputs the feature Xcnn; Feature Xvit input global feature extraction module; The inverted pyramid convolution block outputs the feature Xcnn which is input into the local-global feature fusion classification module.
6. The multi-source remote sensing image classification method based on axial Transformer of hierarchical multimodal feature aggregation according to claim 5 is characterized by: The data X* is input into the pyramid convolution block, and the pyramid convolution block outputs the feature X P conv; the specific process is: The data X* is input into the tenth 1×1 convolution layer which is divided into a group, and the tenth 1×1 convolution layer outputs the feature D; The data X* is input into the eleventh 3×3 convolutional layer divided into two groups, and the eleventh 3×3 convolutional layer outputs the feature E; The data X* is input into the twelfth 5×5 convolutional layer divided into three groups, and the twelfth 5×5 convolutional layer outputs the feature F; The data X* is input into the thirteenth 7×7 convolutional layer divided into four groups, and the thirteenth 7×7 convolutional layer outputs the feature G; Cascade feature D, feature E, feature F, and feature G to get feature X P conv; It is expressed as: in, represents a 2D convolution kernel of size 5, represents a 2D convolution kernel of size 7, represents a 2D convolution kernel of size 3, represents a 2D convolution kernel of size 1, X P conv represents the output of pyramid convolution.
7. The multi-source remote sensing image classification method based on axial Transformer of hierarchical multimodal feature aggregation according to claim 6 is characterized by: The feature X′ P ′conv is sent to the inverted pyramid convolution block, and the inverted pyramid convolution block outputs the feature Xcnn; the specific process is: Data X′ P ′conv input is divided into four groups of four 7×7 convolutional layers, and the fourteenth 7×7 convolutional layer outputs features D′; Data X′ P 'conv input is divided into three groups of fifteenth 5×5 convolutional layer, and the fifteenth 5×5 convolutional layer outputs the feature E'; Data X′ P ′conv input is divided into two groups of sixteenth 3×3 convolutional layers, and the sixteenth 3×3 convolutional layer outputs features F′; Data X′ P ′conv input is divided into a group of seventeenth 1×1 convolutional layers, and the seventeenth 1×1 convolutional layer outputs the feature Xcnn; It is expressed as:
8. The multi-source remote sensing image classification method based on axial Transformer of hierarchical multimodal feature aggregation according to claim 7 is characterized in that: The feature Xvit is input into the global feature extraction module; the specific process is: The global feature extraction module is the multi-head axial attention mechanism module MHAA; For the input sequence Add learnable tokens and integrate positional encoding information to obtain information The information X′vit is mapped through two parallel linear layers to map the query vector Xq and the key vector Xk, that is, n′ represents the length of the spatial sequence after adding the learnable token; n represents the length of the spatial sequence, and c represents the length of the channel sequence; A separate linear layer is used to map the information X′vit to a value vector Xv, that is, This process is represented by X′vit=Fconcat(Xcls,Xvit)+X PE Among them, X PE represents positional encoding, Xcls represents learnable tokens, Fl′inear(·) represents a linear layer with weight sharing; Fconcat represents concatenation; Flinear(·) represents a linear layer; Reshape the query vector Xq and key vector Xk into matrix shape, i.e. h represents the length of the channel sequence in the row direction, and w represents the length of the channel sequence in the column direction; The value vector Xv is split into data of multiple attention heads to form d represents the length of the channel sequence; Then, the column pooling operation is performed on the query vector Xq; The row pooling operation is performed on the key vector Xk; Use the pooled results to build an attention map, and convert the column pooling and row pooling results into normalized attention scores through scaling operations and Softmax functions; Expressed as in, Fa′vg(·) represents the average pooling in the column direction, F ε (·) represents the deformation operation, dq represents the dimension of Xq, Mcol represents the attention map in the column direction, and Fsoftmax(·) represents the Softmax function; Fa″vg(·) represents the average pooling in the row direction, dk represents the dimension of Xk, and Mrow represents the attention map in the row direction; The value vector Xv is matrix multiplied with the attention maps of different scales Mcol and Mrow respectively; the outputs of different attention heads are integrated, the weighted results of the two branches are reshaped into sequences, and the two sequences are merged by matrix addition; layer normalization is used to stabilize the distribution of data and fuse the shallow features from the skip connection to obtain the feature This process is represented as Xmhaa=F LN (Flinear(F ε (Xv×Mrow)+F ε (Xv×Mcol)))+Xv′it Among them, F LN (·) represents layer normalization; The feature Xmhaa is sent to the feedforward perceptron to perform nonlinear transformation on the vector of each position; layer normalization is used to stabilize the distribution of data and fuse the shallow features from the jump connection to obtain the feature; repeat δ times to extract the global feature X′v′it; expressed as X′v′it=F LN (F MLP (Xmhaa)+Xmhaa Among them, F MLP (·) represents a feedforward perceptron.
9. The multi-source remote sensing image classification method based on axial Transformer of hierarchical multimodal feature aggregation according to claim 8, characterized in that: The working process of the local-global feature fusion classification module is as follows: X G =F LN (F MLP (X′v′it)) Xout=η·X G +(1-n)·X L Among them, F MLP (·) represents a multi-layer perceptron, F LN (·) represents layer normalization; represents a 2D convolutional layer of size 1, F BN (·) represents the batch normalization layer, δ represents the ReLU activation function, and Fa″′vg(·) represents the adaptive average pooling layer; η represents the dynamically adjusted weighting coefficient, · represents the Hadamard product; X G represents the feature after global feature processing, X L represents the features after feature Xcnn processing, and Xout represents the output features of the local-global feature fusion classification module.
10. The multi-source remote sensing image classification method based on axial Transformer of hierarchical multimodal feature aggregation according to claim 9, characterized in that: In the step 3, the hyperspectral image data and the light detection and ranging data to be measured are predicted based on the constructed HMAT network model to obtain a classification result; The specific process is: Based on the hyperspectral image data and light detection and ranging data obtained in step 1, the constructed HMAT network model is trained to obtain a trained HMAT network model; Based on the trained HMAT network model, the hyperspectral image data and light detection and ranging data pairs to be tested are predicted to obtain the classification results.
Citation Information
Cited By
Directional image segmentation method and system based on hybrid model
CN120563842A
Retinal blood vessel image segmentation method, device, equipment and medium
CN121053394A