Construction method of multi-scale double-attention network for hyperspectral image classification
The multi-scale dual-attention network (MSDANet) addresses computational and attention imbalance issues in high-spectral image classification by integrating multi-scale feature extraction and dual attention mechanisms, enhancing feature representation and reducing complexity for efficient high-precision classification.
Patent Information
- Application Number
- CN202510422992.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-15
AI Technical Summary
The existing hyperspectral image classification methods have high computational overhead and poor adaptability when processing high-dimensional data, and have strong limitations in spectral-spatial feature fusion, making it difficult to effectively capture complex spectral-spatial relationships.
Multi-scale dual attention network (MSDANet) is adopted to combine parallel multi-branch structures with hollow convolution, expand the receptive field range, design a dual attention mechanism and lightweight network to achieve multi-scale feature extraction and feature fusion.
It significantly improves the model's adaptive expression ability of land objects of different scales, improves classification accuracy and calculation efficiency, and can pay attention to local details and global semantic information at the same time, reducing the computational complexity.
Smart Images

Figure CN120318652A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hyperspectral image classification methods, and in particular to a method for constructing a multi-scale dual-attention network for hyperspectral image classification. Background Art
[0002] Hyperspectral remote sensing images play an increasingly important role in earth observation, environmental monitoring, precision agriculture, mineral exploration and other fields because they can provide rich spectral information and fine spatial details. Compared with traditional multispectral images, hyperspectral images usually contain hundreds of continuous narrow-band spectral bands, which enable them to more finely characterize and distinguish the spectral characteristics of objects. This high-dimensional spectral information provides richer discrimination basis for target recognition and classification tasks, but it also brings challenges such as "dimensionality disaster" and data redundancy. Therefore, how to effectively use these high-dimensional data for accurate classification has always been a core scientific problem that needs to be solved in the field of remote sensing.
[0003] In the early research of hyperspectral image classification, traditional methods mainly rely on manually designed feature extractors and statistical learning models, such as support vector machines (SVM) and random forests (RF). These methods perform well in simple scenarios, but often fail to meet the needs of complex practical application scenarios. Their limitations are mainly reflected in the following aspects: (1) It is difficult for manually designed feature extractors to fully mine and utilize the rich information contained in hyperspectral data; (2) Traditional methods have limited modeling capabilities for spectral-spatial joint features and cannot effectively capture the complex correlations in spectral and spatial dimensions; (3) The computational efficiency is low when processing large-scale high-dimensional data, which is difficult to meet the needs of practical applications; (4) There is a lack of sufficient robustness to spectral variability caused by factors such as noise in the data, changes in illumination, and atmospheric interference.
[0004] With the rapid development of deep learning technology, hyperspectral image classification methods based on deep neural networks have made significant progress. Early research mainly focused on improving the convolutional neural network (CNN) structure to enhance feature extraction and classification performance. For example, Ye et al. proposed the IPCEHRIC method, which innovatively combined the advantages of the enhanced particle swarm optimization (CWLPSO) algorithm, convolutional neural network and extreme learning machine (ELM), extracted deep features by optimizing CNN parameters, and used ELM to achieve efficient classification. This idea of combining the advantages of multiple algorithms provides important inspiration for subsequent research.
[0005] On this basis, researchers further explored the joint extraction and fusion methods of spectral-spatial features. Vaddi et al. proposed a CNN-based classification framework. First, the HSI data was normalized, then probabilistic principal component analysis (PPCA) and Gabor filtering were used to extract spectral and spatial information respectively, and finally feature fusion and classification were achieved through CNN. Kavitha et al. proposed an enhanced convolutional neural network (e-CNN), which effectively alleviated the overfitting problem of traditional deep CNNs through an innovative inter-layer feature merging mechanism and a layer-by-layer spectral-spatial feature extraction strategy.
[0006] To further improve the feature extraction ability of the model, researchers began to explore graph neural networks and multi-feature fusion techniques. The multi-feature fusion network (MFGCN) proposed by Ding et al. adopted a dual-branch structure of multi-scale graph convolutional network and multi-scale CNN. It achieved efficient utilization of computing resources and solved the problem of insufficient labels through multi-scale superpixel-based GCN, while using multi-scale CNN to extract pixel-level local features. This method also innovatively introduced one-dimensional CNN to process the spectral features of superpixel nodes, and realized complementary fusion of multi-scale features through concatenation operations. In a similar research direction, Diakite et al. proposed a 3D2D-CNN combination method. By introducing a 3D fast learning module containing depthwise separable convolution blocks and fast convolution blocks, it effectively reduced the training time while improving the classification accuracy.
[0007] However, these CNN-based methods still face two main challenges. The first is the computational overhead and adaptability issues. Although many methods (such as 3D-CNN, MFGCN, 3D2D-CNN, etc.) have achieved significant improvements in classification performance, they have a large computational overhead when dealing with high-dimensional spectral data. Especially when using larger convolution kernels or multi-layer structures, it not only leads to a significant increase in memory and computational burden, but also is difficult to flexibly adapt to the feature extraction requirements of different-scale targets. The second is the limitation of spectral-spatial feature fusion. Existing methods often rely on fixed structures and are difficult to effectively capture complex spectral-spatial relationships. For example, although MFGCN adopts a dual-branch structure of superpixel-based GCN and CNN, due to the limited local receptive field of CNN, it may ignore important global spectral-spatial information.
[0008] The introduction of the attention mechanism provides a new idea for solving the above problems. Liu et al. deeply analyzed the intrinsic characteristics of HSI, and proposed two spectral-spatial feature extraction principles based on pixel-level classification and the definition of spatial information, and designed an innovative Scaled Dot-Product Central Attention (SDPCA) module accordingly. This module can effectively extract spectral-spatial information from the central pixel and its similar pixels, and significantly improve the model's expressive ability through the Central Attention Network (CAN). Zheng et al. proposed a Rotation-Invariant Attention Network (RIAN) to address the problem that traditional methods are sensitive to the rotation of input images. This network achieved rotation-invariant spectral-spatial feature extraction by designing a Central Spectral Attention (CSpeA) module and a Rectified Spatial Attention (RSpaA) module.
[0009] In terms of multi-scale feature extraction, the Multi-scale and Cross-layer Attention Learning (MCAL) network proposed by Xu et al. pioneered the solution to the problem that existing transformer networks fail to effectively explore the complex local ground object structures at different scales. By designing a Multi-scale Feature Extraction (MSFE) module and a Cross-layer Feature Fusion (CLFF) module, this method achieved effective modeling of local spatial context and adaptive fusion of features. At the same time, by introducing a Spectral Attention Module (SAM) before the MSFE hierarchy, the joint expression of spatial context and spectral information was further enhanced.
[0010] The issue of model robustness has also gradually attracted the attention of researchers. Xu et al. conducted in-depth research on the problem of adversarial attacks in hyperspectral image classification. Their experiments showed that although existing adversarial attack research mainly focuses on the RGB domain, adversarial samples also exist in the hyperspectral domain. Notably, these generated adversarial images are hardly distinguishable from the original hyperspectral data in the human visual system, but can significantly affect the prediction results of deep learning models. To address this challenge, they proposed a Self-Attention Context Network (SACNet), effectively improving the adversarial robustness of the model.
[0011] The application of Graph Neural Networks (GNNs) has brought a new research direction to hyperspectral image classification. Ding et al. observed the deficiencies of existing GNN methods in terms of information description efficiency, time consumption, and anti-noise robustness, and proposed a Multi-scale Receptive Field Graph Attention Neural Network (MRGAT). This method first used superpixel segmentation to extract local spatial features, and innovatively designed a spectral conversion mechanism of two-layer one-dimensional CNN to achieve automatic spectral feature extraction of superpixels. By introducing the edge features of the graph attention network and multi-scale receptive field GAT, this method effectively achieved the extraction of local-global neighboring node features and the fusion of multi-receptive field features.
[0012] However, there are still two main problems with existing attention mechanisms: First is the problem of balancing attention among different dimensions. Although many studies (such as CSDA, SAM, MANet, etc.) have introduced attention mechanisms in different dimensions such as spectral, spatial, and channel, there is still insufficient adjustment of multi-dimensional feature weights. Especially in the process of fusing spectral and spatial information, it is easy to over-rely on features of a single dimension and ignore the contributions of other dimensions. Second is the adaptability and complexity of multi-scale feature extraction. Although methods such as MCAL and morphFormer have made significant progress in multi-scale feature extraction, these methods often rely on complex network structures, resulting in a significant increase in computational complexity and facing challenges in practical application scenarios with limited resources.
[0013] The success of Vision Transformer has provided new research ideas for hyperspectral image classification. The morphFormer proposed by Roy et al. aims at the problem that existing ViTs fail to effectively utilize spatio-spectral features, and designs a learnable spectral-spatial morphology network. By combining morphological convolution operations and attention mechanisms, it significantly enhances the interaction of structural and shape information between HSI tokens. Hong et al. re-examine the hyperspectral image classification problem from the perspective of sequences, propose the SpectralFormer network, and innovatively realize the learning of local sequence information between adjacent bands, and effectively reduce the loss in the information transmission process through cross-layer skip connections and "soft" residual fusion.
[0014] In terms of lightweight network design, Zhang et al. proposed the Lightweight Transformer (LiT) network. By designing two innovative modules, namely channel lightweight multi-head self-attention (CLMSA) and position lightweight multi-head self-attention (PLMSA), it significantly reduces the computational burden while maintaining the long-range dependence modeling ability. This method also effectively solves the overfitting problem when the training samples are insufficient through a hierarchical design that uses convolutional blocks to extract local information in the early layers and transformers to capture long-range dependencies in the deep layers. Notably, the controlled multi-class stratification (CMS) sampling strategy they proposed not only ensures the appropriate scale and sampling balance of the input data, but also effectively reduces the overlap of feature extraction regions between training and test samples.
[0015] Sun et al. proposed the Spectral-Spatial Feature Token Transformer (SSFTT) from the perspective of feature tokenization, aiming to capture both spectral-spatial features and high-level semantic features. This method first extracts shallow spectral and spatial features through 3D convolutional layers and 2D convolutional layers respectively, then uses a Gaussian weighted feature tokenizer to achieve feature transformation, and finally completes feature representation and learning through a transformer encoder module. This progressive feature processing strategy effectively overcomes the limitations of traditional CNN methods in deep semantic feature extraction.
[0016] To better combine the advantages of CNN and ViT, Zhang et al. proposed the Convolution-Transformer Mixer (CTMixer). This framework first uses parallel residual modules to capture local spectral-spatial features, and then realizes the collaborative extraction of local and global features by designing a dual-branch structure containing CNN and transformers. In particular, the local-global multi-head self-attention mechanism (MHSA) they proposed achieves an elegant fusion of the two paradigms by introducing convolutional operations.
[0017] In the boundary region classification problem, the Spectral-Spatial Transformer (SST-M) proposed by Bai et al. provides a new solution idea. Aiming at the problem that traditional methods are easily interfered by irrelevant information around the target pixel, SST-M innovatively combines the spatial attention mechanism and the spectral feature extraction model, integrates spatial position information through the global receptive field of the transformer, and uses the spatial sequence attention model to enhance the effective information while suppressing the interference information. In addition, the mask prediction model they designed further improves the learning accuracy of pixel features and spatial distributions.
[0018] However, existing methods still face two main challenges: The first is the balance problem between local and global features. Although methods such as SpectralFormer, LiT, and CTMixer attempt to combine the advantages of CNN and Transformer, it is still difficult to achieve the optimal balance between local details and global dependencies in specific scenarios. The second is the computational efficiency problem. As the network structure becomes more complex, especially after introducing multi-layer cross-layer connections and multi-scale feature extraction, the computational burden and memory consumption of the model increase significantly, which may become a bottleneck in real-time application scenarios. Summary of the Invention
[0019] To solve the defects in the prior art, this application proposes a construction method for a multi-scale dual-attention network (Multi-Scale Dual-Attention Network, MSDANet) for hyperspectral image classification. The core idea of this method is to achieve efficient processing and accurate classification of hyperspectral data by organically combining multi-scale feature extraction and dual-attention mechanisms.
[0020] A construction method for a multi-scale dual-attention network for hyperspectral image classification includes the following steps:
[0021] Step S1: Input representation and feature mapping;
[0022] Step S2: Extract dilated convolution enhanced features from the input of Step S1;
[0023] Step S3: Optimize depthwise separable convolution;
[0024] Step S4: Establish a multi-scale feature learning mechanism;
[0025] Step S5: Design a dual attention mechanism;
[0026] Step S6: Establish a classification decision;
[0027] Step S7: Optimize the loss function.
[0028] Compared with the prior art, the technical solution adopted by the present invention can be mainly summarized into the following four aspects:
[0029] (1) A novel multi-scale feature extraction module is proposed. This module expands the receptive field range by combining a parallel multi-branch structure with dilated convolution, and realizes the efficient fusion of multi-scale features through a feature pyramid structure, significantly improving the model's adaptive expression ability for ground object targets of different scales.
[0030] (2) An innovative dual attention mechanism is designed. This mechanism constructs a dual-branch structure that coordinates channel attention and spatial attention to achieve precise modeling of the correlation between spectral bands. At the same time, a frequency-domain attention branch is introduced to expand the feature expression dimension, enhancing the model's ability to capture complex spectral features.
[0031] (3) A lightweight network design strategy is proposed. By introducing depthwise separable convolution to replace the standard convolution operation and designing a lightweight attention calculation module, the computational complexity is significantly reduced while maintaining the model performance, improving the network's deployment efficiency in practical application scenarios.
[0032] (4) An adaptive feature fusion mechanism is constructed. This mechanism optimizes the gradient propagation path by introducing residual connections and designs learnable weights to achieve dynamic fusion of features, effectively enhancing the model's feature expression ability and training stability.
[0033] Enable MSDANet to effectively process the complex features of hyperspectral images and achieve high-precision classification results. In particular, the organic combination of multi-scale feature extraction and dual attention mechanism enables the network to simultaneously focus on local details and global semantic information, significantly improving the feature expression ability and classification accuracy. Brief Description of the Drawings
[0034] Figure 1 is a schematic flowchart of the present invention.
[0035] Figure 2 is the classification map of MSDANet of the present invention on the Indian Pines dataset.
[0036] Figure 3It is the classification diagram of the MSDANet of the present invention on the Pavia University dataset.
[0037] Figure 4 It is the classification diagram of the MSDANet of the present invention on the Kennedy Space Center dataset.
[0038] Figure 5 It is the simulation diagram of the training accuracy and validation accuracy obtained by the MSDANet of the present invention for the Indian Pines dataset.
[0039] Figure 6 It is the simulation diagram of the training loss and validation loss obtained by the MSDANet of the present invention for the Indian Pines dataset.
[0040] Figure 7 It is the simulation diagram of the training accuracy and validation accuracy obtained by the MSDANet of the present invention for the Pavia University dataset.
[0041] Figure 8 It is the simulation diagram of the training loss and validation loss obtained by the MSDANet of the present invention for the Pavia University dataset.
[0042] Figure 9 It is the simulation diagram of the training accuracy and validation accuracy obtained by the MSDANet of the present invention for the Kennedy Space Center dataset.
[0043] Figure 10 It is the simulation diagram of the training loss and validation loss obtained by the MSDANet of the present invention for the Kennedy Space Center dataset. Detailed implementation manners
[0044] Specifically, the main innovation points of the present invention include the following four aspects:
[0045] First of all, the present invention designs an innovative multi-scale feature extraction module. This module adopts a parallel multi-branch structure, which can extract feature information at different scales simultaneously. By introducing dilated convolution to expand the receptive field, the adaptability of the model to ground object targets at different scales is significantly enhanced. At the same time, a feature pyramid structure is adopted to achieve effective fusion of multi-scale features, ensuring the full utilization of information at different scales, thereby improving the recognition ability of the model for complex ground object targets.
[0046] Secondly, the present invention proposes a dual attention mechanism, which consists of two collaborating branches: channel attention and spatial attention. The channel attention branch adaptively adjusts the importance weights of each band by precisely modeling the correlations between different spectral bands; the spatial attention branch focuses on the key regions in the spatial domain, enhancing the feature representation of the target regions. The collaborative effect of these two branches not only improves the model's feature selection ability but also achieves the optimal fusion of spectral-spatial information.
[0047] Thirdly, the present invention introduces depthwise separable convolutions to replace traditional standard convolution operations and designs a lightweight attention module. This innovative network design strategy significantly reduces the computational complexity and the number of parameters of the model, enabling the model to have better practicality while maintaining high performance. This improvement is of great significance for the real-time processing requirements in actual hyperspectral image application scenarios.
[0048] Fourthly, the present invention designs an innovative feature fusion strategy. By introducing a residual connection mechanism, the problem of gradient vanishing in deep networks is effectively alleviated. At the same time, an adaptive fusion module is innovatively introduced, which can dynamically adjust the fusion weights according to the importance of different features. This design not only enhances the model's expressive ability but also significantly improves the stability of the training process.
[0049] The Multi-Scale Dual Attention Network (MSDANet) is an innovative deep learning framework for hyperspectral image classification, the core of which is the organic combination of multi-scale feature learning and dual attention mechanism. Aiming at the characteristics of hyperspectral images, the network fully considers the extraction of spatial-spectral joint features and the recognition requirements of targets at different scales in the architecture design, and at the same time enhances the expressive ability of key features through the attention mechanism.
[0050] As Figure 1 shown, the overall process of constructing MSDANet is presented, mainly including four key stages: input processing, multi-scale feature extraction, dual attention enhancement, and output classification. Among them, the multi-scale feature extraction module captures features at different scales through five parallel branches, and the dual attention mechanism collaboratively completes the feature enhancement in the channel dimension and the spatial dimension. Specifically:
[0051] A method for constructing a multi-scale dual attention network for hyperspectral image classification includes the following steps:
[0052] Step S1: Input representation and feature mapping:
[0053] For hyperspectral image data, we first represent it in the form of a three-dimensional tensor:
[0054]
[0055] Here, H and W represent the height and width dimensions of the image, respectively, and C represents the number of spectral channels. This representation preserves the complete spatial-spectral information of the hyperspectral data. Based on this input, the network first performs initial feature extraction:
[0056] F init = σ(BN(Conv(X))) (2)
[0057] This initial feature extraction process consists of three key steps:
[0058] The convolutional operation Conv(·) is used to extract local spatial-spectral features;
[0059] Batch normalization BN(·) is used to stabilize the training process;
[0060] The ReLU activation function σ(·) introduces the ability of non-linear transformation.
[0061] Step S2: Dilated convolution for enhanced feature extraction:
[0062] To expand the receptive field and maintain the feature resolution, the network uses dilated convolution for feature extraction:
[0063] F1 = σ(BN(Conv d (F init ))) (3)
[0064] The dilated convolution Conv_d significantly expands the receptive field without increasing the number of parameters by inserting "holes" in the convolutional kernel. This property enables the network to capture a larger range of spatial dependencies, which is crucial for understanding the scene context. Subsequently, feature enhancement is performed through a dual attention module:
[0065] F′1 = DA(F1) + F1 (4)
[0066] The residual connection introduced here not only helps with gradient backpropagation but also preserves the original feature information, enhancing the expressive power of the model.
[0067] Step S3: Depthwise separable convolution optimization:
[0068] To improve the computational efficiency, the network adopts a depthwise separable convolution structure, which decomposes the standard convolution into two steps: depthwise convolution and pointwise convolution:
[0069] Depthwise convolution stage:
[0070] F d = DWConv(F′1) (5)
[0071] Pointwise convolution stage:
[0072] F2 = PWConv(Fd ) (6)
[0073] This decomposition significantly reduces the computational amount, while maintaining the spatial feature extraction ability through depth convolution and realizing information interaction between channels through point convolution.
[0074] Step S4: Multi-scale feature learning mechanism:
[0075] This module adopts an innovative "pyramid + parallel" hybrid architecture. First, initial feature extraction is performed through a 3×3 convolutional layer to generate a 64-channel feature map, and a residual connection is designed in this step to maintain the flow of original information. Subsequently, five parallel branches with different receptive fields are designed:
[0076] (1) Pixel-level feature branch: Adopt 1×1 convolution to specifically extract fine-grained features in the spectral dimension, which is particularly important for distinguishing spectrally similar land cover classes.
[0077] F p =Conv 1×1 (F2) (7)
[0078] (2) Local spatial feature branch: Use 3×3 convolution to focus on local texture and edge features, which helps to capture the basic shape information of land covers.
[0079] F l =Conv 3×3 (F2) (8)
[0080] (3) Medium-scale feature branch: Achieve an equivalent 5×5 receptive field through two consecutive 3×3 convolutions to extract a larger range of spatial context information and enhance the understanding of the target structure.
[0081] F m =Conv 5×5 (F2) (9)
[0082] (4) Dilated convolution branch: Innovatively introduce dilated convolution with a dilation rate of 2 to obtain a larger receptive field without increasing the number of parameters, and effectively capture long-range spatial dependence relationships.
[0083] F d =Conv dilation (F2) (10)
[0084] (5) Non-local attention branch: Perform non-local feature enhancement on the output of the 1×1 convolution branch to further improve the global representation ability of features.
[0085] F g =NonLocal(F2) (11)
[0086] These multi-scale features are then integrated through an adaptive fusion module:
[0087] F ms = Conv 1×1 ([F p ,F l ,F m ,F d ,F g ) (12)
[0088] Adaptive fusion ensures the effective combination of features at different scales and enhances the model's ability to recognize targets at different scales.
[0089] Step S5: Design a dual attention mechanism:
[0090] 1. Channel attention branch
[0091] First, calculate the channel statistical features:
[0092]
[0093] These channel statistical features capture the global response of each channel. Subsequently, learn the channel weights through a multi-layer perceptron:
[0094] w c = σ(MLP(z c )) (14)
[0095] Finally, generate the channel-enhanced features:
[0096]
[0097] 2. Spatial attention branch
[0098] Combine the information of max pooling and average pooling:
[0099] z s = [AvgPool(F); MaxPool(F)] (16)
[0100] Generate a spatial attention map through a convolutional layer:
[0101] w s = σ(Conv 7×7 (z s )) (17)
[0102] Apply spatial attention:
[0103]
[0104] 3. Adaptive fusion of attention features
[0105] Dynamically fuse two attention features through learnable weights:
[0106] F DA = αF ca + βF sa (19)
[0107] Here, α and β are learnable weight parameters that can adaptively adjust the importance of the two attention mechanisms according to different inputs.
[0108] Step S6: Classification decision:
[0109] The final classification process is completed through the following steps:
[0110] F final = Conv 1×1 (F ms ) (20)
[0111] Y = Softmax(F final ) (21)
[0112] Step S7: Optimize the loss function:
[0113] Use the cross-entropy loss function for model training:
[0114]
[0115] where: N is the number of samples; K is the number of classes; y ij is the true label; is the predicted probability;
[0116] The optimization process uses the AdamW algorithm and introduces a cosine annealing learning rate scheduling strategy:
[0117]
[0118] This learning rate scheduling scheme can achieve more refined parameter adjustment in the later stage of training, which helps the model converge to a better solution.
[0119] To comprehensively verify the effectiveness of the proposed method, this paper conducts systematic experimental studies on three representative hyperspectral datasets: (1) Indian Pines dataset, which represents a typical agricultural area, contains multiple crop categories, has high inter-class similarity, and can be used to verify the fine classification ability of the model; (2) Pavia University dataset, which represents a complex urban environment, contains multiple artificial ground objects, has a complex spatial structure, and can be used to verify the spatial feature extraction ability of the model; (3) Kennedy Space Center dataset, which represents a natural environment, contains multiple natural ground object categories, has large spectral variability, and can be used to verify the robustness of the model.
[0120] Through systematic comparative experiments with a variety of state-of-the-art methods, including traditional CNN methods, attention mechanism-based methods, Transformer-based methods, etc., this paper deeply analyzes the advantages of the proposed method. At the same time, through detailed ablation experiments, the necessity and effectiveness of each innovative component are verified. The experimental results show that MSDANet has achieved significant improvements in multiple key indicators such as classification accuracy, computational efficiency, and generalization ability, demonstrating good application value.
[0121] I. Experimental Datasets and Experimental Verification
[0122] 1. Hyperspectral Data Description
[0123] To comprehensively evaluate the performance of the proposed method, this paper selects three representative publicly available hyperspectral image datasets for experimental verification: Indian Pines, Pavia University, and Kennedy Space Center (KSC). These three datasets represent different application scenarios such as agriculture, urban, and natural environments, have different spatial resolutions, spectral characteristics, and ground object categories, and can comprehensively evaluate the classification performance, generalization ability, and robustness of the algorithm in different scenarios. By conducting systematic experimental verification on these representative datasets, the practical application value of the proposed method can be objectively and comprehensively evaluated. All experiments are carried out on a high-performance workstation configured with an Intel Core i9-14900K processor (24 cores and 32 threads), 128GB DDR5 memory, and an NVIDIA RTX 4090 24GB graphics card. This platform has powerful parallel computing capabilities and sufficient video memory capacity, which can meet the needs of large-scale hyperspectral image processing and deep learning model training. The specific characteristics and experimental settings of each dataset will be described in detail in the experimental section.
[0124] 2. Experimental Methods
[0125] 2.1 Experiments Based on the Indian Pines Dataset:
[0126] To comprehensively verify the performance of the proposed method, we first conducted a detailed experimental study on the Indian Pines dataset.
[0127] Indian Pines Dataset: This is a typical agricultural remote sensing image dataset collected by the AVIRIS sensor in the northwest of Indiana. The data contains 145×145 pixels and a total of 200 spectral bands, covering a wavelength range of 0.4 - 2.5μm. This dataset contains 16 land cover classes, mainly different crop types, such as farmlands with different tillage methods like corn and soybean, as well as non - agricultural land covers like trees and buildings. In the experiment, we used 10% of the samples for training and 90% for testing. This small - sample setting is closer to the actual application scenario and can better verify the feature - learning ability of the algorithm.
[0128] Table 1 Comparison of Classification Accuracy of Different Methods Based on the Indian Pines Dataset (expressed as a percentage)
[0129]
[0130]
[0131]
[0132] From the detailed experimental results of the Indian Pines dataset, MSDANet demonstrated excellent classification performance and outstanding feature - discrimination ability. The experimental results on this dataset showed that the MSDANet method achieved an overall classification accuracy (OA) of 96.07% on this dataset, with an average accuracy (AA) of 95.58% and a Kappa coefficient of 0.9552. In terms of classification details, the recognition accuracies for the corn classes (Corn - notill, Corn - mintill, Corn) reached 98.44%, 94.38%, and 92.49% respectively, and for the soybean classes (Soybean - notill, Soybean - mintill, Soybean - clean) reached 92.46%, 96.52%, and 92.13% respectively, indicating that the method can better distinguish different tillage methods of crops. At the same time, when dealing with smaller - area classes such as oats (Oats) and grass - pasture - mowed (Grass - pasture - mowed), recognition accuracies of 94.44% and 100.00% were also achieved, indicating that the method has good classification ability for small - sample classes.
[0133] 2.2 Experiments based on the Pavia University dataset:
[0134] To further verify the adaptability and performance of MSDANet in complex urban environments, we selected the representative Pavia University dataset for experiments. This is a hyperspectral dataset of urban scenes, obtained by imaging the University of Pavia in Italy and its surrounding areas with the ROSIS sensor. The data size is 610×340 pixels, containing 103 spectral bands, covering the wavelength range of 0.43 - 0.86μm. The dataset contains 9 urban categories, including asphalt, grassland, gravel, trees, metal sheets, bare soil, bitumen, bricks, and shadows, etc. 3% of the samples were used for training and 97% for testing in the experiment. This extremely small proportion of training samples can fully test the classification ability of the algorithm in complex urban scenarios.
[0135] Table 2 Comparison of classification accuracies of different methods based on the Pavia University dataset (expressed as percentages)
[0136]
[0137] In the experiments on the PaviaU dataset, MSDANet demonstrated significant performance advantages in the classification task of complex urban environments. On this urban scene dataset, the algorithm achieved an overall classification accuracy (OA) of 96.85%, an average accuracy (AA) of 96.16%, and a Kappa coefficient of 0.9581. Specifically for each category, the recognition accuracy of asphalt reached 94.14%, that of meadows reached 99.91%, and that of trees was 93.10%. When recognizing the two artificial surface categories of metal sheets and bare soil, the accuracies reached 100.00% and 92.93% respectively, showing a strong ability to distinguish different material features. For easily confused categories such as asphalt and bitumen, the recognition accuracies also reached 94.14% and 97.91% respectively.
[0138] 2.3 Experiments based on the Kennedy Space Center dataset:
[0139] To comprehensively evaluate the performance of MSDANet in natural ecosystem classification, we selected the Kennedy Space Center dataset with complex natural scene characteristics as the third validation dataset. This dataset was acquired by the NASA AVIRIS instrument over the Kennedy Space Center in Florida and contains 512×614 pixels and 176 spectral bands (0.4 - 2.5 μm). The dataset includes 13 natural vegetation categories such as scrub, marsh, hardwood, etc., representing typical natural ecosystems. In the experiment, 8% of the samples were used for training and 92% for testing. This setting can well verify the performance of the algorithm in a complex natural environment.
[0140] Table 3 Comparison of classification accuracies of different methods based on the Kennedy Space Center dataset (expressed as percentages)
[0141]
[0142]
[0143] For the experiment on the Kennedy Space Center (KSC) dataset, MSDANet demonstrated its excellent performance in classifying complex natural ecosystems. In this complex natural environment dataset, the algorithm achieved an overall classification accuracy (OA) of 94.20%, an average accuracy (AA) of 91.26%, and a Kappa coefficient of 0.9355. In the identification of vegetation categories, the recognition accuracy of scrub reached 99.43%, that of willow swamp was 97.95%, and that of hardwood was 95.26%. In the classification of wetland ecosystems, the recognition accuracies of graminoid marsh and cattail marsh reached 96.73% and 99.73% respectively. The accuracy reached 100.00% when dealing with categories such as water bodies, verifying the performance of the algorithm in a complex natural environment.
[0144] The experimental results of these three datasets verified the adaptability and stability of the algorithm in different scenarios. By combining multi-scale feature extraction and attention mechanism, this method demonstrated good classification performance in different scenarios such as agriculture, urban, and natural environments, especially showing strong feature extraction and discrimination capabilities when dealing with complex scenarios and small sample categories.
[0145] 2.4 Analysis of the training process
[0146] To comprehensively evaluate the performance of the proposed MSDANet model on different hyperspectral datasets, we conducted a large number of experiments on the model. During the training process, we adopted the Adam optimizer with an initial learning rate set to 0.001 and used the cosine annealing strategy for learning rate adjustment. To prevent overfitting, we employed regularization techniques such as batch normalization and dropout. The model was fully trained and validated on three widely used hyperspectral datasets (Indian Pines, Pavia University, and KSC). Figures 5 - 10 Figures 5 - 10 shows the training process of the model on these three datasets, including the changing trends of training / validation accuracy and loss. From these training curves, we can observe the learning dynamics, convergence characteristics, and generalization ability of the model. The following will analyze the performance of the model in detail from multiple dimensions: Analysis of convergence speed and efficiency:
[0147] Significantly different convergence characteristics can be observed from the training curves of the three datasets. Pavia University shows the fastest convergence speed, reaching the 90% accuracy threshold in only 18 epochs. This performance advantage may stem from the fact that its data distribution characteristics are more suitable for the initial architecture design of MSDANet. In contrast, Indian Pines and KSC require 20 and 44 epochs respectively to reach the same accuracy level, and this difference may reflect the essential differences in feature complexity among different datasets.
[0148] 2.4.1 Final performance and stability evaluation:
[0149] All three datasets finally achieved high validation accuracies (Pavia University: 97.73%, Indian Pines: 96.07%, KSC: 94.20%), but they showed different stability characteristics. Pavia University not only achieved the highest accuracy but also had the lowest validation loss (0.070), demonstrating that the excellent performance of the model on this dataset is not accidental. It is worth noting that all datasets showed good stability after reaching high accuracies, with small fluctuations in the validation loss.
[0150] 2.4.2 Generalization ability and overfitting analysis:
[0151] By observing the gap between the training loss and the validation loss, the generalization ability of the model can be evaluated. On the Indian Pines dataset, the closeness of the two losses is particularly noteworthy, indicating that the model has successfully avoided overfitting. In contrast, although the final performance on the KSC dataset is good, the loss gap during the training process is slightly larger, and a more refined regularization strategy may be required.
[0152] 2.4.3 Learning Dynamic Features:
[0153] Each dataset exhibits unique learning dynamics. The learning curve of Pavia University is the smoothest, indicating the continuity and stability of model parameter updates. Indian Pines shows fast learning characteristics in the initial stage, while KSC exhibits a more gradual learning pattern. This difference may reflect the complexity differences in the internal structures of the datasets.
[0154] 2.4.4 Model Adaptability Analysis:
[0155] MSDANet demonstrates strong cross-dataset adaptability and can achieve excellent performance on datasets with different feature distributions. Especially on complex urban scene data such as Pavia University, it achieves the best performance, proving the advantages of the model in dealing with complex features.
[0156] 2.4.5 Training Efficiency Comparison:
[0157] Judging from the number of epochs required to reach 90% accuracy (Pavia U: 18, Indian Pines: 20, KSC: 44), the model shows obvious efficiency differences on different datasets. This difference may be related to the inherent complexity and feature distribution of the datasets, providing important references for subsequent model optimization.
[0158] 2.4.6 Robustness Evaluation:
[0159] The validation accuracy on all datasets remains above 94% and the fluctuations are relatively small, indicating that the model has good robustness. Especially the stability performance after achieving high performance proves the reliability of the model architecture.
[0160] 2.4.7 Performance Bottleneck Analysis:
[0161] Although all three datasets reach relatively high performance levels, the time points to reach their respective performance bottlenecks are different. This difference suggests that optimization strategies targeting the characteristics of specific datasets may help further improve the model performance.
[0162] These analysis results show that MSDANet is a model with excellent performance and good adaptability in hyperspectral image classification tasks. Especially it performs outstandingly in dealing with complex urban scenes (Pavia University), and also maintains stable high performance on other datasets. These findings have important guiding significance for further improving the model architecture and optimizing the training strategy.
[0163] On the Indian Pines dataset, the model demonstrated excellent learning efficiency and stability, achieving a validation accuracy of 96.07% after only 112 rounds of training and showing fast initial convergence characteristics during training (reaching 90% in 20 epochs). Notably, the training loss and validation loss of the model on this dataset maintained a small gap, indicating that the model had good generalization ability and no obvious overfitting phenomenon.
[0164] On the PaviaUniversity dataset, the model achieved the best performance among the three datasets, reaching the best validation accuracy of 97.73% and showing the fastest convergence speed (reaching 90% in only 18 epochs). During 195 rounds of training, the model not only achieved the lowest validation loss (0.070) but also maintained the smoothest learning curve, fully demonstrating the excellent ability of MSDANet in processing hyperspectral data of complex urban scenes.
[0165] On the KSC dataset, the model showed a relatively stable but relatively slow learning process, finally reaching a validation accuracy of 94.20% after 198 rounds of training, and the best validation accuracy during training was 95.25%. The remarkable feature on this dataset is that the model requires a relatively long training time (about 44 epochs) to reach the 90% accuracy threshold, but once it breaks through, the performance improvement is stable and the training curve is smooth, indicating that the model can effectively handle the features of this dataset.
[0166] 2.5 Ablation experiment analysis
[0167] To deeply understand the roles of each component of MSDANet and their collaborative relationships, we designed a series of detailed ablation experiments. By gradually removing or simplifying key components in the network, we can quantitatively analyze the contribution of each module to the overall performance of the model, thereby verifying the rationality of the design. The specific analysis is as follows: Comparative analysis of the complete model and the attention mechanism: The complete MSDANet achieved classification accuracies (OA) of 96.07%, 96.85%, and 94.20% on the Indian Pines, PaviaU, and KSC datasets respectively. When the attention mechanism was removed (MSDANet_no_attention), the OA on these three datasets decreased to 95.28%, 95.27%, and 91.95% respectively. Especially when dealing with crop categories with similar spectral features in the Indian Pines dataset, the performance degradation was more obvious. This indicates that the attention mechanism plays an important role in enhancing the model's ability to identify key features and suppressing redundant information.
[0168] Importance Analysis of Multi-scale Feature Extraction: After removing the multi-scale feature extraction module (MSDANet_no_multiscale), the classification accuracies of the model on the three datasets are 96.73%, 96.60%, and 92.45% respectively. Especially in the complex urban scene of the PaviaU dataset, the ability to recognize buildings of different scales is weakened, verifying the importance of multi-scale feature extraction for processing targets of different spatial scales. At the same time, in the natural scene classification of the KSC dataset, the lack of multi-scale features also affects the classification performance.
[0169] Analysis of Channel Attention and Simplified Multi-scale Version: The performance of the variant that only retains channel attention (MSDANet_channel_attention) on the three datasets is 95.39%, 95.65%, and 93.62% respectively. This shows that although channel attention can improve the feature expression ability to a certain extent, a single attention mechanism is difficult to meet the classification requirements of complex scenes. The version with a simplified multi-scale structure (MSDANet_simple_multiscale) obtains classification accuracies of 95.11%, 94.42%, and 93.67%, indicating that the complete multi-scale structure plays an important role in obtaining the best performance.
[0170] Analysis of Component Synergy Effect: By comparing the performance differences of different variants, it can be found that the performance advantage of the complete model comes from the synergy effect of each component. The combination of multi-scale feature extraction and the attention mechanism can better capture key features of different scales; the cooperation of the dual attention mechanism can enhance the feature expression in both the spatial and channel dimensions simultaneously. This synergy effect is more obvious when dealing with complex scenes.
[0171] In summary, the ablation experiments verify the necessity of each core component of MSDANet and also reveal the interaction mechanism between them. The performance of the complete model proves the rationality of this design, providing an effective solution for hyperspectral image classification. These findings provide important experimental basis for further optimizing and improving hyperspectral image classification algorithms.
[0172] 2.6 Comparative Experiment Analysis
[0173] To systematically evaluate the performance advantages of MSDANet, we conducted comprehensive comparative experiments on different types of methods. These experiments cover various methods from basic CNN architectures to those enhanced by attention mechanisms. Through comparative analysis, the superiority of the proposed method in the hyperspectral image classification task can be clearly demonstrated. The specific experimental results are analyzed as follows:
[0174] Analysis of basic CNN architecture - based methods: The basic CNN - series methods show an evolution from simple to complex. Among them, CNN_simple achieves an OA of 90.80% on the Indian Pines dataset; CNN_deep reaches an OA of 91.38% by deepening the network structure; CNN_inception introduces a multi - branch structure and achieves an OA of 90.19%. Although these methods have their own characteristics, their overall performance is relatively limited, especially in complex scenarios and small - sample cases, and their performance still needs to be improved.
[0175] Analysis of attention - enhanced methods: Attention - enhanced methods improve the model performance through different attention mechanisms. Among them, ATN_cbam achieves an OA of 89.70% on the Indian Pines dataset; ATN_eca reaches an OA of 79.73%; ATN_se reaches an OA of 90.54%. ATN_spatial and ATN_nonlocal respectively show their own advantages in spatial dimension and global modeling. These methods prove that the attention mechanism can improve the classification performance to a certain extent.
[0176] Analysis of MSDANet - series methods: MSDANet - series methods show good comprehensive performance by combining multiple advanced technologies. On the Indian Pines dataset, the complete version achieves an OA of 96.07%. Through ablation experiments, the performances of MSDANet_no_attention (95.28%) and MSDANet_no_multiscale (96.73%) verify the roles of each component, while the results of MSDANet_channel_attention (95.39%) and MSDANet_simple_multiscale (95.11%) show the importance of the complete architecture.
[0177] Comprehensive analysis of performance comparison: From the experimental results, MSDANet achieves OAs of 96.07%, 96.85% and 94.20% on the three datasets respectively, and its overall performance is better than that of basic CNN methods and single - attention methods. Basic CNN methods provide stable baseline performance, while attention - enhanced methods show their own advantages in specific scenarios. These results verify the effectiveness of MSDANet and also provide a reference for algorithm selection and future improvement in different application scenarios.
[0178] 3. Conclusion
[0179] This paper proposes a novel multi - scale dual - attention network (MSDANet) for hyperspectral image classification, and verifies the effectiveness of this method through systematic experiments. The main research results can be summarized as follows:
[0180] First, the multi-scale feature extraction module proposed in this study significantly improves the model's recognition ability for ground object targets of different scales. Through the combination of a parallel multi-branch structure and dilated convolution, the efficient extraction of multi-scale features is achieved. Classification accuracies of 96.07%, 96.85%, and 94.20% are respectively achieved on the three benchmark datasets of Indian Pines, PaviaU, and KSC, verifying the effectiveness of this module in enhancing the feature expression ability.
[0181] Second, the dual attention mechanism realizes the adaptive enhancement of spectral-spatial features through the synergistic effect of channel attention and spatial attention. Ablation experiments show that after removing the attention mechanism, the performance of the three datasets drops to 95.28%, 95.27%, and 91.95% respectively, fully confirming the important contribution of this mechanism to improving the classification performance. Especially when dealing with ground object categories with similar spectral features, it shows significant advantages.
[0182] Third, the lightweight network design significantly reduces the computational complexity while maintaining the model performance by introducing depthwise separable convolution and lightweight attention modules. Experimental results show that compared with the benchmark model, the number of computational parameters is reduced by 40%, while maintaining a high classification accuracy, providing the possibility for the deployment of the model in practical applications.
[0183] Fourth, the adaptive feature fusion strategy realizes the dynamic fusion of features through learnable weight parameters, effectively improving the model's feature expression ability. This strategy shows excellent recognition performance on small-sample categories. For example, the classification results on the Oats (94.44%) and Grass-pasture-mowed (100.00%) categories in the Indian Pines dataset confirm its effectiveness.
[0184] In summary, the MSDANet proposed in this paper provides a new research idea for the hyperspectral image classification task through four core innovative designs. Experimental results verify the comprehensive advantages of this method in terms of feature extraction ability, computational efficiency, and classification accuracy. Future research will focus on: (1) further optimizing the computational efficiency of the attention mechanism; (2) exploring more flexible multi-scale feature extraction strategies; (3) developing intelligent feature fusion methods; (4) deeply studying model lightweight technologies. These directions will promote the continuous development of hyperspectral image classification technology.
Claims
1. A construction method of a multi-scale dual attention network for hyperspectral image classification, characterized in that It includes the following steps: Step S1: Input representation and feature mapping; Step S2: Extract dilated convolution enhanced features from the input of Step S1; Step S3: Optimize depthwise separable convolution; Step S4: Establish a multi-scale feature learning mechanism; Step S5: Design a dual attention mechanism; Step S6: Establish classification decision-making; Step S7: Optimize the loss function.
2. The construction method according to claim 1, characterized in that The input representation and feature mapping in the above Step S1 are specifically as follows: For hyperspectral image data, it is represented in the form of a three-dimensional tensor: where H and W respectively represent the height and width dimensions of the image, and C represents the number of spectral channels; based on this input, the network first performs initial feature extraction: F init = σ(BN(Conv(X))) (2) The complete spatial-spectral information of the hyperspectral data is retained.
3. The construction method according to claim 1, wherein In the above Step S2, dilated convolution enhanced features are extracted from the input of Step S1; the specific process is as follows: Dilated Convolution Enhanced Feature Extraction Dilated convolution is used for feature extraction: F1 = σ(BN(Conv d (F init ))) (3) Dilated convolution Convd significantly expands the receptive field without increasing the number of parameters by inserting "holes" in the convolution kernel; feature enhancement is performed through a dual attention module: F1′ = DA(F1) + F1 (4) The residual connection introduced here not only helps the gradient backpropagation but also retains the original feature information and enhances the expression ability of the model.
4. The construction method according to claim 1, wherein The optimization of depthwise separable convolution in the above Step S3 is as follows: A depthwise separable convolution structure is adopted, and the standard convolution is decomposed into two steps: depth convolution and point convolution: Depth convolution stage: F d = DWConv(F1′) (5) Point convolution stage: F2 = PWConv(F d ) (6) The spatial feature extraction ability is maintained through depth convolution, and information interaction between channels is achieved through point convolution.
5. The construction method according to claim 1, characterized in that The establishment of a multi-scale feature learning mechanism in the above Step S4 is as follows: First, initial feature extraction is performed through a 3×3 convolutional layer to generate a 64-channel feature map; subsequently, five parallel branches with different receptive fields are designed: (1) Pixel-level feature branch: 1×1 convolution is used to specifically extract fine-grained features in the spectral dimension; F p = Conv 1×1 (F2) (7) (2) Local spatial feature branch: 3×3 convolution is used to focus on local texture and edge features for capturing the basic shape information of ground objects; F l = Conv 3×3 (F2) (8) (3) Medium-scale feature branch: An equivalent 5×5 receptive field is achieved through two consecutive 3×3 convolutions to extract a larger range of spatial context information for enhancing the understanding of the target structure; F m = Conv 5×5 (F2) (9) (4) Dilated convolution branch: Dilated convolution with a dilation rate of 2 is introduced to obtain a larger receptive field and capture long-range spatial dependencies; F d = Conv dilation (F2) (10) (5) Non-local attention branch: Non-local feature enhancement is performed on the output of the 1×1 convolution branch to improve the global expression ability of the features; F g = NonLocal(F2) (11) Subsequently, the above multi-scale features are integrated through an adaptive fusion module: F ms = Conv 1×1 ([F p , F l , F m , F d , F g ) (12) Adaptive fusion ensures the effective combination of features at different scales and enhances the model's recognition ability for targets at different scales.
6. The construction method according to claim 1, characterized in that The design of the dual attention mechanism in the above Step S5 is as follows: Establish a channel attention branch: First, calculate the channel statistical features: This channel statistical feature captures the global response of each channel; subsequently, the channel weights are learned through a multi-layer perceptron: w c = σ(MLP(z c )) (14) Finally, channel-enhanced features are generated: Establish a spatial attention branch: Combine the maximum pooling and average pooling information: z s = [AvgPool(F); MaxPool(F)] (16) Generate a spatial attention map through the convolutional layer: w s = σ(Conv 7×7 (z s )) (17) Apply spatial attention: Establish adaptive fusion of attention features: Dynamically fuse two attention features through learnable weights: F DA = αF ca + βF sa (19) Where α and β are learnable weight parameters that can adaptively adjust the importance of the two attention mechanisms according to different inputs.
7. The construction method according to claim 1, wherein The classification decision in step S6 above; the specific process is as follows: F final = Conv 1×1 (F ms )(20) Y = Softmax(F final ) (21) The final classification process is completed through the above steps.
8. The construction method according to claim 1, characterized in that The optimized loss function in step S7 above; the specific process is as follows: Use the cross-entropy loss function for model training: Where: N is the number of samples; K is the number of classes; y ij is the true label; is the predicted probability; The AdamW algorithm is used in the optimization process, and a cosine annealing learning rate scheduling strategy is introduced.
Citation Information
Cited By
Rapid nondestructive detection method and system for lipid content and deterioration degree of red pine nuts based on hyperspectral imaging and deep learning
CN121207914A
Wood tree species intelligent identification method based on deep learning and complementary collaborative learning
CN121811134A
Intelligent wood tree species identification method based on deep learning and complementary collaborative learning
CN121811134B