Self-adaptive spectrum attention network for hyperspectral image classification and wave band feature learning method

The self-adaptive spectral attention network (ATN_spectral) addresses the challenges of waveband feature expression and feature fusion in high-spectral image classification by employing a dual attention mechanism and multi-dimensional extraction, achieving superior classification performance and stability across diverse datasets.

CN120318653APending Publication Date: 2025-07-15GUANGZHOU MARITIME INST
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510423330.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

When the existing hyperspectral image classification method processes high-dimensional spectral data, it is difficult to accurately express and select band features, and it is difficult to distinguish between geographic categories with similar spectral features, and the problems of classification stability and feature fusion efficiency in complex scenarios have not been effectively solved.

Method used

The adaptive spectral attention network (ATN_spectral) is designed, and the secondary attention weight generation and multi-dimensional feature extraction strategies are combined with feature cascade and convolutional fusion to realize intelligent extraction of key band information and dynamic weight learning, and cross-entropy loss and Adam optimizer are used to ensure the stability of model training.

Benefits of technology

It significantly improves the accuracy and stability of hyperspectral image classification, and can maintain high-precision classification performance in complex scenarios, especially in the case of similar spectral characteristics and uneven sample distribution, which shows excellent robustness and generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318653A_ABST
    Figure CN120318653A_ABST
Patent Text Reader

Abstract

The invention provides a self-adaptive spectrum attention network for hyperspectral image classification and a wave band feature learning method, namely a novel self-adaptive spectrum attention network. According to the method, intelligent extraction and dynamic weight learning of key wave band information in hyperspectral data are realized by innovatively designing a wave band adaptive attention mechanism and a multi-dimensional feature extraction strategy. The method is mainly characterized by comprising the following steps: designing a secondary attention weight generation network, and realizing accurate modeling of wave band importance; a multi-dimensional wave band feature extraction strategy is provided, and comprehensive characterization of features is achieved through combination of multiple statistics; a deep feature fusion strategy is realized, and original information and enhanced features are effectively combined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hyperspectral image analysis methods, and in particular to an adaptive spectral attention network and a band feature learning method for hyperspectral image classification. Background Art

[0002] In recent years, hyperspectral remote sensing technology has experienced rapid development and wide application. Due to its unique advantages, namely rich spectral information and spatial features, hyperspectral images are playing an increasingly important role in the field of remote sensing. However, the inherent high-dimensional characteristics of hyperspectral data bring huge challenges to image processing and analysis. Especially in practical application scenarios, complex problems such as ground object categories with highly similar spectral features, significant intra-class heterogeneity, and severely unbalanced sample distributions are often encountered. These challenges make it difficult for traditional classification methods to achieve ideal classification results when dealing with hyperspectral data.

[0003] In traditional machine learning methods, support vector machine (SVM) has always been an important direction in hyperspectral image classification research due to its excellent feature mapping ability and strong generalization performance. Kuo et al. proposed an innovative kernel-based feature selection method, and the core innovation of this method lies in the organic integration of the automatic selection mechanism of radial basis function (RBF) parameters and the feature linear combination criterion. This method not only realizes the intelligent selection and priority ranking of feature subsets, but more importantly, can effectively overcome the Hughes phenomenon commonly existing in hyperspectral image classification. To further improve the performance of SVM in hyperspectral image classification, Jain et al. proposed an optimization method combined with self-organizing mapping (SOM). This method adopts a two-stage strategy: first, the SOM network is used to optimize the extraction of important features in the training samples, and then the internal and external pixels in the image are accurately identified by analyzing the posterior probability of pixel intensity. This method significantly improves the accuracy and robustness of classification. Against the backdrop of the rapid development of deep learning technology, researchers began to attempt to apply deep learning architectures to support vector machines. For example, Okwuashi et al. proposed the deep support vector machine (DSVM) framework, and the innovation of this method lies in using multiple independent SVMs as network connection weights and flexibly applying various kernel functions such as exponential RBF, Gaussian RBF, neurons, and polynomials to construct a classification framework with deep learning characteristics.

[0004] With the booming development of deep learning technology, convolutional neural networks (CNNs) have shown great potential in the field of hyperspectral image classification. The Contextual deep CNN proposed by Lee et al. is an important breakthrough. By designing a multi-scale convolutional filter bank, this network effectively fuses spatial and spectral information, enabling in-depth exploration of the local context relationships in images. Under the design of a fully convolutional architecture, the network can intelligently fuse the spatial-spectral feature maps extracted by the multi-scale filter bank, significantly improving the classification performance. To address the common problems of overfitting and vanishing gradients when CNNs process hyperspectral data, Paoletti et al. proposed an improved deep CNN framework. The innovation of this framework lies in introducing a shortcut connection mechanism between the layers of the network, allowing the underlying feature maps to be directly used as the input of the current layer and passing the processed output to the upper-layer network. This design not only realizes the effective combination of spatial-spectral features at different levels, but more importantly, significantly enhances the generalization ability of the network. To solve the problem of difficult acquisition of labeled samples in hyperspectral image classification, Cao et al. proposed a method that combines the active learning strategy with a deep learning model. This method adopts an iterative optimization strategy: first, a CNN model is trained using a limited number of labeled samples, then the most informative pixels are selected through the active learning mechanism for labeling and network fine-tuning, and finally, a Markov random field is introduced to achieve smooth optimization of the class labels. This method effectively reduces the dependence on a large number of labeled samples. Considering the particularity of hyperspectral data processing, Luo et al. proposed an innovative HSI-CNN framework, the core idea of which is to convert one-dimensional spectral-spatial features into a two-dimensional matrix for processing. Specifically, this method first extracts the spectral-spatial features of the target pixel and its neighborhood, and then, through a carefully designed feature recombination strategy, systematically stacks the one-dimensional feature maps into a two-dimensional matrix, and finally inputs it into a standard CNN network for classification. This feature recombination method significantly improves the network's processing ability for high-dimensional data. To solve the problem of insufficient training data commonly existing in hyperspectral image classification, Li et al. proposed a novel data augmentation method - pixel block pairs (PBPs). This method extracts PBP features through a deep CNN network and combines a decision fusion mechanism for label assignment. This data augmentation strategy not only effectively expands the training samples, but also maintains the authenticity and diversity of the samples, providing a new idea for solving the small-sample classification problem. In the field of semi-supervised learning, Liu et al. proposed an innovative semi-supervised CNN framework, which features adding jump connection parameters between the encoder and decoder layers in a clever way to make the network more adaptable to semi-supervised learning scenarios. By simultaneously optimizing the supervised and unsupervised loss functions, this method not only solves the problem of limited labeled samples, but also can automatically learn the inherent features of complex hyperspectral data structures, significantly enhancing the generalization ability of the model.

[0005] To overcome the limitations of a single CNN structure, Guo et al. proposed the Deep Collaborative Attention Network (CACNN). This network innovatively combines the advantages of 2D-CNN and 3D-CNN, and realizes the deep fusion of the two features and the spatial attention mechanism through the NonLocalBlock. The network also cleverly uses a lightweight dense block-like Conv_Block to extract relevant information from the feature map, and adopts a deep multi-layer feature fusion strategy to enhance the ability to extract spectral-spatial information. This method of multi-modal feature fusion significantly improves the classification performance.

[0006] Inspired by the attention mechanism in the human visual system, Mei et al. proposed an innovative spectral-spatial attention network. The uniqueness of this network lies in using a recurrent neural network (RNN) with an attention mechanism to learn the correlation between consecutive spectral bands, while using a CNN with an attention mechanism to focus on significant features and the spatial correlation between adjacent pixels. The design of this dual attention mechanism effectively improves the model's perception ability of key features. Haut et al. proposed a hyperspectral image classification technique driven by visual attention. By cleverly introducing an attention mechanism into the ResNet structure, it realizes a better representation of spectral-spatial information in the data. The innovation of this method lies in designing an efficient mask calculation mechanism, which can accurately screen the most valuable features extracted by the network for the classification task, significantly improving the classification accuracy. In terms of network structure optimization, Roy et al. proposed the Attention Adaptive Spectral-Spatial Kernel Improved Residual Network (A2S2K-ResNet). This network realizes the joint feature extraction of selective 3D convolutional kernels through improved 3D ResBlocks, and innovatively introduces an efficient feature recalibration (EFR) mechanism, effectively improving the classification performance of the network. This strategy of adaptive feature extraction significantly enhances the feature expression ability of the model. The Spatial Attention Guided Residual Attention Network (SpaAG-RAN) proposed by Li et al. is another important breakthrough. This network contains three core modules: the Spatial Attention Module (SpaAM), the Spectral Attention Module (SpeAM), and the Spectral-Spatial Feature Extraction Module (SSFEM). Through the SpaAM based on spectral similarity and the design of an innovative activation function, the network can accurately capture the features of relevant spatial regions. At the same time, by guiding band selection and feature extraction through a spatial attention mask and introducing a spatial consistency loss function to ensure the accurate discrimination of features, the classification performance is significantly improved. In the research on dense connection networks, Fang et al.

[16] proposed an end-to-end 3D dense convolutional network with a spectral attention mechanism (MSDN-SA). This network realizes the synchronous extraction of spectral-spatial features of different scales through 3D dilated convolutions and densely connects all 3D feature maps. In particular, by introducing a spectral attention mechanism, the discriminability of spectral features is significantly enhanced, and the classification accuracy is improved.

[0007] With the Transformer architecture achieving breakthrough progress in the field of computer vision, researchers have begun to apply it to hyperspectral image classification tasks. Yang et al. proposed the Hyperspectral Image Transformer (HiT) classification network, which effectively captures subtle spectral differences and local spatial context information by cleverly embedding convolutional operations in the Transformer structure. This network contains two key innovative modules: the spectral adaptive 3D convolutional projection module and the Conv-Permutator, which significantly improve the efficiency and accuracy of feature extraction. The Spatial-Spectral Transformer (SST) classification framework proposed by He et al. aims to address the inherent limitations of CNNs in dealing with long-range dependencies in hyperspectral images. This framework innovatively uses CNNs to extract spatial features, employs an improved dense connection Transformer to capture sequential spectral relationships, and completes the final classification through a multi-layer perceptron. In particular, by introducing strategies such as dynamic feature augmentation (SST-FA), transfer learning (T-SST), and label smoothing (T-SST-L), problems such as overfitting and limited training samples are effectively solved. In terms of the combination of active learning and Transformer, Zhao et al. proposed a classification framework based on multi-attention Transformer and adaptive superpixel segmentation active learning (MAT-ASSAL). This framework realizes the modeling of long-range context dependencies through the self-attention module of the Transformer, captures local features using the appearance attention module, and combines the active learning strategy based on adaptive superpixel segmentation to effectively solve the problem of limited labeled samples. In the field of self-supervised learning, Scheibenreif et al. explored the application of masked image reconstruction in Transformer models for hyperspectral remote sensing images. The research team significantly improved the feature learning ability of the model by pre-training on a large-scale unlabeled dataset of the EnMAP satellite and adopting innovative strategies such as block image embedding, spatial-spectral self-attention, and spectral position encoding. To address the problem of local feature enhancement, Zou et al. proposed the Local Enhanced Spectral-Spatial Transformer (LESSFormer). This network realizes the efficient conversion of hyperspectral images to adaptive spectral-spatial tokens through the careful design of two key modules, the HSI2Token module and the local enhanced Transformer encoder, and simultaneously enhances local features and preserves long-range information through a simple and effective attention masking mechanism. In terms of cross-domain few-shot learning, Li et al. proposed the Spectral Coordinate Transformer (SCFormer). This network significantly improves the generalization ability of feature representation by embedding dense spectral coordinate blocks in the encoder and combining rotational position encoding.Meanwhile, the research team innovatively designed two masking modes, namely random masking and sequence masking, and designed an intra-domain loss function based on the theory of orthogonal complement space projection. An inter-domain loss function was constructed through the Wasserstein distance to achieve effective domain alignment.

[0008] Although these studies have made remarkable progress, they still face the following key challenges when dealing with high-dimensional hyperspectral data: First, the precise expression and selection of band features have not been effectively solved. Existing methods usually adopt a single feature expression method, which is difficult to comprehensively capture the rich band information in hyperspectral data. Especially when dealing with ground object categories with similar spectral features, it is still challenging to accurately distinguish subtle spectral differences and adaptively select key bands. Second, the classification stability problem in complex scenarios is prominent. Hyperspectral images often face complex situations such as unbalanced sample distribution and large intra-class heterogeneity. Existing methods often perform unstably when dealing with these scenarios. Especially for categories with a small number of samples, the generalization ability and robustness of the model still need to be improved. Third, the effectiveness and computational efficiency of feature fusion still exist. Existing fusion strategies often adopt simple feature splicing or weighted average methods, which fail to fully explore the complementary information of different-level features and also face the balance problem of feature redundancy and computational efficiency. The existence of these challenges severely restricts the further development and practical application of hyperspectral image classification technology, and new solutions are urgently needed. Summary of the Invention

[0009] To address the above problems, the present invention proposes a novel Adaptive Spectral Attention Network (ATN_spectral), namely an adaptive spectral attention network and band feature learning method for hyperspectral image classification. This method realizes the intelligent extraction of key band information and dynamic weight learning in hyperspectral data by innovatively designing a band adaptive attention mechanism and a multi-dimensional feature extraction strategy. The main innovation points include: designing a two-level attention weight generation network to achieve precise modeling of band importance; proposing a multi-dimensional band feature extraction strategy to comprehensively represent features through the combination of multiple statistics; realizing a deep feature fusion strategy to effectively combine the original information and enhanced features.

[0010] An adaptive spectral attention network and band feature learning method for hyperspectral image classification includes the following steps:

[0011] Step S1: Perform data representation and preprocessing on hyperspectral image data, and after standardizing the preprocessed data, perform recombination;

[0012] Step S2: Perform multi-dimensional band feature extraction on the recombined data, including band feature decoupling, multi-dimensional statistical feature extraction, and feature fusion representation;

[0013] Step S3: Establish an adaptive spectral attention mechanism, including attention weight generation and feature adaptive weighting;

[0014] Step S4: Deep feature fusion and classification, including multi-layer feature fusion and class prediction;

[0015] Step S5: Optimization strategy, including designing a loss function and parameter optimization.

[0016] The algorithm proposed by the present invention has the following theoretical innovations:

[0017] 1. Multi-dimensional band feature extraction: Through the combination of three statistics, namely mean, root mean square, and maximum value, a comprehensive representation of band features is achieved.

[0018] 2. Adaptive attention mechanism: A two-level attention weight generation mechanism is designed to achieve dynamic weight learning for key bands.

[0019] 3. Deep feature fusion: Through feature concatenation and convolutional fusion, the effective combination of original information and enhanced features is ensured.

[0020] 4. Multi-level optimization strategy: Combining cross-entropy loss and Adam optimizer ensures the stability and convergence of model training.

[0021] The organic combination of these innovation points enables the model to better adapt to the characteristics of hyperspectral data, improve classification accuracy, and has strong practical value and theoretical significance. Description of the Drawings

[0022] Figure 1 is the flowchart of the present invention.

[0023] Figure 2 is the classification map of the ATN_spectral of the present invention on the Indian Pines dataset.

[0024] Figure 3 is the classification map of the ATN_spectral of the present invention on the PaviaUniversity dataset.

[0025] Figure 4 is the classification map of the ATN_spectral of the present invention on the Kennedy Space Center dataset.

[0026] Figure 5 is the classification map of the ATN_spectral of the present invention on the Salinas Valley dataset.

[0027] Figure 6It is the simulation diagram of the training accuracy and validation accuracy obtained from the Indian Pines dataset based on the ATN_spectral of the present invention.

[0028] Figure 7 It is the simulation diagram of the training loss and validation loss obtained from the Indian Pines dataset based on the ATN_spectral of the present invention.

[0029] Figure 8 It is the simulation diagram of the training accuracy and validation accuracy obtained from the Pavia University dataset based on the ATN_spectral of the present invention.

[0030] Figure 9 It is the simulation diagram of the training loss and validation loss obtained from the Pavia University dataset based on the ATN_spectral of the present invention.

[0031] Figure 10 It is the simulation diagram of the training accuracy and validation accuracy obtained from the Kennedy Space Center dataset based on the ATN_spectral of the present invention.

[0032] Figure 11 It is the simulation diagram of the training loss and validation loss obtained from the Kennedy Space Center dataset based on the ATN_spectral of the present invention.

[0033] Figure 12 It is the simulation diagram of the training accuracy and validation accuracy obtained from the Salinas Valley dataset based on the ATN_spectral of the present invention.

[0034] Figure 13 It is the simulation diagram of the training loss and validation loss obtained from the Salinas Valley dataset based on the ATN_spectral of the present invention. Detailed implementation manners

[0035] The following details the specific solution of the present invention with reference to the accompanying drawings:

[0036] In order to systematically elaborate the implementation mechanism of the ATN_spectral network, Figure 1 the architecture design and data flow process of the method are shown in detail. The ATN_spectral network proposed by the present invention mainly consists of five core modules: input preprocessing, band feature extraction, adaptive spectral attention enhancement, feature fusion, and classification prediction, forming a complete end-to-end learning framework.

[0037] During the forward propagation of the network, the input processing module first performs standardization and dimensional reorganization operations on the original hyperspectral data (H×W×L), laying a foundation for subsequent feature extraction. The band feature extraction module realizes a comprehensive representation of band information by designing a combination strategy of multidimensional statistics. In particular, this module integrates three complementary statistical features: mean, variance, and maximum value, to capture different statistical characteristics of band data.

[0038] In the feature enhancement stage, the adaptive spectral attention module dynamically generates attention weights through a two-layer transformation network, realizing the intelligent identification and highlighting of key bands. Subsequently, the feature fusion module adopts a strategy combining feature splicing and 1×1 convolution to effectively integrate features while retaining the original information. Finally, the classification module completes the conversion from high-dimensional features to class probabilities through global average pooling and multi-layer feature mapping.

[0039] The training process of the entire network is uniformly managed by the optimization module, and end-to-end optimization of parameters is achieved by combining the cross-entropy loss function and the Adam optimizer (with an initial learning rate set to 0.001). At the same time, an early stopping strategy with patience set to 50 is introduced to prevent overfitting. This modular design not only improves the model's feature learning ability but also ensures the reliability and stability of the classification results.

[0040] An adaptive spectral attention network and a band feature learning method for hyperspectral image classification, comprising the following steps:

[0041] Step S1: Perform data representation and preprocessing on the hyperspectral image data, and after standardizing the preprocessed data, perform reorganization.

[0042] The specific process is as follows:

[0043] In hyperspectral remote sensing, image data has spatial-spectral dual characteristics and is mathematically expressed by a three-dimensional tensor:

[0044]

[0045] where H and W respectively represent the height and width of the image in the spatial dimension, and L represents the number of bands in the spectral dimension; the pixel at each spatial position (i, j) contains the complete spectral response curve information at that position. This representation method completely retains the spatial structure features and spectral dimension information of the hyperspectral data.

[0046] To eliminate the scale differences between different bands and improve the stability and convergence speed of model training, the original data is first standardized:

[0047]

[0048] Among them, μ represents the global mean of the data, and σ represents the standard deviation. This normalization process has the following advantages: eliminating the dimensional differences between different bands; adjusting the data distribution to a similar scale range; helping to accelerate the training convergence process of the model; and improving the numerical stability of the model.

[0049] To adapt to the batch processing mechanism of the deep learning framework, the standardized data is reorganized into a four-dimensional tensor structure:

[0050]

[0051] Where: B represents the batch size, P×P represents the size of the image block in the spatial dimension, and L maintains the original number of spectral bands.

[0052] This reorganization strategy not only maintains the spatial local correlation of the data but also makes full use of the parallel computing power of modern deep learning frameworks to improve the training efficiency.

[0053] Step S2: Extract multi-dimensional band features from the reorganized data, including band feature decoupling, multi-dimensional statistical feature extraction, and feature fusion representation.

[0054] The specific process is as follows:

[0055] Band feature decoupling:

[0056] First, perform feature separation at the band level to decouple the high-dimensional spectral data into independent band feature sets:

[0057]

[0058] This decoupling enables the model to independently analyze and process the feature information of each band, laying a foundation for subsequent feature extraction; Multi-dimensional statistical feature extraction: Establish three complementary statistical features:

[0059] Mean feature extraction: The mean feature reflects the overall energy level and baseline characteristics of the band and can effectively characterize the main spectral response characteristics of the band;

[0060] Root mean square feature calculation: The root mean square feature describes the fluctuation intensity of the band energy and can capture the internal change pattern and energy distribution characteristics of the band;

[0061] Maximum value feature extraction: m i = max h,w f i (h, w) The maximum value feature is used to capture the peak response of the band and can effectively identify the key bands with significant features;

[0062] Feature fusion representation:

[0063]

[0064] Combine the above three statistical features to form a complete band feature representation vector.

[0065] This multi-dimensional feature combination has the following advantages: providing a multi-dimensional expression of band features; capturing complementary information in different statistical dimensions; enhancing the feature expression ability and distinctiveness.

[0066] Step S3: Establish an adaptive spectral attention mechanism, including attention weight generation and feature adaptive weighting.

[0067] The specific process is as follows:

[0068] Attention weight generation, including a secondary attention weight generation mechanism:

[0069] First, perform feature dimensionality reduction and non-linear transformation:

[0070] A1 = ReLU(W1F stat + b1) (6)

[0071] Implement the preliminary transformation of features through the learnable weight matrix W1 and bias term b1. The ReLU activation function introduces non-linearity to enhance the feature expression ability of the model;

[0072] Subsequently, generate the final attention weights through the second transformation and Sigmoid activation:

[0073] A2 = σ(W2A1 + b2) (7)

[0074] Where: W2 and b2 are the parameters of the second-layer transformation, and σ(·) represents the Sigmoid activation function, which normalizes the weight values to the [0, 1] interval. A2 is the importance score for each band;

[0075] Feature adaptive weighting: Adaptively weight the generated attention weights with the original features:

[0076] F att = F band ⊙ A2 (8)

[0077] Where ⊙ represents element-wise multiplication, realizing the adaptive enhancement of features.

[0078] This weighting mechanism has the following characteristics: dynamically adjusting the importance of different bands; enhancing the feature expression of key bands; suppressing the interference effects of irrelevant bands.

[0079] Step S4: Deep feature fusion and classification, including multi-layer feature fusion and class prediction.

[0080] The specific process is as follows:

[0081] Multi-layer feature fusion

[0082] Using the original features and attention-weighted features, adopt the feature concatenation and convolution fusion strategy:

[0083] F cat = [F band ; F att (9)

[0084] F fuse = ReLU(Conv 1×1 (F cat )) (10)

[0085] The advantage of this fusion method is that it retains the complete information of the original features; realizes information interaction between channels through 1×1 convolution; and the ReLU activation introduces non-linearity to enhance feature expression.

[0086] Class prediction

[0087] Global feature extraction:

[0088] F global = GAP(F fuse ) (11)

[0089] Global average pooling compresses the spatial dimension and extracts the global feature representation;

[0090] Multi-layer feature mapping:

[0091] Z1 = ReLU(W3F global + b3) (12)

[0092] Z2 = Dropout(Z1, p = 0.5) (13)

[0093] Realize the non-linear mapping of the feature space through the fully connected layer, and the Dropout layer prevents overfitting;

[0094] Final classification:

[0095] Y = Softmax(W4Z2 + b4) (14)

[0096] The Softmax function converts the output into a class probability distribution.

[0097] Step S5: Optimization strategy, including designing the loss function and parameter optimization.

[0098] The specific process is as follows:

[0099] Design the loss function

[0100] The cross - entropy loss function is used to measure the difference between the prediction result and the true label:

[0101]

[0102] Among them: C represents the total number of categories, y i is the true label, is the class probability predicted by the model;

[0103] Parameter optimization

[0104] The overall optimization objective is expressed as:

[0105]

[0106] Among them, θ represents all the trainable parameters of the model; the Adam optimizer is used for parameter update, and an early stopping strategy is introduced to prevent overfitting.

[0107] I. Experimental dataset and experimental verification

[0108] 1. Hyperspectral dataset

[0109] To comprehensively evaluate the performance and universality of the proposed method, four representative publicly available hyperspectral image datasets are selected in this invention for systematic verification: Indian Pines, Pavia University, Kennedy SpaceCenter (KSC), and Salinas Valley. These datasets cover different application scenarios: The Indian Pines dataset is collected from the agricultural area of Indiana and contains 16 ground object categories, mainly used to evaluate the classification ability of algorithms in complex crop scenarios. The Pavia University dataset is obtained from the campus of the University of Pavia in Italy and covers 9 urban ground object categories, suitable for verifying the performance of algorithms in urban environments. The Kennedy Space Center dataset is collected from the Kennedy Space Center in the United States and contains 13 natural ecological categories, used to test the classification effect of algorithms in natural environments. The Salinas Valley dataset comes from the Salinas Valley in California and contains 16 precision agriculture categories, mainly used to evaluate the performance of algorithms in precision agriculture scenarios.

[0110] These datasets have different spatial resolutions, spectral characteristics, and land cover class distributions, providing an ideal test platform for comprehensively evaluating the classification performance, generalization ability, and robustness of algorithms. All experiments were conducted on a high-performance workstation configured with an Intel Core i9-14900K processor (24 cores and 32 threads), 128GB of DDR5 memory, and an NVIDIA RTX 4090 24GB graphics card. This computing platform has sufficient parallel computing power and video memory capacity to effectively support large-scale hyperspectral data processing and deep learning model training.

[0111] By selecting four standard datasets with different characteristics, the present invention constructs a comprehensive evaluation framework. These datasets cover various typical scenarios such as agriculture, urban, natural ecology, and precision agriculture, with diverse and challenging data characteristics, providing an ideal verification platform for systematically evaluating the performance and generalization ability of the proposed method.

[0112] 2. Experimental Methods

[0113] To ensure the reproducibility and fairness of the experiments, the present invention uses a fixed-ratio random stratified sampling strategy to divide the training sets for each dataset. Specifically, 10% of the samples from the Indian Pines and Kennedy Space Center datasets are used for training, while 3% and 2% of the samples from the PaviaU and Salinas datasets are used for training, respectively. These training ratios are chosen mainly considering the following factors: maintaining comparability with existing studies, validating the performance of the model in small-sample scenarios, considering the cost of obtaining labeled data in practical applications, and ensuring the representational integrity of different class samples. The experimental results are analyzed as follows:

[0114] 2.1. Experiments Based on the Indian Pines Dataset:

[0115] First, in the experiments on the Indian Pines agricultural dataset, the ATN_spectral method demonstrated excellent classification performance on this agricultural area dataset, achieving an overall accuracy (OA) of 96.51%, an average accuracy (AA) of 92.06%, and a Kappa coefficient of 96.02%. Looking at the class accuracy distribution, significant high accuracies were achieved in major crop and mixed land cover classes such as "Corn" (99.53%), "Soybeans-clean-till" (98.13%), "Bldg-grass-tree" (97.98%), etc., with relatively lower accuracy only in the "Oats" class (50.00%), which may be due to the small number of samples in this class. This method shows obvious advantages over other methods in dealing with complex agricultural scenarios, especially in distinguishing crop classes with similar spectral characteristics.

[0116] Table 1 Comparison of classification accuracies of different methods based on the Indian Pines dataset (expressed as percentages)

[0117]

[0118]

[0119] 2.2. Experiments based on the Pavia University dataset:

[0120] Next, in the experiment on the Pavia University urban dataset, the ATN_spectral method achieved an OA of 96.01%, an AA of 92.15%, and a Kappa coefficient of 94.69%, demonstrating excellent urban land cover classification ability. This method obtained extremely high classification accuracies in categories such as "Metal Sheets" (100%), "Meadows" (99.34%), and "Bare Soil" (96.27%), but had a relatively low accuracy in the "Bitumen" category (69.61%), which might be due to a certain degree of spectral feature confusion of this category with other artificial materials. Overall, this method still maintained stable and excellent classification performance in the complex urban environment.

[0121] Table 2 Comparison of classification accuracies of different methods based on the Pavia University dataset (expressed as percentages)

[0122]

[0123]

[0124] 2.3. Experiments based on the Kennedy Space Center dataset:

[0125] Furthermore, in the experiment on the Kennedy Space Center ecological dataset, for this complex wetland ecosystem dataset, the ATN_spectral method achieved excellent results with an OA of 97.08%, an AA of 95.54%, and a Kappa coefficient of 96.75%. Notably, this method reached 100% classification accuracy in categories such as "Hardwood swamp", "Spartina marsh", "Water", and "Salt marsh", and although the accuracy in the "Cabbage palm / oak" category was relatively low (70.48%), it was still better than most of the comparison methods. This indicates that this method can effectively handle the fine discrimination of different vegetation types in complex natural ecosystems.

[0126] Table 3 Comparison of Classification Accuracy of Different Methods Based on Kennedy Space Center Dataset (Expressed in Percentage)

[0127]

[0128] 2.4. Experiments Based on Salinas Valley Dataset:

[0129] Finally, in the experiments on the Salinas Valley precision agriculture dataset, the ATN_spectral method achieved an OA of 94.47%, an AA of 97.39%, and a Kappa coefficient of 93.87% on this agricultural area dataset. The classification results showed that 100% accuracy was achieved in multiple crop categories including "Broccoli-green-weeds-2", "Fallow", "Stubble", "Celery", and "Lettuce-romaine-5wk", and it was only relatively weak in the "Grapes-untrained" category (79.12%), mainly due to the large internal heterogeneity of this category. These results fully demonstrated the classification ability of this method for crops at different growth stages in precision agriculture scenarios.

[0130] Table 4 Comparison of Classification Accuracy of Different Methods Based on Salinas Valley Dataset (Expressed in Percentage)

[0131]

[0132]

[0133] A comprehensive analysis of the classification results of all methods on four typical hyperspectral datasets shows the following conclusions: The ATN_spectral method achieves excellent classification performance using only 2%-10% of the training samples. The ATN_spectral method achieves optimal or near-optimal classification performance on all datasets, with an overall accuracy (OA) of 96.51% for Indian Pines, 96.01% for Pavia University, 97.08% for Kennedy Space Center, and 94.47% for Salinas Valley. The average accuracy (AA) is 92.06%, 92.15%, 95.54%, and 97.39%, respectively. The Kappa coefficient is 96.02%, 94.69%, 96.75%, and 93.87%, respectively, which are significantly better than other comparison methods. Although the traditional CNN series methods performed well on some data sets (such as cnn_deep achieved 91.52% OA on Indian Pines), the performance fluctuated greatly and the accuracy dropped significantly in complex scenes (such as cnn_efficientnetv2 had the worst performance on multiple data sets); the methods that introduced the attention mechanism (such as ATN_cbam and ATN_nonlocal) have been significantly improved compared with the traditional CNN, but there are still some limitations, such as ATN_eca and ATN_se have unstable performance when dealing with complex ground object categories, which may be due to their failure to fully utilize the special properties of hyperspectral data. In contrast, the ATN_spectral method not only maintains a leading position in overall accuracy, but also shows excellent performance in the classification of specific categories, especially when dealing with categories with similar spectral features (such as different crops in Indian Pines and different vegetation types in Kennedy Space Center), it can maintain a high classification accuracy, which fully proves the advancement and universality of this method in hyperspectral image classification tasks. This excellent classification performance is mainly due to its innovations in feature extraction and attention mechanism design, which enables it to better adapt to the characteristics of hyperspectral data and provides an effective solution for the fine classification of hyperspectral remote sensing images.

[0134] Through in-depth analysis of experimental results and systematic evaluation of method design, the ATN_spectral method demonstrates significant advantages and innovative contributions: First, the method achieves optimal or near-optimal classification performance on four representative hyperspectral datasets, with the overall accuracy (OA) reaching over 94% in all cases, which fully proves its excellent classification ability and method stability; Second, the method shows excellent discrimination ability when dealing with complex land cover classes with similar spectral features (such as crops, urban features, wetland vegetation, etc.), which benefits from its innovative algorithm design, mainly including the following four aspects: (1) A band adaptive attention mechanism is proposed. By designing a two-level attention weight generation network (as shown in Equations 5 and 6), it realizes dynamic and precise modeling of the importance of different bands, effectively improving the accuracy of feature selection; (2) A multi-dimensional band feature extraction strategy is designed. Through the combination of three complementary statistics, namely the mean feature (Equation 13), the root mean square feature (Equation 14), and the maximum value feature (Equation 15), it achieves a comprehensive representation of band features and enhances the expressive power of features; (3) A deep feature fusion strategy is implemented. By combining feature concatenation (Equation 8) and convolutional fusion (Equation 9), it effectively preserves the original feature information while enhancing feature expression; (4) A multi-level optimization strategy is adopted, combining the cross-entropy loss function (Equation 14) and the Adam optimizer to ensure the stability and convergence of model training. These innovative designs enable the ATN_spectral method to better adapt to the characteristics of hyperspectral data, improve classification accuracy while maintaining strong generalization ability and robustness, providing a new paradigm and solution for the study of fine classification of hyperspectral remote sensing images. The organic combination of these innovation points not only promotes the development of hyperspectral image classification technology but also provides important reference value for the design of deep learning methods in related fields.

[0135] II. Analysis of the Training Process

[0136] Experimental results on four standard datasets show that the proposed ATN_spectral model exhibits different training characteristics. On the Indian Pines dataset, the experimental results show that the model has significant fast learning ability: within the first 20 training epochs, the classification accuracy rapidly increases from 40.14% to 82.23%. Subsequently, the model enters a stable optimization stage. Although there are slight fluctuations during the 80 - 120 epochs, the overall trend remains upward, and finally reaches the optimal validation accuracy of 96.79% at the 135th epoch. It is worth noting that the loss function shows a stable downward trend, gradually decreasing from the initial value of 2.10 to 0.13, and the training loss and validation loss maintain a reasonable gap, confirming the good generalization performance of the model.

[0137] In the experiment of KSC dataset, the model showed excellent convergence characteristics. It only takes 15 training cycles to achieve a classification accuracy of more than 85%, which is significantly faster than other datasets. The training process showed extremely high stability, and the fluctuation of verification accuracy was effectively controlled within the range of ±0.3%, which fully verified the robustness of the model. In the 159th cycle, the model achieved a verification accuracy of 97.08%, which is the best result among all test datasets. In addition, the loss function showed an ideal exponential decay characteristic and tended to be stable in the later stage.

[0138] For the PaviaU dataset, the experimental results show that the model has excellent learning stability. The classification accuracy gradually increased from the initial 60.33% to 96.01%, and the entire training process showed the smoothest learning curve. In particular, the difference between the training accuracy and the verification accuracy was always maintained within 1%, which strongly confirmed the generalization ability of the model. The loss function converged to the range of 0.14-0.16 in the late stage of training, further verifying the optimization stability of the model.

[0139] The most obvious stage-by-stage training characteristics were observed in the experiments on the Salinas dataset. Specifically, the classification accuracy in the rapid learning stage (1-20 cycles) increased from 47.32% to 83.17%; the stable optimization stage (20-80 cycles) achieved a continuous improvement of about 10%; and the fine-tuning stage (>80 cycles) finally reached a verification accuracy of 95.60%. The experimental results show that the loss function presents a step-by-step decline feature, and the dynamic adjustment of the learning rate can bring significant performance improvements in each stage.

[0140] Performance Summary Across Datasets

[0141]

[0142] 1. Training process and convergence analysis:

[0143] Experimental results show that the proposed ATN_spectral model exhibits significant stage characteristics during the training process. In the early stage of training (epoch 1-20), the model exhibits rapid learning ability, and the classification accuracy is significantly improved from the baseline level (40%-60%) to 80%-85%. In the subsequent stable optimization stage (epoch 20-80), the performance growth slows down but continues to improve. In the convergence stage (epoch>80), the model reaches a stable state on all data sets, and the fluctuation of the verification accuracy is effectively controlled within the range of ±0.5%, which fully confirms that the method has good convergence characteristics.

[0144] 2. Analysis of loss function behavior:

[0145] The evolution process of the loss function exhibits typical exponential decay characteristics. In the initial stage of training, the loss value rapidly drops from 1.8 - 2.1 to 0.4 - 0.5, indicating that the model can effectively capture the basic features of the data. In the subsequent optimization stage, the loss value continues to steadily decline to the range of 0.2 - 0.3. It is worth noting that the training loss and the validation loss always maintain a moderate gap (≤0.03) and finally stably converge to the range of 0.13 - 0.16. This phenomenon indicates that the model has excellent generalization performance and effectively avoids the overfitting problem.

[0146] 3. Analysis of the learning rate strategy:

[0147] The present invention adopts a cosine annealing learning rate scheduling strategy, and the initial learning rate is set to 1e-3. This method effectively prevents the model from falling into local optima by periodically adjusting the learning rate. Experiments show that this adaptive learning rate mechanism can play different roles at different stages of training: maintaining a relatively high learning rate in the initial stage to promote rapid convergence, achieving fine optimization through progressive decay in the middle stage, and using a smaller learning rate in the later stage to ensure convergence stability. The results of quantitative analysis show that this strategy significantly improves the optimization efficiency and the final performance of the model.

[0148] 4. Comprehensive evaluation of the model performance:

[0149] The experimental results on four benchmark datasets fully verify the effectiveness of the proposed method. The model achieves a validation accuracy of over 95% on all datasets (Indian Pines: 96.79%, KSC: 97.08%, PaviaU: 96.01%, Salinas: 95.60%). Further quantitative evaluation shows that the model obtains excellent Kappa coefficients (>0.94) and macro F1 scores (>92%) on each dataset. Particularly noteworthy is that the performance difference between the training set and the validation set is controlled within 2%, and this result strongly confirms the generalization ability of the model. At the same time, the model shows significant computational efficiency advantages, and the average training time per round is only 0.12 - 0.15 seconds. These experimental results strongly confirm the superiority of the proposed method in the hyperspectral image classification task.

[0150] The results of the training process analysis show that the ATN_spectral model exhibits good convergence characteristics and learning stability on different datasets. The training process of the model can be clearly divided into three stages: rapid learning, stable optimization, and fine tuning, and the loss function shows an ideal convergence trend. In particular, through the application of the cosine annealing learning rate scheduling strategy, the optimization efficiency and the final performance of the model are effectively improved, providing a stable and reliable solution for the hyperspectral image classification task.

[0151] III. Overall experimental analysis

[0152] 1. Overall performance analysis:

[0153] Quantitative evaluation on four representative hyperspectral datasets shows that the proposed ATN_spectral method exhibits significant advantages in classification performance. This method achieved overall accuracies (OA) of 96.51%, 96.01%, 97.08%, and 94.47% on the Indian Pines, Pavia University, Kennedy Space Center, and Salinas Valley datasets respectively, and the corresponding Kappa coefficients reached 96.02%, 94.69%, 96.75%, and 93.87%. This result is significantly better than existing benchmark methods, fully demonstrating the effectiveness and advancement of the proposed method.

[0154] 2. Performance evaluation of traditional CNN methods

[0155] Experimental results show that traditional CNN architectures have significant performance limitations in hyperspectral image classification tasks. Although methods such as cnn_deep achieved an acceptable accuracy of 91.52% in specific scenarios (such as the Indian Pines dataset), they are unstable in complex scenarios. For example, the overall accuracy of cnn_simple significantly decreased to 79.72% on the Salinas Valley dataset. Notably, deep network structures (such as cnn_resnet and cnn_efficientnetv2) show severe performance degradation when dealing with some challenging classes. For instance, the classification accuracy on the Gravel class of the Pavia University dataset was only 9.77%, highlighting the inherent defects of traditional CNN architectures in processing high-dimensional spectral data.

[0156] 3. Evaluation of attention mechanism methods

[0157] Although the improved methods based on the attention mechanism have enhanced the classification performance to some extent, they still fail to fully address the core challenges of hyperspectral classification. Although ATN_cbam and ATN_nonlocal achieved an overall accuracy of over 90% on multiple datasets, they are still insufficient in dealing with complex scenarios. For example, ATN_eca completely fails on specific classes of the Pavia University dataset, and ATN_se is unstable when dealing with classes with similar spectral features. These phenomena indicate that simply introducing the attention mechanism is difficult to effectively cope with the inherent characteristics of hyperspectral data.

[0158] 4. Class accuracy and stability evaluation

[0159] In - depth analysis shows that the ATN_spectral method has outstanding advantages in dealing with the key challenges of hyperspectral classification tasks. For complex scenarios such as unbalanced sample distribution and similar spectral features, this method demonstrates excellent classification performance and remarkable stability. Especially when dealing with categories with a small number of samples, traditional methods generally experience a significant decline in performance, while the ATN_spectral method can maintain stable and high - precision classification results, which fully verifies the outstanding contribution of this method in enhancing classification robustness and generalization ability, providing a new technical paradigm for the fine classification of hyperspectral remote - sensing images.

[0160] Systematic experimental analysis of four typical hyperspectral datasets shows that the proposed ATN_spectral method has achieved significant breakthroughs in key indicators such as classification accuracy, performance stability, and class discrimination ability. The excellent performance of this method mainly stems from its innovative contributions in feature - extraction strategies and attention - mechanism design, enabling it to effectively capture and utilize the inherent features of hyperspectral data. Especially when dealing with complex scenarios of unbalanced sample distribution and classes with highly similar spectral features, this method still maintains stable and high - precision classification performance, fully demonstrating its superiority in feature expression and class recognition. These experimental results not only verify the technological advancement of the proposed method but also provide a robust and practical solution for the fine - classification task of hyperspectral remote - sensing images, having important theoretical and practical value for promoting the research and development in this field.

[0161] Systematic experimental analysis of four typical hyperspectral datasets shows that the proposed ATN_spectral method has achieved significant breakthroughs in key indicators such as classification accuracy, performance stability, and class discrimination ability. The excellent performance of this method mainly stems from its innovative contributions in feature - extraction strategies and attention - mechanism design, enabling it to effectively capture and utilize the inherent features of hyperspectral data. Especially when dealing with complex scenarios of unbalanced sample distribution and classes with highly similar spectral features, this method still maintains stable and high - precision classification performance, fully demonstrating its superiority in feature expression and class recognition. These experimental results not only verify the technological advancement of the proposed method but also provide a robust and practical solution for the fine - classification task of hyperspectral remote - sensing images, having important theoretical and practical value for promoting the research and development in this field.

[0162] IV. Conclusion

[0163] The present invention proposes a hyperspectral image classification method based on an adaptive spectral attention mechanism. Through systematic theoretical analysis and experimental verification, the following conclusions are drawn:

[0164] The experimental results show that even under the condition of extremely low training sample ratios (2% - 10%), the proposed method can still maintain stable high-precision classification performance. The proposed ATN_spectral method has achieved excellent classification performance on four representative hyperspectral datasets, with the overall accuracy exceeding 94% in all cases, fully verifying the effectiveness and universality of the method. Especially when dealing with complex scenarios (such as unbalanced sample sizes, similar spectral features, etc.), it shows significant advantages, providing a new technical paradigm for the fine classification of hyperspectral images.

[0165] The second-level attention weight generation network and multi-dimensional band feature extraction strategy innovatively designed in this invention effectively solve the performance limitations of traditional methods when dealing with high-dimensional spectral data. The experimental results show that this method can accurately capture key band information and achieve dynamic weight learning, significantly improving the classification accuracy and model stability.

[0166] Through in-depth analysis of the training process, it is found that the proposed method has excellent convergence characteristics and generalization ability. The training curves on the four datasets all show an ideal convergence trend, with the validation loss and training loss maintaining a reasonable gap, fully confirming the robustness of the model.

[0167] The comparative experimental results show that compared with traditional CNN methods and basic attention mechanism methods, the ATN_spectral method proposed in this invention has significant advantages in terms of classification accuracy, performance stability, and computational efficiency, providing new ideas and methods for the research on hyperspectral remote sensing image classification.

[0168] Future research directions include: further optimizing the model structure to improve computational efficiency, exploring more efficient feature fusion strategies, and extending this method to other remote sensing image analysis tasks.

Claims

1. An adaptive spectral attention network and band feature learning method for hyperspectral image classification, characterized in that It includes the following steps: Step S1: Perform data representation and preprocessing on the hyperspectral image data. After standardizing the preprocessed data, reorganize it; Step S2: Extract multi-dimensional band features from the reorganized data, including band feature decoupling, multi-dimensional statistical feature extraction, and feature fusion representation; Step S3: Establish an adaptive spectral attention mechanism, including attention weight generation and feature adaptive weighting; Step S4: Deep feature fusion and classification, including multi-layer feature fusion and class prediction; Step S5: Optimization strategy, including designing a loss function and parameter optimization.

2. The adaptive spectral attention network and band feature learning method for hyperspectral image classification according to claim 1, wherein The specific process of the above Step S1 is as follows: In hyperspectral remote sensing, image data has spatial-spectral dual characteristics and is mathematically expressed by a three-dimensional tensor: where H and W represent the height and width of the image in the spatial dimension respectively, and L represents the number of bands in the spectral dimension; the pixel at each spatial position (i, j) contains the complete spectral response curve information at that position; To eliminate the scale differences between different bands and improve the stability and convergence speed of model training, first perform standardization processing on the original data: where μ represents the global mean of the data, and σ represents the standard deviation; To adapt to the batch processing mechanism of the deep learning framework, reorganize the standardized data into a four-dimensional tensor structure: where: B represents the batch size, P×P represents the size of the image block in the spatial dimension, and L maintains the original number of spectral bands.

3. The adaptive spectral attention network and band feature learning method for hyperspectral image classification according to claim 2, characterized in that The specific process of extracting multi-dimensional band features from the reorganized data in the above Step S2 is as follows: Band feature decoupling: First, perform feature separation at the band level to decouple the high-dimensional spectral data into an independent set of band features: This decoupling enables the model to independently analyze and process the feature information of each band, laying a foundation for subsequent feature extraction; Multi-dimensional statistical feature extraction: Establish three complementary statistical features: Mean feature extraction: The mean feature reflects the overall energy level and baseline characteristics of the band, and can effectively characterize the main spectral response characteristics of the band; Root Mean Square Feature Calculation: The root mean square feature describes the fluctuation intensity of the band energy and can capture the internal change patterns and energy distribution characteristics of the band; Maximum value feature extraction: m i = max h,w f i (h, w) The maximum value feature is used to capture the peak response of the band and can effectively identify the key bands with significant features; Feature fusion representation: Combine the above three statistical features to form a complete band feature representation vector.

4. The adaptive spectral attention network and band feature learning method for hyperspectral image classification according to claim 3, characterized in that The specific process of establishing an adaptive spectral attention mechanism in the above Step S3 is as follows: Attention weight generation, including a secondary attention weight generation mechanism: First, perform feature dimensionality reduction and non-linear transformation: A1 = ReLU(W1F stat + b1) (6) Implement the preliminary transformation of the feature through the learnable weight matrix W1 and the bias term b1, and introduce non-linearity through the ReLU activation function to enhance the feature expression ability of the model; Subsequently, generate the final attention weight through the second transformation and the Sigmoid activation: A2 = σ(W2A1 + b2) (7) where: W2 and b2 are the parameters of the second-layer transformation, σ(·) represents the Sigmoid activation function, normalizes the weight value to the [0,1] interval, and A2 is the importance score of each band; Feature adaptive weighting: Adaptively weight the generated attention weight with the original feature: F att = F band ⊙A2 (8) where ⊙ represents element-wise multiplication, realizing the adaptive enhancement of the feature.

5. The adaptive spectral attention network and band feature learning method for hyperspectral image classification according to claim 4, wherein The specific process of deep feature fusion and classification in the above Step S4 is as follows: Multi-layer feature fusion Utilize the original feature and the attention-weighted feature, and adopt a feature concatenation and convolution fusion strategy: F cat = [F band ; F att (9) Class prediction Global feature extraction: F global = GAP(F fuse ) (11) Use global average pooling to compress the spatial dimension and extract the global feature representation; Multi-layer feature mapping: Z1 = ReLU(W3F global + b3) (12) Z2 = Dropout(Z1, p = 0.5) (13) Realize the non-linear mapping of the feature space through the fully connected layer, and the Dropout layer prevents overfitting; Finally Classification: Y = Softmax(W4Z2 + b4) (14) The Softmax function converts the output into a categorical probability distribution.

6. The adaptive spectral attention network and band feature learning method for hyperspectral image classification according to claim 5, characterized in that The optimization strategy in step S5 above is as follows: Design the loss function Use the cross-entropy loss function to measure the difference between the prediction result and the true label: Among them: C represents the total number of categories, and y i is the true label, is the category probability predicted by the model; Parameter optimization The overall optimization objective is expressed as: where θ represents all the trainable parameters of the model; the Adam optimizer is used for parameter update, and an early stopping strategy is introduced to prevent overfitting.

Citation Information

Cited By

  • Hyperspectral image-based rice bakanae disease bacteria-carrying seed detection method and device

    CN121207889A