Construction method of adaptive temperature-guided hybrid attention network for hyperspectral image classification

The ATN-Hybrid network addresses the limitations of single-mode attention mechanisms in high-spectral image classification by combining deterministic and probabilistic attention with learnable fusion, improving feature selection and fusion strategies to enhance classification performance in complex environments with limited training data.

CN120318651APending Publication Date: 2025-07-15GUANGZHOU MARITIME INST
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510422304.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing hyperspectral image classification methods have insufficient flexibility in feature selection, insufficient optimization of feature fusion strategy, and prominent adaptability problems in small sample scenarios, making it difficult to effectively deal with complex object categories and spectral feature similarities.

Method used

An adaptive temperature-guided hybrid attention network (ATN-Hybrid) is designed, using a deterministic-probability hybrid attention mechanism and a multi-level feature fusion strategy based on learnable fusion coefficients. Through explicit feature transformation of deterministic branches and temperature-guided soft attention mechanism of probabilistic branches, the robustness and adaptability of feature selection are achieved.

Benefits of technology

It significantly improves the robustness and adaptability of hyperspectral image classification, especially maintains excellent classification performance under the conditions of similar complex geographic categories and spectral characteristics, and can maintain high-precision classification under limited training samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318651A_ABST
    Figure CN120318651A_ABST
Patent Text Reader

Abstract

The invention relates to a construction method of an adaptive temperature guided mixed attention network for hyperspectral image classification, and provides a novel mixed attention network for hyperspectral image classification. A deterministic-probabilistic mixed attention mechanism is designed, and a traditional single-mode attention mechanism is expanded into a double-branch structure. The deterministic branch generates stable feature weights through explicit feature transformation and normalization, and the probabilistic branch introduces a soft attention mechanism of temperature parameter adjustment to realize flexible feature selection. The robustness and adaptability of feature selection are remarkably improved through the complementary effect of the two mechanisms. Meanwhile, a multi-level feature fusion strategy based on a learnable fusion coefficient is innovatively designed, and adaptive adjustment of feature importance is realized by dynamically balancing contributions of two types of attention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hyperspectral image classification methods, and in particular to a method for constructing an adaptive temperature-guided hybrid attention network for hyperspectral image classification. Background Art

[0002] Hyperspectral remote sensing images exhibit great application potential in fields such as ground object recognition, environmental monitoring, and precision agriculture due to their rich spectral dimension information. Compared with traditional remote sensing images, each pixel point in a hyperspectral image contains complete spectral curve information, providing important spectral feature bases for the precise recognition and classification of ground objects. However, the inherent high-dimensional characteristics of hyperspectral data also bring many challenges: high data redundancy, strong correlations between bands, and problems of feature expression under limited training samples, etc., which all pose great challenges to the hyperspectral image classification task (Li et al., 2024).

[0003] In recent years, researchers have explored solutions to the hyperspectral image classification problem from different perspectives. These studies can be mainly divided into the following directions:

[0004] Early hyperspectral image classification research mainly focused on the improvement and optimization of traditional machine learning methods. These methods usually carried out research from multiple aspects such as feature extraction, classifier design, and post-processing optimization. In terms of feature extraction and classifier design, Tarabalka et al. (2010) proposed a spectral-spatial classification method combining probabilistic support vector machine (SVM) and Markov random field (MRF). This method first uses probabilistic SVM to obtain the class probability distribution at the pixel level, and then models the spatial dependence relationship between pixels through MRF, introducing spatial smoothing constraints while maintaining the discriminability of spectral information, effectively improving the spatial continuity of classification. However, this method has a large computational overhead when dealing with large-scale data. To solve the problem of excessive computational load of SVM in multi-class classification scenarios, Hosseini et al. (2011) proposed a hierarchical two-stage classification framework. This method first uses a computationally efficient maximum likelihood classifier for preliminary classification to obtain the class probability distribution of each pixel, and then constructs a tree-structured SVM classifier based on classification uncertainty for fine classification only in high-probability classes. However, this method is highly dependent on the initial classification result and is easily affected by noise. Aiming at the limitations of the traditional random forest algorithm in high-dimensional data processing, Tong et al. (2021) proposed the spectral-spatial deep random forest (SSDRF) method. This method combines two spatial feature extraction strategies: fixed-size image patches are used to capture local texture information, and shape-adaptive superpixels are used to maintain the integrity of the target boundary. However, this method is still sensitive to the selection of superpixel segmentation parameters.

[0005] With the development of deep learning technology, hyperspectral image classification methods based on deep neural networks have demonstrated powerful feature learning capabilities and classification performance. These methods mainly conduct research from aspects such as network architecture design, optimization strategy improvement, and small sample adaptability.

[0006] The introduction of the attention mechanism provides a new feature selection paradigm for hyperspectral image classification. Related research mainly focuses on aspects such as attention module design, multi-attention fusion, and the combination of attention and traditional architectures. In terms of the design of basic attention modules, Hang et al. (2020) proposed a dual-branch attention CNN architecture. This method designed a spectral attention sub-network and a spatial attention sub-network respectively: the spectral attention sub-network highlights key spectral features by learning band weights, while the spatial attention sub-network identifies important spatial regions through an adaptive weight map. The classification results of the two sub-networks are fused through learnable weighting coefficients to achieve adaptive adjustment of feature importance. This design improves the model performance while providing good interpretability, but the dual-branch structure increases the computational complexity. In terms of the collaborative application of multi-attention mechanisms, Zhang et al. (2023) proposed the MATNet network that combines multi-attention mechanisms and Transformer. This method designed three key modules: the spatial attention module is used to highlight important spatial regions, the channel attention module is responsible for selecting discriminative feature channels, and the tokenizer module realizes the semantic-level representation of ground object categories. In particular, this method also proposed a polynomial label smoothing loss function (Lpoly) to improve the model generalization ability by adaptively adjusting the smoothing degree of different datasets. This collaborative design of multi-attention significantly improves the expressive ability of features, but the training difficulty of the model also increases accordingly. Yang et al. (2021) proposed the cross-attention spectral-spatial network (CASSN) to address the problem of CNN being sensitive to image rotation. The core innovation of this method lies in the design of a cross-spectral attention component, which generates band weights by establishing the correlation relationship between bands, effectively suppressing the influence of redundant bands. At the same time, the cross-spatial attention component generates spectral-spatial features under the guidance of the pixel to be classified, improving the discriminative ability of features, but still needs to be optimized in terms of computational efficiency. Xue et al. (2021) proposed the hierarchical residual attention network (HResNetAM) from the perspective of hierarchical feature learning. This method extracts multi-scale spatial and spectral features through a hierarchical residual network structure, and at the same time uses the attention mechanism to set adaptive weights for features at different scales. This design expands the receptive field of the network and also realizes the adaptive selection of features, but the number of model parameters is large. In terms of the innovative application of the attention mechanism, Sun et al. (2023) proposed the large kernel spectral-spatial attention network (LKSSAN). This method addresses the problem that existing Transformer methods ignore the spatial attributes of hyperspectral images by designing a spectral-spatial attention module (SSAM) to maintain the 3D structure of the data. At the same time, long-range 3D feature dependencies are extracted by introducing large kernel attention (LKA) and convolutional feed-forward network (CFF). This design effectively improves the model's ability to model long-range dependency relationships, but has higher requirements for hardware resources.

[0007] With the in-depth research, researchers have gradually realized that the single feature learning strategy is difficult to fully express the complex characteristics of hyperspectral data. Therefore, they began to explore the organic combination of multiple feature learning methods. Ghaderizadeh et al. (2021) proposed a hybrid CNN architecture from the perspective of network structure design. The core of this method is to combine the advantages of 3D fast learning blocks and 2D CNNs. Among them, the 3D fast learning block contains depthwise separable convolution blocks and fast convolution blocks, which are used to efficiently extract spectral-spatial features. This design significantly reduces the computational overhead while maintaining high performance. In terms of hybrid feature learning in specific application scenarios, Zhao et al. (2022) proposed an innovative hybrid convolutional network architecture for wheat seed classification. This method designs a dual-path feature extraction strategy: spectral features of regions of interest (ROIs) are specifically extracted through 1D convolution, while spatial features are directly analyzed from hyperspectral images using 2D convolution. Feature fusion is achieved in the fully connected layer of the network. This method enhances the correlation between features by designing a feature interaction module. However, there is still room for optimization in the automatic selection of ROIs. Pan et al. (2023) proposed a hyperspectral image classification network based on a multi-scale hybrid network and attention mechanism. This method designs three sub-networks: a spectral-spatial feature extraction network, a spatial pyramid inverse network, and a classification network, which are responsible for multi-scale feature extraction, redundant information reduction, and feature fusion classification respectively. Combining the multi-scale fusion network and attention mechanism enables effective capture of local and global features, achieving efficient utilization of computing resources while enhancing the feature expression ability. However, the computational overhead is still large when processing large-scale images. Dong et al. (2022) proposed a weighted feature fusion model (WFCG) of convolutional neural network and graph attention network. This method reduces the computational complexity through a superpixel-based GAT encoding and decoding module, enhances the pertinence of feature extraction by a CNN combined with an attention mechanism, and realizes feature adaptive fusion using learnable weight coefficients. This design effectively enhances the respective advantages of GAT and CNN, and effectively solves the small sample classification problem. However, the model training process is complex and requires fine-tuning of parameters. Bhatti et al. (2023) proposed a multi-feature fusion method (MFFCG) of 3D-CNN and graph attention network. This method combines the advantages of 3D-CNN and GAT encoding and decoding modules, designs two optimized GAT models for fusion with different levels of 3D-CNN, and forms MFFCG-1 and MFFCG-2 variants adapted to different application scenarios. This multi-level feature fusion strategy significantly enhances the feature expression ability of the model, but also increases the complexity and training difficulty of the model.Islam et al. (2024) proposed the Hybrid-2DNET hybrid classification method, which realizes the dimensionality reduction and optimization of features through three key steps: firstly, the factor analysis method is used for preliminary feature extraction, then feature selection is carried out based on the minimum redundancy-maximum relevance (mRMR) criterion, and finally a 2D wavelet convolutional neural network is used to further optimize the features. This multi-stage feature processing strategy effectively reduces the dimensionality of the data while retaining key information, but the selection of parameters in each stage is relatively complex.

[0008] Through a systematic analysis of existing hyperspectral image classification methods, it is found that the current research still faces the following several key scientific problems:

[0009] Firstly, there are obvious deficiencies in the flexibility of feature selection in existing methods. Although the introduction of the attention mechanism provides a new paradigm for feature selection, most existing methods adopt a single-mode attention mechanism, such as spatial attention or channel attention. This single-mode design is difficult to meet the requirements of both the stability and adaptability of feature selection simultaneously. Especially when dealing with highly similar ground object categories, a single-mode attention mechanism often fails to capture fine-grained feature differences, resulting in a decline in classification performance.

[0010] Secondly, the optimization of the feature fusion strategy still needs to be further explored. Most existing studies adopt basic fusion methods such as simple feature concatenation or weighted averaging, lacking a dynamic evaluation and adjustment mechanism for the importance of different types of features. This static fusion strategy cannot adapt to the dynamic changes in the importance of features in different scenarios. Especially in complex surface environments, the contribution degree of different types of features to the classification result often changes with the scenario. Therefore, how to design a more flexible and adaptive feature fusion strategy has become an urgent problem to be solved.

[0011] Thirdly, the adaptability problem in small-sample scenarios is still prominent. When the number of training samples is limited, existing methods often struggle to extract discriminative enough feature representations. This problem is particularly significant in hyperspectral image classification because obtaining high-quality labeled samples often requires a large amount of human and time costs. Although existing research has attempted to alleviate this problem through methods such as transfer learning or data augmentation, how to maintain the generalization ability of the model under limited sample conditions remains an important challenge.

[0012] The existence of these scientific problems not only restricts the popularization of hyperspectral image classification methods in practical applications but also provides important innovation space for future research. The present invention will propose a new solution to address the above problems. Summary of the Invention

[0013] Aiming at the limitations of existing methods, the present invention proposes a novel attention-based hybrid network (ATN-Hybrid) for hyperspectral image classification. The core innovation of this method lies in the design of a deterministic-probabilistic hybrid attention mechanism, which extends the traditional single-mode attention mechanism to a dual-branch structure. The deterministic branch generates stable feature weights through explicit feature transformation and normalization, while the probabilistic branch introduces a soft attention mechanism regulated by a temperature parameter to achieve flexible feature selection. The complementary effect of the two mechanisms significantly improves the robustness and adaptability of feature selection. At the same time, the present invention innovatively designs a multi-level feature fusion strategy based on learnable fusion coefficients, which adaptively adjusts the feature importance by dynamically balancing the contributions of the two attentions.

[0014] A method for constructing an adaptive temperature-guided hybrid attention network for hyperspectral image classification, comprising the following steps:

[0015] Step S1: Feature extraction and preprocessing;

[0016] Step S2: Establish a hybrid attention mechanism;

[0017] Step S3: Adaptive feature fusion;

[0018] Step S4: Make a classification decision and output the land cover classification result.

[0019] The adaptive temperature-guided hybrid attention network for hyperspectral image classification constructed by the present invention, compared with the prior art:

[0020] The deterministic-probabilistic hybrid attention mechanism is proposed for the first time. The dual-branch attention structure is constructed by the mutual complement of the explicit channel mapping of the deterministic branch and the temperature-guided soft attention of the probabilistic branch. The deterministic branch uses 1×1 convolution and Sigmoid activation function to generate stable feature weights, while the probabilistic branch introduces a temperature parameter and Softmax function to achieve soft assignment of feature importance. The complementary effect of the two mechanisms improves the flexibility and robustness of feature selection.

[0021] The multi-level feature fusion strategy based on learnable fusion coefficients is innovatively designed, and the dynamic balance of deterministic and probabilistic attention weights is achieved by introducing learnable fusion coefficients. This strategy adopts a multi-level fusion design of "feature concatenation - 1×1 convolution - batch normalization", which introduces attention-enhanced features while retaining the original feature information, realizes the adaptive adjustment of feature importance, and provides richer and more expressive feature representations.

[0022] Extensive experimental validations were carried out on four publicly available hyperspectral datasets, covering different application scenarios such as agricultural, urban, and natural environments. The experimental results show that the method proposed in the present invention can still maintain excellent classification performance under the condition of limited training samples, especially showing significant advantages when dealing with complex ground object categories with similar spectral features. Description of the Drawings

[0023] Figure 1 is a schematic flow chart of the construction method of the present invention.

[0024] Figure 2 is the classification map of the ATN-Hybrid of the present invention on the Indian Pines dataset.

[0025] Figure 3 is the classification map of the ATN-Hybrid of the present invention on the Pavia University dataset.

[0026] Figure 4 is the classification map of the ATN-Hybrid of the present invention on the Kennedy Space Center dataset.

[0027] Figure 5 is the classification map of the ATN-Hybrid of the present invention on the Salinas Valley dataset.

[0028] Figure 6 is the simulation diagram of the training accuracy and validation accuracy obtained by the ATN-Hybrid of the present invention for the Indian Pines dataset.

[0029] Figure 7 is the simulation diagram of the training loss and validation loss obtained by the ATN-Hybrid of the present invention for the Indian Pines dataset.

[0030] Figure 8 is the simulation diagram of the training accuracy and validation accuracy obtained by the ATN-Hybrid of the present invention for the Pavia University dataset.

[0031] Figure 9 is the simulation diagram of the training loss and validation loss obtained by the ATN-Hybrid of the present invention for the Pavia University dataset.

[0032] Figure 10 is the simulation diagram of the training accuracy and validation accuracy obtained by the ATN-Hybrid of the present invention for the Kennedy Space Center dataset.

[0033] Figure 11It is a simulation diagram of the training loss and validation loss obtained by the ATN-Hybrid based on the present invention for the Kennedy Space Center dataset.

[0034] Figure 12 It is a simulation diagram of the training accuracy and validation accuracy obtained by the ATN-Hybrid based on the present invention for the Salinas Valley dataset.

[0035] Figure 13 It is a simulation diagram of the training loss and validation loss obtained by the ATN-Hybrid based on the present invention for the Salinas Valley dataset. Detailed implementation manners

[0036] The technical solutions of the present invention will be described in detail below with reference to the accompanying drawings:

[0037] The ATN-Hybrid hybrid attention network proposed by the present invention is dedicated to solving the problems of feature expression and band selection in hyperspectral image classification. Hyperspectral images have rich spectral dimensional information, and each pixel contains a complete spectral curve, but not all bands are equally important for the classification task. To effectively utilize this information, the present invention designs a novel hybrid attention mechanism that combines deterministic and probabilistic attention strategies to achieve adaptive enhancement of hyperspectral features.

[0038] The innovation of the present invention lies in the design of the hybrid attention mechanism. This mechanism includes two complementary branches: the deterministic branch generates stable weight representations for each feature channel through explicit feature transformation and normalization operations; the probabilistic branch introduces a soft attention mechanism adjusted by a temperature parameter to achieve adaptive allocation of feature importance by adjusting the temperature coefficient. The present invention will introduce in detail the design principles and implementation methods of each key module.

[0039] A construction method for an adaptive temperature-guided hybrid attention network for hyperspectral image classification includes the following steps:

[0040] Step S1: Feature extraction and preprocessing module

[0041] In hyperspectral image processing, the original data usually has the problem of inconsistent scales, and there are large differences in the numerical ranges between different bands. Therefore, standardization processing needs to be carried out first. For the input hyperspectral image:

[0042]

[0043] where H and W represent the spatial dimensions, and L represents the number of spectral bands. Each band is normalized:

[0044]

[0045] Among them, μ l and σ l respectively represent the mean and standard deviation of the l-th band. The normalization operation can eliminate the scale differences between different bands, making the subsequent feature learning more stable and efficient.

[0046] To adapt to the deep learning framework, the data needs to be reorganized into the form of a 4D tensor:

[0047]

[0048] where B is the batch size and C is the number of feature channels. This reorganization maintains the spatial locality of the data and facilitates subsequent convolutional operations.

[0049] Step S2: Design of the hybrid attention mechanism

[0050] In the ATN-Hybrid model proposed in the present invention, the hybrid attention mechanism works collaboratively through two branches to enhance the feature selection ability in hyperspectral image classification. Specifically, the deterministic branch generates stable feature weights through explicit feature transformation and normalization, while the probabilistic branch assigns dynamic weights to each feature through a temperature-guided soft attention mechanism. These two attention mechanisms play different roles in the feature selection process.

[0051] To further optimize the feature fusion process, we introduce a learnable fusion coefficient, which is used to dynamically balance the contributions of the two attention mechanisms. Specifically, the fusion coefficient is adaptively adjusted according to the characteristics of different input features, thereby determining the weight allocation between the deterministic branch and the probabilistic branch. Specifically, when features have different importance in different channels or spatial dimensions, the fusion coefficient allocates weights between the two attention branches, enabling the network to adjust its attention strategy according to the specific needs of the features.

[0052] The temperature-guided soft attention mechanism can achieve the importance assignment of features by adjusting the temperature parameter T. At high temperatures (T > 1), the temperature mechanism tends to evenly distribute attention, retaining more feature information; while at low temperatures (T < 1), the temperature mechanism strengthens the high-response features, focusing on the key feature dimensions. On this basis, the learnable fusion coefficient can flexibly adjust the contributions of these two mechanisms, enabling the network to automatically optimize this balance during training to adapt to different input features and data distributions.

[0053] In this way, the ATN-Hybrid model can not only optimize features globally but also adaptively adjust the intensity of attention in local feature selection, thereby improving its adaptability to complex scenarios.

[0054] Step S2.1 Deterministic Attention Branch

[0055] Step S2.1.1 Cross-Channel Feature Mapping Stage

[0056] The deterministic branch generates deterministic weights for each feature channel through explicit feature transformation and normalization operations. Its calculation process is as follows:

[0057] F d = Conv 1×1 (X) (4)

[0058] The linear combination between feature channels is achieved through 1×1 convolution, reducing the computational complexity while maintaining the feature expression ability. The design effectively utilizes the correlation between adjacent bands in hyperspectral data, realizes information interaction between channels, and enhances the feature expression ability.

[0059] Step S2.1.2 Feature Distribution Standardization Stage

[0060] Batch normalization is performed on the features to improve training stability:

[0061] F d_bn = BN(F d ) (5)

[0062] Batch normalization effectively reduces internal covariate shift and enhances model stability by dynamically adjusting the mean and variance of the feature distribution. The normalized feature distribution not only accelerates model convergence but also plays a regularization role, improving the model's generalization ability.

[0063] Step S2.1.3 Deterministic Weight Generation Stage

[0064] Deterministic weights are generated through the Sigmoid function:

[0065] W d = σ(F d_bn ) (6)

[0066] where The saturation characteristic of the Sigmoid activation function provides a natural boundary constraint for weight generation, effectively suppressing the influence of outliers. The smooth non-linear transformation ensures the continuity of the weights, avoids mutation phenomena, and provides a stable feature selection mechanism.

[0067] Step S2.2 Probabilistic Attention Branch

[0068] Step S2.2.1 Feature Extraction

[0069] Feature extraction also uses 1×1 convolution:

[0070] F p = Conv1×1 (X) (7)

[0071] Step S2.2.2 Distributed Standardization

[0072] Batch Normalization processing:

[0073] F p_bn = BN(F p ) (8)

[0074] Step S2.2.3 Temperature Adjustment

[0075] Introduce the temperature parameter T to adjust the distribution:

[0076] F p_t = F p_bn / T (9)

[0077] A large temperature parameter (T > 1) weakens the differences between feature responses, generates a more uniform attention distribution, and is beneficial to maintaining more feature information. A small temperature parameter (0 < T < 1) strengthens the differences in feature responses, makes the attention more concentrated in the high-response regions, and highlights key features

[0078] Step S2.2.4 Probability Weight Generation

[0079] Use Softmax to generate probability weights:

[0080] W p = softmax(F p_t ) (10)

[0081] Step S3 Adaptive Feature Fusion

[0082] The feature fusion module adopts a three-step design strategy to effectively integrate features through feature concatenation, non-linear transformation, and normalization operations. The core of this module lies in introducing learnable fusion coefficients to dynamically balance the contributions of the deterministic and probabilistic branches. It specifically includes the following three key steps:

[0083] Step S3.1 Adaptive Weight Fusion Strategy

[0084] First, perform the adaptive fusion of mixed weights. The core of the adaptive fusion is to introduce a learnable fusion coefficient α. The fusion coefficient provides the model with adaptive capabilities in different scenarios and enhances the generality of the algorithm

[0085] W = αW d +(1 - α)W p (11)

[0086] Here, the fusion coefficient α is a learnable parameter that can adaptively adjust the proportion of the two types of attention according to different data.

[0087] Step S3.2 Feature Enhancement Mechanism

[0088] After weight fusion, the original features are enhanced through the attention mechanism, achieving importance weighting of the features, highlighting the contribution of key information, and preserving the integrity of spatial information through channel-wise weighting. The features are enhanced using the attention weights:

[0089] Y = X ⊙ W (12)

[0090] Step S3.3 Multi-level Feature Fusion Strategy

[0091] The final feature fusion adopts a three-step design idea: Feature concatenation preserves the complete information of the original features and the enhanced features, avoiding information loss; 1×1 convolution realizes information recombination between channels, providing opportunities for feature interaction; batch normalization and ReLU activation introduce non-linear transformations, enhancing the feature expression ability. The final feature fusion is achieved through feature concatenation and non-linear transformation:

[0092] F c = [X; Y] (13)

[0093] F f = Conv 1×1 (F c ) (14)

[0094] F = ReLU(BN(F f )) (15)

[0095] This multi-level fusion design not only preserves the original feature information but also introduces attention-enhanced features, providing richer feature expressions.

[0096] Step S4: Make classification decisions and output the land cover classification results. The specific process is as follows:

[0097] Step S4.1 Model training process:

[0098] For each epoch loop:

[0099] Forward propagation to calculate the loss:

[0100] Backward propagation to update the parameters:

[0101] Evaluate on the validation set: acc val = evaluate(model, X val )

[0102] Early stopping check: Stop training if the validation set accuracy has not improved for N consecutive epochs;

[0103] Learning rate adjustment: lr = scheduler.step().

[0104] Step S4.1 Classification decision:

[0105] Fully connected layer mapping: Z = FC(F);

[0106] Class prediction: P = argmax(softmax(Z)) is the final ground object classification result map.

[0107] To systematically elaborate on the implementation process of the ATN-Hybrid algorithm, the input, output, and key processing steps of the algorithm are described in detail below. The algorithm takes the original hyperspectral image and its corresponding ground object class labels as input, and through a series of operations such as data preprocessing, model initialization, deterministic-probabilistic dual-branch attention mechanism, and adaptive feature fusion, finally outputs high-precision ground object classification results.

[0108]

[0109]

[0110]

[0111] I. Algorithm Advantage Analysis

[0112] 1. Hybrid attention mechanism design

[0113] The present invention creatively proposes a deterministic-probabilistic hybrid attention mechanism. The deterministic branch generates stable feature weights through explicit channel mapping and normalization, ensuring the reliability of feature selection; the probabilistic branch realizes the adaptive allocation of feature importance through a temperature-guided soft attention mechanism, providing flexible feature selection ability. The complementary effect of the two mechanisms significantly improves the robustness and adaptability of feature selection.

[0114] 2. Adaptive feature fusion strategy

[0115] An innovative multi-level feature fusion strategy based on learnable fusion coefficients is designed. Through a three-step design of "feature concatenation - non-linear transformation - normalization", while retaining the original feature information, attention-enhanced features are introduced to realize the dynamic optimization of feature importance. This strategy provides an effective technical solution for the processing of complex spectral features.

[0116] These innovative designs construct a theoretically complete feature learning framework, providing new methodological ideas for hyperspectral image classification. The hybrid attention mechanism expands the expression ability of the traditional attention mechanism, and the adaptive fusion strategy provides a more flexible way of feature integration.

[0117] II. Experimental Datasets and Experimental Verification

[0118] 1. Hyperspectral Datasets

[0119] To evaluate the performance of the proposed method, the present invention selects four representative publicly available hyperspectral datasets for experimental verification:

[0120] The Indian Pines dataset was collected from the Indian Pine experimental area in northwestern Indiana and acquired by the AVIRIS sensor. The dataset contains 145×145 pixels and a total of 200 spectral channels (wavelength range: 400 - 2500 nm). The land cover classes mainly include different types of crops, forest land, and other natural vegetation, with a total of 16 land cover classes. In the experiment, only 9% of the samples are used for model training. This setting of a low training sample ratio is closer to the actual application scenario and helps to verify the practicality of the algorithm.

[0121] The Pavia University dataset was collected by the ROSIS sensor in the area of the University of Pavia, Italy. The data contains 610×340 pixels and 103 spectral channels (wavelength range: 430 - 860 nm), mainly including 9 typical urban land cover classes such as buildings, roads, and vegetation. The present invention uses 3% of the samples for training. This strict experimental setting poses higher requirements for the feature learning ability of the algorithm.

[0122] The Kennedy Space Center (KSC) dataset was acquired by the NASA AVIRIS sensor over the Kennedy Space Center in Florida. The data coverage area is 512×614 pixels and contains 176 spectral bands (wavelength range: 400 - 2500 nm), with a total of 13 natural vegetation and wetland classes. In the experiment, 8% of the samples are used for training to evaluate the classification performance of the algorithm in a complex natural environment.

[0123] The Salinas Valley dataset was also collected by the AVIRIS sensor and is located in California. The data size is 512×217 pixels and contains 204 spectral bands (after removing 20 water vapor absorption bands), with a total of 16 different types of crop classes. The present invention only uses 3% of the samples for training, which not only tests the feature extraction ability of the algorithm but also verifies its potential in precision agriculture applications.

[0124] These datasets cover different application scenarios such as agriculture, urban areas, and natural environments, with different spatial resolutions, spectral characteristics, and land cover classes, and can comprehensively evaluate the performance of classification algorithms. Especially when using a lower training sample ratio, the practical application value of the algorithm can be more reflected.

[0125] 2. Experimental Methods

[0126] In this experiment, a variety of common convolutional neural networks (CNNs) and attention mechanisms were selected as comparison methods to evaluate the effectiveness of the proposed ATN-Hybrid method. The following is a brief introduction to each comparison method:

[0127] CNN_deep: The deep convolutional neural network (CNN_deep) is a basic convolutional neural network architecture designed to extract spatial features of images through multiple layers of convolutional operations. This method is suitable for processing high-dimensional image data but lacks in-depth mining of different spectral information in hyperspectral image classification.

[0128] CNN_inception: The Inception model performs multi-scale feature extraction by introducing convolutional kernels of different sizes, thus enhancing the diversity and expressiveness of the model. It is suitable for capturing multiple scale features of images and has good performance in image classification tasks.

[0129] CNN_resnet: ResNet (Residual Networks) solves the problem of gradient disappearance in the training of deep networks by introducing residual connections, enabling the network to be deeper and more effective in learning features. In hyperspectral image classification, ResNet exhibits strong feature extraction capabilities, especially when dealing with high-dimensional data.

[0130] ATN_cbam: The convolutional block attention module (CBAM) combines channel attention and spatial attention to adaptively adjust the feature importance of each channel and spatial position. This method enhances the network's ability to focus on key features through two independent attention modules.

[0131] ATN_eca: The efficient channel attention (ECA) mechanism quickly models the dependencies between channels through one-dimensional convolution without dimensionality reduction operations, thus improving efficiency and maintaining good performance. In hyperspectral image classification, ECA exhibits strong modeling capabilities for channel features.

[0132] ATN_nonlocal: The Non-local attention mechanism models global dependencies through self-attention operations, captures remote spatial information in images, and thus enhances the model's ability to process complex scenes. It is suitable for the classification of hyperspectral images, especially when dealing with long-range dependent features.

[0133] ATN_bam: The boundary attention module (BAM) focuses on the boundary information of images, aiming to enhance the model's perception ability of object boundaries. This method can more precisely identify the boundaries of ground object categories by optimizing boundary features, especially suitable for fine-grained classification tasks.

[0134] The selection of these comparison methods covers classic convolutional neural network (CNN) architectures, attention mechanisms, and network structures that combine both, thus providing a series of powerful benchmarks to help us evaluate the advantages of the ATN-Hybrid method in hyperspectral image classification.

[0135] To comprehensively evaluate the performance of the proposed method, the present invention uses a limited proportion of training samples for experimental verification. This experimental setting with a low proportion of training samples is of great significance: on the one hand, it is often difficult to obtain a large number of labeled samples in practical applications, and such a setting is closer to the actual application scenario; on the other hand, the classification performance under the condition of limited samples can better reflect the feature learning ability and generalization performance of the algorithm. Based on the above experimental setting, the present invention conducts a systematic evaluation of the ATN-Hybrid method, and the specific result analysis is as follows:

[0136] 3. Experiments based on the Indian Pines dataset:

[0137] Combined with Table 1 (Comparison of classification accuracies of different methods based on the Indian Pines dataset (expressed as percentages)), for the Indian Pines dataset, the ATN-Hybrid method shows significant classification advantages. Compared with other comparison methods, its overall accuracy (96.01%), average accuracy (93.25%), and Kappa coefficient (95.45%) are all optimal. Notably, this method performs outstandingly in dealing with crop categories with spectral similarity. For example, it achieves 100% classification accuracy in categories such as "Corn", "Grass-pasture-mowed", and "Woods"; it also achieves significant improvements in the classification of crop categories such as "Corn-no-till" (95.23%) and "Soybeans-min-till" (98.52%), indicating that the ATN-Hybrid method can effectively capture and distinguish subtle spectral differences in complex agricultural scenarios. In contrast, other methods generally show varying degrees of performance degradation in these challenging categories, especially in categories with a small number of samples (such as "Alfalfa" and "Oats") where the classification accuracies are generally low.

[0138] Table 1 Comparison of classification accuracies of different methods based on the Indian Pines dataset (expressed as percentages)

[0139]

[0140] 4. Experiments based on the Pavia University dataset:

[0141] Combined with Table 2 (Comparison of classification accuracies of different methods based on the Pavia University dataset (expressed as percentages)), on the Pavia University dataset, the ATN-Hybrid method also demonstrated excellent classification performance, achieving an overall accuracy of 95.23%, an average accuracy of 91.71%, and a Kappa coefficient of 93.67%. This method showed strong feature discrimination ability when dealing with urban land cover classes. In particular, it achieved a perfect classification of 100% in the "Metal Sheets" class and maintained relatively high classification accuracies in major urban land cover types such as "Asphalt" (97.95%) and "Meadows" (98.49%). It is worth noting that in the "Gravel" class with complex texture features, although the classification accuracy was relatively low (61.35%), it was still better than some of the comparison methods, which reflects that this method still has certain advantages when dealing with objects with complex spatial-spectral features.

[0142] Table 2 Comparison of classification accuracies of different methods based on the Pavia University dataset (expressed as percentages)

[0143]

[0144] 5. Experiments based on the Kennedy Space Center dataset:

[0145] Combined with Table 3 (Comparison of classification accuracies of different methods based on the Kennedy Space Center dataset (expressed as percentages)), the experimental results for the Kennedy Space Center dataset further verified the effectiveness of the ATN-Hybrid method. This method achieved the best performance on this dataset, with an overall accuracy of 97.37%, an average accuracy of 95.76%, and a Kappa coefficient of 97.07%. It performed particularly well in the classification of natural vegetation classes. For example, it achieved a classification accuracy of 100% in the "Scrub" and "Spartina marsh" classes and was also significantly better than other comparison methods in complex vegetation types such as "Slash pine" (91.22%) and "Oak / broadleaf hammock" (85.31%). This indicates that the ATN-Hybrid method can effectively handle land cover classes with similar spectral features in natural ecosystems, demonstrating its superiority in fine-grained land cover classification tasks.

[0146] Table 3 Comparison of classification accuracies of different methods based on the Kennedy Space Center dataset (expressed as percentages)

[0147]

[0148] 6. Experiments based on the Salinas Valley dataset:

[0149] Combined with Table 4 (Comparison of classification accuracies of different methods based on the Salinas Valley dataset (expressed as percentages)), for the Salinas Valley dataset, the ATN-Hybrid method continued to maintain excellent classification performance, with an overall accuracy of 95.99%, an average accuracy as high as 97.92%, and a Kappa coefficient reaching 95.54%. In terms of agricultural vegetation classification, the method achieved perfect classification for multiple crop categories. For example, the classification accuracies of "Broccoli-green-weeds-1", "Stubble", and "Lettuce-romaine-6wk" all reached 100%. Especially when dealing with crops at different growth stages with similar spectral characteristics (such as lettuce at various stages), the ATN-Hybrid method demonstrated excellent discrimination ability, highlighting the potential of this method in precision agriculture applications. Even in challenging categories such as "Grapes-untrained" (86.41%) and "Vineyard-untrained" (95.01%), the method maintained relatively stable classification performance.

[0150] Table 4 Comparison of classification accuracies of different methods based on the Salinas Valley dataset (expressed as percentages)

[0151]

[0152]

[0153] III. Analysis of the training process

[0154] 1. Analysis of the training process and convergence:

[0155] Such as Figure 6 、 8As shown in Figures 10 and 12, the ATN-Hybrid model demonstrates excellent convergence characteristics on each dataset. The training process can be clearly divided into three stages: rapid learning (1 - 30 epochs), stable optimization (31 - 100 epochs), and fine-tuning (>100 epochs). In the initial stage, the model exhibits strong feature learning ability, and the training accuracy increases rapidly (e.g., from 36.23% to over 80% on the Indian Pines dataset); subsequently, it enters the stable optimization stage, and the validation accuracy steadily increases to the range of 90% - 95%; finally, in the fine-tuning stage, the model successfully reaches validation accuracies of 96.01% (Indian Pines), 95.23% (Pavia University), 97.37% (Kennedy Space Center), and 95.99% (Salinas Valley), and the training process shows extremely high stability, with the fluctuation range controlled within 0.5%.

[0156] 2. Analysis of the behavior of the loss function:

[0157] The dynamic change of the loss function shows typical exponential decay characteristics, and the training loss and validation loss maintain a high degree of synchronization. As Figure 7 , 9 As shown in Figures 11 and 13, in the initial stage, the loss value rapidly drops from above 2.0 to below 0.5, reflecting the efficient learning ability of the model; in the middle stage, the loss function shows a steady downward trend with a significantly reduced fluctuation range; in the final stage, the loss value stably converges to the range of 0.1 - 0.2, and the gap between the training loss and the validation loss is maintained within a small range (<0.05). This characteristic strongly proves that the model has excellent generalization performance and stability.

[0158] 3. Comprehensive evaluation of the model performance:

[0159] The ATN-Hybrid model demonstrates excellent performance in all evaluation metrics: in terms of classification accuracy, the validation accuracies of all four datasets exceed 95%, and the Kennedy Space Center dataset reaches the highest of 97.37%; in terms of stability, the gap between the training and validation accuracies is controlled within 1%, indicating that the model has excellent generalization ability; in terms of reliability, the Kappa coefficients of all datasets exceed 0.93, verifying the high credibility of the classification results; in addition, the model shows strong robustness on datasets with different complexities, and the performance difference is controlled within 2.5%, fully demonstrating the generality and practical value of this method.

[0160] IV. Overall experimental analysis

[0161] Through comprehensive analysis of the experimental results on four benchmark datasets, the ATN-Hybrid method shows significant advantages in overall performance. On the Indian Pines, Kennedy Space Center, and Pavia University datasets, this method is significantly better than the comparative methods, especially outstanding in dealing with complex scenarios and under the condition of low training samples. Although it has comparable performance with the ATN_bam method on the Salinas Valley dataset, considering the obvious advantages of this method on other datasets, it fully verifies the effectiveness of the hybrid attention mechanism in improving the generalization ability and adaptability of the model.. By comparing and analyzing the classification performance of different methods on the four datasets, significant performance differences can be observed among the methods:

[0162] 1. Traditional CNN-based methods:

[0163] As a basic deep learning architecture, the CNN_deep method performs okay in simple scene classification. For example, it achieves 100% classification accuracy in the regular farmland categories ("Broccoli-green-weeds-1" and "Stubble") of the Salinas Valley dataset. However, it has obvious limitations in dealing with complex scenarios. Especially for small sample categories such as "Alfalfa" (21.43%) and "Oats" (38.89%) in the Indian Pines dataset, the classification accuracy is relatively low, indicating the insufficiency of its feature extraction ability.

[0164] 2. Improved CNN architecture methods:

[0165] By introducing a multi-scale feature extraction mechanism, the CNN_inception method performs well in urban scene classification. For example, it obtains an overall accuracy of 95.90% on the Pavia University dataset. However, it is still insufficient in dealing with land cover categories with similar spectral features. For example, the accuracy of the "Slash pine" category on the Kennedy Space Center dataset only reaches 74.32%. Although the CNN_resnet method adopts a deeper network structure, its classification performance is unstable, and significant performance degradation occurs in some categories (such as "Metal Sheets" on Pavia University, 33.56%).

[0166] 3. Single attention mechanism methods:

[0167] ATN_cbam improves the overall classification performance by combining spatial and channel attention mechanisms, but there are still deficiencies in the fine feature discrimination. For example, it only reaches an accuracy of 29.17% in the "Gravel" category of the Pavia University dataset. ATN_eca adopts a lightweight channel attention strategy, which improves the computational efficiency, but the overall classification accuracy is relatively low. For example, it only reaches an overall accuracy of 79.17% in the Indian Pines dataset.

[0168] 4. Complex attention mechanism methods:

[0169] ATN_nonlocal enhances feature representation by establishing long-range dependencies and performs well in some complex scenarios. For example, it reaches 98.48% in the "Asphalt" category of the Pavia University dataset. However, it only reaches an accuracy of 32.43% in the "Slash pine" category of the Kennedy Space Center dataset, showing the limitations of its generalization ability. ATN_bam improves the feature extraction ability through a branched attention structure and achieves an overall accuracy of over 90% on multiple datasets. However, it still faces challenges when dealing with categories with severe spectral aliasing.

[0170] 5. Hybrid attention mechanism methods:

[0171] The proposed ATN-Hybrid method in the present invention achieves optimal or near-optimal classification performance on all datasets by fusing deterministic and probabilistic attention strategies. The overall accuracy of this method on datasets such as Indian Pines (96.01%), Kennedy Space Center (97.37%), Pavia University (95.23%), and Salinas Valley (95.99%) is significantly better than other methods. Especially when dealing with complex feature categories with similar spectral characteristics, it shows significant advantages, fully demonstrating the effectiveness of the hybrid attention strategy.

[0172] Through comparative analysis, it can be seen that different types of classification methods have their own characteristics. However, the ATN-Hybrid method successfully solves the key problems in hyperspectral image classification and achieves an overall improvement in classification performance through an innovative hybrid attention mechanism and a multi-level feature fusion strategy. This provides a new technical paradigm for hyperspectral remote sensing image classification.

[0173] V. Conclusion

[0174] The proposed ATN-Hybrid method in the present invention makes the following key innovative contributions to the field of hyperspectral image classification:

[0175] Innovative Hybrid Attention Mechanism: The present invention pioneered the deterministic-probabilistic hybrid attention mechanism, which complements each other through the explicit channel mapping of the deterministic branch and the temperature-guided soft attention of the probabilistic branch, providing a more comprehensive feature selection strategy. Experimental results show that this hybrid mechanism has significant advantages in processing complex spectral features. For example, the classification accuracy of the "Slash pine" category in the Kennedy Space Center dataset reaches 91.22%, showing a significant improvement compared to traditional single-attention methods (such as 79.05% of ATN_cbam); in small-sample categories such as "Oats" in the Indian Pines dataset, the classification accuracy is improved from 38.89% of traditional methods to 94.44%.

[0176] Adaptive Feature Fusion: An innovative multi-level feature fusion strategy based on learnable fusion coefficients is designed to achieve dynamic adaptive adjustment of attention weights. While retaining the original feature information, this design effectively enhances the expression of key features. On the Pavia University dataset, for the "Bricks" category with complex texture features, the proposed method achieves a classification accuracy of 95.63%, significantly superior to the baseline method; in the classification of crops at different growth stages in the Salinas Valley dataset, such as the "Lettuce-romaine" series categories, high-precision classifications above 96% are achieved for all of them.

[0177] Excellent Classification Performance: Comprehensive experimental verification on four benchmark datasets shows that the proposed method still maintains excellent performance under the condition of limited training samples. Specifically, using only 9% of the training samples in the Indian Pines dataset achieves an overall accuracy of 96.01%, using 3% of the training samples in the Pavia University dataset obtains a classification accuracy of 95.23%, using 8% of the training samples in the Kennedy Space Center dataset achieves an accuracy of 97.37%, and using 3% of the training samples in the Salinas Valley dataset reaches an accuracy of 95.99%. These results fully demonstrate the reliability and effectiveness of the proposed method in practical applications.

[0178] The theoretical significance of the present invention lies in the first proposed concept of the deterministic-probabilistic hybrid attention mechanism, opening up a new research direction for feature selection and representation learning in hyperspectral image analysis. The temperature-guided attention mechanism and the adaptive fusion strategy provide innovative solutions for dealing with the high dimensionality and complex spectral relationships of hyperspectral data. Experimental results prove that this hybrid mechanism can effectively handle the key challenges in hyperspectral image classification, including spectral feature similarity, insufficient training samples, and other problems.

[0179] Future research can further explore the application of the hybrid attention mechanism in other remote sensing tasks and incorporate more spatial context information to enhance the processing ability of complex scenes.

Claims

1. A construction method of an adaptive temperature-guided hybrid attention network for hyperspectral image classification, characterized in that It includes the following steps: Step S1: Feature extraction and preprocessing; Step S2: Establish a hybrid attention mechanism; Step S3: Adaptive feature fusion; Step S4: Make a classification decision and output the land cover classification result.

2. The construction method of the adaptive temperature-guided hybrid attention network for hyperspectral image classification according to claim 1, wherein The feature extraction and preprocessing in the above step S1 are specifically as follows: In hyperspectral image processing, first, standardization processing is required. For the input hyperspectral image: where H and W are the spatial dimensions, and L is the number of spectral bands. Normalize each band: Among them, μ l and σ l respectively represent the mean and standard deviation of the l-th band; the normalization operation can eliminate the scale differences between different bands; To adapt to the deep learning framework, the data is reorganized into a 4D tensor form: where B is the batch size and C is the number of feature channels.

3. The construction method of the adaptive temperature-guided hybrid attention network for hyperspectral image classification according to claim 1, characterized in that The establishment of the hybrid attention mechanism in the above step S2 is specifically as follows: Step S2.1 Deterministic attention branch Cross-channel feature mapping: The deterministic branch generates deterministic weights for each feature channel through explicit feature transformation and normalization operations; The calculation process is as follows: F d = Conv 1×1 (X) (4) Realize the linear combination between feature channels through 1×1 convolution, reducing the computational complexity while maintaining the feature expression ability; Normalize the feature distribution: Perform batch normalization on the features to improve the training stability: F d_bn = BN(F d ) (5) Batch normalization reduces the internal covariate shift by dynamically adjusting the mean and variance of the feature distribution; Generate deterministic weights: Generate deterministic weights through the Sigmoid activation function: W d = σ(F d_bn ) (6) Among them The saturation characteristic of the Sigmoid activation function provides a natural boundary constraint for weight generation, suppressing the influence of outliers; Step S2.2 Probabilistic attention branch Feature extraction: Feature extraction uses 1×1 convolution: F p = Conv 1×1 (X) (7) Distribution standardization: Batch normalization processing: F p_bn = BN(F p ) (8) Temperature adjustment: Introduce the temperature parameter T to adjust the distribution: F p_t = F p_bn / T (9) A large temperature parameter, i.e., T>1, will weaken the differences between feature responses, generate a more uniform attention distribution, and be able to retain more feature information; A small temperature parameter, i.e., 0<T<1, strengthens the differences in feature responses, makes the attention more concentrated on the high-response region, and highlights the key features; Probability weight generation W p = softmax(F p_t ) (10) Use Softmax to generate probability weights.

4. The construction method of the adaptive temperature-guided hybrid attention network for hyperspectral image classification according to claim 1, characterized in that The adaptive feature fusion in the above step S3 is specifically as follows: Step S3.1 Adaptive weight fusion strategy First, perform adaptive fusion of the hybrid weights. Introduce a learnable fusion coefficient α. The fusion coefficient provides the model's adaptive ability in different scenarios and enhances the generality of the algorithm W = αW d +(1 - α)W p (11) Among them, the fusion coefficient α is a learnable parameter that can adaptively adjust the proportion of the two attentions according to different data; Step S3.2 Feature enhancement mechanism After weight fusion, enhance the original features through the attention mechanism, realize importance weighting of the features, highlight the contribution of key information, and retain the integrity of the spatial information through channel-wise weighting. Apply attention weights to enhance the features: Y = X⊙W (12) Step S3.3 Multi-level feature fusion strategy Feature concatenation retains the complete information of the original features and enhanced features, avoiding information loss; 1×1 convolution realizes information recombination between channels, providing an opportunity for feature interaction; batch normalization and ReLU activation introduce non-linear transformations, enhancing the feature expression ability; F c = [X; Y] (13) F f = Conv 1×1 (F c )(14) F = ReLU(BN(F f )) (15) Realize the final feature fusion through feature concatenation and non-linear transformation.

5. The construction method of the adaptive temperature-guided hybrid attention network for hyperspectral image classification according to claim 1, characterized in that The classification decision in the above step S4 and output of the land cover classification result are specifically as follows: Step S4.1 Model training process: Each epoch loop: Forward propagation to calculate the loss: Backpropagation to update parameters: Validation set evaluation: acc val = evaluate(model, X val ); Early stopping check: Stop training if the validation set accuracy does not improve for N consecutive epochs; Learning rate adjustment: lr = scheduler.step(); Step S4.1 Classification decision: Fully connected layer mapping: Z = FC(F); Class prediction: P = argmax(softmax(Z)) is the final ground object classification result map.

Citation Information

Cited By

  • Carrier phase hopping detection method and device and storage medium

    CN120742359A