Hyperspectral image accurate classification method based on improved multi-scale attention and Transform network

Through the improved multi-scale attention and Transformer network, combined with dynamic spatial attention units, multi-core fusion attention modules and cross-attention mechanism, the feature redundancy and information dispersion problems in hyperspectral image classification are solved, and high-precision space-spectral feature fusion and multi-scale modeling are achieved, which improves classification performance.

CN120047816AActive Publication Date: 2025-05-27HAINAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411967702.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-27
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

The prior art has problems of feature redundancy and information overdispersion caused by the multi-head self-attention mechanism in hyperspectral image classification, making it difficult to effectively perform multi-scale feature extraction and spatial-spectral feature fusion.

Method used

Using an improved multi-scale attention and Transformer network, multi-core fusion attention module and cross-attention mechanism are used to achieve multi-scale feature extraction and deep fusion through dynamic spatial attention units, multi-core fusion attention modules and cross-attention mechanisms, combined with dynamic weight convolution and different-scale convolution kernels.

Benefits of technology

The classification accuracy of hyperspectral images is significantly improved, which is superior to existing methods, and has excellent performance on multiple public data sets, especially in complex geographic classification and uneven category distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047816A_ABST
    Figure CN120047816A_ABST
Patent Text Reader

Abstract

The invention discloses a hyperspectral image accurate classification method based on improved multi-scale attention and a Transform network. The method comprises the following steps: firstly, extracting space information of a dynamic space attention unit; then multi-scale feature extraction is carried out by a multi-core fusion attention module; and finally, the information fusion is promoted by a cross attention Swin Transform module. Specifically, the method comprises the following steps: inputting features subjected to multi-scale attention processing into LayerNorm; then, a multi-head attention mechanism is introduced, so that the network can process features of different scales; multi-head cross attention is introduced, different scale features are well fused by the network, and MLP is added after a multi-head attention mechanism for integrating and extracting features; according to the method, the problems of feature redundancy and excessive information dispersion possibly caused by a multi-head self-attention mechanism in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning in remote sensing image classification, and particularly relates to a hyperspectral image precise classification method based on improved multi-scale attention and Transformer network. Background Technique

[0002] Hyperspectral remote sensing technology has been widely applied in recent years in fields such as earth science, environmental monitoring, agriculture, and urban planning. Compared with traditional multi-spectral images, hyperspectral images can provide hundreds of continuous bands, which provides richer information for more detailed object recognition and classification. In HSI, each pixel represents the reflection spectral characteristics of the target surface in multiple bands, which makes HSIC have significant advantages in many application scenarios, especially in the fields of land cover classification and environmental monitoring.

[0003] HSIC is a technology for accurately classifying land covers in images according to the spectral characteristics of pixels. Since HSI contains a large amount of band information, traditional image classification methods often face challenges of dimensionality disaster and redundant information when processing HSI, resulting in poor classification accuracy. In recent years, with the booming development of machine learning and deep learning technologies, more and more advanced algorithms have been introduced into HSIC. For example, methods such as support vector machine (SVM), random forest (RF), convolutional neural network (CNN), recurrent neural network (RNN), and long short-term memory network (LSTM) have all made significant progress in the classification accuracy and efficiency of HSI. Especially deep learning methods, relying on their ability of feature extraction and end-to-end training, have become important tools for solving the HSIC problem.

[0004] Huang et al. proposed a method using one-dimensional convolutional neural network (1D-CNN) to classify and identify textile fiber HSI. Zhao et al. used two-dimensional convolutional neural network (2D-CNN) to extract high-level spatial features of HSI and stacked the spatial and spectral features to achieve HSIC. Kanthi et al. proposed a 3D deep feature extraction CNN model that can simultaneously utilize the spectral and spatial information in HSI for HSIC and achieved good classification performance. Ge et al. combined 2D-CNN and 3D-CNN and designed a HSI classification model, and achieved good results on four public HSI datasets.

[0005] The Transformer model was first used to handle natural language processing (NLP) tasks. Leveraging the self-attention mechanism and excellent long-range dependency modeling capabilities, the Transformer model has achieved great success. The self-attention mechanism of the Transformer can dynamically focus on the key parts in the data, demonstrating powerful feature extraction capabilities in fields such as NLP and computer vision (CV). With the introduction of the Vision Transformer (ViT), the Transformer model has been introduced into image classification tasks and shown excellent performance. Different from traditional CNNs, the Transformer can directly capture the global features of images without relying on convolutional kernels, making it particularly suitable for processing high-dimensional data. In HSIC, the introduction of the Transformer provides a new approach to solving the problem of spatial-spectral feature fusion.

[0006] Mei et al. proposed a Group-Aware Hierarchical Transformer (GAHT) for HSIC, which solved the problem of over-dispersion in feature extraction by multi-head self-attention. Yang et al. proposed a method of embedding convolutional operations into the transformer structure to capture subtle spectral differences and convey local spatial context information, improving the classification performance. To solve the problem that it is difficult for CNN to extract deep semantic features, Sun et al. proposed a spectral–spatial feature tokenization transformer (SSFTT) model to extract spectral-spatial features and high-level semantic features in HSI, and achieved good classification results on three standard datasets. Zhang et al. proposed an HSIC model combining multiple attention and Transformer, which solved the problem that the network is easily affected by surrounding redundant information during the training phase, resulting in inaccurate feature extraction and poor model generalization ability. First, they used spatial attention (SA) and channel attention (CA) to focus on more important information parts, then used the tokenizer module to perform semantic-level representation of different types of ground objects, and then used the Transformer encoder module for deep semantic feature extraction. The results show that the model performs well in extracting the spatial-spectral features of HSI and understanding the semantic level. To improve the performance of traditional HSIC tasks, Huang et al. proposed a Transformer based on spectral-spatial vision foundation model (SS-VFMT). Secondly, based on SS-VFMT, to solve the generalized zero-shot classification task, they proposed a Transformer based on spectral-spatial vision language (SS-VLFMT), providing a new solution idea for HSI zero-shot classification. Guo et al. utilized the multi-attention mechanism in Swin-Transformer to make full use of rich discriminative information and designed an end-to-end network, further improving the classification performance of HSI.

[0007] Although these methods have further improved the performance of the model, there are still some difficult problems to be solved, which are as follows.

[0008] 1. Hyperspectral images usually contain features at different scales. Traditional CNNs have advantages in multi-scale feature extraction, while native Transformers perform limitedly in this aspect. Although Swin Transformer achieves multi-scale modeling to a certain extent through hierarchical window attention mechanism, there are still deficiencies in multi-scale feature fusion of hyperspectral images. 2. Hyperspectral images contain rich spatial and spectral information, which are highly correlated. However, there are still challenges in achieving deep fusion of spatial and spectral features in the Transformer structure. 3. The multi-head self-attention mechanism introduces powerful feature expression ability in hyperspectral image classification, but too many attention heads may lead to feature redundancy and over-dispersion of information, affecting the discriminative ability of the model. For example, GAHT has been improved to solve this problem, but the phenomena of feature dispersion and redundancy still exist. Summary of the Invention

[0009] The object of the present invention is to provide an accurate classification method for hyperspectral images based on improved multi-scale attention and Transformer network, which solves the problem that the multi-head self-attention mechanism in the prior art may lead to feature redundancy and over-dispersion of information.

[0010] The technical solution adopted by the present invention is an accurate classification method for hyperspectral images based on improved multi-scale attention and Transformer network, which specifically comprises the following steps:

[0011] Step 1: Extract spatial information by the dynamic spatial attention unit.

[0012] Step 2: Extract multi-scale features by the multi-core fusion attention module.

[0013] Step 3: Improve information fusion by the cross-attention Swin Transformer module.

[0014] The characteristics of the present invention also lie in that

[0015] Step 1 specifically comprises the following steps:

[0016] Step 1.1: Perform normalization processing on the input.

[0017] Step 1.2: Extract features from the normalized data by convolution.

[0018] Step 1.3: Extract spatial features by using dynamic weights and different convolutions.

[0019] Step 1.3 specifically comprises the following steps:

[0020] First, given the input feature x ∈ R B×C×L, where B, C, and L represent the batch size, number of channels, and feature length of the input hyperspectral image respectively. First, perform normalization and channel expansion operations on the input features to obtain the expanded feature representation as shown in Equation (1):

[0021] z = Conv1(LayerNorm(x)) (1)

[0022] Among them, z represents the features extracted by convolution. Split z along the channel dimension into two sub-features a and b, and calculate the dynamic weights as shown in Equation (2) and Equation (3).

[0023] a, b = chunk(z, 2) (2)

[0024] w = Softmax(ReLU(Conv1(AdaptiveAvgPool(a)))) (3)

[0025] Among them, w represents the dynamic weight. Perform a weighted convolution operation on b, and combine with a to extract spatial features, as shown in Equation (4):

[0026] y = b ⊙ DWConv1(a ⊙ w) (4)

[0027] Among them, y represents the features extracted after the dynamic weight. ⊙ represents the element-wise dot product operation, and DWConv1 represents the depthwise separable convolution, which is used to capture spatial information;

[0028] Finally, output the finally extracted spatial features through 1D convolution and residual connection, as shown in Equation (5):

[0029] y' = Conv2(y) ⊙ scale + x (5)

[0030] Among them, y' represents the extracted spatial features.

[0031] Step 2 is specifically carried out according to the following steps:

[0032] Step 2.1: Perform normalization processing on the input;

[0033] Step 2.2: Extract features using convolution kernels of different scales on the normalized data;

[0034] Step 2.3: Fuse features of different scales.

[0035] Step 2.1 is specifically carried out according to the following steps:

[0036] First, perform layer normalization operation on the input x, and the calculation is as shown in Equation (6)

[0037] x' = LayerNorm(x) (6)

[0038] where x’ represents the result after normalization;

[0039] Step 2.2 is specifically carried out according to the following steps:

[0040] Parallel multi-scale feature extraction is performed through convolutions with different-scale convolutional kernels, and the specific calculation is shown in formula (7):

[0041] x k = Conv k (x′), k ∈ {3, 5, 7} (7)

[0042] where x k represents the features extracted by different convolutional kernels, and k represents the size of the convolutional kernel.

[0043] All the features extracted above are concatenated and then channel fusion is performed, and the specific calculation is shown in formula (8):

[0044] x fused = Fusion(concat(x 3 , x 5 , x 7 )) (8)

[0045] where x fused represents the fused features;

[0046] Finally, the spectral features are obtained by combining the scale factor and residual connection:

[0047] y = x fused ⊙ scale + x (9)

[0048] where y represents the extracted spectral features.

[0049] Step 2.3 is specifically carried out according to the following steps:

[0050] The hybrid scale attention module HSA extracts multi-scale spatial-spectral features through the dynamic spatial attention unit DASU and the multi-kernel fusion attention module MKFA respectively, and captures the information at different scales, as shown in formula (10):

[0051]

[0052] y spectral represents the spectral features extracted by the DSAU module, and y spatial represents the spatial features extracted by the MKFA module;

[0053] Then, the two types of features are weighted and fused to output the final features, as shown in formula (11):

[0054] y = y spectral + yspatial (11)

[0055] Among them, y represents the finally extracted feature.

[0056] Step 3 is specifically carried out according to the following steps:

[0057] Step 3.1: Input the features processed by multi-scale attention into LayerNorm;

[0058] Step 3.2: Introduce the multi-head attention mechanism to enable the network to process features of different scales;

[0059] Step 3.3: Introduce multi-head cross-attention to enable the network to better fuse features of different scales

[0060] Step 3.4: Add an MLP after the multi-head attention mechanism for integrating and extracting features;

[0061] Step 3.5: Introduce a residual connection after the MLP to reduce the computational complexity of the model.

[0062] Step 3.1 is specifically carried out according to the following steps:

[0063] First, standardize the input feature x that has undergone multi-scale attention so that the features are similar in distribution for each channel, facilitating subsequent feature extraction operations, as specifically shown in formula (12):

[0064]

[0065] Among them, is the feature after standardization, x ∈ R L×B×D , L represents the length of the feature sequence, B represents the batch size, D is the feature dimension, and the dimension of the feature obtained after standardization is the same as that of the input x;

[0066] Step 3.2 is specifically carried out according to the following steps:

[0067] In the multi-head self-attention module, calculate the long-range dependence relationship of the input features through the self-attention mechanism. The calculation process of self-attention is as shown in formula (13):

[0068]

[0069] Among them, respectively represent the query, key, and value vectors, and these vectors are respectively obtained by the linear projection matrices W Q , W K , W V obtained, d k is the feature dimension of each attention head, and the calculation is as shown in formula (14):

[0070] dk = D / num_heads (14)

[0071] After MHSA, the obtained output is calculated as shown in Equation (15):

[0072]

[0073] where W O ∈ R D×D is the output projection matrix. Finally, the feature representation ability is enhanced through residual connection, and the calculation is shown in Equation (16):

[0074] x 1 = x + MHSA(LayerNorm(x)) (16).

[0075] Step 3.3 is specifically carried out according to the following steps:

[0076] A cross-attention module is proposed to further integrate the information between different features. By applying the cross-attention mechanism to the input features after normalization again, the calculation is shown in Equation (17):

[0077] x 2 = x 1 + CrossAttention(LayerNorm(x 1 ), LayerNorm(x 1 ), LayerNorm(x 1 )) (17)

[0078] x 1 represents the features extracted by the multi-head attention mechanism, and x 2 represents the features extracted by the cross-attention mechanism;

[0079] where the calculation method of CrossAttention is the same as that of MHSA, but it is used to capture the interaction relationships between different-level features, thereby enhancing the spatial-spectral information fusion of features.

[0080] Step 3.4 is specifically carried out according to the following steps:

[0081] The MLP module is used to further model the non-linear relationship of features, and the calculation process is shown in Equation (18):

[0082] MLP(x 2 ) = W 2 · Dropout(GELU(W 1 · x 2 + b 1 )) + b 2 (18)

[0083] MLP(x 2 ) represents the features extracted by the MLP;

[0084] Among them, W 1 ∈R D×mlp_dim , W 2 ∈R mlp_dim×D is the weight matrix, b 1 ∈R mlp_dim , b 2 ∈R D is the bias, GELU is the activation function, Dropout is the random inactivation operation to prevent overfitting. Finally, the output of the MLP is shown in Equation (19):

[0085] y = x 2 + MLP(LayerNorm(x 2 )) (19).

[0086] y represents the spectral-spatial features of the finally extracted hyperspectral image, which are used for the accurate classification of hyperspectral images.

[0087] The beneficial effect of the present invention is that, based on the accurate hyperspectral image classification method of improved multi-scale attention and Transformer network, the MSA2T-Net framework is proposed. This is the first hyperspectral image classification model that combines multi-scale attention mechanism, dynamic spatial attention and cross-attention, and can simultaneously achieve multi-scale modeling, deep fusion of spatial-spectral features and feature redundancy optimization. A large number of experiments have been carried out on four representative hyperspectral datasets (Pavia, PaviaU, Houston2013 and Salinas). The proposed MSA2T-Net significantly outperforms the existing SOTA methods in terms of overall accuracy (OA), average accuracy (AA) and Kappa coefficient and other indicators. Description of the Drawings

[0088] Figure 1 is the network model MSA2T-Net that combines multi-scale attention and improved Transformer;

[0089] Figure 2 is the DSAU model structure;

[0090] Figure 3 is the MKFA model structure;

[0091] Figure 4 is the CASTA model structure;

[0092] Figure 5 is the Pavia dataset;

[0093] Figure 6 is the Houston2013 dataset;

[0094] Figure 7 is the PaviaU dataset;

[0095] Figure 8 is the Salinas dataset;

[0096] Figure 9 are the F1-Scores of different datasets;

[0097] Figure 10 are the results of different training-testing ratios. Detailed implementation manners

[0098] The present invention will be described in detail below with reference to the accompanying drawings and specific implementation manners.

[0099] The precise hyperspectral image classification method of the present invention based on improved multi-scale attention and Transformer network, combined with Figure 1 , specifically follows the following steps:

[0100] Step 1: Spatial information extraction by the Dynamic Spatial Attention Unit (DSAU);

[0101] Combined with Figure 2 , Step 1 specifically follows the following steps:

[0102] Step 1.1: Normalize the input;

[0103] Step 1.2: Extract features from the normalized data by convolution;

[0104] Step 1.3: Use dynamic weights and different convolutions to complete spatial feature extraction.

[0105] Step 1.3 specifically follows the following steps:

[0106] Dynamic Spatial Attention Unit (DSAU)

[0107] First, given the input feature x ∈ R B×C×L , where B, C, and L respectively represent the batch size, number of channels, and feature length of the input hyperspectral image. First, perform standardization and channel expansion operations on the input feature to obtain the expanded feature representation as shown in formula (1):

[0108] z = Conv1(LayerNorm(x)) (1)

[0109] where z represents the feature extracted by convolution, split z along the channel dimension into two sub-features a, b, and calculate the dynamic weights as shown in formula (2) and formula (3).

[0110] a, b = chunk(z, 2) (2)

[0111] w = Softmax(ReLU(Conv1(AdaptiveAvgPool(a)))) (3)

[0112] Wherein, w represents the dynamic weight, performs a weighted convolution operation on b, and combines with a to extract spatial features, as shown in formula (4):

[0113] y = b ⊙ DWConv1(a ⊙ w) (4)

[0114] Wherein, y represents the features extracted after the dynamic weight, ⊙ represents the element-wise dot product operation, and DWConv1 represents the depthwise separable convolution, which is used to capture spatial information;

[0115] Finally, the final extracted spatial features are output through 1D convolution and residual connection, as shown in formula (5):

[0116] y' = Conv2(y) ⊙ scale + x (5)

[0117] Where y' represents the extracted spatial features.

[0118] Step 2, Multi-Kernel Fusion Attention Module (MKFA) multi-scale feature extraction;

[0119] Combine Figure 3 , and the specific steps of step 2 are as follows:

[0120] Step 2.1, Normalize the input;

[0121] The specific steps of step 2.1 are as follows:

[0122] First, perform a layer normalization operation on the input x, and the calculation is as shown in formula (6)

[0123] x' = LayerNorm(x) (6)

[0124] Where x' represents the result after normalization;

[0125] Step 2.2, Extract features using convolution kernels of different scales for the normalized data;

[0126] The specific steps of step 2.2 are as follows:

[0127] Perform parallel multi-scale feature extraction through convolution with convolution kernels of different scales, and the specific calculation is as shown in formula (7):

[0128] x k = Conv k(x′), k ∈ {3, 5, 7} (7)

[0129] where x k represents the features extracted by different convolutional kernels, and k represents the convolutional kernel size.

[0130] Concatenate all the features extracted above, and then perform channel fusion. The specific calculation is shown in Equation (8):

[0131] x fused = Fusion(concat(x 3 , x 5 , x 7 )) (8)

[0132] where x fused represents the fused features;

[0133] Finally, combine the scale factor and residual connection to obtain the spectral features:

[0134] y = x fused ⊙ scale + x (9)

[0135] where y represents the extracted spectral features.

[0136] Step 2.3: Fuse features of different scales.

[0137] Step 2.3 is specifically carried out according to the following steps:

[0138] To effectively integrate spatial and spectral information, a new attention mechanism called HSA is proposed. By combining the advantages of MKFA and DSAU, it can achieve a balance in multi-scale feature extraction and dynamic spatial attention mechanism, thus effectively enhancing the spatial-spectral feature representation ability.

[0139] The hybrid scale attention module HSA extracts multi-scale spatial-spectral features through the dynamic spatial attention unit DASU and the multi-kernel fusion attention module MKFA respectively, and captures information at different scales, as shown in Equation (10):

[0140]

[0141] y spectral represents the spectral features extracted by the DSAU module, and y spatial represents the spatial features extracted by the MKFA module;

[0142] Then, weight and fuse the two types of features and output the final features, as shown in Equation (11):

[0143] y = y spectral + y spatial (11)

[0144] Among them, y represents the finally extracted feature.

[0145] Step 3: The Cross-Attention Swin Transformer module (CASTB) enhances information fusion;

[0146] Combined with Figure 4 , the specific steps of Step 3 are as follows:

[0147] Step 3.1: Input the features processed by multi-scale attention into LayerNorm;

[0148] The specific steps of Step 3.1 are as follows:

[0149] Normalization (LayerNormalization)

[0150] First, normalize the input feature x that has undergone multi-scale attention, so that the features are similar in the distribution of each channel, facilitating subsequent feature extraction operations, as shown in formula (12):

[0151]

[0152] Among them, is the feature after normalization, x ∈ R L×B×D , L represents the length of the feature sequence, B represents the batch size, D is the feature dimension, and the dimension of the feature obtained after normalization is the same as that of the input x;

[0153] Step 3.2: Introduce the multi-head attention mechanism to enable the network to process features of different scales;

[0154] The specific steps of Step 3.2 are as follows:

[0155] Multi-head self-attention

[0156] In the multi-head self-attention module, the long-range dependence relationship of the input features is calculated through the self-attention mechanism. The calculation process of self-attention is shown in formula (13):

[0157]

[0158] Among them respectively represent the query, key, and value vectors, and these vectors are respectively obtained by the linear projection matrices W Q , W K , W V obtained, d k is the feature dimension of each attention head, and the calculation is shown in formula (14):

[0159] d k = D / num_heads (14)

[0160] After MHSA, the obtained output is calculated as shown in formula (15):

[0161]

[0162] Among them, W O ∈R D×D is the output projection matrix. Finally, the feature expression ability is enhanced through residual connection, and the calculation is shown in formula (16):

[0163] x 1 = x + MHSA(LayerNorm(x)) (16).

[0164] Step 3.3: Introduce multi-head cross-attention to enable the network to better fuse features of different scales

[0165] Step 3.3 is specifically carried out according to the following steps:

[0166] Multi-head cross-attention

[0167] The cross-attention module is proposed to further integrate the information between different features. By applying the cross-attention mechanism to the input features after re-normalization, the calculation is shown in formula (17):

[0168] x 2 = x 1 + CrossAttention(LayerNorm(x 1 ), LayerNorm(x 1 ), LayerNorm(x 1 )) (17)

[0169] x 1 represents the features extracted by the multi-head attention mechanism, and x 2 represents the features extracted by the cross-attention mechanism;

[0170] Among them, the calculation method of CrossAttention is the same as that of MHSA, but it is used to capture the interaction relationship between features at different levels, so as to improve the spatial-spectral information fusion of features.

[0171] Step 3.4: Add an MLP after the multi-head attention mechanism to integrate the extracted features;

[0172] Step 3.4 is specifically carried out according to the following steps:

[0173] Multi-layer perceptron (MLP)

[0174] The MLP module is used to further model the non - linear relationship of features, and the calculation process is shown in formula (18):

[0175] MLP(x 2 )=W 2 ·Dropout(GELU(W 1 ·x 2 +b 1 ))+b 2 (18)

[0176] MLP(x 2 ) represents the features extracted by the MLP;

[0177] Among them, W 1 ∈R D×mlp_dim ,W 2 ∈R mlp_dim×D is the weight matrix, b 1 ∈R mlp_dim ,b 2 ∈R D is the bias, GELU is the activation function, Dropout is the random inactivation operation to prevent overfitting. Finally, the output through the MLP is shown in formula (19):

[0178] y=x 2 +MLP(LayerNorm(x 2 )) (19).

[0179] y represents the spectral - spatial features of the finally extracted hyperspectral image, which is used for the accurate classification of hyperspectral images.

[0180] Step 3.5: Introduce a residual connection after the MLP to reduce the computational complexity of the model.

[0181] Example 1

[0182] The accurate hyperspectral image classification method based on the improved multi - scale attention and Transformer network of the present invention, combined with Figure 1 , specifically follows the following steps:

[0183] Step 1: Extract spatial information by the dynamic spatial attention unit (DSAU);

[0184] Step 2: Extract multi - scale features by the multi - kernel fusion attention module (MKFA);

[0185] Step 3: Improve information fusion by the cross - attention Swin Transformer module (CASTB);

[0186] Example 2

[0187] The precise hyperspectral image classification method based on improved multi-scale attention and Transformer network, combined with Figure 1 , specifically follows the following steps:

[0188] Step 1: Spatial information extraction by the Dynamic Spatial Attention Unit (DSAU);

[0189] Combined with Figure 2 , the specific steps of Step 1 are as follows:

[0190] Step 1.1: Normalize the input;

[0191] Step 1.2: Extract features from the normalized data through convolution;

[0192] Step 1.3: Use dynamic weights and different convolutions to complete spatial feature extraction.

[0193] The specific steps of Step 1.3 are as follows:

[0194] Dynamic Spatial Attention Unit (DSAU)

[0195] First, given the input feature x ∈ R B×C×L , where B, C, and L represent the batch size, number of channels, and feature length of the input hyperspectral image respectively. First, perform standardization and channel expansion operations on the input feature to obtain the expanded feature representation as shown in formula (1):

[0196] z = Conv1(LayerNorm(x)) (1)

[0197] where z represents the feature extracted by convolution. Split z along the channel dimension into two sub-features a and b, and calculate the dynamic weights as shown in formula (2) and formula (3).

[0198] a, b = chunk(z, 2) (2)

[0199] w = Softmax(ReLU(Conv1(AdaptiveAvgPool(a)))) (3)

[0200] where w represents the dynamic weight. Perform a weighted convolution operation on b, and combine with a to extract spatial features, as shown in formula (4):

[0201] y = b ⊙ DWConv1(a ⊙ w) (4)

[0202] where y represents the feature extracted after the dynamic weight. ⊙ represents the element-wise dot product operation, and DWConv1 represents the depthwise separable convolution, which is used to capture spatial information;

[0203] Finally, the final extracted spatial features are output through 1D convolution and residual connection, as shown in Equation (5):

[0204] y′ = Conv2(y) ⊙ scale + x (5)

[0205] where y’ represents the extracted spatial features.

[0206] Step 2, Multi-Kernel Fusion Attention Module (MKFA) multi-scale feature extraction;

[0207] Step 3, Cross-Attention Swin Transformer Module (CASTB) to enhance information fusion.

[0208] Example 3

[0209] The precise hyperspectral image classification method based on improved multi-scale attention and Transformer network of the present invention combines Figure 1 , and specifically follows the following steps:

[0210] Step 1, Dynamic Spatial Attention Unit (DSAU) spatial information extraction;

[0211] Combined with Figure 2 , the specific steps of Step 1 are as follows:

[0212] Step 1.1, Normalize the input;

[0213] Step 1.2, Extract features from the normalized data through convolution;

[0214] Step 1.3, Use dynamic weights and different convolutions to complete spatial feature extraction.

[0215] The specific steps of Step 1.3 are as follows:

[0216] Dynamic Spatial Attention Unit (DSAU)

[0217] First, given the input feature x ∈ R B×C×L , where B, C, and L respectively represent the batch size, number of channels, and feature length of the input hyperspectral image. First, perform standardization and channel expansion operations on the input feature to obtain the expanded feature representation as shown in Equation (1):

[0218] z = Conv1(LayerNorm(x)) (1)

[0219] where z represents the features extracted by convolution. Split z along the channel dimension into two sub-features a and b, and calculate the dynamic weights as shown in Equation (2) and Equation (3).

[0220] a, b = chunk(z, 2) (2)

[0221] w = Softmax(ReLU(Conv1(AdaptiveAvgPool(a)))) (3)

[0222] Wherein, w represents the dynamic weight, performs a weighted convolution operation on b, and extracts spatial features in combination with a, as shown in formula (4):

[0223] y = b ⊙ DWConv1(a ⊙ w) (4)

[0224] Wherein, y represents the features extracted after the dynamic weight, ⊙ represents the element-wise dot product operation, and DWConv1 represents the depthwise separable convolution, which is used to capture spatial information;

[0225] Finally, the finally extracted spatial features are output through 1D convolution and residual connection, as shown in formula (5):

[0226] y′ = Conv2(y) ⊙ scale + x (5)

[0227] Where y’ represents the extracted spatial features.

[0228] Step 2, Multi-Kernel Fusion Attention Module (MKFA) multi-scale feature extraction;

[0229] Combined with Figure 3 , the specific steps of step 2 are as follows:

[0230] Step 2.1, Normalize the input;

[0231] The specific steps of step 2.1 are as follows:

[0232] First, perform a layer normalization operation on the input x, and the calculation is as shown in formula (6)

[0233] x′ = LayerNorm(x) (6)

[0234] Where x’ represents the result after normalization;

[0235] Step 2.2, Extract features using convolution kernels of different scales for the normalized data;

[0236] The specific steps of step 2.2 are as follows:

[0237] Perform parallel multi-scale feature extraction through convolution with convolution kernels of different scales, and the specific calculation is as shown in formula (7):

[0238] x k = Conv k (x′), k ∈ {3, 5, 7} (7)

[0239] where x k represents the features extracted by different convolutional kernels, and k represents the convolutional kernel size.

[0240] Concatenate all the features extracted above, and then perform channel fusion. The specific calculation is shown in Equation (8):

[0241] x fused = Fusion(concat(x 3 , x 5 , x 7 )) (8)

[0242] where x fused represents the fused features;

[0243] Finally, combine the scale factor and residual connection to obtain the spectral features:

[0244] y = x fused ⊙ scale + x (9)

[0245] where y represents the extracted spectral features.

[0246] Step 2.3: Fuse features of different scales.

[0247] Step 2.3 is specifically carried out according to the following steps:

[0248] To effectively integrate spatial and spectral information, a new attention mechanism called HSA is proposed. By combining the advantages of MKFA and DSAU, it can achieve a balance in multi-scale feature extraction and dynamic spatial attention mechanism, thus effectively enhancing the spatial-spectral feature representation ability.

[0249] The hybrid-scale attention module HSA extracts multi-scale spatial-spectral features through the dynamic spatial attention unit DASU and the multi-kernel fusion attention module MKFA respectively, and captures information at different scales, as shown in Equation (10):

[0250]

[0251] y spectral represents the spectral features extracted by the DSAU module, and y spatial represents the spatial features extracted by the MKFA module;

[0252] Then, the two types of features are weighted and fused to output the final features, as shown in Equation (11):

[0253] y = y spectral + y spatial (11)

[0254] where y represents the finally extracted features.

[0255] Step 3. Cross-attention Swin Transformer module (CASTB) improves information fusion;

[0256] Combined with Figure 4 , the specific steps of step 3 are as follows:

[0257] Step 3.1. Input the features processed by multi-scale attention into LayerNorm;

[0258] Step 3.2. Introduce the multi-head attention mechanism to enable the network to process features of different scales;

[0259] Step 3.3. Introduce multi-head cross-attention to enable the network to better fuse features of different scales

[0260] Step 3.4. Add MLP after the multi-head attention mechanism to integrate and extract features;

[0261] Step 3.5. Introduce a residual connection after MLP to reduce the computational complexity of the model.

[0262] Example 4

[0263] The high-spectral image precise classification method based on improved multi-scale attention and Transformer network of the present invention, combined with Figure 1 , specifically follows the following steps:

[0264] Step 1. Spatial information extraction by the dynamic spatial attention unit (DSAU);

[0265] Combined with Figure 2 , the specific steps of step 1 are as follows:

[0266] Step 1.1. Normalize the input;

[0267] Step 1.2. Extract features by convolution on the normalized data;

[0268] Step 1.3. Use dynamic weights and different convolutions to complete spatial feature extraction.

[0269] The specific steps of step 1.3 are as follows:

[0270] Dynamic spatial attention unit (DSAU)

[0271] First, given the input feature x∈R B×C×L , where B, C, and L represent the batch size, number of channels, and feature length of the input high-spectral image respectively. First, perform standardization and channel expansion operations on the input feature to obtain the expanded feature representation as shown in formula (1):

[0272] z = Conv1(LayerNorm(x)) (1)

[0273] Among them, z represents the features extracted by convolution. z is split into two sub-features a and b along the channel dimension, and the dynamic weights are calculated as shown in formulas (2) and (3).

[0274] a, b = chunk(z, 2) (2)

[0275] w = Softmax(ReLU(Conv1(AdaptiveAvgPool(a)))) (3)

[0276] Among them, w represents the dynamic weight. A weighted convolution operation is performed on b, and spatial features are extracted in combination with a, as shown in formula (4):

[0277] y = b ⊙ DWConv1(a ⊙ w) (4)

[0278] Among them, y represents the features extracted after the dynamic weight. ⊙ represents the element-wise dot product operation, and DWConv1 represents the depthwise separable convolution, which is used to capture spatial information;

[0279] Finally, the finally extracted spatial features are output through 1D convolution and residual connection, as shown in formula (5):

[0280] y′ = Conv2(y) ⊙ scale + x (5)

[0281] Among them, y’ represents the extracted spatial features.

[0282] Step 2: Multi-core fusion attention module (MKFA) multi-scale feature extraction;

[0283] Combine Figure 3 , and the specific steps of step 2 are as follows:

[0284] Step 2.1: Normalize the input;

[0285] The specific steps of step 2.1 are as follows:

[0286] First, perform a layer normalization operation on the input x, and the calculation is as shown in formula (6)

[0287] x′ = LayerNorm(x) (6)

[0288] Among them, x’ represents the result after normalization;

[0289] Step 2.2: Extract features using convolution kernels of different scales for the normalized data;

[0290] The specific steps of step 2.2 are as follows:

[0291] Parallel multi-scale feature extraction is performed through convolutions with convolution kernels of different scales. The specific calculation is shown in formula (7):

[0292] x k = Conv k (x′), k ∈ {3, 5, 7} (7)

[0293] where x k represents the features extracted by different convolution kernels, and k represents the convolution kernel size.

[0294] All the features extracted above are concatenated and then channel fusion is performed. The specific calculation is shown in formula (8):

[0295] x fused = Fusion(concat(x 3 , x 5 , x 7 )) (8)

[0296] where x fused represents the fused features;

[0297] Finally, the spectral features are obtained by combining the scale factor and residual connection:

[0298] y = x fused ⊙ scale + x (9)

[0299] where y represents the extracted spectral features.

[0300] Step 2.3: Fuse features of different scales.

[0301] Step 2.3 is specifically carried out according to the following steps:

[0302] To effectively integrate spatial and spectral information, a new attention mechanism called HSA is proposed. By combining the advantages of MKFA and DSAU, a balance can be achieved in multi-scale feature extraction and dynamic spatial attention mechanism, thus effectively enhancing the spatial-spectral feature representation ability.

[0303] The hybrid scale attention module HSA extracts multi-scale spatial-spectral features through the dynamic spatial attention unit DASU and the multi-kernel fusion attention module MKFA respectively, and captures information at different scales, as shown in formula (10):

[0304]

[0305] y spectral represents the spectral features extracted by the DSAU module, and y spatial represents the spatial features extracted by the MKFA module;

[0306] Then, the two features are weighted and fused to output the final feature, as shown in Equation (11):

[0307] y = y spectral + y spatial (11)

[0308] Where y represents the finally extracted feature.

[0309] Step 3. Cross-Attention Swin Transformer Module (CASTB) to enhance information fusion;

[0310] Combined with Figure 4 , Step 3 is specifically carried out according to the following steps:

[0311] Step 3.1. Input the feature processed by multi-scale attention into LayerNorm;

[0312] Step 3.1 is specifically carried out according to the following steps:

[0313] Normalization (LayerNormalization)

[0314] First, normalize the input feature x that has undergone multi-scale attention to make the distribution of features similar in each channel, facilitating subsequent feature extraction operations, as shown in Equation (12):

[0315]

[0316] Where is the feature after normalization, x ∈ R L×B×D , L represents the length of the feature sequence, B represents the batch size, D is the feature dimension, and the dimension of the feature obtained after normalization is the same as that of the input x;

[0317] Step 3.2. Introduce the multi-head attention mechanism to enable the network to process features of different scales;

[0318] Step 3.3. Introduce multi-head cross-attention to enable the network to better fuse features of different scales

[0319] Step 3.4. Add MLP after the multi-head attention mechanism to integrate and extract features;

[0320] Step 3.5. Introduce a residual connection after MLP to reduce the computational complexity of the model.

[0321] Embodiment 5

[0322] The high-spectral image precise classification method based on the improved multi-scale attention and Transformer network of the present invention, combined with Figure 1 , is specifically carried out according to the following steps:

[0323] Step 1: Spatial information extraction by the Dynamic Spatial Attention Unit (DSAU);

[0324] Combined with Figure 2 , Step 1 is specifically carried out according to the following steps:

[0325] Step 1.1: Normalize the input;

[0326] Step 1.2: Extract features by convolution on the normalized data;

[0327] Step 1.3: Extract spatial features using dynamic weights and different convolutions.

[0328] Step 1.3 is specifically carried out according to the following steps:

[0329] Dynamic Spatial Attention Unit (DSAU)

[0330] First, given the input feature x ∈ R B×C×L , where B, C, and L represent the batch size, number of channels, and feature length of the input hyperspectral image respectively. First, perform standardization and channel expansion operations on the input feature to obtain the expanded feature representation as shown in Equation (1):

[0331] z = Conv1(LayerNorm(x)) (1)

[0332] where z represents the feature extracted by convolution. Split z along the channel dimension into two sub-features a and b, and calculate the dynamic weights as shown in Equation (2) and Equation (3).

[0333] a, b = chunk(z, 2) (2)

[0334] w = Softmax(ReLU(Conv1(AdaptiveAvgPool(a)))) (3)

[0335] where w represents the dynamic weight. Perform a weighted convolution operation on b, and combine with a to extract spatial features, as shown in Equation (4):

[0336] y = b ⊙ DWConv1(a ⊙ w) (4)

[0337] where y represents the feature extracted after the dynamic weight. ⊙ represents the element-wise dot product operation, and DWConv1 represents the depthwise separable convolution, which is used to capture spatial information;

[0338] Finally, output the finally extracted spatial feature through 1D convolution and residual connection, as shown in Equation (5):

[0339] y' = Conv2(y) ⊙ scale + x (5)

[0340] Among them, y’ represents the extracted spatial features.

[0341] Step 2: The multi-core fusion attention module (MKFA) performs multi-scale feature extraction;

[0342] Combine Figure 3 , and the specific steps of Step 2 are as follows:

[0343] Step 2.1: Normalize the input;

[0344] The specific steps of Step 2.1 are as follows:

[0345] First, perform layer normalization on the input x, and the calculation is as shown in formula (6)

[0346] x′ = LayerNorm(x) (6)

[0347] Among them, x’ represents the result after normalization;

[0348] Step 2.2: Extract features from the normalized data using convolution kernels of different scales;

[0349] Step 2.3: Fuse features of different scales.

[0350] Step 3: The cross-attention Swin Transformer module (CASTB) enhances information fusion;

[0351] Combine Figure 4 , and the specific steps of Step 3 are as follows:

[0352] Step 3.1: Input the features processed by multi-scale attention into LayerNorm;

[0353] The specific steps of Step 3.1 are as follows:

[0354] Normalize (LayerNormalization)

[0355] First, normalize the input features x that have undergone multi-scale attention, so that the features are similar in distribution for each channel, facilitating subsequent feature extraction operations, specifically as shown in formula (12):

[0356]

[0357] Among them, is the feature after normalization, x ∈ R L×B×D , L represents the length of the feature sequence, B represents the batch size, D is the feature dimension, and the feature dimension obtained after normalization is the same as the input x;

[0358] Step 3.2: Introduce the multi-head attention mechanism to enable the network to process features of different scales;

[0359] The specific steps of Step 3.2 are as follows:

[0360] Multi-head self-attention

[0361] In the multi-head self-attention module, the long-range dependence relationship of the input features is calculated through the self-attention mechanism. The calculation process of self-attention is shown in Equation (13):

[0362]

[0363] where, respectively represent the query, key, and value vectors, and these vectors are obtained by the linear projection matrices W Q , W K , W V respectively. d k is the feature dimension of each attention head, and the calculation is shown in Equation (14):

[0364] d k = D / num_heads (14)

[0365] After passing through the MHSA, the obtained output is calculated as shown in Equation (15):

[0366]

[0367] where, W O ∈ R D×D is the output projection matrix. Finally, the feature expression ability is enhanced through the residual connection, and the calculation is shown in Equation (16):

[0368] x 1 = x + MHSA(LayerNorm(x)) (16).

[0369] Step 3.3: Introduce the multi-head cross-attention to enable the network to better fuse features of different scales

[0370] The specific steps of Step 3.3 are as follows:

[0371] Multi-head cross-attention

[0372] The proposed cross-attention module aims to further integrate the information between different features. By applying the cross-attention mechanism to the input features after re-normalization, the calculation is shown in Equation (17):

[0373] x 2 = x 1 + CrossAttention(LayerNorm(x1 ), LayerNorm(x 1 ), LayerNorm(x 1 )) (17)

[0374] x 1 represents the features extracted by the multi-head attention mechanism, and x 2 represents the features extracted by the cross-attention mechanism;

[0375] Among them, the calculation method of CrossAttention is the same as that of MHSA, but it is used to capture the interaction relationships between features at different levels, thereby enhancing the spatial-spectral information fusion of features.

[0376] Step 3.4. Add an MLP after the multi-head attention mechanism to integrate the extracted features;

[0377] Step 3.5. Introduce a residual connection after the MLP to reduce the computational complexity of the model.

[0378] Example 6

[0379] Figure 1 shows a schematic diagram of the proposed method. Now test this method.

[0380] The objects of the experimental test are four standard hyperspectral public datasets: Pavia, WHU-Hi-HanChuan (HanChuan), WHU-Hi-HongHu (HongHu), and XuZhou. Figure 5 shows the true color map (left) and the ground truth (right) of the Pavia dataset, Figure 6 shows the true color map (top) and the ground truth (bottom) of the Houston2013 dataset, Figure 7 shows the true color map (left) and the ground truth (right) of the PaviaU dataset, Figure 8 shows the true color map (left) and the ground truth (right) of the Salinas dataset.

[0381] The proposed algorithm is implemented by Python 12 and torch 2.4.1. The hardware used for training is an NVIDIA GeForce RTX 3060Ti GPU, x64, Win11.

[0382] To compare the performance of various classification algorithms, three evaluation criteria commonly used in HSI classification tasks were adopted: overall classification accuracy (OA), average accuracy (AA), and kappa coefficient. In addition to comparing with OA, AA, and kappa, we also used F1-Score to compare the effectiveness of the proposed method. We further compared the methods with different training samples to evaluate whether the proposed method is better for low samples or only accurate for higher training samples. The dataset categories were divided at a ratio of 10% and 90%.

[0383] In this experiment, four hyperspectral datasets (Pavia, Houston2013, PaviaU, Salinas) were used to comprehensively compare a variety of mainstream classification methods (such as SVM, KNN, RF, 2D-CNN, SACNet, SSFCN, ViT, SF, MF) with the proposed method. As shown in Tables 1-4, the results indicate that the proposed method has significant advantages in classification performance. In the Pavia dataset, the overall accuracy (OA) of the proposed method reached 0.988, significantly superior to traditional methods (such as 0.978 for SVM and 0.982 for RF) and some deep learning methods (such as 0.921 for SACNet), and the classification accuracies for class 7 and class 8 reached 0.9986 and 1.000 respectively, showing the best performance; on the Houston2013 dataset, the OA of the proposed method was 0.902, showing a significant improvement compared to RF (0.877) and 2D-CNN (0.864), especially outstanding in the classification of complex ground objects (such as class 3 and class 14), with accuracies of 0.996 and 0.988 respectively; in the PaviaU dataset, the OA of the proposed method reached 0.919, and the AA and Kappa values were 0.892 and 0.893 respectively, and the classification accuracies for class 1 and class 8 were 0.952 and 0.999 respectively, significantly superior to other methods; in the Salinas dataset, the OA of the proposed method reached 0.915, and the AA and Kappa values were 0.950 and 0.905*, and it was particularly outstanding in dealing with difficult-to-classify categories (such as class 4 and 9), with accuracies reaching 0.989 and 0.965, surpassing other comparison methods. Through comprehensive comparative analysis, the proposed HSIC method in this paper performs excellently in terms of OA, AA, and Kappa, especially having significant advantages in the classification of complex ground objects and the case of uneven class distribution, fully demonstrating the superiority and stability of this algorithm.

[0384] Table 1. Quantitative results of the Pavia dataset

[0385] ClassNo. SVM KNN RF 2D-CNN SACNet SSFCN ViT SF MF Proposed 0 1.000 1.000 1.000 0.995 0.957 0.983 0.875 0.828 0.854 1.000 1 0.967 0.962 0.966 0.885 0.852 0.924 0.935 0.832 0.975 0.956 2 0.818 0.825 0.852 0.937 0.834 0.923 0.731 0.474 0.393 0.916 3 0.808 0.782 0.806 0.982 0.954 0.961 0.950 0.854 0.867 0.873 4 0.923 0.950 0.951 0.973 0.660 0.795 0.981 0.957 0.940 0.978 5 0.923 0.928 0.950 0.931 0.843 0.965 0.893 0.774 0.615 0.959 6 0.925 0.933 0.955 0.967 0.904 0.900 0.846 0.576 0.563 0.973 7 0.996 0.997 0.995 0.967 0.948 0.979 0.820 0.691 0.673 0.998 8 1.000 1.000 0.999 0.896 0.794 0.907 0.936 0.990 0.957 1.000 OA(%) 0.978 0.979 0.982 0.972 0.921 0.962 0.902 0.812 0.837 0.988 AA(%) 0.927 0.940 0.948 0.948 0.861 0.926 0.885 0.775 0.760 0.964 Kappa(%) 0.969 0.970 0.975 0.951 0.863 0.910 0.869 0.743 0.780 0.983

[0386] Table 2. Quantitative results of the Houston2013 dataset

[0387]

[0388]

[0389] Table 3. Quantitative results of the PaviaU dataset

[0390]

[0391]

[0392] Table 4. Quantitative results of the Salinas dataset

[0393] ClassNo. SVM KNN RF 2D-CNN SACNet SSFCN ViT SF MF Proposed 0 1.000 1.000 0.998 0.988 0.989 0.985 0.855 0.932 0.759 0.993 1 0.991 0.989 0.996 0.993 0.997 0.941 0.926 0.874 1.000 0.993 2 0.907 0.931 0.967 1.000 0.966 0.958 0.882 0.812 0.923 0.924 3 0.975 0.982 0.985 0.993 0.964 0.957 0.888 0.985 0.986 0.982 4 0.981 0.984 0.989 0.944 0.946 0.822 0.880 0.720 0.901 0.989 5 1.000 1.000 0.999 0.991 0.979 0.993 0.937 1.000 0.983 0.998 6 0.984 0.994 1.000 0.996 0.979 0.985 0.896 0.999 0.995 1.000 7 0.697 0.727 0.793 0.765 0.639 0.605 0.910 0.856 0.789 0.794 8 0.988 0.988 0.986 0.994 0.985 0.989 0.895 1.000 1.000 0.992 9 0.899 0.911 0.926 0.960 0.928 0.893 0.925 0.941 0.951 0.965 10 0.885 0.923 0.943 0.997 1.000 0.999 0.957 0.980 0.998 0.942 11 0.957 0.954 0.972 0.958 0.951 0.472 0.949 1.000 1.000 0.971 12 0.911 0.935 0.959 0.971 0.957 0.737 0.949 0.997 0.987 0.954 13 0.979 0.961 0.964 0.971 0.973 0.974 0.908 0.973 0.959 0.956 14 0.838 0.647 0.791 0.678 0.795 0.493 0.751 0.305 0.635 0.771 15 0.989 0.991 0.983 0.983 0.965 0.969 0.948 0.976 0.955 0.987 OA(%) 0.886 0.881 0.917 0.918 0.878 0.799 0.902 0.836 0.883 0.915 AA(%) 0.928 0.935 0.951 0.961 0.938 0.861 0.922 0.897 0.9226 0.950 Kappa(%) 0.872 0.868 0.907 0.951 0.928 0.841 0.940 0.815 0.869 0.905

[0394] To comprehensively evaluate the reliability of the proposed model, the F1-Score of four types of datasets was analyzed, as Figure 9 shown. It can be concluded from the figure that the F1-Score of class 11 in the Salinas dataset is greater than 0.65, and the F1-Score of each other class is greater than 0.8. In the PaviaU dataset, the F1-Score of each class except class 2 is greater than 0.8. In the Pavia dataset, the F1-Score value of almost each class is greater than 0.9, fully demonstrating the effectiveness of the algorithm. For the Houston2013 dataset, the F1-Score of class 12 is greater than 0.6, and the rest of the classes are greater than 0.8. Through comprehensive analysis, the algorithm has good classification performance and strong generalization ability for HSIC.

[0395] To evaluate the performance of the proposed method under different training-test sample ratios, five different training-test ratios (0.1 - 0.9, 0.15 - 0.85, 0.2 - 0.8, 0.25 - 0.75, 0.3 - 0.7) were selected for experiments. The results are as Figure 9 shown. As the proportion of training samples increases, the overall accuracy (OA), average accuracy (AA), and Kappa coefficient of the proposed method all show a stable upward trend, fully indicating that more training data helps to improve the classification performance. At a low training sample ratio (0.1 - 0.9), the OA, AA, and Kappa coefficient of the proposed method are significantly better than other comparison methods, indicating its strong robustness in the case of insufficient samples; while at a high training ratio (0.3 - 0.7), the classification performance tends to saturate, further verifying the generalization ability and consistency of the method. Under different training-test ratios, the proposed method shows the characteristics of stable performance and smooth change, indicating its wide applicability and superiority in hyperspectral classification tasks.

Claims

1. An accurate hyperspectral image classification method based on improved multi-scale attention and Transformer network, characterized by: Follow these steps: Step 1: Extract spatial information of dynamic spatial attention unit; Step 2: Multi-scale feature extraction of multi-core fusion attention module; Step 3: Cross-attention Swin Transformer module improves information fusion.

2. The method for accurate classification of hyperspectral images based on improved multi-scale attention and Transformer network according to claim 1, characterized in that: The step 1 specifically follows the following steps: Step 1.1, normalize the input; Step 1.2, perform convolution on the normalized data to extract features; Step 1.3: Use dynamic weights and different convolutions to complete spatial feature extraction.

3. The hyperspectral image accurate classification method based on improved multi-scale attention and Transformer network according to claim 2 is characterized in that: The step 1.3 specifically follows the following steps: First, given the input feature x∈R B×C×L , where B, C and L represent the batch size, number of channels and feature length of the input hyperspectral image respectively. First, the input features are standardized and channel expanded to obtain the extended feature representation as shown in formula (1): z=Conv1(LayerNorm(x)) (1) Where z represents the feature extracted by convolution, z is divided into two sub-features a and b along the channel dimension, and the dynamic weights are calculated as shown in formula (2) and formula (3); a,b=chunk(z,2) (2) w=Softmax(ReLU(Conv1(AdaptiveAvgPool(a)))) (3) Where w represents the dynamic weight, and a weighted convolution operation is performed on b to extract spatial features in combination with a, as shown in formula (4): y=b⊙DWConv1(a⊙w) (4) Among them, y represents the features extracted after dynamic weighting, ⊙ represents the element-by-element dot product operation, and DWConv1 represents the depth-wise separable convolution, which is used to capture spatial information; Finally, the final extracted spatial features are output through 1D convolution and residual connection, as shown in formula (5): y′=Conv2(y)⊙scale+x (5) Where y' represents the extracted spatial features.

4. The method for accurate classification of hyperspectral images based on improved multi-scale attention and Transformer network according to claim 3, characterized in that: The step 2 specifically follows the following steps: Step 2.1, normalize the input; Step 2.2, perform feature extraction on the normalized data using convolution kernels of different scales; Step 2.3: Fusion of features at different scales.

5. The method for accurate classification of hyperspectral images based on improved multi-scale attention and Transformer network according to claim 4, characterized in that: The step 2.1 specifically follows the following steps: First, perform layer normalization on the input x, as shown in formula (6): x′=LayerNorm(x) (6) Where x' represents the normalized result; The step 2.2 specifically follows the following steps: Parallel multi-scale feature extraction is performed through convolution of convolution kernels of different scales. The specific calculation is shown in formula (7): x k =Conv k (x′),k∈{3,5,7} (7) where x k Represents the features extracted by different convolution kernels, and k represents the size of the convolution kernel; All the features extracted above are concatenated and then channel fusion is performed. The specific calculation is shown in formula (8): x fused =Fusion(concat(x3,x5,x7)) (8) where x fused Indicates fusion features; Finally, the scale factor and residual connection are combined to obtain the spectral characteristics: y=x fused ⊙scale+x (9) Where y represents the extracted spectral features.

6. The method for accurate classification of hyperspectral images based on improved multi-scale attention and Transformer network according to claim 5, characterized in that: The step 2.3 specifically follows the following steps: The hybrid scale attention module HSA extracts multi-scale spatial-spectral features through the dynamic spatial attention unit DASU and the multi-core fusion attention module MKFA to capture information at different scales, as shown in formula (10): y spectral Represents the spectral features extracted by the DSAU module, y spatial Represents the spatial features extracted by the MKFA module; Then the two features are weighted fused and the final feature is output, as shown in formula (11): y=y spectral +y spatial (11) Among them, y represents the final extracted feature.

7. The method for accurate classification of hyperspectral images based on improved multi-scale attention and Transformer network according to claim 6, characterized in that: The step 3 specifically follows the following steps: Step 3.1, input the features after multi-scale attention processing into LayerNorm; Step 3.2: Introduce a multi-head attention mechanism to enable the network to process features of different scales; Step 3.3: Introduce multi-head cross attention to allow the network to better integrate features of different scales Step 3.4: Add MLP after the multi-head attention mechanism to integrate the extracted features; Step 3.5: After MLP, residual connection is introduced to reduce the computational complexity of the model.

8. The method for accurate classification of hyperspectral images based on improved multi-scale attention and Transformer network according to claim 7, characterized in that: The step 3.1 specifically follows the following steps: First, the input feature x after multi-scale attention is standardized so that the distribution of the features in each channel is similar, which is convenient for subsequent feature extraction operations, as shown in formula (12): in, is the standardized feature, x∈R L×B×D , L represents the length of the feature sequence, B represents the batch size, and D represents the feature dimension. The feature dimension obtained after standardization is the same as the input x; The step 3.2 specifically follows the following steps: In the multi-head self-attention module, the long-range dependency of the input features is calculated through the self-attention mechanism. The calculation process of self-attention is shown in formula (13): in, denote the query, key, and value vectors, respectively, which are represented by the linear projection matrix W Q ,W K ,W V Get,d k For the feature dimension of each attention head, the calculation is shown in formula (14): d k =D / num_heads (14) After MHSA, the output calculation is shown in formula (15): Among them, W O ∈R D×D is the output projection matrix. Finally, the feature expression capability is enhanced through residual connection. The calculation is shown in formula (16): x1=x+MHSA(LayerNorm(x)) (16).

9. The method for accurate classification of hyperspectral images based on improved multi-scale attention and Transformer network according to claim 8, characterized in that: The step 3.3 specifically follows the following steps: The cross-attention module is proposed to further integrate the information between different features. By normalizing the input features again, the cross-attention mechanism is applied to the normalized input features. The calculation is shown in formula (17): x2=x1+CrossAttention(LayerNorm(x1),LayerNorm(x1),LayerNorm(x1)) (17) x1 represents the features extracted by the multi-head attention mechanism, and x2 represents the features extracted by the cross-attention mechanism; Among them, the calculation method of CrossAttention is the same as that of MHSA, but it is used to capture the interactive relationship between features at different levels, thereby improving the spatial-spectral information fusion of features.

10. The method for accurate classification of hyperspectral images based on improved multi-scale attention and Transformer network according to claim 9, characterized in that: The step 3.4 specifically follows the following steps: The MLP module is used to further model the nonlinear relationship of features. The calculation process is shown in formula (18): MLP(x2)=W2·Dropout(GELU(W1·x2+b1))+b2 (18) MLP(x2) represents the features extracted by MLP; Where W1∈R D×mlp_dim ,W2∈R mlp_dim×D is the weight matrix, b1∈R mlp_dim ,b2∈R D is the bias, GELU is the activation function, and Dropout is a random inactivation operation to prevent overfitting. Finally, the output of MLP is shown in formula (19): y=x2+MLP(LayerNorm(x2)) (19) y represents the spectral-spatial features of the final extracted hyperspectral image, which is used for accurate classification of the hyperspectral image.

Citation Information

Patent Citations

  • Hyperspectral remote sensing image classification method based on hybrid convolutional neural network

    CN115909052A

  • Ground feature classification method based on spectral space fusion Transform feature extraction

    CN116229153A

  • Hyperspectral image classification method combined with spatial pyramid attention mechanism

    CN116843975A

  • Hyperspectral image classification method based on high-order interactive convolutional network

    CN118506112A

  • Method for classifying hyperspectral images on basis of adaptive multi-scale feature extraction model

    US20230252761A1