Hyperspectral Image Classification Method Based on Transformer-Enhanced Dual-Stream Complementary Convolutional Neural Network

Through a dual-stream complementary convolutional neural network enhanced based on Transformer, combining spectral and spatial feature extraction and weight feature complementarity, the problem of insufficient local-global feature extraction in hyperspectral image classification is solved, and more efficient feature capture and classification accuracy is achieved.

CN117975268BActive Publication Date: 2025-07-11QIQIHAR UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410134030.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-31
Publication Date
2025-07-11
Estimated Expiration
2044-01-31

AI Technical Summary

Technical Problem

The existing hyperspectral image classification methods are difficult to fully utilize local and global feature dependencies, resulting in insufficient feature extraction, especially the methods based on Transformer and CNN architectures in local-global feature extraction.

Method used

A two-stream complementary convolutional neural network based on Transformer enhancement is adopted to extract spectral and spatial features through SpeFES and SpaFES respectively. Combined with the spectrum-spatial weight feature complementary module, the weight features between the spectrum and spatial features are supplemented, and the token with semantic features is generated, and the classification results are finally generated through the classification module.

Benefits of technology

Effectively capturing spectral and spatial characteristics improves the classification accuracy and stability of hyperspectral images, especially in finite labeled samples and complex scenarios, improving classification performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117975268B_ABST
    Figure CN117975268B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image classification methods, and the steps are as follows: First, a spectral and spatial feature extraction stream is proposed, both of which include a hybrid convolution block and an attention module composed of Transformers; the hybrid convolution block aims to mine high-resolution features based on local dependencies, and the attention module is used to extract key features in the spectral and spatial channels based on global dependencies; then spectral and spatial feature tokens are derived from SpeFES and SpaFES respectively; the spectral-spatial weight feature complementary module is used to make full use of the features in the tokens through the spectral weights of the spatial tokens and the spatial weights of the spectral tokens; finally, these feature tokens with different semantic features are fed into the classification module to generate classification results. The present invention solves the problem that most methods based on convolutional neural networks and Transformers cannot make full use of the spectral and spatial features in hyperspectral images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image classification methods, and specifically to a hyperspectral image classification method based on a dual-stream complementary convolutional neural network enhanced by Transformer. Background Art

[0002] The development of remote sensing technology has facilitated hyperspectral images with rich spectral and spatial features in hundreds to thousands of spectral bands. Hyperspectral images have been widely applied in various practical application fields including precision agriculture, military target detection, and mineral exploration. In these applications, the classification of hyperspectral images plays a crucial role by assigning each pixel in the hyperspectral image to a specific category based on its spectral and spatial features. Therefore, how to extract and utilize spectral and spatial features for classification tasks has attracted extensive attention.

[0003] In recent years, the Transformer architecture has performed excellently in the field of computer vision. The Transformer mainly relies on multiple self-attention mechanisms and effectively obtains global dependencies by studying the relationships between internal components of the sequence. Recent studies have shown that the Transformer has the ability to model content-related global-scale interactions and can adjust its perception field to highlight prominent components, thereby learning discriminative feature representations. Methods for hyperspectral image classification based on the Transformer can be mainly divided into two categories: pure Transformer architectures and hybrid CNN-Transformer architectures. Pure Transformer architecture methods use the Transformer to process HSI data. For example, a spectral-spatial Transformer network includes spectral and spatial sequence branches. The spatial sequence branch is used to learn fine-grained spatial features, while the spectral Transformer branch is designed to extract spectral features and model the dependencies within the spectral sequence. A spectral Swin Transformer network for HSI classification proposes a new method for dealing with HSI data scenarios. The network uses group attention to learn feature dependencies and considers the sliding window calculation of long-range information between different windows. Classification results show that the spectral Swin Transformer network exhibits better performance. Although pure Transformer methods can capture global-scale feature dependencies, they tend to ignore local-scale feature dependencies. Methods based on the Transformer and its variants can extract features with various dependencies. In particular, methods combining the Transformer and the CNN architecture can capture features with local-global scale dependencies. A new dual-branch Swin Transformer method combines a ViT-based branch and a CNN-based branch. The ViT-based branch aims to enhance the connection between adjacent windows, while the CNN-based branch aims to learn features from various feature maps. A new residual local-global spectral-spatial feature Transformer network, in which shallow and deep features are gradually derived through convolution and the Transformer. At the same time, multi-scale Transformer layers are used to further extract local-global features. A new method combining the Transformer and the CNN. A hierarchical 2D CNN dense network is used to learn spatial information, and an improved compressed axis Transformer is used to model global-scale feature dependencies. Although methods based on hybrid Transformer and CNN architectures can better extract local-global feature dependencies, they still have limited ability to fully utilize local-global spectral features and spatial information. Summary of the Invention

[0004] The purpose of the present invention is to provide a hyperspectral image classification method based on a Transformer-enhanced dual-stream complementary convolutional neural network to solve the problems raised in the above background art.

[0005] To achieve the above purpose, a hyperspectral image classification method based on a Transformer-enhanced dual-stream complementary convolutional neural network (i.e., the TECCNet method) includes the following steps:

[0006] Step S1, extracting key features in the high resolution and spectral-spatial channels of the hyperspectral image, and using SpeFES and SpaFES to extract spectral and spatial features respectively;

[0007] The SpeFES includes a spectral mixed convolution module and a spectral attention mechanism composed of a Transformer encoder, and the high-resolution feature map extracted by the spectral mixed convolution module is fed into the spectral attention mechanism; the SpaFES includes a spatial attention mechanism composed of a Transformer encoder and a spatial mixed convolution module, and the mixed convolution module uses a simplified 3D convolution kernel;

[0008] The mixed convolution module consists of five convolutional layers, the five convolutional layers include four 3D convolutional layers and one 3D transposed convolutional layer, and the 3D transposed convolutional layer is located in the middle of the other four 3D convolutional layers, and the feature extraction process of the mixed convolution module is expressed as follows:

[0009]

[0010] f(σ(X m+4 )); θ) = ω(σ(X m+4 )) * k m+1 + S m+1

[0011] σ(X m+4 ) = X m+1 + f(X m+1 ; θ) + ε(X m+2 ; θ) + f(X m+3 ; θ);

[0012] Where X m+1 represents the 3D cube input to the (m + 1)-th layer, k m+1 , S m+1 respectively represent the convolution kernel and stride of the (m + 1)-th layer, f(X; θ) represents the 3D convolution operation on X, and ε(X; θ) represents the 3D transposed convolution operation on X;

[0013] The feature maps A after processing by the spectral mixed convolution module and the spatial mixed convolution module spe∈R c×h×w and B spa ∈R c×h×w Preprocessing of Tokenization needs to be performed separately to obtain spectral Tokens and spatial Tokens with 64 output channels, and then the spectral Tokens (spectral markers) and spatial Tokens (spatial markers) are fed into the spectral and spatial attention mechanisms respectively to obtain spectral Tokens and spatial Tokens with key features based on global dependencies;

[0014] Step S2, a spectral-spatial weight feature complementary module is used to supplement the spectral and spatial weight features between the spectral feature Tokens and spatial feature Tokens, so as to make full use of the spectral and spatial features; and at the same time, spectral and spatial semantic features are extracted;

[0015] The specific supplementation process of the spectral-spatial weight feature complementary module is as follows:

[0016]

[0017]

[0018]

[0019]

[0020] ; where represents the weight parameter matrix, and C() represents the complementary function;

[0021] The spectral-spatial weight feature complementary module calculates the expectation E(x cls ) of each element in each t i Token and uses it as the semantic feature of the corresponding Token. The specific process is as follows:

[0022]

[0023] where X class belongs to the set of elements of the c-th class, and n class represents the number of elements in the c-th class;

[0024] The output of the spectral-spatial weight feature complementary module is shown as follows:

[0025]

[0026]

[0027] where represents the spectral Token with semantic features, Represents a spatial Token with semantic features;

[0028] Step S3, using a classification module to generate a classification result.

[0029] Preferably: The 3D convolutional layer includes 3D convolution operations, batch normalization operations, and the Mish activation function. In the 3D transposed convolutional layer, 3D transposed convolution, batch normalization operations, and the Mish activation function are used. The mathematical formula of the Mish activation function is as follows:

[0030] f(x) = x tan h(softplus(x)) = x tan h(ln(1 + e x ))

[0031] where x represents the input of the activation function, and Mish is in the range of [≈ -0.31, ∞].

[0032] Preferably: Before obtaining Tokens containing different types of features with a channel dimension of 64 from the feature maps processed by the spectral mixture convolution module and the spatial mixture convolution module, these 3D patches need to be tokenized. The specific process is to add clsToken (represented as spe in the spectral feature map i and spa in the spatial feature map i) and posToken (represented as spe in the spectral feature map i and spa in the spatial feature map i) to A spe ∈R hw×64 and B spa ∈R hw×64 respectively, to obtain spectral feature Tokens and spatial feature Tokens, thereby effectively embedding the spectral feature Tokens and spatial feature Tokens into the spectral and spatial Transformer encoding mechanisms;

[0033] The spectral and spatial Transformer encoding mechanisms are as follows:

[0034]

[0035]

[0036] Preferably: The supplement process of the spectral-spatial weight feature complementary module is that the spectral-spatial weight feature complementary module copies the processed and Tokens from the spectral feature Tokens and spatial feature Tokens respectively, merges them with each other, and then constructs two new and Tokens use the self-attention mechanism to complement each other and the weight features in, which are used as the classification information of spectral tokens and spatial tokens respectively.

[0037] Preferably, the semantic feature extraction process of the spectral-spatial weight feature complementary module is as follows: for each element x i ∈t cls in, perform softmax processing on each element x i , and then assign a probability vector p i ∈R i to each x class , where class represents the number of categories contained in the ground to be classified. According to the probability vector p i , calculate the expected value E(x i ) of each x i .

[0038] S(x i ) = p i

[0039] E(x i ) = x i ·p i

[0040] where S(x i ) represents the softmax function, and E(x i ) represents the expectation function. For all elements x i ∈t cls , its corresponding expected value matrix is expressed as U∈R C×class , where C is the total number of elements. Calculate E(x cls ) for each t i token based on the expected value matrix U and use it as the semantic feature of the corresponding token.

[0041] Preferably, the classification process of step S3 is as follows: the spectral and spatial feature tokens obtained based on step S2 are respectively processed through the Reshape operation, and the spectral and spatial feature tokens are input into the average pooling layer with batch normalization and Mish activation. Use the concatenation operation to merge the spectral and spatial feature tokens along the channel dimension, fuse them into a feature token with spectral-spatial semantic features, and obtain the final classification result through the fully connected layer.

[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0043] To fully capture spectral and spatial features, the present invention proposes a dual pipeline for spectral and spatial extraction, namely SpeFES and SpaFES. SpeFES is used to extract spectral features, which includes a hybrid convolution module and a spectral attention mechanism composed of Transformer encoders. The spectral hybrid convolution module aims to extract high-resolution spectral features while avoiding information loss. In addition, the spectral hybrid convolution module uses a simplified 3D convolution kernel to operate within the spectral dimension and reduce the overall parameters of the network. Subsequently, the high-resolution feature map extracted by the spectral hybrid convolution module is fed into the spectral attention mechanism, enabling the network to capture key features based on global dependencies. Similar to SpeFES, SpaFES is used to extract spatial features, including a spatial attention mechanism and a spatial hybrid convolution module composed of Transformer encoders. Then, spectral and spatial feature tokens are respectively extracted from the spectral and spatial feature extraction streams. In addition, a spectral-spatial weight feature complementary module is used after the dual pipeline. This module first extracts the spectral weight features contained in the spatial classification Tokens and the spatial weight features contained in the spectral classification Tokens, and then supplements the extracted weight features with the original spectral and spatial semantic features respectively. Subsequently, this module generates spectral semantic features for the spectral feature classification Tokens and spatial semantic features for the spatial feature classification Tokens. Finally, these Tokens with different semantic features are fed into a classification module including an average pooling layer, a Mish activation function, a batch normalization layer, and a linear layer to generate classification results. The cross-entropy loss function is used to optimize the learnable parameters in the network. Description of the Drawings

[0044] Figure 1 Schematic diagram of the detailed structure of the present invention;

[0045] Figure 2 Schematic diagram of the structure of the hybrid convolution module of the present invention;

[0046] Figure 3 Schematic diagram of the structure of the spectral-spatial weight feature complementary module of the present invention;

[0047] Figure 4 Graph of the change in overall accuracy based on the sample ratio of the present invention and various comparison methods on four hyperspectral image datasets;

[0048] Figure 5 Classification results of the present invention on input patches of different sizes;

[0049] Figure 6 Graph of the change in OA value of the network of the present invention for different patches;

[0050] Figures 7 - 10Full-pixel classification map for the present invention and various comparison methods;

[0051] Figure 11 Table 4 of the present invention;

[0052] Figure 12 Table 5 of the present invention;

[0053] Figure 13 Table 6 of the present invention;

[0054] Figure 14 Table 7 of the present invention. Detailed implementation manners

[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0056] Embodiment

[0057] Please refer to Figure 1 , the figure is a preferred implementation manner in the present invention, a hyperspectral image classification method based on a dual-stream complementary convolutional neural network enhanced by Transformer, including the following steps:

[0058] Step S1, extraction of key features in the high-resolution and spectral-spatial channels of the hyperspectral image, using SpeFES and SpaFES to extract spectral and spatial features respectively;

[0059] The SpeFES includes a spectral mixing convolution module and a spectral attention mechanism composed of a Transformer encoder, and the high-resolution feature map extracted by the spectral mixing convolution module is fed into the spectral attention mechanism; the SpaFES includes a spatial attention mechanism composed of a Transformer encoder and a spatial mixing convolution module, and the mixing convolution module uses a simplified 3D convolution kernel to reduce the overall parameters of the network;

[0060] The mixing convolution module is composed of five convolutional layers, the five convolutional layers include four 3D convolutional layers and one 3D transposed convolutional layer, and the 3D transposed convolutional layer is located in the middle of the other four 3D convolutional layers, and the feature extraction process of the mixing convolution module is expressed as follows:

[0061]

[0062] f(σ(X m+4 )); θ) = ω(v(X m+4 )) * km+1 +S m+1

[0063] σ(X m+4 ) = X m+1 +f(X m+1 ; θ)+ε(X m+2 ; θ)+f(X m+3 ; θ);

[0064] where X m+1 represents the 3D cube input to the (m + 1)-th layer, k m+1 , S m+1 respectively represent the convolution kernel and stride of the (m + 1)-th layer, f(X; θ) represents the 3D convolution operation on X, and ε(X; θ) represents the 3D transposed convolution operation on X;

[0065] The feature maps A spe ∈R c×h×w and B spa ∈R c×h×w processed by the spectral mixing convolution module and the spatial mixing convolution module also need to be preprocessed by Token (marking) respectively to obtain the spectral Token and spatial Token with an output channel of 64, and then the spectral Token (spectral mark) and spatial Token (spatial mark) are fed into the spectral and spatial attention mechanisms respectively to obtain the spectral Token and spatial Token with key features based on global dependence; and the spectral and spatial feature marks are extracted from the spectral and spatial feature extraction streams respectively;

[0066] Step S2, use the spectral-spatial weight feature complementary module to supplement the spectral and spatial weight features between the spectral feature Token and the spatial feature Token, and simultaneously extract the spectral and spatial semantic features;

[0067] The specific supplementary process of the spectral-spatial weight feature complementary module is as follows:

[0068]

[0069]

[0070]

[0071]

[0072] where, represents the weight parameter matrix, and C() represents the complementary function;

[0073] The spectral-spatial weight feature complementary module calculates each t based on the matrix U clsThe expected E(x i ) for each element in the Token and the specific process of using it as the semantic feature of the corresponding Token are as follows:

[0074]

[0075] where X class belongs to the set of elements of the c-th class, and n class represents the number of elements in the c-th class;

[0076] The output of the spectral-spatial weight feature complementary module is shown as follows:

[0077]

[0078]

[0079] where, represents the spectral Token with semantic features, represents the spatial Token with semantic features;

[0080] Step S3, use the classification module to generate classification results.

[0081] In this embodiment, in order to fully capture spectral and spatial features, the present invention proposes a dual pipeline for spectral and spatial extraction, namely spectral feature excavation stream (SpeFES) and spatial feature excavation stream (SpaFES). SpeFES is used to extract spectral features, including a spectral mixing convolution module and a spectral attention mechanism composed of Transformer encoders. The spectral mixing convolution module is designed to extract high-resolution spectral features while avoiding information loss. In addition, the spectral mixing convolution module uses a simplified 3D convolution kernel for operations within the spectral channels and to reduce the overall parameters of the network. Subsequently, the high-resolution feature map extracted by the spectral mixing convolution module is fed into the spectral attention mechanism, enabling the network to capture the global dependencies of spectral features. Similar to SpeFES, SpaFES is used to extract spatial features, including a spatial attention mechanism and a spatial mixing convolution module composed of Transformer encoders. Then, spectral and spatial feature tokens are respectively extracted from the spectral and spatial feature extraction streams. In addition, a spectral-spatial weight feature complementary module is used after the dual pipeline. This module first extracts the spectral weight features contained in the spatial classification Token and the spatial weight features contained in the spectral classification Token, and then supplements the extracted weight features with the original spectral and spatial semantic features respectively. Subsequently, this module generates spectral semantic features for the spectral feature classification Token and spatial semantic features for the spatial feature classification Token. Finally, these Tokens with different semantic features are fed into a classification module including an average pooling layer, a Mish activation function, a batch normalization layer, and a linear layer to generate classification results. The cross-entropy loss function is used to optimize the learnable parameters in the network. The whole process is as follows:

[0082]

[0083] where Z xi denotes the set of labeled pixels, and Z T denotes the set of labeled samples, denotes the probability output by the network. The present invention applies the Adam optimizer algorithm to update the parameters of the network. In addition, a dynamic learning rate and an early stopping mechanism are respectively introduced into the framework of the network to promote network convergence and reduce the training time cost; cls Token: class token or classification token.

[0084] Furthermore, the proposed hybrid convolution module of the present invention consists of five convolutional layers, including four 3D convolutional layers and one 3D transposed convolutional layer. The 3D transposed convolutional layer is located in the middle of these four 3D convolutional layers, which is used to extract the local dependence of features. At the same time, the addition of the 3D transposed convolutional layer helps the hybrid convolution module to extract excellent high-resolution features. Adding residual connections in the hybrid convolution block helps to transfer gradient information from the first layer to the last convolutional layer, thus effectively preventing the occurrence of gradient vanishing in the stacking of extremely deep convolutional layers.

[0085] Embodiment 1

[0086] Refer to Figure 2 , in the hybrid convolution module, a 3D convolutional layer with a convolution kernel of (a×a×d) is applied in this embodiment, and a 3D transposed convolutional layer with a convolution kernel size of (1×1×1) is applied, which is beneficial to the extraction of high-resolution features. In addition, when extracting spectral features and spatial features respectively, the parameters in the 3D convolutional layer are different. In the hybrid convolution module, the residual connection F(X v ; θ) and a connection σ(X v ) are introduced. The feature extraction process of the hybrid convolution module can be expressed as follows:

[0087]

[0088] f(σ(X m+4 ); θ) = ω(σ(X m+4 )) * k m+1 + S m+1

[0089] σ(X m+4 ) = X m+1 + f(X m+1 ; θ) + ε(X m+2 ; θ) + f(X m+3 ; θ);

[0090] where X m+1 represents the 3D cube input to the (m + 1)-th layer, k m+1 , S m+1 respectively represent the convolution kernel and stride of the (m + 1)-th layer, f(X; θ) represents the 3D convolution operation on X, and ε(X; θ) represents the 3D transposed convolution operation on X;

[0091] At the same time, the 3D convolutional layer includes 3D convolution operation, batch normalization operation, and Mish activation function. In the 3D transposed convolutional layer, transposed convolution, batch normalization operation, and Mish activation function are used. A non-monotonic Mish activation function is introduced to accelerate the training process of the network. The mathematical formula of the Mish activation function is as follows:

[0092] f(x) = x tanh(softplus(x)) = x tanh(ln(1 + e x ))

[0093] where x represents the input of the activation function, and Mish is in the range of [≈ -0.31, ∞).

[0094] In this embodiment, a 3D patch with a size of (h×w×d) is randomly selected from the Pavia University scene data and then fed into the spectral mixture convolution module to mine spectral features, where d represents the depth of the 3D patch, and h and w represent the height and width of the 3D patch respectively. In the spectral mixture convolution module, a convolution kernel of (1×1×7) is used in the first layer, the stride is set to (1×1×2), and the number of filters is set to 48. In the fourth layer of the spectral mixture convolution module, a convolution kernel of (1×1×7) is used, the stride is set to (1×1×1), the padding is set to (0×0×3), and the number of filters is 12. The middle 3D transposed convolution layer configures both the convolution kernel size and the stride to (1×1×1). In the last layer, a convolution kernel with a size of (1×1×band) is used, the stride is set to (1×1×1), and the number of filters is 48. Subsequently, one-shot connections are used to connect the feature maps generated by the first four layers through the channel dimension, thereby obtaining a feature map with a size of (11×11×49, 96). Through the last layer of the spectral mixture convolution block, a feature map with a size of (11×11×49, 48) is obtained. Finally, residual connections are used to combine the output feature map of the first layer with the output feature map of the fifth layer to generate the output of the spectral mixture convolution module. In addition, an additional convolution layer is used after the spectral mixture convolution module to further mine spectral information, and the convolution kernel size and stride are (1×1×band) and (1×1×1) respectively. Finally, a spectral feature map with a size of (11×11×49, 48) is obtained.

[0095] Similarly, a 3D patch with a size of (11×11×band) is fed into the spatial mixture convolution block for extracting spatial features. During the spatial feature extraction process, the spatial convolution kernel size in the first layer is set to (1×1×band), the convolution kernel sizes in the second and fourth layers are set to (3×3×1), and the spatial convolution kernel size in the fifth layer is set to (1×1×1). Except for the different convolution kernel sizes mentioned above, the other parameter settings in the spatial CTC mixture convolution module are basically the same as those in the spectral mixture convolution module. Tables 1 and 2 show the entire process of mining spectral and spatial features.

[0096] Table 1

[0097] The detailed information of the spectral hybrid convolution (hc) block.

[0098]

[0099] Table 2

[0100] The detailed information of the spatial hybrid convolution (hc) block

[0101]

[0102]

[0103] To better embed the patches obtained after processing by the hybrid convolution module into the Transformer encoder, these patches need to be tokenized.

[0104] First, two feature maps A spe ∈R c×h×w and B spa ∈R c×h×w from the spectral hybrid convolution module and the spatial hybrid convolution module respectively need to be pre - processed, that is, use 2D transposed convolution operation to convert the channel sizes of these two feature maps to appropriate dimension sizes, which is a hyper - parameter. Here, the output dimension is set to 64. After processing A spe ∈R c×h×w and B spa ∈R c×h×w respectively, tokens with a channel dimension of 64 will be obtained, which contain different types of features. During the tokenization process of these 3D patches, it includes adding the clsToken (denoted as spe in the spectral feature map i and spa in the spatial feature map) and posToken (denoted as spe in the spectral feature map i and spa in the spatial feature map) to A spe ∈R hw×64 and B spa ∈R hw×64 respectively, and are used for classification tasks and marking position information respectively. Then, spectral feature tokens and spatial feature tokens are obtained respectively, so as to effectively embed these two feature tokens into the spectral and spatial Transformer encoding mechanisms:

[0105]

[0106]

[0107] Example 2

[0108] In this example, considering that the semantic information of each category in the hyperspectral image is different and the probability of each category is also different, the semantic features of the hyperspectral image (HSI) are used to improve the classification performance. In the present invention, a spectral-spatial weight feature complementary module is adopted to complement the spectral semantic features and the spatial features with each other, and at the same time, labels with spectral semantic features and spatial semantic features are obtained.

[0109] Furthermore, after being processed by the spectral and spatial attention mechanisms, the and Tokens with classification features can be used for classification, and they complement the extracted features with the original spectral and spatial semantic features respectively. As Figure 3 shown, this module copies and Tokens from the spectral Token and the spatial Token respectively, merges them with each other and then constructs two new and Tokens respectively. Then, the self-attention mechanism is used to complement the and weight features in them, which are respectively used as the classification features of the spectral and spatial Tokens. The specific complement process is as follows:

[0110]

[0111]

[0112]

[0113]

[0114] where represents the weight parameter matrix, and C() represents the complementary function.

[0115] To extract the semantic features of the feature Token, first, each element x i ∈t cls in x i is processed by softmax, and then a probability vector p i is assigned to each x i ∈R class , where class represents the number of categories contained in the ground to be classified. According to the probability vector p i , each xi The expected value E(x i ).

[0116] S(x i ) = p i

[0117] E(x i ) = x i ·p i

[0118] where S(x i ) represents the softmax function, and E(x i ) represents the expectation function. For all elements x i ∈t cls , its corresponding expected matrix expression is U ∈ R C×class , where C is the total number of elements. In this embodiment, the E(x cls ) of each t i Token will be calculated based on the expected matrix U and used as the semantic feature of the corresponding Token. The process is as follows:

[0119]

[0120] where X class is the set of elements belonging to the c-th class, and n class represents the number of elements in the c-th class.

[0121] Furthermore, the output of the spectral-spatial weight feature complementary module is as follows:

[0122]

[0123]

[0124] where, represents the spectral Token with semantic features, represents the spatial Token with semantic features. The spectral-spatial weight feature complementary module complements the extracted spectral and spatial features with the original spectral and spatial features respectively. At the same time, it also generates spectral and spatial feature Tokens with semantic features, thus enhancing the classification performance.

[0125] In this embodiment, after being processed by the above modules, the two semantic feature Tokens obtained are respectively processed by the Reshape operation to facilitate inputting the two feature Tokens into the average pooling layer with batch normalization and the Mish activation. Then, the concatenation operation is used to merge the two feature Tokens along the channel dimension, effectively fusing them into a feature Token with spectral-spatial semantic features. Finally, the final classification result is obtained through the fully connected layer. The specific process is shown in Table 3 below:

[0126] Table 3The detailed information of classification block

[0127]

[0128] Example 3

[0129] Refer to Figure 4 , 1%, 2%, 3%, 4%, and 5% of the labeled samples on the Pavia University, WHU-Hi-Honghu, University of Houston, and Kennedy Space Center datasets were randomly selected in the present invention (these ratios are shown on the Sample axis of the 3D color mapping surface graph in Figure 4 ). Additionally, due to hardware limitations, the spectral dimension of the Houston scene data was reduced to 45. In Figure 4 , the variation of the overall accuracy (OA) of various methods with different training sample percentages on four hyperspectral image datasets is described in this paper.

[0130] Generally, increasing the proportion of the training sample size can provide more discriminative features for the data-driven network, thus helping to improve the quality of model classification. At the same time, the training time and cost of the network are also considered. Specifically, as can be seen from Figure 4 , the red contour represents the maximum region of the overall accuracy (OA), and the purple contour represents the minimum region of the overall accuracy (OA). It can be clearly seen from the four changing graphs in Figure 4 that as the training sample ratio increases, the overall accuracy (OA) of these algorithms also increases, which is consistent with the previous research results. In Figure 4The phenomena observed in (a), (b), (c), and (d) are that when the sample percentage is 5%, the colors corresponding to most networks are close to the dark red area of the 3D color mapping surface map. When the sample ratio is 1%, by comparing various classification methods, it can be observed that the overall accuracy (OA) of FDSSC, OSDN, MVAHN, and TECCNet is higher than other methods. This indicates that these methods have greater feature extraction capabilities under limited labeled samples. At the same time, it can be seen that the overall accuracy (OA) of TECCNet increases with the growth of samples and maintains stable classification results, while being superior to other methods. This result once again verifies the effectiveness and practicality of the TECCNet network feature extraction under limited labeled samples.

[0131] Generally speaking, the patch size of hyperspectral images will affect the information carried by the patch itself. A certain amount of information is helpful for the network to achieve satisfactory classification results. Therefore, in this embodiment, the OA values of the network when using different patches are compared on four datasets, that is, the network classification accuracies when the patch sizes are 5×5, 7×7, 9×9, 11×11, and 13×13 as inputs, and the comparison is made taking OA as an example, as Figure 6 shown. Obviously, with the increase of the patch size, the OA values show an upward trend on the four hyperspectral datasets. When the patch size is 11×11, OA reaches the maximum value. Then when the patch size is 13×13, OA begins to decline. Especially in the Kennedy Space Center dataset, compared with the patch of size 11×11, when the patch size is 13×13, OA is significantly lower. At the same time, with the increase of the patch size, the number of network parameters also increases. In Figure 5 , taking the datasets of the University of Pavia, WHU-Hi-Honghu, the University of Houston, and the Kennedy Space Center as examples, the classification results on input patches of different sizes are shown. In the experiments of this article, for the sake of consistency, the patch sizes of these four datasets are uniformly configured as 11×11.

[0132] In this embodiment, four commonly used hyperspectral datasets are used to verify the practicality and robustness of the network of the present invention, namely the University of Pavia dataset, the University of Houston dataset, the WHU-Hi-Honghu dataset, and the Kennedy Space Center dataset:

[0133] (1) University of Pavia data: The spatial size of the University of Pavia (PU) image is 610×340 pixels, and the spatial resolution of each pixel is about 1.3 meters. The University of Pavia image contains 103 bands and includes nine categories. In Table 4, the sample numbers of each category used in the experiment in the training, validation, and test sets are given;

[0134] (2) WHU-Hi-Honghu: The WHU-Hi-Honghu (Honghu) image has a spatial size of 940×475 pixels, including 270 bands and 22 classes. However, in the experiment, due to equipment limitations, only 16 classes were selected. The selected classes and the corresponding number of samples are shown in Table 5;

[0135] (3) University of Houston data: The University of Houston (Houston) image has a spatial size of 349×1905 pixels, including 144 bands and 15 classes. The corresponding number of training, validation, and test labeled samples are listed in Table 6;

[0136] (4) Kennedy Space Center data: The Kennedy Space Center (KSC) image has a spatial size of 512×614 pixels. After removing water absorption and low signal-to-noise ratio bands, only 176 spectral bands in the range of 400 - 2500 nm were used for research. The Kennedy Space Center includes 13 classes, and the corresponding number of training, validation, and test samples are shown in Table 7.

[0137] To evaluate the effectiveness of the network, the present invention compares TECCNet with ten different hyperspectral image classification methods, which are TBPFA, SSGCA, SSFTT, OSDN, HDDA, SSRN, MVAHN, MDFFA, FDSSC, and HResNet. To more clearly understand the classification results of each method, the present invention will use OA, AA, and Kappa, as well as the classification results of each class to evaluate their performance. All experiments will be performed on a small-scale deep learning base station equipped with 128GB of DDR4 RAM and 8 NVIDIA GeForce RTX 2080Ti graphics processing units with 11GB of memory. The software environment is CUDA version 11.0, PyTorch 2.0.1, and Python 3.8.

[0138] The ten methods are summarized as follows:

[0139] (1) SSRN: SSRN uses two consecutive residual blocks to mine spectral and spatial features. Among them, a convolutional kernel with a size of (1×1×7) is used in the spectral feature extraction residual block, and a convolution with a size of (3×3×8) is used in the spatial feature extraction residual block. In addition, batch normalization processing is used in the network as a regularization function to improve the classification results;

[0140] (2) FDSSC: FDSSC is a spectral-spatial network designed specifically for hyperspectral image classification. It adopts a deep structure, uses a convolutional kernel of size (9×9×b) to extract spectral features, and simultaneously uses a convolutional kernel of size (7×7×7) to extract spatial features, thereby learning advanced spectral and spatial features;

[0141] (3) SSGCA: SSGCA is a network that integrates channel and global context. It consists of two parallel branches, namely the spectral branch and the spatial branch. In the spectral branch, channel global context attention is used to explore the interrelationships between channels. In the spatial branch, position global context attention is used to explore the relationships between positions. Finally, the spectral and spatial features are merged to facilitate hyperspectral image classification;

[0142] (4) HResNet: HResNet is a two-branch network designed specifically for fine-grained hyperspectral image classification. It uses HResNet blocks to extract multi-scale features and applies an attention mechanism to automatically calibrate spectral and spatial features at different scales. In addition, in the spectral HResNet block, the convolutional kernel size is (7×7×49), and in the spatial HResNet block, the convolutional kernel size is (7×7×1);

[0143] (5) MDFFA: MDFFA uses a 3D multi-scale feature module for efficient feature extraction, which includes introducing a fine-grained multi-scale receptive field. Then, a 3D two-branch feature interaction module is used to promote feature reuse in the previous layers. Finally, a 3D spatial-channel attention module is used to innovatively modify the previous weight distribution in the channel and spatial dimensions, enhancing the network's ability to represent features in hyperspectral image classification;

[0144] (6) HDDA: HDDA is a hybrid dense network with a dual mechanism for simultaneously extracting spectral-spatial features, which are carried out in 3D and 2D spaces respectively. In HDDA, a residual dual attention mechanism is constructed to enhance the features extracted in the channel and spatial dimensions respectively;

[0145] (7) OSDN: OSDN consists of two different feature extraction branches: one is the spectral branch and the other is the spatial branch. In the spectral and spatial branches, one-time dense blocks are used to independently extract spectral and spatial features respectively. In the spectral branch, a channel-only attention mechanism is used to emphasize important channel features, while in the spatial branch, a spatial-only attention mechanism is used to point to regions with more significant features. In addition, residual connections are added to the network to address problems related to performance gradient disappearance and saturation in the network;

[0146] (8) TBPFA: TBPFA is a dual convolutional neural network that integrates a full attention mechanism with polarization effects, effectively combining CNN and self-attention mechanisms for hyperspectral image classification. This network uses two independent branches to extract spectral and spatial features. In addition, a one-time connection is added to the network to facilitate network training;

[0147] (9) SSFTT: SSFTT utilizes a combination of a 3D convolutional layer and a 2D convolutional layer to extract local features. Subsequently, a Tokenization module weighted by a Gaussian distribution is used to convert local spectral-spatial features into Tokenized semantic features. Finally, the CNN and Transformer architectures are combined to provide semantic features for accurate classification of hyperspectral images;

[0148] (10) MVAHN: MVAHN proposes a hybrid network architecture that integrates CNN and Transformer components for local and global feature extraction. In addition, they develop a graph convolutional module to enhance the classification performance of the network;

[0149] To ensure the fairness of the experiment, unified input sizes and consistent variables are adopted. Specifically, the spatial size of the input hyperspectral image patches is (11×11), the number of image patches input to the network each time is 32, the number of training epochs is 200, the initial learning rate is set to 0.0003. In addition, the Adam optimizer is used with a decay rate of (0.9, 0.999). An early stopping mechanism is used during training. If the loss on the validation set remains unchanged for approximately 20 epochs, the training process will switch to the testing phase.

[0150] To evaluate the effectiveness of the network proposed in the present invention, three commonly used evaluation metrics are used in this embodiment, namely overall accuracy (OA), average accuracy (AA), and Kappa coefficient (Kappa), as well as the accuracy of each class. OA refers to the ratio of the number of target objects with true labels and predicted labels in the test set to the total number of target objects in the test dataset. AA is the expected value of the classification accuracy of all class targets in the test set. Kappa reflects the level of consistency between the true labels and predicted labels of the targets in the test set.

[0151] This embodiment aims to evaluate the effectiveness of TECCNet compared with CNN-based methods, CNN-Attention-based methods and CNN-Transformer-based methods, where the CNN-based methods include SSRN and FDSSC, the CNN-Attention-based methods include SSGCA, HResNet, MDFFA, HDDA, OSDN, TBPFA, and the CNN-Transformer-based methods include SSFTT and MVAHN. These methods are trained and evaluated using randomly sampled training sets and test sets. The overall accuracy (OA), average accuracy (AA), Kappa coefficient (Kappa) of these methods on four data sets, as well as the accuracy of each category, and their standard deviations are recorded in Tables 8 to 11 (the values ​​in Tables 8 to 11 are the average values ​​and corresponding standard deviations in five iterations). At the same time, the full pixel classification maps of each method are shown in Figure 7 , 8 , 9 and 10, for all the compared methods and our proposed method, all the training samples in each dataset account for 1% of the total dataset.

[0152] Table 8

[0153] The accuracy of classification for various methods on the PU.

[0154]

[0155]

[0156] Table 9The accuracy of classification for various methods on the Honghu.

[0157]

[0158] Table 10The accuracy of classification for various methods on theHouston.

[0159]

[0160]

[0161] Table 11The accuracy of classification for various methods on the KSC.

[0162]

[0163] First, to verify the performance of each module inside the TECCNet network, experiments were conducted on the PU dataset. The classification results regarding the PU dataset are shown in Table 8. In terms of the overall accuracy (OA), TECCNet is respectively about 22.75%, 0.01%, 2.8%, 4.86%, 13.27%, 2.42%, 0.56%, 5.4%, 21.41% and 0.95% higher than SSRN, FDSSC, SSGCA, HResNet, MDFFA, HDDA, OSDN, TBPFA, SSFTT, MVAH. This is because TECCNet uses a hybrid convolution module to jointly extract high-resolution spectral and spatial features, and then applies a spectral-spatial weight feature complementary module to supplement the spectral and spatial feature Tokens generated by the Transformer encoder, thereby preventing the loss of spectral and spatial information and making full use of spectral and spatial features at the same time; compared with other methods, the OA value of SSFTT is relatively low, mainly due to the fact that its feature extraction process only relies on 3D convolutional layers and 2D convolutional layers, resulting in insufficient feature extraction. Although the Transformer encoder is used, the insufficient information extraction still leads to unsatisfactory classification results. MVAHN combines CNN and Transformer, aiming to extract various types of features to achieve higher classification results. Methods such as OSDN, HDDA, SSGCA and HResNet are all dual-branch networks, which contribute to sufficient feature extraction and thus obtain excellent classification results. In addition, methods such as SSRN, TBPFA and FDSSC use multiple stacked convolutional layers, which also helps with feature extraction and classification. In addition, in different backgrounds of feature extraction, the lack of appropriate feature fusion and classification methods may also lead to inconsistent final classification results. Considering the OA value of the PU dataset, currently both TECCNet and OSDN have demonstrated the robustness of their networks as well as effective feature extraction and fusion. From Figure 7 it can be seen that except for SSRN, MDFFA and SSFTT, all classification methods generated smooth classification maps, while salt-and-pepper noise appeared in the classification maps of these three methods. At the same time, it is worth noting that in all these classification maps, the classification map generated by the network of the present invention Figure 7 (l) has the highest similarity with the ground truth label Figure 7 (a).

[0164] Furthermore, to verify the effectiveness and robustness of the TECCNet network in different scenarios, experiments were conducted in the complex WHU-Hi-Honghu scenario and the University of Houston scenario. As can be seen from Table 9 and Table 10, the overall accuracy (OA), average accuracy (AA), and Kappa coefficient values of TECCNet in WHU-Hi-Honghu and the University of Houston are higher than those of other methods. From the classification results of WHU-Hi-Honghu, it can be seen that in complex and diverse scenarios, except for SSRN, MDFFA, and SSFTT, the OA of more than half of the comparison methods exceeds 90%, while TECCNet achieved the highest OA value, namely 97.91%. In the University of Houston dataset, the OA value of TECCNet is still the highest among other methods, namely 90.30%. Especially for classes C3, C6, C13, and C14 in the University of Houston dataset, the TECCNet network achieved the highest values. At the same time, in Figure 8 and Figure 9 , the classification map of TECCNet is the clearest and smoothest compared to the comparison methods.

[0165] To verify the performance of TECCNet in the case of small samples, the KSC dataset was used for experiments. Table 11 lists the classification results of all methods in the context of small-sample training data. As can be seen from Table 7, the training dataset consists of 7, 2, 2, 2, 1, 2, 1, 4, 5, 4, 4, 5, and 9 samples randomly sampled from classes C1 to C13 respectively. Obviously, due to limited training data, the classification accuracy of almost all methods shows a relatively low level. It is obvious that TECCNet is superior to other methods in terms of overall accuracy (OA), average accuracy (AA), and Kappa coefficient. Another important phenomenon is that compared with other methods, except for the MVAHN method, TECCNet has an obvious advantage when the training set is limited, which also proves that MVAHN performs well under limited labeled samples. In fact, the smaller the number of the training set, the higher the OA value, and the more significant the feature extraction advantage of TECCNet under small-sample conditions. For example, in the Kennedy Space Center scenario, although the training datasets of classes C5 and C7 each have only 1 sample, TECCNet is still superior to other methods in terms of overall accuracy. According to the classification results in Figure 10 , it further proves that TECCNet performs well in terms of overall classification efficiency.

[0166] The above content is a further detailed description of the present invention in combination with specific embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as falling within the protection scope determined by the claims submitted for the present invention.

Claims

1. A hyperspectral image classification method based on a Transformer-enhanced dual-stream complementary convolutional neural network, characterized in that, The method includes the following steps: Step S1, extraction of key features in the high - resolution and spectral - spatial channels of the hyperspectral image, using SpeFES and SpaFES to extract spectral and spatial features respectively; The SpeFES includes a spectral mixing convolution module and a spectral attention mechanism composed of Transformer encoders, and the high - resolution feature map extracted by the spectral mixing convolution module is fed into the spectral attention mechanism; the SpaFES includes a spatial attention mechanism composed of Transformer encoders and a spatial mixing convolution module, and the mixing convolution module uses a simplified 3D convolution kernel; The mixing convolution module consists of five convolutional layers, the five convolutional layers include four 3D convolutional layers and one 3D transposed convolutional layer, and the 3D transposed convolutional layer is located in the middle of the other four 3D convolutional layers, and the feature extraction process of the mixing convolution module is expressed as follows: f(σ(X m+4 )); θ) = ω(σ(X m+4 )) * k m+1 + S m+1 σ(X m+4 ) = X m+1 + f(X m+1 ; θ) + ε(X m+2 ; θ) + f(X m+3 ; θ); where X m+1 represents the 3D cube input to the (m + 1)-th layer, k m+1 , S m+1 represent the convolutional kernel and stride of the (m + 1)-th layer respectively, f(X; θ) represents the 3D convolutional operation on X, and ε(X; θ) represents the 3D transposed convolutional operation on X; The feature maps A spe ∈R c×h×w and B spa ∈R c ×h×w also need to be preprocessed by tokenization respectively to obtain spectral tokens and spatial tokens with 64 output channels. Then, the spectral tokens and spatial tokens are fed into the spectral and spatial attention mechanisms respectively to obtain spectral tokens and spatial tokens with key features based on global dependencies; Step S2, using a spectral - spatial weight feature complementary module to supplement the spectral and spatial weight features between the spectral feature Token and the spatial feature Token; and simultaneously extracting spectral and spatial semantic features; The specific supplement process of the spectral - spatial weight feature complementary module is as follows: Among them, represents the weight parameter matrix, and C() represents the complementary function; The spectral-spatial weight feature complementary module calculates each t based on matrix U cls The expected E(x of each element in the Token i ) and taking it as the semantic feature of the corresponding Token is as follows: where X class is the set of elements belonging to class c, and n class represents the number of elements in class c; The output of the spectral - spatial weight feature complementary module is shown as follows: Among them, represents a spectral Token with semantic features, represents a spatial Token with semantic features; Step S3, using a classification module to generate a classification result.

2. The hyperspectral image classification method based on the Transformer-enhanced dual-stream complementary convolutional neural network according to claim 1, wherein: The 3D convolutional layer includes 3D convolution operation, batch normalization operation and Mish activation function, and in the 3D transposed convolutional layer, 3D transposed convolution, batch normalization operation and Mish activation function are used. The mathematical formula of the Mish activation function is as follows: f(x) = x tanh(softplus(x)) = x tanh(ln(1 + e x )) where x represents the input of the activation function, and Mish is in the range of [-0.31, ∞].

3. The hyperspectral image classification method based on the Transformer-enhanced dual-stream complementary convolutional neural network according to claim 1, wherein: Before obtaining the Tokens with different types of features with a channel dimension of 64 from the feature maps processed by the spectral mixing convolution module and the spatial mixing convolution module, these 3D patches need to be tokenized. The specific process is to add the clsToken and to A spe ∈R hw×64 and B spa ∈R hw ×64 respectively to obtain the spectral feature Token and the spatial feature Token, so as to effectively embed the spectral feature Token and the spatial feature Token into the spectral and spatial Transformer encoding mechanisms; The spectral and spatial Transformer encoding mechanism is as follows:

4. The hyperspectral image classification method based on the Transformer-enhanced dual-stream complementary convolutional neural network according to claim 1, characterized in that: The complementary process of the spectral-spatial weight feature complementary module is that the spectral-spatial weight feature complementary module copies the and Tokens after the processing of the spectral and spatial attention mechanisms from the spectral feature Tokens and spatial feature Tokens respectively, merges them with each other, and then constructs two new and Tokens respectively. Each uses the self-attention mechanism to complement the weight features in and , and they are used as the classification information of the spectral Token and the spatial Token respectively. and Token from the spectral feature Tokens and spatial feature Tokens respectively, merges them with each other, and then constructs two new and Tokens respectively. and Token, each uses the self-attention mechanism to complement each other and in the weight features, and they are used as the classification information of the spectral Token and the spatial Token respectively.

5. The hyperspectral image classification method based on the Transformer-enhanced dual-stream complementary convolutional neural network according to claim 1, wherein: The semantic feature extraction process of the spectral-spatial weight feature complementary module is as follows: for each element x i ∈t cls in it, perform softmax processing on x i , and then assign a probability vector p i to each x i ∈R class , where class represents the number of categories included in the ground to be classified. According to the probability vector p i , calculate the expected value E(x i ) of each x i . S(x i ) = p i E(x i ) = x i ·p i where S(x i ) represents the softmax function, and E(x i ) represents the expectation function. For all elements x i ∈t cls , its corresponding expectation matrix is expressed as U∈R C×class , where C is the total number of elements. Calculate E(x cls ) for each t i Token and use it as the semantic feature of the corresponding Token.

6. The hyperspectral image classification method based on the Transformer-enhanced dual-stream complementary convolutional neural network according to claim 1, characterized in that: The classification process in step S3 is that the spectral and spatial feature Tokens obtained based on step S2 are respectively processed by Reshape operation, and the spectral and spatial feature Tokens are input into an average pooling layer with batch normalization and Mish activation. The connection operation is used to merge the spectral and spatial feature Tokens along the channel dimension, fuse them into a feature Token with spectral - spatial semantic features, and obtain the final classification result through a fully - connected layer.

Citation Information

Patent Citations

  • Hyperspectral image classification method based on complementary integrated Transform network

    CN115205590A

  • Hyperspectral image classification method based on double-branch multi-scale Transform network

    CN117456263A