Hyperspectral Image Classification Method Based on Spectral-Spatial Attention Mechanism of Dual-Branch Network

By adopting a dual-branch network and attention mechanism in hyperspectral image classification, the problem of insufficient feature extraction under finite samples is solved, and efficient feature representation and accurate classification results are achieved.

CN117218429BActive Publication Date: 2025-05-27ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311178205.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-13
Publication Date
2025-05-27
Estimated Expiration
2043-09-13

AI Technical Summary

Technical Problem

In the case of finite samples, the effective feature extraction of hyperspectral images is insufficient, resulting in low classification accuracy and large model complexity and computational volume of traditional methods.

Method used

The spectral-spatial attention mechanism based on the dual-branch network is adopted to reduce the dimensionality through principal component analysis (PCA), combining spectral subnetwork and spatial subnetwork, spectral and spatial features are extracted respectively, and feature extraction capabilities are enhanced through the attention mechanism.

Benefits of technology

The effective feature representation of hyperspectral images is improved under limited samples, ensuring classification accuracy, and reducing the number of parameters and calculation amount of the network model, achieving better classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117218429B_ABST
    Figure CN117218429B_ABST
Patent Text Reader

Abstract

The present invention relates to a hyperspectral image classification method based on a spectral-spatial attention mechanism of a dual-branch network, including: inputting hyperspectral image data cubes; using principal component analysis to reduce the spectral dimension of the hyperspectral image, and for the pixels to be classified, encapsulating them into adjacent region blocks as the input of the spectral sub-network; creating an adjacent region around each specific pixel to collect spatial information and using it as the input of the spatial sub-network; obtaining one-dimensional spectral features; obtaining one-dimensional spatial features; fusing and balancing the one-dimensional spectral features and the one-dimensional spatial features through a fusion layer, and using a softmax regression layer to predict the probability distribution of each type of ground object. By taking the hyperspectral image as the research object, the present invention improves the effective feature representation of the hyperspectral image under the premise of limited samples, ensuring both the classification accuracy and reducing the number of parameters and the computational amount of the network model; fully utilizing the spectral information and the spatial information to obtain better classification results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing, feature extraction, spectral feature selection, and visual attention, and in particular to a hyperspectral image classification method based on a spectral-spatial attention mechanism of a dual-branch network. Background Art

[0002] Hyperspectral image classification is an important task in the field of remote sensing image processing and analysis. With its rich spatial and spectral features, it has become a research hotspot in the remote sensing field. However, the main difficulty in hyperspectral image classification lies in the huge amount of data, strong band correlation, and at the same time, the redundancy and noise of high-dimensional data, resulting in low classification accuracy.

[0003] Traditional machine learning methods for hyperspectral image classification, such as support vector machine (SVM), k-nearest neighbor (KNN), etc., have some limitations when dealing with hyperspectral image classification. These methods often require manual feature extraction and cannot fully utilize the spectral and spatial information of hyperspectral images.

[0004] With the development of deep learning, convolutional neural network (CNN) has been introduced into hyperspectral image classification. CNN has good feature learning ability and has achieved good performance in hyperspectral image classification tasks. Networks with one-dimensional and two-dimensional convolutional layers are widely used in hyperspectral image classification. The one-dimensional network method takes the spectrum as the input and uses the spectral information to learn features, but does not utilize the original spatial features of the image. Two-dimensional CNN is used to process spatial information. However, two-dimensional CNN cannot extract good discriminative features from the spectral dimension. Similarly, using the depth feature extraction of three-dimensional CNN for HSI classification can extract spectral and spatial information simultaneously. Although the classification performance is improved, due to more parameters, it may lead to overfitting and increase the computational cost. At the same time, when dealing with hyperspectral images, CNN cannot fully utilize the spatial information between pixels, while the attention mechanism can help the network focus on the key parts of the feature space, more effectively allocate limited attention resources to the core area, select more important features, and thus improve the performance of visual tasks. Therefore, researchers have begun to explore introducing visual attention modules into CNN to further improve the performance of hyperspectral image classification. Different types of attention mechanisms, such as spatial attention mechanism and channel attention mechanism, are introduced into CNN to improve the network's perception ability of spectral and spatial features.

[0005] In recent years, dual-branch CNN networks have been proposed to handle hyperspectral image classification tasks. The model is divided into two branches, using different attention modules to extract spatial information and spectral information respectively, achieving good classification results. Although the introduction of the attention module can effectively improve the information extraction ability of the model, it will also increase the computational complexity of the model.

[0006] In summary, compared with traditional machine learning methods, the above methods have more advantages in hyperspectral image classification and have strong generalization ability. However, how to extract effective feature representations of hyperspectral images under limited samples has become a key part in the research of hyperspectral image classification. During the extraction process of hyperspectral images, a large amount of redundant information and the imbalance between different labeled samples greatly reduce the classification performance of hyperspectral images. Therefore, how to obtain more features under limited samples still deserves further research. Summary of the Invention

[0007] To solve the problems of insufficient extraction of effective features of hyperspectral images and complex classification method models under limited samples, the purpose of the present invention is to provide a hyperspectral image classification method based on a spectral-spatial attention mechanism of a dual-branch network, which can ensure the classification accuracy and reduce the number of parameters and computational complexity of the network model under the premise of limited samples.

[0008] To achieve the above purpose, the present invention adopts the following technical solutions: A hyperspectral image classification method based on a spectral-spatial attention mechanism of a dual-branch network, the method includes the following steps in sequence:

[0009] (1) Input the hyperspectral image data cube I ∈ R H×W×B , where H, W, and B respectively represent the length, width, and spectral dimension of I;

[0010] (2) Use principal component analysis PCA to reduce the spectral dimension B of the hyperspectral image to c. After dimensionality reduction, for the pixel P to be classified, encapsulate it into an adjacent region block X spa as the input of the spectral sub-network; at the same time, create an adjacent region X spa around each pixel P to collect spatial information and use it as the input of the spatial sub-network. The size of the adjacent region X spa is w × w × c, and w × w represents the spatial size;

[0011] (3) Input the adjacent region block X spe into the spectral sub-network, and after passing through three three-dimensional convolutional layers, one two-dimensional convolutional layer, one channel attention mechanism module, and three fully connected layers, finally obtain one-dimensional spectral features;

[0012] (4) Input the adjacent region X spaThe input space sub-network passes through a spatial attention residual module, two two-dimensional convolutional layers, two max pooling layers, and a fully connected layer, and finally obtains one-dimensional spatial features;

[0013] (5) The one-dimensional spectral features and one-dimensional spatial features are fused and balanced through a fusion layer, and a softmax regression layer is used to predict the probability distribution of each type of ground object.

[0014] The specific steps of step (3) are as follows:

[0015] (3a) Input the adjacent region block X spe into the spectral sub-network, and perform three-dimensional convolution operations on the adjacent region block X spe in sequence. The sizes of the three-dimensional convolution kernels are 3×3×7, 3×3×5, and 5×5×3 respectively. Add a bias term and use the ReLU activation function for activation. Through three-dimensional convolution, a feature cube covering spectral spatial information is finally generated;

[0016] (3b) Rearrange the feature cube as the input of the two-dimensional convolutional layer, and use a 3×3 convolutional kernel to perform convolution operations on the rearranged feature cube to obtain a 64-channel feature map, and then use the ReLU activation function for activation;

[0017] (3c) Introduce a channel attention mechanism module after the two-dimensional convolutional layer. The size of the feature map activated by the ReLU activation function is X spe ∈R s×s×d , where s×s represents the spatial size and d represents the number of channels; the feature map activated by the ReLU activation function undergoes global average pooling and is compressed into a feature vector [1, 1, d]; then, through a fully connected layer, the pooling result is mapped to the size of the original number of feature channels that is where ratio is the ratio. Activate the output of this fully connected layer using the ReLU activation function, and then map the output back to the size of the original number of feature channels through another fully connected layer, and convert it into a normalized weight vector restricted between 0 and 1 through the Sigmoid activation function; finally, multiply the feature map activated by the ReLU activation function with the normalized weight channel by channel, and expand the obtained weighted feature map into a one-dimensional vector;

[0018] (3d) Use three fully connected layers with 1024, 512, and 256 neurons respectively to perform linear transformations on the one-dimensional vector, and use the ReLU activation function for activation, and finally obtain one-dimensional spectral features.

[0019] The specific steps of step (4) are as follows:

[0020] (4a) Input the adjacent region Xspa The input space sub-network adjusts X spa in shape to meet the training requirements of the space sub-network;

[0021] (4b) The spatial attention residual module is located at the head of the space sub-network. The adjusted X spa is input into the spatial attention residual module to obtain a feature map. Let the input feature X spa ∈R w×w×c , where w×w represents the spatial size and c represents the spectral dimension. The input feature is normalized using a BN layer, the normalized feature is activated using the ReLU function, and a 1×1 convolutional kernel is used to reduce the number of channels of the feature map to 4. The feature map is normalized again using a BN layer, the normalized feature map is activated using the ReLU function, and a 3×3 convolutional kernel is used to perform a convolution operation on the feature map. A spatial attention mechanism is added to the convolved feature map to obtain the weighted attention feature x spa . The input feature map and the weighted attention feature x spa are merged in the channel dimension to obtain a new feature input f(x). A spatial attention mechanism is added before the concatenation operation to construct the spatial attention residual module, which is expressed as:

[0022] y = [[x, f(x)], f([x, f(x)])]

[0023] where x and y represent the input feature and output feature of the spatial attention residual module respectively; [·] is the concatenation operation, and f(·) is the composite function operation, including the normalization layer - ReLU activation function - two-dimensional convolution 1×1, i.e., BN - ReLU - Conv1×1, the normalization layer - ReLU activation function - two-dimensional convolution 3×3, i.e., BN - ReLU - Conv3×3, and the spatial attention mechanism;

[0024] (4c) A 5×5 convolutional kernel is used to perform a two-dimensional convolution operation on the feature map to obtain a 32-channel feature map, and a max-pooling operation is performed to reduce the size of the feature map to half of the original. A 5x5 convolutional kernel is used again to perform a two-dimensional convolution operation and a max-pooling operation;

[0025] (4d) The output of the second max-pooling layer is unfolded into a one-dimensional tensor, passed through a fully connected layer operation, weights and biases are used therein, and then passed through the ReLU activation function. Overfitting is reduced by using a dropout layer. A linear transformation and a ReLU activation operation are performed on the output of the dropout layer to obtain a feature vector, and a linear transformation is performed on it to obtain the output of the space sub-network, i.e., the one-dimensional spatial feature.

[0026] Step (5) specifically includes the following steps:

[0027] (5a) Process the one-dimensional spectral features and one-dimensional spatial features, and use ReLU for activation operation processing;

[0028] (5b) Connect the last fully-connected layer in the spectral sub-network and the last fully-connected layer in the spatial sub-network to form a new fully-connected layer, fuse the one-dimensional spectral features and one-dimensional spatial features, and utilize both spectral correlation and spatial correlation to extract joint spectral-spatial features;

[0029] (5c) Predict the probability distribution of each type of ground object by using a softmax regression layer.

[0030] As can be seen from the above technical solutions, the beneficial effects of the present invention are as follows: First, the spectral sub-network and spatial sub-network in the present invention constitute a hyperspectral image classification method. By taking the hyperspectral image as the research object, it improves the effective feature representation of the hyperspectral image under the premise of limited samples, ensuring both classification accuracy and reducing the number of parameters and computational complexity of the network model; Second, the hyperspectral image classification method of the present invention adopts a dual-branch network to respectively perform spectral and spatial correlation learning, making full use of spectral information and spatial information to obtain better classification results; Third, the present invention respectively adds different attention mechanisms to the dual-branch network. The spectral sub-network embedded with the channel attention mechanism module serves as a spectral feature learner, and the spatial sub-network embedded with the spatial attention residual module serves as a spatial feature learner, more effectively allocating limited attention resources to the core area, selecting more important features, and improving the information extraction ability of the model; Fourth, in the present invention, the spatial attention mechanism is combined with the residual network, and the weighted attention feature map is fused with the original feature map to extract a more informative feature representation, greatly improving the performance and accuracy of the classification model. Description of the Drawings

[0031] Figure 1 It is the overall architecture diagram of the present invention;

[0032] Figure 2 It is the basic structure diagram of the channel attention mechanism module embedded in the present invention;

[0033] Figure 3 It is the basic structure diagram of the spatial attention residual module embedded in the present invention;

[0034] Figure 4 It is the bar chart of the influence of the attention mechanism on the overall accuracy;

[0035] Figure 5 It is the bar chart of the influence of the attention mechanism on the average accuracy;

[0036] Figure 6Bar chart of the influence of the attention mechanism on the Kappa coefficient;

[0037] Figure 7 Classification diagram of six classification methods for the Pavia Center dataset;

[0038] Figure 8 Classification diagram of six classification methods for the Pavia University dataset;

[0039] Figure 9 Classification diagram of six classification methods for the Salinas dataset. Detailed implementation manners

[0040] As Figure 1 shown, a hyperspectral image classification method based on a dual-branch network with spectral-spatial attention mechanism, the method includes the following steps in sequence:

[0041] (1) Input the hyperspectral image data cube I ∈ R H×W×B , where H, W, and B respectively represent the length, width, and spectral dimension of I;

[0042] (2) Use principal component analysis PCA to reduce the spectral dimension B of the hyperspectral image to c. After reduction, for the pixel P to be classified, encapsulate it into an adjacent region block X spe as the input of the spectral sub-network; at the same time, create an adjacent region X spa around each pixel P to collect spatial information and use it as the input of the spatial sub-network. The size of the adjacent region X spa is w × w × c, and w × w represents the spatial size;

[0043] (3) Input the adjacent region block X spe into the spectral sub-network. After passing through three three-dimensional convolutional layers, one two-dimensional convolutional layer, one channel attention mechanism module, and three fully connected layers, finally obtain a one-dimensional spectral feature;

[0044] (4) Input the adjacent region X spa into the spatial sub-network. After passing through one spatial attention residual module, two two-dimensional convolutional layers, two max pooling layers, and one fully connected layer, finally obtain a one-dimensional spatial feature;

[0045] (5) Through the fusion layer, fuse and balance the one-dimensional spectral feature and the one-dimensional spatial feature, and use the softmax regression layer to predict the probability distribution of each type of ground object.

[0046] The step (3) specifically includes the following steps:

[0047] (3a) Input the adjacent region block X spe into the spectral sub-network, and perform operations on the adjacent region block Xspe Perform 3D convolution operations in sequence. The sizes of the 3D convolution kernels are 3×3×7, 3×3×5, and 5×5×3 respectively. Add a bias term and use the ReLU activation function for activation. Through 3D convolution, a feature cube covering spectral spatial information is finally generated;

[0048] (3b) Rearrange the feature cube as the input of the 2D convolutional layer. Use a 3×3 convolutional kernel to perform convolution operations on the rearranged feature cube to obtain a feature map with 64 channels, and then use the ReLU activation function for activation;

[0049] (3c) Introduce a channel attention mechanism module after the 2D convolutional layer. The size of the feature map activated by the ReLU activation function is X spe ∈R s×s×d , where s×s represents the spatial size and d represents the number of channels; the feature map activated by the ReLU activation function undergoes global average pooling and is compressed into a feature vector [1, 1, d]; then, through a fully connected layer, the pooling result is mapped to the size of the original number of feature channels That is where ratio is the ratio. Activate the output of this fully connected layer using the ReLU activation function, and then map the output back to the size of the original number of feature channels through another fully connected layer, and convert it into a normalized weight vector restricted between 0 and 1 through the Sigmoid activation function; finally, multiply the feature map activated by the ReLU activation function with the normalized weights channel by channel, and expand the obtained weighted feature map into a one-dimensional vector;

[0050] (3d) Use three fully connected layers with 1024, 512, and 256 neurons respectively to perform linear transformation on the one-dimensional vector, and use the ReLU activation function for activation to finally obtain one-dimensional spectral features.

[0051] The specific steps of step (4) include the following steps:

[0052] (4a) Input the adjacent region X spa into the spatial sub-network, and re-adjust the shape of X spa to meet the training requirements of the spatial sub-network;

[0053] (4b) The spatial attention residual module is located at the beginning of the spatial sub-network. Input the adjusted X spa into the spatial attention residual module to obtain a feature map; assume the input feature X spa ∈R w×w×c, where \(w\times w\) represents the spatial size, \(c\) represents the spectral dimension. The input features are normalized using the BN layer, the ReLU function is used to activate the normalized features, and a \(1\times1\) convolutional kernel is used to reduce the number of channels of the feature map to 4. The feature map is normalized again using the BN layer, the ReLU function is used to activate the normalized feature map, and a \(3\times3\) convolutional kernel is used to perform a convolution operation on the feature map. A spatial attention mechanism is added to the convolved feature map to obtain the weighted attention feature \(x\). spa , the input feature map and the weighted attention feature \(x\). spa are merged in the channel dimension to obtain a new feature input \(f(x)\). A spatial attention mechanism is added before the concatenation operation to construct a spatial attention residual module, which is expressed as:

[0054] y = [[x, f(x)], f([x, f(x)])]

[0055] where \(x\) and \(y\) represent the input feature and output feature of the spatial attention residual module respectively; [·] is the concatenation operation, and \(f(·)\) is the composite function operation, including the normalization layer - ReLU activation function - two-dimensional convolution \(1\times1\) i.e., BN - ReLU - Conv1×1, the normalization layer - ReLU activation function - two-dimensional convolution \(3\times3\) i.e., BN - ReLU - Conv3×3, and the spatial attention mechanism;

[0056] (4c) A \(5\times5\) convolutional kernel is used to perform a two-dimensional convolution operation on the feature map to obtain a 32-channel feature map. A max-pooling operation is performed to reduce the size of the feature map to half of the original. A \(5\times5\) convolutional kernel is used again to perform a two-dimensional convolution operation and a max-pooling operation;

[0057] (4d) The output of the second max-pooling layer is unfolded into a one-dimensional tensor, passed through a fully connected layer operation, weights and biases are used therein, and then passed through the ReLU activation function. Overfitting is reduced by using a dropout layer. A linear transformation and a ReLU activation operation are performed on the output of the dropout layer to obtain a feature vector, and a linear transformation is performed on it to obtain the output of the spatial subnetwork, i.e., the one-dimensional spatial feature.

[0058] The specific steps of step (5) include the following steps:

[0059] (5a) The one-dimensional spectral feature and the one-dimensional spatial feature are processed, and the ReLU is used for activation operation processing;

[0060] (5b) The last fully-connected layer in the spectral sub-network is connected to the last fully-connected layer in the spatial sub-network to form a new fully-connected layer, which fuses one-dimensional spectral features and one-dimensional spatial features, and simultaneously utilizes spectral correlation and spatial correlation to extract joint spectral-spatial features;

[0061] (5c) The probability distribution of each type of ground object is predicted by using a softmax regression layer.

[0062] As Figure 2 shown, the channel attention mechanism module enhances or suppresses different channels for different tasks by modeling the importance of each feature's channel, and is divided into two parts: squeeze and excitation. First, the global spatial information is compressed, then feature learning is performed in the channel dimension to form the importance of each channel, and finally different weights are assigned to each channel through the excitation part.

[0063] As Figure 3 shown, this figure details the structure of the spatial attention residual module. The adjusted X spa is input into the spatial attention residual module to obtain a more informative feature map. The input feature map is enhanced through the spatial attention mechanism, and the weighted attention feature map is fused with the original feature map to extract a more informative feature representation. Applying the spatial attention mechanism to the ResNet residual network greatly improves the performance and accuracy of the HSI model.

[0064] Figure 4 、 5 、6 are the effects of the attention mechanism on the overall accuracy, average precision, and Kappa coefficient respectively. From Figure 4 、 5As can be seen from Figure 6, the overall classification of the channel attention mechanism module, i.e., the SE module, is better than the non-attention mechanism, and the spatial attention dense module is better than the SE module. The effect of embedding the SE module and the spatial attention dense module at the same time is better. The embedding of the two modules has obvious improvements on the three datasets of Pavia Center, Pavia University, and Salinas, i.e., PC, PU, ​​and SV. For the PC dataset, the use of the SE module or the spatial attention dense module alone has greatly improved the overall accuracy, average accuracy, and Kappa coefficient of PC. For the PU dataset, the use of the SE module or the spatial attention dense module alone has a significant improvement on its average accuracy, but the overall accuracy and Kappa coefficient are not very good, and it is difficult to play an effective role. For the SV dataset, the use of the SE module or the spatial attention dense module alone has a significant improvement on its overall accuracy, average accuracy, and Kappa coefficient. The experimental results show that the simultaneous use of the SE module and the spatial attention dense module can make full use of spectral information and spatial information, and can achieve good results in hyperspectral image classification.

[0065] Tables 1, 2, and 3 show the classification accuracy of the six classification methods for the Pavia Center, Pavia University, and Salinas datasets, respectively.

[0066] Table 1 Accuracy of six classification methods for Pavia Center dataset (%)

[0067]

[0068]

[0069] Table 2 Accuracy of six classification methods for Pavia University dataset (%)

[0070]

[0071] Table 3 Accuracy of six classification methods for the Salinas dataset (%)

[0072]

[0073]

[0074] Figure 7 , 8, 9 are the classification diagrams of six classification methods for Pavia Center, Pavia University, and Salinas datasets respectively. Among them, (a) is the false color image, (b) is the true ground object distribution, (c) is CNN, (d) is ACNN, (e) is SSAN, (f) is Hybrid-SN, (g) is DBSMA, and (h) is DBSSAN. Three typical publicly available hyperspectral datasets, Pavia Center, Pavia University, and Salinas, were selected for experiments. Five classification methods, CNN, ACNN, SSAN, Hybrid-SN, and DBSMA, were used to compare with the DBSSAN of the present invention, and the training samples used by each method were exactly the same. It can be seen from Tables 1, 2, and 3 that the overall accuracy and Kappa coefficient of the method used in the present invention reach the highest, and the classification accuracy of each category is also very excellent.

[0075] In Figure 7 , most areas of (c) are incompletely classified, and there are many misclassification phenomena. The misclassification phenomenon of (d) is also relatively serious, and there are many noise points. The misclassification phenomenon of (e) is significantly improved, but the classification effect is uneven. The classification effect of (f) is not smooth, and the boundaries of each category are not clear enough. Most areas of (g) are relatively completely classified, and the classification effect of (h) is more uniform and accurate, and the boundaries of each category are clearer and closer to the true ground object distribution map.

[0076] In Figure 8 , the ground object classification effect of (c) is poor, and the boundaries of each category are blurred. The classification effect of each area of (d) has been significantly improved, but the boundaries of each category are blurred. The classification effect of (e) is not smooth, and there are many noise points. The classification effect of each area of (f) is good, but there are some noise points. The misclassification phenomenon of (g) is significantly improved, and there are fewer noise points. Most areas of (h) are completely classified, and the classification effect of each ground object is very good, preserving the integrity of the object.

[0077] In Figure 9 , the speckle noise phenomenon of (c) is relatively serious, and there are many misclassifications. The speckle noise phenomenon in the classification effect diagram of (d) is significantly more. Most areas of (e) are relatively completely classified, but the speckle noise phenomenon is more. The classification of each ground object in (f) is relatively good, but there is still a speckle noise phenomenon. The boundaries of each category in (g) are clear, and the speckle noise phenomenon is significantly reduced. The classification effect of (h) is more uniform and smooth, and there are fewer noise points. It can be seen that the classification effect of each ground object of the present invention is very good, preserving the integrity of the object, having great advantages in capturing unique features of different categories, and the classification effect is smoother, with significantly fewer noise points, and having better classification performance and robustness compared with other methods.

[0078] In summary, the spectral sub-network and the spatial sub-network in the present invention constitute a hyperspectral image classification method. By taking the hyperspectral image as the research object, the effective feature representation of the hyperspectral image is improved under the premise of limited samples, which not only ensures the classification accuracy but also reduces the number of parameters and the computational complexity of the network model. The hyperspectral image classification method of the present invention uses a dual-branch network to separately perform spectral and spatial correlation learning, fully utilizing spectral information and spatial information to obtain better classification results.

Claims

1. A hyperspectral image classification method based on a spectral-spatial attention mechanism of a dual-branch network, the method comprising the following steps in sequence: (1) Input the hyperspectral image data cube \(I\in\mathbb{R}\) H×W×B , where \(H\), \(W\), and \(B\) represent the length, width, and spectral dimension of \(I\), respectively; (2) Use principal component analysis (PCA) to reduce the spectral dimension B of the hyperspectral image to c. For the pixel P to be classified, encapsulate it into an adjacent region block X spe as the input of the spectral sub-network; at the same time, create an adjacent region X around each pixel P spa to collect spatial information and use it as the input of the spatial sub-network. The adjacent region X spa has a size of w×w×c, where w×w represents the spatial size; (3) Input the adjacent regional block X spe into the spectral sub-network. After passing through three 3D convolutional layers, a feature cube covering spectral spatial information is generated. Rearrange the feature cube and input it into the 2D convolutional layer to obtain a feature map. Input the feature map into the channel attention mechanism module to obtain a weighted feature map and expand it into a one-dimensional vector. Finally, obtain the one-dimensional spectral feature through three fully connected layers; (4) Adjacent region X spa is input into the spatial sub-network, and a feature map is obtained through a spatial attention residual module. Through two two-dimensional convolutional layers and two max pooling layers, two-dimensional convolutional operations and max pooling operations are performed on the feature map, reducing the size of the feature map to half of the original. Then, two-dimensional convolutional operations and max pooling operations are performed again. The output of the second max pooling layer undergoes a fully connected layer operation to finally obtain a one-dimensional spatial feature; (5) Fuse and balance the one-dimensional spectral feature and the one-dimensional spatial feature through a fusion layer, and use a softmax regression layer to predict the probability distribution of each type of ground object.

2. The hyperspectral image classification method based on the spectral-spatial attention mechanism of a dual-branch network according to claim 1, characterized in that: The specific steps of the step (3) are as follows: (3a) Input the adjacent region block X spe into the spectral sub-network, and perform 3D convolution operations on the adjacent region block X spe successively. The sizes of the 3D convolution kernels are 3×3×7, 3×3×5, and 5×5×3 respectively. Add the bias term and use the ReLU activation function for activation. Through 3D convolution, finally generate a feature cube covering spectral spatial information; (3b) Rearrange the feature cube as the input of a two-dimensional convolutional layer, perform a convolutional operation on the rearranged feature cube using a 3×3 convolutional kernel to obtain a feature map with 64 channels, and then use a ReLU activation function for activation; (3c) After the two-dimensional convolutional layer, a channel attention mechanism module is introduced. The size of the feature map after activation using the ReLU activation function is X spe ∈R s×s×d , where s×s represents the spatial size and d represents the number of channels; the feature map after activation using the ReLU activation function undergoes global average pooling and is compressed into a feature vector [1, 1, d]; then, through a fully connected layer, the pooling result is mapped to the size of the original feature channels That is where ratio is the ratio. The output of this fully connected layer is activated using the ReLU activation function, and then through another fully connected layer, the output is mapped back to the size of the original feature channels and transformed into a normalized weight vector restricted between 0 and 1 through the Sigmoid activation function; finally, the feature map after activation using the ReLU activation function is multiplied by the normalized weight channel by channel, and the weighted feature map is obtained and unfolded into a one-dimensional vector; (3d) Perform a linear transformation on the one-dimensional vector using three fully connected layers with 1024, 512, and 256 neurons in sequence, and use a ReLU activation function for activation to finally obtain a one-dimensional spectral feature.

3. The hyperspectral image classification method based on the spectral-spatial attention mechanism of a dual-branch network according to claim 1, characterized in that: The specific steps of the step (4) are as follows: (4a) Input the adjacent region X spa into the spatial sub-network and re-adjust the shape of X spa to meet the training requirements of the spatial sub-network; (4b) The spatial attention residual module is located at the beginning of the spatial sub-network and inputs the adjusted X spa into the spatial attention residual module to obtain a feature map. Let the input feature be X spa ∈R w×w×c , where w×w represents the spatial size and c represents the spectral dimension. The input feature is normalized using a BN layer, the normalized feature is activated using the ReLU function, and the number of channels of the feature map is reduced to 4 using a 1×1 convolutional kernel. The feature map is again normalized using a BN layer, the normalized feature map is activated using the ReLU function, and the feature map is convolved using a 3×3 convolutional kernel. A spatial attention mechanism is added to the convolved feature map to obtain the weighted attention feature x spa . The input feature map and the weighted attention feature x spa are merged in the channel dimension to obtain a new feature input f(x). A spatial attention mechanism is added before the concatenation operation to construct the spatial attention residual module, which is expressed as: y = [[x, f(x)], f([x, f(x)])] where x and y respectively represent the input feature and the output feature of the spatial attention residual module; [·] is a concatenation operation, and f(·) is a composite function operation, including a normalization layer - ReLU activation function - two-dimensional convolution 1×1, i.e., BN - ReLU - Conv1×1, a normalization layer - ReLU activation function - two-dimensional convolution 3×3, i.e., BN - ReLU - Conv3×3, and a spatial attention mechanism; (4c) Perform a two-dimensional convolution operation on the feature map using a 5×5 convolutional kernel to obtain a feature map with 32 channels, perform a max pooling operation to reduce the size of the feature map to half of the original, and perform a two-dimensional convolution operation again using a 5×5 convolutional kernel and a max pooling operation; (4d) Expand the output of the second max pooling layer into a one-dimensional tensor, perform a fully connected layer operation, use weights and biases therein, and then pass through a ReLU activation function. Reduce overfitting by using a random inactivation Dropout layer; perform a linear transformation and a ReLU activation operation on the output of the random inactivation Dropout layer to obtain a feature vector, perform a linear transformation on it to obtain the output of the spatial sub-network, i.e., a one-dimensional spatial feature.

4. The hyperspectral image classification method based on the spectral-spatial attention mechanism of a dual-branch network according to claim 1, characterized in that: The specific steps of the step (5) are as follows: (5a) Process the one-dimensional spectral feature and the one-dimensional spatial feature, and perform an activation operation using ReLU; (5b) Connect the last fully connected layer in the spectral sub-network with the last fully connected layer in the spatial sub-network to form a new fully connected layer, fuse the one-dimensional spectral feature and the one-dimensional spatial feature, and extract joint spectral-spatial features by utilizing spectral correlation and spatial correlation. (5c)Predict the probability distribution of each type of ground object by using a softmax regression layer.

Citation Information

Patent Citations

  • Hyperspectral remote sensing image classification method based on attention joint network

    CN115564996A

  • Hyperspectral image classification method based on double-branch spatial-spectral global feature extraction network

    CN116563606A