Hyperspectral image classification method based on M-HybridSN-Attention

The hyperspectral data set is expanded through the Mixup algorithm and the CBAM attention module is introduced into the HybridSN network, which solves the problem of insufficient overfitting and feature extraction capabilities in hyperspectral image classification, and achieves higher classification accuracy and computing efficiency.

CN114187469BActive Publication Date: 2025-05-13HARBIN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111340017.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-12
Publication Date
2025-05-13
Estimated Expiration
2041-11-12

AI Technical Summary

Technical Problem

There are problems in hyperspectral image classification with overfitting and insufficient 3DCNN feature extraction capabilities, especially when there are fewer training samples for hyperspectral data labeling.

Method used

The virtual samples are generated through the Mixup algorithm to expand the data set and reduce overfitting; the CBAM attention module is added to the 3DCNN part of the HybridSN network to enhance feature extraction capabilities.

Benefits of technology

It effectively alleviates the overfitting phenomenon during the training process, improves feature extraction capabilities, and improves the accuracy and calculation speed of hyperspectral image classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present invention discloses a hyperspectral image classification method based on M-HybridSN-Attention. Firstly, a Mixup algorithm is used to construct a hyperspectral image virtual data set to expand the original data. The amount of data after expansion is twice the amount of the original data, which greatly alleviates the overfitting phenomenon caused by the small sample size of the hyperspectral image; secondly, the network structure of the 3DCNN part in HybridSN is improved, and a convolution block attention module is added between the three-dimensional convolution layer and the Relu layer to enhance the discriminative features in the spectrum and space and suppress the non-discriminative features, thereby improving the role of the discriminative features in recognition; then a 2DCNN network is applied to distinguish the spatial information in different spectral bands without losing a large amount of spectral information, thereby ensuring the integrity of the hyperspectral data information; finally, the obtained spectral-spatial features are sent to the SoftMax classifier to obtain the final classification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of hyperspectral image classification, and mainly to a hyperspectral image classification method based on M-HybridSN-Attention. Background Art

[0002] Remote sensing technology is a technology that detects and identifies targets by sensing electromagnetic waves, visible light, and infrared rays reflected or radiated by the target from a long distance. Hyperspectral remote sensing is a new type of earth observation technology that emerged in the 1980s. It is a technology that obtains a lot of very narrow and spectrally continuous image data in the ultraviolet, visible light, near-infrared, and mid-infrared regions of the electromagnetic spectrum. The hyperspectral imaging spectrometer provides tens to hundreds of narrow-band spectral information for each pixel, and can generate a spectral curve between the spectral band and the spectral value. Therefore, the hyperspectral image not only contains two-dimensional spatial information, but also contains a large amount of spectral information. Different substances behave differently under different band spectral signals. According to the differences in spectral curves, we can classify different substances in hyperspectral images.

[0003] With the continuous development of deep learning, a hyperspectral image classification method based on deep learning has been proposed. This method has the advantage of automatically learning deep features of images, and can effectively extract more representative features in hyperspectral images, making breakthrough progress in hyperspectral image classification. In order to better utilize the spatial feature information of hyperspectral images, hyperspectral image classification algorithms based on convolutional neural networks (CNN) are widely used. Among them, when applying two-dimensional convolutional neural networks (2DCNN) for hyperspectral image classification, since 2DCNN is a shallow network, the original hyperspectral image needs to be reduced in dimension before training, which will lead to the loss of detail information and poor classification effect. Therefore, a hyperspectral image classification algorithm based on three-dimensional convolutional neural network (3DCNN) was proposed. 3DCNN is a classification network that constructs joint spectral-spatial information, which can effectively utilize the spatial information of the image and obtain deep spectral information at the same time. However, 3DCNN has disadvantages such as many network training parameters and slow calculation speed. HybridSN is a hyperspectral image classification network that combines 2DCNN and 3DCNN. It overcomes the problem that 2DCNN cannot obtain deep image features from the spectral dimension and the computational redundancy caused by the deeper 3DCNN network.

[0004] However, the above network models are all trained on the original hyperspectral data set, and the dimension of hyperspectral images is high and the number of labeled training samples is small, so the dimensionality disaster and Hughes phenomenon are prone to occur in the classification of hyperspectral images. In addition, as the number of network layers increases and the number of weight learning iterations increases, overfitting is prone to occur during training. In addition, when using 3DCNN for hyperspectral image classification, 3DCNN will consider both discriminative information that is beneficial to the classification of objects and non-discriminative information that interferes with the classification of objects in the feature extraction stage, which leads to insufficient feature extraction and ultimately unsatisfactory classification results. Summary of the invention

[0005] The technical problem to be solved by the present invention is to provide a hyperspectral image classification method based on M-HybridSN-Attention to address the training overfitting phenomenon caused by the small number of labeled training samples in hyperspectral data, as well as the insufficient feature extraction capability of 3DCNN and computational redundancy.

[0006] The technical solution adopted by the present invention to solve the technical problem mainly includes the following steps:

[0007] In the present invention, the original hyperspectral data is first processed by the Mixup algorithm to construct virtual samples, and the virtual samples are mixed with the original data to form a new data set. At this time, the number of samples in the new data set is twice that of the original data set. The new data set is sent to the network for training. As the number of samples increases, the overfitting phenomenon of network training caused by the small sample characteristics of hyperspectral data is effectively alleviated.

[0008] Secondly, this paper improves the network structure of HybridSN by adding a convolutional block attention module (CBAM) between each three-dimensional convolution layer and Relu layer in its 3DCNN network, using a total of 3 modules. After the feature map passes through CBAM, the pixel values ​​in the spectral and spatial dimensions of the feature map are given weights. Through the continuous learning of the network during the training process, the corresponding weights are close to 1 for the discriminative features of the spectral and spatial dimensions that are beneficial to classification; and the corresponding weights are close to 0 for the non-discriminative features of the spectral and spatial dimensions that have a negative impact on classification, so as to strengthen the discriminative features and suppress the non-discriminative features, improve the feature extraction ability of the network, and reduce the computational redundancy to a certain extent, which can ultimately improve the classification accuracy and computational speed of the network.

[0009] The beneficial effects of the present invention are as follows: by applying the Mixup algorithm, the number of hyperspectral data samples is increased, and the overfitting phenomenon occurring during the training process is effectively alleviated; by adding CBAM between the three-dimensional convolutional layer and the Relu layer in the 3DCNN part of the HybridSN network, that is, adding an attention mechanism in the spectral and spatial dimensions, after training, weights can be assigned to different spatial and spectral positions of the feature graph, thereby distinguishing between discriminative features and non-discriminative features, and improving the overall feature extraction capability; and then applying the 2DCNN network to distinguish spatial information in different spectral bands without losing a large amount of spectral information, thereby ensuring the integrity of hyperspectral data information and ultimately improving the accuracy of hyperspectral image classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Attached Figure 1 This is a flow chart of a method for hyperspectral image classification based on the M-HybridSN-Attention network disclosed in the present invention.

[0011] Attached Figure 2 A flowchart for the overall implementation.

[0012] Attached Figure 3 This is the 3DCNN network structure diagram of the embedded attention mechanism proposed in this invention.

[0013] Attached Figure 4 This is a diagram of the overall network architecture proposed by the present invention. DETAILED DESCRIPTION

[0014] The present invention is not limited by the following embodiments, and specific implementation methods can be determined based on the technical solutions and actual conditions of the present invention.

[0015] Combined with Figure 1 , which is a flow chart of a method for hyperspectral image classification based on M-HybridSN-Attention network disclosed in the present invention, specifically comprising the following steps:

[0016] A1. Build a sample quantity increase network based on Mixup, apply the Mixup algorithm to expand the sample operation of the original hyperspectral data, obtain a new data set after expansion, and then send the new data set to the subsequent network for training, where:

[0017] The specific steps to build a Mixup-based sample size increase network include:

[0018] B1. Obtaining the original hyperspectral data set:

[0019] In the present invention, for the input data, the hyperspectral original data set is directly used without any dimensionality reduction processing, which can protect the data structure and obtain complete image information.

[0020] B2. Process the original data set to obtain a random data set:

[0021] In this step, the original hyperspectral data set is randomly shuffled according to the spatial position index to form a random data set.

[0022] B3. Mixup the original data set and the random data set to obtain a virtual data set:

[0023] In this step, the Mixup method is a general neighborhood distribution. In order to solve the small sample problem of hyperspectral images, this paper introduces the Mixup method into the hyperspectral image classification method. The specific formula is as follows:

[0024]

[0025]

[0026] Among them, (x i ,y i ), (x j ,y j ) are two samples randomly selected from the training data, λ is the weight of the Mixup method, λ obeys Beta distribution, and its value range is λ∈[0,1]. Mixup produces a smooth transition between two classes. In the classification of hyperspectral images, it can solve the problem of small sample size of hyperspectral images and the problem that the same category of materials are affected by different radiation at different locations.

[0027] Next, we take out the spatial neighborhood pixel samples at the same spatial position in the original dataset and the random dataset, perform linear weighting operations on them according to weights λ and (1-λ), and finally get virtual samples. In this paper, λ is set to 0.5.

[0028] B4. Mix the virtual sample set with the original data set to obtain a new data set with expanded sample capacity:

[0029] The virtual samples are mixed with the spatial neighborhood pixel samples taken from the original data set to form a new data set. At this time, the number of samples in the new data set is twice that of the original data set.

[0030] A2. Construct a hyperspectral image classification network based on M-HybridSN-Attention. On the basis of applying the Mixup algorithm for data expansion, the HybridSN network is used as the basic model, and the attention mechanism is added to the 3DCNN part of the network to form M-HybridSN-Attention. The specific steps include:

[0031] B5. Add CBAM between each Conv3D and Relu layer in the internal 3DCNN network:

[0032] CBAM consists of two parts, namely the spectral attention mechanism module and the spatial attention mechanism module. Among them, the spectral attention recognizes and learns the feature information of the spectral bands of the hyperspectral data; the spatial attention recognizes and learns the spatial information of the hyperspectral data. Specifically:

[0033] In the application of spectral attention for data training, the feature map is first input, and its size is h×w×c. The input feature map is subjected to space-based maximum pooling (MaxPool) and average pooling (AvgPool) operations respectively to extract richer high-level features, and two feature maps of size 1×1×c are obtained, where c is the spectral dimension. These two feature maps are then sent to the same shared multi-layer perceptron (Shared MLP) network for further feature extraction. In this network, the number of neurons in the first layer is C / r, where r is the reduction rate, the purpose is to reduce parameter overhead, the activation function is set to Relu, and the number of neurons in the second layer is C. After passing through this parameter-sharing Shared MLP, the output size of the feature map remains unchanged. The two feature maps are element-wise added and sent to the sigmoid function for operation. At this time, each spectrum is assigned its own weight, and finally a spectral attention module of size 1×1×c is obtained. Then, the spectral attention module is multiplied with the original input feature map by the corresponding pixel values ​​to obtain a spectral attention output feature map of size h×w×c, whose pixel values ​​are reassigned according to the weights of each spectrum. The specific calculation process is shown in formula (3):

[0034]

[0035] Among them, M c (F) represents the spectral attention module of size 1×1×c, F represents the input feature map, σ represents the sigmoid function, W0∈R C / r×C , W1∈R C×C / r .

[0036] In the application of spatial attention for data training, the obtained spectral attention output feature map is subjected to channel-based maximum pooling (MaxPool) and average pooling (AvgPool), that is, weights are assigned to different spatial positions to obtain two feature maps of size h×w×1. The two pooled feature maps are then convolved with a convolution window of 7×7. After channel splicing, a feature map of size h×w×1 is obtained. Finally, after the sigmoid function, the spatial attention module is obtained. At this time, different spatial positions have their own weights. The spatial attention module is multiplied by the pixel values ​​of the corresponding positions of the spectral attention output feature map to obtain a feature map of size h×w×c, and its pixel values ​​are reassigned according to the weights of each spatial position. The specific calculation process is shown in formula (4):

[0037]

[0038] Among them, M s (F) represents the spatial attention module, F represents the spectral attention output feature map, σ represents the sigmoid function, and f 7×7 Represents a convolution operation with a convolution window size of 7×7.

[0039] Therefore, after the feature graph has been processed by the spectral attention mechanism and the spatial attention mechanism, each piece of information in the feature graph cube has its own weight according to its spectral and spatial position. Through continuous learning during the training process, the corresponding weights of the discriminative features of the spectral and spatial dimensions that are beneficial to classification are close to 1; and the corresponding weights of the non-discriminative features of the spectral and spatial dimensions that have a negative impact on classification are close to 0.

[0040] B6. Forming a HybridSN-Attention network with attention mechanism:

[0041] On the basis of the internal 3DCNN, a CBAM is added between each Conv3D and Relu layer, for a total of 3, to ensure full training of the attention mechanism. For the specific implementation steps, please refer to the attached Figure 3 .

[0042] B7. Get the classification result through SoftMax classifier.

[0043] A3. Input the image to be classified into the hyperspectral classification network to obtain the classification result.

[0044] Combined with Figure 2 , which is the overall implementation flow chart. Combining the above-mentioned image classification methods, the overall implementation process can be divided into the following steps:

[0045] First, the raw hyperspectral data are input into the network;

[0046] Secondly, the Mixup algorithm is introduced to perform data expansion operations;

[0047] Secondly, CBAM is added to the internal 3DCNN basic model to form a HybridSN-Attention network combined with the attention mechanism;

[0048] Secondly, the SoftMax classifier is used to classify the objects;

[0049] Secondly, obtain the classification results;

[0050] Secondly, train the classification model and make further judgments based on the classification accuracy. When the classification accuracy does not meet the preset conditions, perform network training until a network with higher classification accuracy is obtained for subsequent steps.

[0051] Finally, the network is evaluated using the test set.

[0052] Combined with Figure 3 , which is a 3DCNN network structure diagram of the embedded attention mechanism proposed in the present invention, and its specific structure is as follows:

[0053] Conv3D represents a three-dimensional convolution layer. The present invention sets three layers in total, each layer has 128, 192, and 256 convolution kernels, and the sizes of the three convolution kernels are 2×2×32, 2×2×16, and 2×2×10, respectively. A CBAM attention mechanism module is embedded behind each Conv3D. After the feature map is processed by CBAM, the Relu activation function is used to accelerate training and nonlinear processing. Average Pooling 3D represents a three-dimensional convolution pooling layer, and its step size is set to 2×2×2, which compresses the features while reducing the amount of calculation. The Dropout regularization method is used to randomly inactivate neurons in the hidden layer. The inactivation rate is set to 50% in this paper, which can alleviate the overfitting phenomenon to a certain extent. Conv2D represents a two-dimensional convolution layer. This paper sets a total of one layer, with a total of 288 convolution kernels, and the convolution kernel size is 2×2×1. Finally, the obtained results are input into the fully connected (FC, FullyConnected) layer, and then the Softmax activation function is used for classification operation to obtain the category label.

[0054] Combined with Figure 4 , which is the overall network architecture diagram proposed by the present invention.

[0055] The present invention discloses a method for classifying hyperspectral images based on M-HybridSN-Attention, which aims at the training overfitting phenomenon caused by the small sample characteristics of the hyperspectral image itself, and the problem of insufficient 3DCNN feature extraction capability when using HybridSN for training. By introducing the Mixup algorithm, the sample capacity is increased to alleviate overfitting; CBAM is embedded in the original 3DCNN network, and the attention mechanism is added to both the spectral and spatial dimensions. After continuous training of the network, it can effectively distinguish between discriminative features and non-discriminative features, thereby improving the feature extraction capability of the network, and then the features trained by the attention mechanism are sent to the 2DCNN network in the HybridSN, which can distinguish the spatial information in different spectral bands without losing a lot of spectral information, thereby ensuring the integrity of the hyperspectral data information, and finally effectively improving the image classification accuracy, which is suitable for long-term promotion and application.

[0056] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make more forms without departing from the scope of protection of the purpose of the present invention and the claims, which are all within the protection of the present invention.

Claims

1. A hyperspectral image classification method based on M-HybridSN-Attention, the implementation method is as follows: Firstly, the original hyperspectral data are randomly shuffled according to the spatial position to obtain a random data set. Then, the spatial neighborhood pixel samples in the original data set and the spatial neighborhood pixel samples in the random data set are linearly weighted according to the corresponding spatial positions based on the Mixup algorithm to obtain a virtual data set. Then, the virtual data set is mixed with the original data set to obtain a new data set with an expanded sample capacity. At this time, the sample capacity of the new data set is twice that of the original data set. Next, the new data set is sent to the subsequent network for training. The subsequent network consists of a HybridSN network embedded with an attention mechanism. That is, on the basis of the HybridSN network, a CBAM is added between each Conv3D layer and the Relu layer in its 3DCNN. A total of 3 CBAMs are added to form a HybridSN-Attention network that combines HybridSN with the attention mechanism. The SoftMax classifier is used for object classification. Combined with the previous Mixup algorithm, a hyperspectral image classification method based on M-HybridSN-Attention is formed.

2. The hyperspectral image classification method based on M-HybridSN-Attention according to claim 1, characterized in that: HybridSN consists of two parts, 2DCNN and 3DCNN. CBAM is embedded in 3DCNN, and CBAM consists of two parts, spectral attention and spatial attention. After continuous training of the network, each information in the feature map cube is given its own weight according to its spectrum and spatial position. Among them, for the discriminative features of the spectral and spatial dimensions that are beneficial to classification, the corresponding weight is close to 1; for the non-discriminative features of the spectral and spatial dimensions that have a negative impact on classification, the corresponding weight is close to 0. This can effectively improve the feature extraction ability of 3DCNN, and then apply the 2DCNN network to distinguish the spatial information in different spectral bands without losing a lot of spectral information, ensuring the integrity of hyperspectral data information and increasing the image classification accuracy of the network.

Citation Information

Patent Citations

  • Hyperspectral remote sensing image classification method based on dense residual three-dimensional convolutional neural network

    CN111368896A

  • Hyperspectral image classification method based on deep learning

    CN112101467A