A Hyperspectral Image Classification Method Based on Deep Learning Combining Spatial and Spectral Information
By introducing the dual-branch multi-scale null spectral feature fusion and self-attention mechanism in the hyperspectral image classification model, the problems of missing features and insufficient extraction capabilities during the feature extraction process are solved, and high-precision hyperspectral image classification is achieved.
Patent Information
- Application Number
- CN202410640445.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-05-22
AI Technical Summary
The hyperspectral image classification model has problems such as lack of feature and insufficient extraction ability during feature extraction, especially in the case of limited samples, which leads to low classification accuracy.
A hyperspectral image classification model based on deep learning based on dual-branch multi-scale null spectral feature fusion and self-attention mechanism is used to convert hyperspectral image data into compact representations through data preprocessing, and feature extraction and fusion are used using dense dual-branch pyramid feature extraction module and channel-space attention module.
It significantly improves the accuracy of hyperspectral image classification, alleviates the problems of feature loss and insufficient extraction ability, and improves the model's extraction ability of significant features.
Smart Images

Figure CN118587482B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of hyperspectral image classification, and particularly relates to a spatial-spectral joint hyperspectral image classification method based on deep learning. Background Art
[0002] Hyperspectral remote sensing is a multi-dimensional information acquisition technology integrating imaging technology and spectral technology, which simultaneously detects the two-dimensional spatial information and one-dimensional spectral information of ground object targets, so as to obtain hyperspectral images with high spectral resolution and numerous bands. Hyperspectral images can reflect the nearly continuous spectral characteristic curves of ground objects, contain rich spatial and spectral information, and are a comprehensive carrier of various information. Hyperspectral image classification is the basis of various hyperspectral applications. Its main goal is to determine the pixels in the image, so as to realize the automatic recognition of ground object categories and provide services for other departments. From classical machine learning theories such as SVMs, KNN, etc., to deep neural networks, deep learning model frameworks represented by convolutional neural networks have been widely used in hyperspectral image classification and have achieved remarkable achievements. Compared with traditional methods, convolutional neural networks do not require manual feature extraction, but automatically obtain high-level features of data in the training set through the combination of continuous convolutional blocks, enabling the classification model to better express the characteristics of the data set itself, avoiding the process of manually designing features through expert prior knowledge, and greatly improving the classification accuracy.
[0003] Due to the special data characteristics of hyperspectral images themselves, when using convolutional neural networks to build network models, the following several challenges are still faced:
[0004] (1) The spectral dimension of hyperspectral data has hundreds of band values, and the information between bands is redundant, resulting in high data dimensionality and thus rich spectral information. However, when the number of available labeled samples is limited, this high data dimensionality characteristic will bring the curse of dimensionality problem (usually also known as the Hughes phenomenon), that is, as the number of bands participating in the operation increases, the classification accuracy first increases and then decreases.
[0005] (2) The annotation cost of HSI samples is relatively high, resulting in insufficient labeled samples, which causes the model to overfit and leads to poor generalization performance.
[0006] (3) Due to the functions of convolutional kernels and pooling layers, as the network deepens continuously, inevitable problems such as feature loss and insufficient feature extraction will occur.
[0007] Generally speaking, the features extracted as the network depth increases gradually move from low-level features to high-level features. However, due to the objective limitations of convolutional neural networks, there will be a certain degree of feature loss problems, and it will also lead to problems such as difficult model training and vanishing gradients. Researchers have tried to use different methods to build efficient and high-precision models. Although some achievements have been made, there are still various defects. For example, in the hyperspectral image classification method based on Siamese network, using PCA to reduce the dimension of the original hyperspectral image will inevitably cause the loss of spatial-spectral information, resulting in low classification performance; in a hyperspectral image classification method and system based on 3D U-Net for spectral-spatial information fusion, as the network depth increases, the required training cost also increases continuously. And due to the objective limitations of convolution and pooling, inevitable feature loss and information loss occur, and its classification performance also decreases accordingly; in the hyperspectral image classification method combining spatial pyramid attention mechanism, PCA is still used to reduce the dimension of the original hyperspectral image, resulting in inevitable information loss. And a large number of pooling layers are used in the model for feature fusion, but the pooling layer will cause inevitable feature loss. At the same time, the model pays different attentions to features during training, resulting in important features that may be ignored. Therefore, the classification effect is not satisfactory.
[0008] Traditional residual connections or dense connections are used to alleviate the problems of feature loss and insufficient feature extraction ability, but the effect is still not good. The reason is that the model pays different attentions to features during training. To solve this problem, the attention mechanism has gradually become a hot topic in the application of hyperspectral image classification. Researchers have proposed some hyperspectral image classification models based on the attention mechanism, such as the dense convolutional neural network based on feedback attention and the dual-branch multi-attention mechanism network based on convolutional block attention module. Although these methods have good effects, they do not specifically solve the problems of feature loss and insufficient feature extraction. Especially in the case of limited samples, feature loss has a very large impact on the HIC effect, and satisfactory classification results cannot be obtained.
[0009] Based on this background, the present invention proposes a spectral-spatial joint hyperspectral image classification method based on deep learning. Summary of the Invention
[0010] The purpose of the present invention is to propose a hyperspectral image classification model based on deep learning with dual-branch multi-scale spectral-spatial feature fusion and self-attention mechanism for the situation of feature loss and insufficient feature extraction ability in the feature extraction process of hyperspectral image classification models based on traditional convolutional neural networks. Compared with the prior art, it can significantly improve the classification accuracy.
[0011] To achieve the above object, the present invention provides a hyperspectral image classification method based on deep learning, including:
[0012] Perform data preprocessing on the hyperspectral image data;
[0013] Utilize a convolutional neural network to establish an original hyperspectral image classification model based on dual-branch multi-scale spatial-spectral feature fusion and self-attention mechanism;
[0014] Based on the preprocessed hyperspectral image data, train the original hyperspectral image classification model to obtain a hyperspectral image classification model;
[0015] Utilize the hyperspectral image classification model to classify the pixel categories of the overall hyperspectral image.
[0016] Optionally, performing data preprocessing on the hyperspectral image data includes:
[0017] Convert the hyperspectral image data into a compact representation through attention coefficient weighting;
[0018] Slice the image data converted into a compact representation into several patches; wherein, each patch includes the spectral information of the pixel to be classified and also contains the spatial information of the pixel within a preset distance around it.
[0019] Optionally, converting the hyperspectral image data into a compact representation through attention coefficient weighting includes:
[0020] Obtain the Hellinger distance between the bands in the hyperspectral image data;
[0021] Based on the Hellinger distance, obtain the attention coefficients of the bands;
[0022] Convert the attention coefficients of all bands into a column vector representation;
[0023] Convert the column vector representation into the compact representation.
[0024] Optionally, the Hellinger distance is:
[0025]
[0026] wherein, b i and b j respectively represent the i-th band and the j-th band of the original hyperspectral image, b ik represents the k-th pixel in the i-th band, n represents the total number of pixels in the original hyperspectral image, and H(i, j) represents the Hellinger distance between b i and b j ;
[0027] The attention coefficient is:
[0028]
[0029] Among them, y i represents the attention coefficient of the i-th band, N represents the number of bands of the original hyperspectral image, m represents the m-th band, and H(i, m) represents the Hellinger distance between the i-th band and the m-th band;
[0030] The compact representation is:
[0031]
[0032] Among them, X represents the hyperspectral image data, represents element-wise multiplication, ba is represented as a column vector, ba = (y1, y2,..., y N ) T , y N represents, and T represents.
[0033] Optionally, the original hyperspectral image classification model includes: a dense double-branch pyramid feature extraction module, a spatial feature extractor, and a classifier;
[0034] The dense double-branch pyramid feature extraction module is used to extract multi-scale spatial-spectral feature information of the hyperspectral image data;
[0035] The spatial feature extractor is used to generate a final feature map according to the scale spatial-spectral feature information;
[0036] The classifier is used to classify the pixel categories of the hyperspectral image based on the final feature map.
[0037] Optionally, the dense double-branch pyramid feature extraction module includes: a spatial-spectral joint feature extraction branch and an inter-spectral feature extraction branch;
[0038] The spatial-spectral joint feature extraction branch is used to extract spatial and inter-spectral feature information of the hyperspectral image data;
[0039] The inter-spectral feature extraction branch is used to extract spectral feature information of the hyperspectral image data;
[0040] The dense double-branch pyramid feature extraction module performs channel-level fusion on the spatial and inter-spectral feature information and the spectral feature information to obtain the multi-scale spatial-spectral feature information.
[0041] Optionally, the spatial-spectral joint feature extraction branch includes 3D convolution blocks with successively dense-connected convolution kernel sizes of 7×3×3, 5×3×3, and 3×3×3.
[0042] Optionally, the inter-spectral feature extraction branch includes 3D convolution blocks with densely connected convolution kernels of sizes 7×1×1, 5×1×1, and 3×1×1 in sequence.
[0043] Optionally, the spatial feature extractor adopts a channel-spatial attention module;
[0044] The channel-spatial attention module includes: a channel attention module and a spatial attention module;
[0045] The channel attention module uses GAP to aggregate convolution features and one-dimensional convolution for local cross-channel interaction, and then passes through the Sigmoid activation function. The output feature map after passing through the Sigmoid activation function is multiplied by the input feature map to obtain the output feature map of the channel attention module. Among them, the output feature map after passing through the Sigmoid activation function refers to the feature map after passing through the overall channel-spatial attention module, and the input feature map refers to the original feature map input into the channel-spatial attention module:
[0046] The spatial attention module takes the output feature map of the channel attention module as the input. First, it performs channel-based global maximum pooling and global average pooling, connects the respective outputs in the channel dimension, then performs a convolution operation to reduce the number of channels and output a feature map. Finally, the output feature map with the reduced number of channels is multiplied by the feature map input into the spatial attention module to generate the final feature map.
[0047] Optionally, during the construction of the spatial feature extractor, a deep feature fusion strategy is adopted;
[0048] The deep feature fusion strategy is: upsample the high-level feature map generated by the overall spatial feature extractor to the same size as the input low-level feature map, and then perform channel-level fusion. The high-level feature map refers to the feature map generated after passing through the overall spatial feature extractor.
[0049] The present invention has the following beneficial effects:
[0050] In the data preprocessing stage, the hyperspectral image data is transformed into a compact representation through attention coefficient weighting to suppress the correlation between bands, thereby reducing information redundancy and improving data utilization; a dual-branch multi-scale spatio-spectral information feature fusion model that can perform high-efficiency spatio-spectral joint feature extraction and spatial information extraction is established, which can greatly alleviate the problems of feature loss and insufficient feature extraction ability; a new channel-spatial attention mechanism for enhancing the significant feature extraction ability of the model is proposed, which greatly improves the model's ability to extract significant features; a spatial information extraction module based on the deep feature fusion strategy is proposed, which fully enhances the feature reuse ability and alleviates the feature loss problem caused by deepening the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0052] Figure 1 It is a schematic flow chart of a hyperspectral image classification method based on deep learning according to an embodiment of the present invention;
[0053] Figure 2 It is a model block diagram of hyperspectral image classification based on dual-branch multi-scale spatio-spectral feature fusion and attention mechanism according to an embodiment of the present invention;
[0054] Figure 3 It is a schematic diagram of a dense dual-branch pyramid feature extraction module according to an embodiment of the present invention;
[0055] Figure 4 It is a schematic diagram of a channel-spatial attention module according to an embodiment of the present invention;
[0056] Figure 5 It is a pseudo-color schematic diagram of the Indian Pines dataset according to an embodiment of the present invention;
[0057] Figure 6 It is a schematic diagram of the true ground object labels of the Indian Pines dataset according to an embodiment of the present invention;
[0058] Figure 7 It is a schematic diagram of the classification effect drawn by different classification methods according to an embodiment of the present invention; among them, (a) is SVM, (b) is SSRN, (c) is A2S2K-ResNet, (d) is HyBridSN, (e) is SSDGL, (f) is RSSGL, (g) is LANet, and (h) is the classification model proposed in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.
[0060] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0061] As Figure 1 shown, this embodiment proposes a hyperspectral image classification method based on deep learning, including:
[0062] Perform data preprocessing on the hyperspectral image data;
[0063] Utilize a convolutional neural network to establish an original hyperspectral image classification model based on dual-branch multi-scale spatio-spectral feature fusion and an attention mechanism;
[0064] Based on the preprocessed hyperspectral image data, train the original hyperspectral image classification model to obtain a hyperspectral image classification model;
[0065] Use the hyperspectral image classification model to classify the pixel categories of the overall hyperspectral image.
[0066] Furthermore, the data preprocessing of the hyperspectral image data includes:
[0067] Convert the hyperspectral image data into a compact representation through attention coefficient weighting;
[0068] Divide the image data converted into a compact representation into several patches; wherein, each patch includes the spectral information of the pixel to be classified and also contains the spatial information of the pixel within a preset distance around it.
[0069] Furthermore, converting the hyperspectral image data into a compact representation through attention coefficient weighting includes:
[0070] Obtain the Hellinger distance between the bands in the hyperspectral image data;
[0071] Based on the Hellinger distance, obtain the attention coefficients of the bands;
[0072] Convert the attention coefficients of all bands into a column vector representation;
[0073] Convert the column vector representation into a compact representation.
[0074] Specifically, in this embodiment, perform data preprocessing on the hyperspectral image data, that is, reduce the information redundancy caused by the strong correlation between the bands of the original hyperspectral data and at the same time process the hyperspectral image into a size suitable for model operation, including steps such as a full-band attention module and slicing the hyperspectral three-dimensional data block. Subsequently, divide the preprocessed data into training set data and validation set data.
[0075] The spatial size of the original hyperspectral data cube is W×H, and the number of spectral bands is B, then it can be represented as C W×H×B , use the full-band attention module to reduce the redundant information brought by the high similarity between the bands. After passing through the attention module, the hyperspectral image can be denoted as U W×H×B , denote U W×H×BSplit into S patches of size c, where M is the predefined neighborhood size. The patch category is determined by the category of the central pixel. The patch contains not only the spectral information of the pixel to be classified but also the spatial information of the pixel within a certain distance around it. Each patch can be represented as P M×M×B , the model takes P M×M×B as input;
[0076] To solve the problem of information redundancy caused by strong inter-band correlation and alleviate the negative impact of redundant information on the training process, this embodiment proposes a full-band attention module. The idea is to measure the similarity between bands by calculating the Hellinger distance between bands. The Hellinger distance is a statistical distance metric used to measure the difference between two probability distributions. Based on the similarity, attention coefficients are inferred. The attention coefficients reflect the similarity of each band relative to other bands. By introducing these attention coefficients, the original hyperspectral image data is weighted, emphasizing the bands with lower similarity and suppressing the bands with higher similarity, thereby achieving the suppression of strong inter-band correlation. In other words, this embodiment transforms the original hyperspectral image data into a compact representation by weighting with attention coefficients. This compact representation helps to reduce information redundancy, improve the utilization efficiency of data, and at the same time, it is easier to capture the important features of the original hyperspectral image data during the training process.
[0077] Let b i , b j be the i-th band and the j-th band of the original hyperspectral image, and their Hellinger distance is denoted as H(i, j), which is given by Equation (1):
[0078]
[0079] where, b ik represents the k-th pixel in the i-th band, b jk represents the k-th pixel in the j-th band, and n represents the total number of pixels in the original hyperspectral image. The attention coefficient of the i-th band is denoted as y i , which is calculated by Equation (2):
[0080]
[0081] where, N represents the number of bands in the original hyperspectral image. In this embodiment, the output full-band attention coefficients are represented as a column vector ba = (y1, y2,..., y N ) T . Denote the original hyperspectral image data as X, then its i-th band is X i , and X is denoted as X = {X1, X2,..., X N}, and it is transformed into a compact representation through Equation (3).
[0082]
[0083] Among them, represents element multiplication.
[0084] Furthermore, the original hyperspectral image classification model includes: a dense double-branch pyramid feature extraction module, a spatial feature extractor, and a classifier;
[0085] The dense double-branch pyramid feature extraction module is used to extract multi-scale spatial-spectral feature information of hyperspectral image data;
[0086] The spatial feature extractor is used to generate the final feature map according to the scale spatial-spectral feature information;
[0087] The classifier is used to classify the pixel categories of the hyperspectral image based on the final feature map.
[0088] Furthermore, the dense double-branch pyramid feature extraction module includes: a spatial-spectral joint feature extraction branch and an inter-spectral feature extraction branch;
[0089] The spatial-spectral joint feature extraction branch is used to extract the spatial and inter-spectral feature information of hyperspectral image data;
[0090] The inter-spectral feature extraction branch is used to extract the spectral feature information of hyperspectral image data;
[0091] The dense double-branch pyramid feature extraction module performs channel-level fusion on the spatial and inter-spectral feature information and the spectral feature information to obtain multi-scale spatial-spectral feature information.
[0092] Specifically, in this embodiment, a convolutional neural network is used as the basic component of the model to establish a hyperspectral image classification model based on double-branch multi-scale spatial-spectral feature fusion and attention mechanism, as Figure 2 shown.
[0093] The hyperspectral image classification model based on double-branch multi-scale spatial-spectral feature fusion and attention mechanism includes: a dense double-branch pyramid feature extraction module, as Figure 3 shown, which includes a spatial-spectral joint feature extraction branch and an inter-spectral feature extraction branch. The spatial-spectral joint feature extraction branch is used to extract the spatial and inter-spectral information of the input hyperspectral data block, and it is composed of 3D convolutional blocks as the basic components. Combining with the dense connection technology, as shown in Equation (4), 3D convolutional blocks with convolutional kernel sizes of 7×3×3, 5×3×3, and 3×3×3 are used for multi-scale spatial-spectral information joint extraction, and its detailed architecture is shown in Table 1. Among them, the use of dense connection alleviates the problem of gradient disappearance, strengthens feature propagation, and greatly enhances feature reuse;
[0094] x i= f([x0, x1, …, x i-1 ) (4)
[0095] In Equation (4), x i represents the output of the i-th layer. The i-th layer receives the feature maps [x0, x1, …, x i-1 from the previous i - 1 layers. f(·) is defined as a composite function of three consecutive operations: batch normalization (BN), LeakyReLU, and 3D convolution.
[0096] Table 1 Structure Table of Spatial-Spectral Joint Feature Extraction Branch
[0097]
[0098] The spectral feature extraction branch is also composed of 3D convolution blocks with sizes of 7×1×1, 5×1×1, and 3×1×1 in sequence, combined with the dense connection technology, as shown in Table 2. Since the sizes of the feature maps output by each layer are different, the output feature maps of each layer are passed through max pooling to adapt to the size of the densely connected feature maps. In the process of spatial-spectral joint feature extraction, the dense double-branch pyramid feature extraction module uses dense pyramid convolution to extract multi-scale spatial-spectral feature information from the samples; in the spectral branch, the dense pyramid convolution layer is used to extract spectral features, and then the high-level feature maps generated by the two branches are fused at the channel level. After the feature maps are generated, Reshape is used to adjust the dimensions of the feature maps, and then they are fed into the subsequent spatial feature extractor.
[0099] Table 2 Structure Table of Spectral Feature Extraction Branch
[0100]
[0101] Furthermore, the spatial feature extractor adopts a channel-spatial attention module;
[0102] The channel-spatial attention module includes: a channel attention module and a spatial attention module;
[0103] The channel attention module uses GAP to aggregate convolutional features and one-dimensional convolution for local cross-channel interaction, and then passes through the SigMoid activation function. The output feature map after passing through the SigMoid activation function is multiplied by the input feature map to obtain the output feature map of the channel attention module. Among them, the output feature map after passing through the SigMoid activation function refers to the feature map after passing through the overall channel-spatial attention module, and the input feature map refers to the original feature map input into the channel-spatial attention module:
[0104] The spatial attention module takes the output feature map of the channel attention module as input. First, it performs channel-based global maximum pooling and global average pooling, concatenates the respective outputs along the channel dimension, then performs a convolution operation to reduce the number of channels and output a feature map. Finally, it multiplies the output feature map with the reduced number of channels by the feature map input into the spatial attention module to generate the final feature map.
[0105] The above-mentioned spatial feature extractor is composed of 3 2D convolution blocks as basic components, and the detailed settings are shown in Table 3. The kernel sizes of the convolutions are 7×7, 5×5, and 3×3 in sequence. During the construction of the spatial feature extractor, in order to enhance the model's ability to extract significant features, a brand-new channel-spatial attention module is adopted, as Figure 4 shown. The channel-spatial attention module abandons the process of reducing the dimension of the feature map in the traditional attention mechanism and designs the attention mechanism through local cross-channel interaction. The channel-spatial attention module consists of two independent sub-modules, namely the channel attention module and the spatial attention module. The channel attention module uses GAP to aggregate convolutional features, uses a one-dimensional convolution to achieve local cross-channel interaction, sets the kernel size of the convolution to 3, then passes through the SigMoid activation function, and finally multiplies the output feature map by the input feature map to obtain the output feature map of the channel attention module. Among them, the output feature map specifically refers to the feature map after passing through the overall channel-spatial attention module, and the input feature map refers to the original feature map input into the channel-spatial attention module. The spatial attention module takes the feature map generated by the channel attention module as the input feature map of the spatial attention module. First, it performs channel-based global maximum pooling and global average pooling, concatenates the respective outputs along the channel dimension, then performs a convolution operation with a kernel size of 7×7 to reduce the number of channels to 1, and finally multiplies the output feature map by the input feature map to generate the final feature map. The output feature map refers to the feature map after the convolution reduces the number of channels as mentioned above, and the input feature map refers to the output feature map of the channel attention module, that is, the feature map input into the spatial attention module. As shown in Eqs. (5) and (6).
[0106] M C (F) = σ(C1D3(GAP(F))) · F (5)
[0107] In Eq. (5), F is the input feature map, σ(·) is the Sigmoid activation function, C1D3 is a one-dimensional convolution block with a kernel size of 3, and GAP is global average pooling.
[0108] M S (F′) = σ(C2D 7×7 ([GAP(F′); GMP(F′)])) · F′ (6)
[0109] In Equation (6), F′ is the input feature map, GMP(·) is global max pooling, and C2D 7×7 is a 2D convolutional block with a convolutional kernel size of 7×7.
[0110] During the construction of the model's spatial feature extractor, a deep feature fusion strategy is adopted. The high-level feature map generated by the overall spatial feature extractor is upsampled to the same size as the input low-level feature map, and then channel-level fusion is performed. The high-level feature map refers to the feature map generated after passing through the overall spatial feature extractor. Specifically, it is the feature map generated by a three-layer 2D convolutional block, the corresponding batch normalization layer (BN), and the rectified linear unit (ReLU) after passing through the fused channel-spatial attention module; the low-level feature map refers to the input feature map of the spatial feature extractor.
[0111] Specifically, the high-level feature map generated by the spatial feature extractor is upsampled to the same size as the input low-level feature map, and then channel-level fusion is performed, which greatly alleviates the feature loss caused by convolution. Specifically: during the construction of the above-mentioned spatial feature extractor, the extracted high-level feature map is deconvolved to expand its size to the same size as the input low-level feature map, and then channel-level connection is performed, that is, the upsampled high-level feature map is fused with the low-level feature map. This fusion process not only makes full use of the rich semantic information of the high-level feature map but also combines the fine-grained features of the low-level feature map, thereby constructing a more comprehensive and detailed feature representation.
[0112] Table 3 Spatial Feature Extractor Architecture Table
[0113]
[0114]
[0115] In this embodiment, two fully connected layers and a Softmax classification loss function are used to build the classifier. Among them, the maximum number of parameters in the first fully connected layer is 1568, and the number of nodes in the last fully connected layer is the same as the number of categories in the corresponding given dataset, which is 16. In addition, in order to effectively deal with the model overfitting situation caused by a large number of model parameters and few training samples, Dropout is used to achieve a regularization effect to some extent and alleviate the occurrence of overfitting; ReLU is selected as the non-linear activation function; a BN layer is added after each layer of convolution to make the training easier and accelerate convergence.
[0116] The model is trained using the training set data to obtain a trained model.
[0117] The computing environment of the embodiments of the present invention runs on an Intel Core i7-13700KF processor and an NVIDIA GeForce RTX 4090 graphics card. The programming language used is Python. The classification network is built using PyTorch, with PyCharm as the compiler. The batch size is set to 32, Adam is used as the optimizer, and the learning rate is set to 0.001.
[0118] S4. Use the validation set to evaluate the model performance, automatically determine the pixel categories of the overall hyperspectral image, and complete the classification task.
[0119] Preferably, when using the validation set data to finally evaluate the model, performance evaluation metrics are used to evaluate the classification performance of the model. The performance evaluation metrics mainly include the overall classification accuracy (OA, Overall Accuracy) as shown in Equation (7), the average accuracy (AA, Average Accuracy) as shown in Equation (8), and the Kappa coefficient as shown in Equation (9). Among them, OA calculates the ratio of correctly classified samples to all samples in the test set. AA is used to determine the average classification accuracy of all classes, and it can also well evaluate the classification results of classes with fewer training samples. The Kappa coefficient calculated based on the confusion matrix represents the degree of agreement between the classification labels obtained by the classification model and the ground truth labels.
[0120]
[0121] To illustrate the effectiveness of this embodiment, the following experimental examples are disclosed:
[0122] First, the hyperspectral dataset in the experiment is: Indian Pines dataset, as Figure 5 and Figure 6 shown. This scene was collected by an airborne visible / infrared imaging spectrometer (AVIRIS) at the IP test site in the northwest of India. It contains 145×145 pixels and 224 spectral bands, with a wavelength range of 0.4 - 2.5 μm. Due to water absorption and damage, 24 of these spectral bands have been removed. The reference ground objects are designed as 16 vegetation categories, and not all vegetation categories are mutually exclusive (i.e., there are phenomena such as spectral mixing, different objects with the same spectrum, and the same object with different spectra), and the sample sizes of some categories are highly unbalanced. The spatial resolution of this dataset is approximately 20 meters per pixel.
[0123] As shown in Table 4, the OA of the hyperspectral image classification model proposed in the embodiment of the present invention is 99.76%, which is 6.91%, 5.41%, 7.58%, 2.92%, 3.39% and 4.36% higher than that of other deep learning-based methods respectively, showing very excellent classification performance. In addition, the classification accuracy of each category exceeds 98%, which is better than other deep learning-based classification models, such as Figure 7 shown, where Figure 7 (a) is SVM, Figure 7 (b) is SSRN, Figure 7 (c) is A2S2K-ResNet, Figure 7 (d) is HyBridSN, Figure 7 (e) is SSDGL, Figure 7 (f) is RSSGL, Figure 7 (g) is LANet, Figure 7 (h) is the classification algorithm proposed in this example.
[0124] Table 4 Classification accuracies of different methods on the dataset
[0125]
[0126] Therefore, this embodiment proposes a hyperspectral image classification model with a high-precision double-branch multi-scale feature aggregation and self-attention mechanism. The model achieves high-efficiency extraction of spatial-spectral joint feature information and spectral inter-feature information by using dense double-branch pyramid connections, greatly reducing feature loss and enhancing the model's context information extraction ability; a channel-spatial attention module ECBAM is proposed, which greatly improves the model's ability to extract significant features; a spatial information extraction module based on a deep feature fusion strategy is proposed, which fully enhances the feature reuse ability and alleviates the feature loss problem caused by deepening the model.
[0127] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A spatial-spectral combined hyperspectral image classification method based on deep learning, characterized in that: include: Perform data preprocessing on hyperspectral image data; Using convolutional neural networks, a raw hyperspectral image classification model based on dual-branch multi-scale spatial-spectral feature fusion and attention mechanism is established; Based on the preprocessed hyperspectral image data, the original hyperspectral image classification model is trained to obtain a hyperspectral image classification model; Using the hyperspectral image classification model, classifying the overall hyperspectral image pixel categories; Data preprocessing of hyperspectral image data includes: The hyperspectral image data is converted into a compact representation by weighting with attention coefficients; The image data converted into a compact representation is divided into a number of patches; each patch includes the spectral information of the pixel to be classified and also includes the spatial information of the pixel within a preset distance around it; The hyperspectral image data is converted into a compact representation weighted by the attention coefficient, including: Get the Hellinger distance between bands in hyperspectral image data; Based on the Hellinger distance, obtaining the attention coefficient of the band; Convert the attention coefficient of the whole band into a column vector representation; Convert the column vector representation into the compact representation; The original hyperspectral image classification model includes: a dense double-branch pyramid feature extraction module, a spatial feature extractor and a classifier; The dense dual-branch pyramid feature extraction module is used to extract multi-scale spatial-spectral feature information of hyperspectral image data; The spatial feature extractor is used to generate a final feature map according to the scale space-spectral feature information; The classifier is used to classify the hyperspectral image pixel categories based on the final feature map; The dense dual-branch pyramid feature extraction module includes: a space-spectrum joint feature extraction branch and an inter-spectrum feature extraction branch; The spatial-spectral joint feature extraction branch is used to extract spatial and inter-spectral feature information of the hyperspectral image data; The inter-spectral feature extraction branch is used to extract spectral feature information of the hyperspectral image data; The dense dual-branch pyramid feature extraction module performs channel-level fusion of spatial and inter-spectral feature information and spectral feature information to obtain the multi-scale spatial-spectral feature information; The spatial-spectral joint feature extraction branch includes densely connected 3D convolution blocks with convolution kernel sizes of 7×3×3, 5×3×3, and 3×3×3 respectively; The inter-spectral feature extraction branch includes densely connected 3D convolution blocks with convolution kernel sizes of 7×1×1, 5×1×1, and 3×1×1, respectively.
2. The spatial-spectral combined hyperspectral image classification method based on deep learning according to claim 1 is characterized in that: The Hellinger distance is: Among them, b i , b j They represent the i-th band and the j-th band of the original hyperspectral image, respectively, and b ik represents the kth pixel in the ith band, b jk represents the kth pixel in the jth band, n represents the total number of pixels in the original hyperspectral image, and H(i,j) represents b i , b j Hellinger distance between ; The attention coefficient is: Among them, y i represents the attention coefficient of the i-th band, N represents the number of bands of the original hyperspectral image, m represents the m-th band, and H(i,m) represents the Hellinger distance between the i-th band and the m-th band; The compact representation is: Where X represents the hyperspectral image data, represents element multiplication, ba is a column vector representation, ba=(y1,y2,…,y N ) T ,y N represents the attention coefficient of the Nth band, and T represents the transpose of the matrix.
3. The spatial-spectral combined hyperspectral image classification method based on deep learning according to claim 1 is characterized in that: The spatial feature extractor adopts a channel-spatial attention module; The channel-spatial attention module includes: a channel attention module and a spatial attention module; The channel attention module uses GAP aggregated convolution features and one-dimensional convolution for local cross-channel interaction, and then passes through the Sigmoid activation function to multiply the output feature map after the Sigmoid activation function with the input feature map to obtain the output feature map of the channel attention module, wherein the output feature map after the Sigmoid activation function refers to the feature map after the overall channel-spatial attention module, and the input feature map refers to the feature map of the original input into the channel-spatial attention module: The spatial attention module takes the feature map output by the channel attention module as input, first performs channel-based global maximum pooling and global mean pooling, connects the respective outputs in the channel dimension, then performs a convolution operation, outputs the feature map after reducing the number of channels, and finally multiplies the feature map output after reducing the number of channels with the feature map input into the spatial attention module to generate the final feature map.
4. The spatial-spectral combined hyperspectral image classification method based on deep learning according to claim 3 is characterized in that: In the process of constructing the spatial feature extractor, a deep feature fusion strategy is adopted; The deep feature fusion strategy is: upsampling the high-level feature map generated by the overall spatial feature extractor to the same size as the input low-level feature map, and then performing channel-level fusion. The high-level feature map refers to the feature map generated after passing through the overall spatial feature extractor.
Citation Information
Patent Citations
Monte carlo characteristics dimension reduction method for small-sample hyperspectral image
CN102663438A
Hyperspectral image classification method and system based on double-branch multi-attention convolutional neural network
CN113159189A