Hyperspectral Image Classification Method Based on Multi-Scale Dense Connection and Feature Aggregation Network

By constructing a multi-scale dense connection and feature aggregation network, the problems of insufficient spectral-spatial feature representation and unutilized multi-level features in hyperspectral image classification are solved, and higher feature expression capabilities and classification accuracy are achieved.

CN116580252BActive Publication Date: 2025-07-01XIDIAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310718315.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-16
Publication Date
2025-07-01
Estimated Expiration
2043-06-16

AI Technical Summary

Technical Problem

In the existing hyperspectral image classification methods, there are problems such as insufficient spectral-spatial feature representation, a large amount of redundant information, and the multi-level features being underutilized.

Method used

Build a multi-scale dense connection and feature aggregation network, including a spectral-spatial feature extraction module, a multi-scale feature extraction module and a multi-level feature aggregation module. The scale features of different targets are extracted through the hollow convolution layer, multi-scale branch extracts multi-scale spatial features, and aggregates different levels of features through cross-attention enhancement.

Benefits of technology

It effectively improves the model's feature expression ability and classification accuracy, and solves the problems of insufficient spectral-spatial feature representation, excessive redundant information, and unused multi-level features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116580252B_ABST
    Figure CN116580252B_ABST
Patent Text Reader

Abstract

The present invention proposes a hyperspectral image classification method based on a multi-scale dense connection and feature aggregation network, and its implementation steps are as follows: constructing a spectral-spatial feature extraction module; constructing a multi-scale feature extraction module; constructing a multi-level feature aggregation module; constructing a multi-scale dense connection and feature aggregation network; training the multi-scale dense connection and feature aggregation network using the generated training set; classifying the hyperspectral image. The present invention can more comprehensively capture the spectral-spatial features of hyperspectral images, and make full use of the spatial features and different-level features of hyperspectral images, thereby effectively solving problems such as insufficient spectral-spatial feature representation, sample imbalance, and a large amount of redundant information in existing hyperspectral image classification methods, and improving the performance of hyperspectral image classification to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and further relates to a hyperspectral image classification method based on a multi-scale dense connection and feature aggregation network in the technical field of hyperspectral image classification. By classifying the types of ground objects in hyperspectral images, the present invention can be applied to technical fields such as agriculture, forestry, geological exploration, and urban planning. Background Art

[0002] A hyperspectral image is a three-dimensional image obtained by simultaneously imaging ground object targets within an image spatial range at spectral bands of different wavelengths through a hyperspectral imaging instrument. Compared with ordinary color images, hyperspectral images can provide more discriminative information for more accurate ground object target recognition and classification. However, hyperspectral image classification technology also faces the challenge of how to improve classification accuracy under the condition of limited training samples.

[0003] Deep learning-based hyperspectral image classification methods can automatically learn features in images and improve classification accuracy by increasing the network depth. Therefore, models such as stacked autoencoder (SAE), deep belief network (DBN), convolutional neural network (CNN), and recurrent neural network (RNN) have been widely applied to hyperspectral image classification tasks. Among them, CNN achieves the best performance due to its excellent feature extraction ability. However, due to the high spatial and spectral resolutions of hyperspectral images, CNN still has problems of insufficient extraction of spatial and spectral features. How to fully extract spatial and spectral information in hyperspectral images is the key to improving classification accuracy.

[0004] 2D-CNN can extract the spatial features of hyperspectral images but cannot extract spectral features; while 3D-CNN can extract both spatial and spectral features, but has a large amount of computation. Combining 2D-CNN and 3D-CNN can further improve classification performance.

[0005] Roy et al. proposed a hyperspectral image classification method of a hybrid spectral convolutional network (HybridSN) that combines 3D-CNN and 2D-CNN in their published paper "HybridSN: Exploring 3-D–2-D CNN Feature Hierarchy for Hyperspectral Image Classification" (IEEE Transactions on Geoscience and Remote Sensing, 2020: 277-281). This method combines the advantages of two-dimensional and three-dimensional convolutions. In the designed network structure, three layers of three-dimensional convolutions are first used, then one layer of two-dimensional convolution is used, and finally two layers of fully connected layers and a Softmax layer are stacked. This not only gives play to the advantages of three-dimensional convolution and fully extracts spectral-spatial features, but also avoids the complexity of the model caused by using only three-dimensional convolution. However, the shortcoming of this method is that using only a single-scale convolution kernel can only capture features of a specific size, resulting in insufficient representation of spectral-spatial features and inability to adapt to other targets of different sizes, thus limiting the classification of targets of different sizes in hyperspectral images and affecting the classification accuracy of the entire image.

[0006] In addition, different spectral bands and spatial pixels contribute differently to the classification result. Highlighting the bands and pixels rich in effective information through the attention mechanism can significantly improve the overall accuracy of classification.

[0007] Xidian University proposed a hyperspectral image classification method based on the attention mechanism and weight sharing in its patent document "Hyperspectral Image Classification Method Based on Attention Mechanism and Weight Sharing" (Patent Application No.: ZL202110399194.7, Authorization Announcement No.: CN 113095409 B). The feature extraction network constructed by this method contains three feature extraction branches with shared weights and the same structure. The three branches respectively input neighborhood blocks of different scale sizes to extract multi-scale features. And by adding an attention mechanism to the feature extraction network, the model can pay more attention to important features. However, the shortcoming of this method is that it does not consider the multi-level features of CNN, resulting in a large amount of redundant information and the multi-level features not being fully utilized, thus limiting the full utilization of hyperspectral image features and affecting the average accuracy of the entire image classification. Summary of the Invention

[0008] The object of the present invention is to propose a hyperspectral image classification method based on a multi-scale dense connection and feature aggregation network aiming at the deficiencies of the above-mentioned existing technologies, aiming to solve the problems of insufficient representation of spectral-spatial features, a large amount of redundant information, and multi-level features not being fully utilized in the existing hyperspectral image classification methods.

[0009] To achieve the object of the present invention, the following technical ideas are proposed: The present invention constructs a multi-scale dense connection and feature aggregation network including a spectral-spatial feature extraction module, a multi-scale feature extraction module, and a multi-level feature aggregation module. Since the spectral-spatial feature extraction module can make full use of the dilated convolutional layer to extract the scale features of different targets in the hyperspectral image, the problem of insufficient spectral-spatial feature representation caused by single-scale convolution in the prior art is solved. Due to the multi-scale feature extraction module constructed by the present invention, multi-scale spatial features in the entire image are extracted using multi-scale branches, information flows between the shallow layer and the deep layer using residual branches, and cross-attention is used to enhance the feature fusion of the two branches, thereby solving the problems of insufficient spatial feature extraction and a large amount of redundant information in the prior art. Since the multi-level feature aggregation module constructed by the present invention can aggregate different-level features extracted by multiple multi-scale feature extraction modules, the problem of insufficient utilization of multi-level features in the prior art is solved.

[0010] To achieve the above object, the specific steps of the present invention are as follows:

[0011] Step 1, construct a spectral-spatial feature extraction module:

[0012] Build a spectral-spatial feature extraction module, the structure of which includes: a first convolutional layer, a first normalization layer, a first activation function layer, a second convolutional layer, a second normalization layer, a second activation function layer, a third convolutional layer, a third normalization layer, a third activation function layer, a concatenate layer, and a fourth convolutional layer; wherein, the first convolutional layer, the first normalization layer, the first activation function layer, the second convolutional layer, the second normalization layer, the second activation function layer, the third convolutional layer, the third normalization layer, the third activation function layer, the concatenate layer, and the fourth convolutional layer are cascaded; at the same time, the first activation function layer and the second activation function layer are also connected in parallel with the concatenate layer respectively; the convolutional kernel sizes of the first to third convolutional layers are all set to 3×3×3, the numbers of convolutional kernels are set to 8, 16, and 24 respectively, the convolutional kernel size of the fourth convolutional layer is set to 1×1×1, and the number of convolutional kernels is set to 16; the dilation rates of the first to third convolutional layers are set to 1, 2, and 3 respectively; the first to third normalization layers all adopt batch normalization functions, and the activation function layers all adopt rectified linear unit functions;

[0013] Step 2, construct a multi-scale feature extraction module:

[0014] Build a multi-scale feature extraction module, whose structure includes: a multi-scale branch, a residual branch, and a cross-attention layer; among them, the multi-scale branch and the residual branch are respectively connected to the cross-attention layer; the multi-scale branch includes: a first convolutional layer, a second convolutional layer, a second normalization layer, a second activation function layer, a first fusion layer, a third convolutional layer, a third normalization layer, a third activation function layer, a second fusion layer, a fourth convolutional layer, a fourth normalization layer, a fourth activation function layer, a third fusion layer, a fifth convolutional layer, a fifth normalization layer, a fifth activation function layer, a concatenate layer, and a sixth convolutional layer; among them, the first convolutional layer, the second convolutional layer, the second normalization layer, the second activation function layer, the first fusion layer, the third convolutional layer, the third normalization layer, the third activation function layer, the second fusion layer, the fourth convolutional layer, the fourth normalization layer, the fourth activation function layer, the third fusion layer, the fifth convolutional layer, the fifth normalization layer, the fifth activation function layer, the concatenate layer, and the sixth convolutional layer are cascaded; the first convolutional layer is connected in parallel with the first fusion layer, the second fusion layer, the third fusion layer, and the concatenate layer respectively, the second activation function layer is connected in parallel with the second fusion layer, the third fusion layer, and the concatenate layer respectively, the third activation function layer is connected in parallel with the third fusion layer and the concatenate layer respectively, and the fourth activation function layer is connected in parallel with the concatenate layer; set the convolutional kernel size of the first convolutional layer to 1×1, set the convolutional kernel sizes of the second to fifth convolutional layers to 3×3, the number of convolutional kernels to 16 in sequence, the dilation rates to 1, 2, 3, 4 respectively, and set the convolutional kernel size of the sixth convolutional layer to 1×1; the second to fifth normalization layers all use the batch normalization function, and the activation function layers all use the rectified linear unit function; the residual branch is cascaded by a seventh convolutional layer, a seventh normalization layer, and a seventh activation function layer in sequence; set the convolutional kernel size of the seventh convolutional layer to 3×3; the seventh normalization layer uses the batch normalization function, and the activation function layer uses the rectified linear unit function;

[0015] Step 3, construct a multi-level feature aggregation module:

[0016] Build a multi-level feature aggregation module, whose structure includes: input feature 1, input feature 2, input feature 3, input feature 4, the first fusion layer, the first convolutional layer, the second fusion layer, the second convolutional layer, the third fusion layer, the third convolutional layer, aggregated feature 1, aggregated feature 2, aggregated feature 3, aggregated feature 4, the concatenate layer, and the fourth convolutional layer; among them, input feature 1, the third fusion layer, the third convolutional layer, and aggregated feature 1 are cascaded, input feature 2, the second fusion layer, the second convolutional layer, and aggregated feature 2 are cascaded, input feature 3, the first fusion layer, the first convolutional layer, and aggregated feature 3 are cascaded, input feature 4 and aggregated feature 4 are cascaded, the concatenate layer is respectively connected to aggregated feature 1, aggregated feature 2, aggregated feature 3, and aggregated feature 4, the first fusion layer is connected to input feature 4, the second fusion layer is connected to the first convolutional layer, and the third fusion layer is connected to the second convolutional layer; set the convolutional kernel sizes of the first to fourth convolutional layers to 3×3;

[0017] Step 4, construct a multi-scale dense connection and feature aggregation network:

[0018] Build a multi-scale dense connection and feature aggregation network, whose structure includes: a spectral-spatial feature extraction module, a first multi-scale feature extraction module, a second multi-scale feature extraction module, a third multi-scale feature extraction module, a multi-level feature aggregation module, and a linear layer; among them, the spectral-spatial feature extraction module, the first multi-scale feature extraction module, the second multi-scale feature extraction module, the third multi-scale feature extraction module, the multi-level feature aggregation module, and the linear layer are cascaded; the multi-level feature aggregation module is respectively connected in parallel with the spectral-spatial feature extraction module, the first multi-scale feature extraction module, and the second multi-scale feature extraction module; the linear layer is successively cascaded by a global average pooling layer, a first fully connected layer, a first activation function layer, a second fully connected layer, and a second activation function layer; the first activation function layer uses a rectified linear unit function, and the second activation function uses a softmax function;

[0019] Step 5, generate a training set:

[0020] Select a hyperspectral remote sensing image with a size of 145×145×200 and a spatial resolution of 20m, which includes 200 available spectral bands, 16 types of land cover, and 10249 labeled samples; perform preprocessing on the hyperspectral remote sensing image in sequence, including mean-variance normalization, principal component analysis for dimensionality reduction, and zero-value edge padding operations; generate neighborhood blocks for the preprocessed hyperspectral image, and randomly sample all neighborhood blocks; form a training set with all the sampled neighborhood blocks and their corresponding labels;

[0021] Step 6, train the multi-scale dense connection and feature aggregation network:

[0022] Input the training set into the multi-scale dense connection and feature aggregation network, calculate the loss value between the predicted label vector and the true label vector using the cross-entropy loss function, and iteratively update the network parameters using the gradient descent method until the loss function of the network converges to obtain a trained multi-scale dense connection and feature aggregation network;

[0023] Step 7, classify the hyperspectral image:

[0024] Adopt the same method as in Step 5 to process the hyperspectral image to be classified; input the processed hyperspectral image into the trained multi-scale dense connection and feature aggregation network, and output the class labels of each pixel in the hyperspectral image to complete the classification of the hyperspectral image.

[0025] The present invention has the following advantages compared with the prior art:

[0026] First, the present invention constructs a spectral-spatial feature extraction module, overcomes the deficiency of insufficient spectral-spatial feature representation in the prior art, realizes the extraction of spectral-spatial features at different scales, enables the present invention to capture the spectral-spatial features of hyperspectral images more comprehensively, and thus effectively improves the feature expression ability of the model.

[0027] Second, the present invention constructs a multi-scale feature extraction module, overcomes the deficiencies of insufficient spatial feature extraction and a large amount of redundant information in the prior art, realizes multi-scale spatial feature extraction and cross-attention enhanced feature fusion, enables the present invention to make more full use of the spatial features of hyperspectral images, and thus effectively improves the classification accuracy of hyperspectral images.

[0028] Third, the present invention constructs a multi-level feature aggregation module, overcomes the deficiency of insufficient utilization of multi-level features in the prior art, realizes the aggregation of different-level features extracted by multiple multi-scale feature extraction modules, enables the present invention to make full use of the complementarity and correlation of different-level features, and thus further enhances the feature expression ability and classification performance of the model. Description of the Drawings

[0029] Figure 1 is the flowchart of the present invention;

[0030] Figure 2 is the structural schematic diagram of the spectral-spatial feature extraction module constructed by the present invention;

[0031] Figure 3 is the structural schematic diagram of the multi-scale feature extraction module constructed by the present invention;

[0032] Figure 4 is the structural schematic diagram of the multi-level feature aggregation module constructed by the present invention;

[0033] Figure 5 It is a schematic diagram of the multi-scale dense connection and feature aggregation network structure constructed by the present invention. Specific embodiments

[0034] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments.

[0035] Refer to Figure 1 for a further detailed description of the specific implementation steps of the present invention.

[0036] The multi-scale dense connection and feature aggregation network constructed by the present invention consists of a spectral-spatial feature extraction module, a first multi-scale feature extraction module, a second multi-scale feature extraction module, a third multi-scale feature extraction module, a multi-level feature aggregation module, and a linear layer; wherein, the multi-scale feature extraction module includes a multi-scale branch, a residual branch, and a cross-attention layer.

[0037] Step 1: Construct a spectral-spatial feature extraction module.

[0038] Refer to Figure 2 for a further description of the spectral-spatial feature extraction module constructed by the present invention.

[0039] Build a spectral-spatial feature extraction module, whose structure includes: a first convolutional layer, a first normalization layer, a first activation function layer, a second convolutional layer, a second normalization layer, a second activation function layer, a third convolutional layer, a third normalization layer, a third activation function layer, a concatenate layer, and a fourth convolutional layer; wherein, the first convolutional layer, the first normalization layer, the first activation function layer, the second convolutional layer, the second normalization layer, the second activation function layer, the third convolutional layer, the third normalization layer, the third activation function layer, the concatenate layer, and the fourth convolutional layer are cascaded; at the same time, the first activation function layer and the second activation function layer are also connected in parallel with the concatenate layer respectively; set the convolutional kernel sizes of the first to third convolutional layers to 3×3×3, and the numbers of convolutional kernels to 8, 16, 24 respectively, set the convolutional kernel size of the fourth convolutional layer to 1×1×1, and the number of convolutional kernels to 16, and set the dilation rates of the first to third convolutional layers to 1, 2, 3 respectively; the first to third normalization layers all adopt batch normalization functions, and the activation function layers all adopt rectified linear unit functions.

[0040] Step 2: Construct a multi-scale feature extraction module.

[0041] Refer to Figure 3 for a further description of the multi-scale feature extraction module constructed by the present invention.

[0042] Build a multi-scale feature extraction module, whose structure includes: a multi-scale branch, a residual branch, and a cross-attention layer; among them, the multi-scale branch and the residual branch are respectively connected to the cross-attention layer; the multi-scale branch includes: a first convolutional layer, a second convolutional layer, a second normalization layer, a second activation function layer, a first fusion layer, a third convolutional layer, a third normalization layer, a third activation function layer, a second fusion layer, a fourth convolutional layer, a fourth normalization layer, a fourth activation function layer, a third fusion layer, a fifth convolutional layer, a fifth normalization layer, a fifth activation function layer, a concatenate layer, and a sixth convolutional layer; among them, the first convolutional layer, the second convolutional layer, the second normalization layer, the second activation function layer, the first fusion layer, the third convolutional layer, the third normalization layer, the third activation function layer, the second fusion layer, the fourth convolutional layer, the fourth normalization layer, the fourth activation function layer, the third fusion layer, the fifth convolutional layer, the fifth normalization layer, the fifth activation function layer, the concatenate layer, and the sixth convolutional layer are cascaded; the first convolutional layer is connected in parallel with the first fusion layer, the second fusion layer, the third fusion layer, and the concatenate layer respectively, the second activation function layer is connected in parallel with the second fusion layer, the third fusion layer, and the concatenate layer respectively, the third activation function layer is connected in parallel with the third fusion layer and the concatenate layer respectively, and the fourth activation function layer is connected in parallel with the concatenate layer; set the convolution kernel size of the first convolutional layer to 1×1, set the convolution kernel sizes of the second to fifth convolutional layers to 3×3, set the number of convolution kernels to 16 in sequence, set the dilation rates to 1, 2, 3, 4 respectively, and set the convolution kernel size of the sixth convolutional layer to 1×1; the second to fifth normalization layers all use the batch normalization function, and the activation function layers all use the rectified linear unit function; the residual branch is cascaded by a seventh convolutional layer, a seventh normalization layer, and a seventh activation function layer in sequence; set the convolution kernel size of the seventh convolutional layer to 3×3; the seventh normalization layer uses the batch normalization function, and the activation function layer uses the rectified linear unit function.

[0043] Step 3, construct a multi-level feature aggregation module.

[0044] Refer to Figure 4 , and make a further description of the multi-level feature aggregation module constructed by the present invention.

[0045] Build a multi-level feature aggregation module, whose structure includes: input feature 1, input feature 2, input feature 3, input feature 4, first fusion layer, first convolutional layer, second fusion layer, second convolutional layer, third fusion layer, third convolutional layer, aggregated feature 1, aggregated feature 2, aggregated feature 3, aggregated feature 4, concatenate layer, fourth convolutional layer; among them, input feature 1, the third fusion layer, the third convolutional layer and aggregated feature 1 are cascaded, input feature 2, the second fusion layer, the second convolutional layer and aggregated feature 2 are cascaded, input feature 3, the first fusion layer, the first convolutional layer and aggregated feature 3 are cascaded, input feature 4 and aggregated feature 4 are cascaded, the concatenate layer is respectively connected to aggregated feature 1, aggregated feature 2, aggregated feature 3 and aggregated feature 4, the first fusion layer is connected to input feature 4, the second fusion layer is connected to the first convolutional layer, and the third fusion layer is connected to the second convolutional layer; the convolutional kernel sizes of the first to fourth convolutional layers are all set to 3×3.

[0046] Step 4, construct a multi-scale dense connection and feature aggregation network.

[0047] Refer to Figure 5 , and make a further description of the multi-scale dense connection and feature aggregation network constructed by the present invention.

[0048] Build a multi-scale dense connection and feature aggregation network, whose structure includes: spectral-spatial feature extraction module, first multi-scale feature extraction module, second multi-scale feature extraction module, third multi-scale feature extraction module, multi-level feature aggregation module and linear layer; among them, the spectral-spatial feature extraction module, the first multi-scale feature extraction module, the second multi-scale feature extraction module, the third multi-scale feature extraction module, the multi-level feature aggregation module and the linear layer are cascaded; the multi-level feature aggregation module is respectively connected in parallel with the spectral-spatial feature extraction module, the first multi-scale feature extraction module and the second multi-scale feature extraction module; the linear layer is successively cascaded by a global average pooling layer, a first fully connected layer, a first activation function layer, a second fully connected layer, and a second activation function layer; the first activation function layer uses a rectified linear unit function, and the second activation function uses a softmax function.

[0049] Step 5, generate a training set.

[0050] Step 5.1, Data preprocessing; In the embodiment of the present invention, a hyperspectral remote sensing image with a size of 145×145×200 and a spatial resolution of 20m is selected, which contains 200 available spectral bands, 16 types of land cover, and 10,249 labeled samples; the hyperspectral remote sensing image is preprocessed by performing mean-variance normalization, principal component analysis for dimensionality reduction, and zero-value edge padding operations in sequence; neighborhood blocks are generated for the preprocessed hyperspectral image, and all neighborhood blocks are randomly sampled; all the sampled neighborhood blocks and their corresponding labels are combined to form a training set;

[0051] For the mean-variance normalization, first calculate the mean and standard deviation of each band in the hyperspectral remote sensing image; then, after subtracting the mean of each band from all pixel values of each band, divide by the standard deviation of that band to complete the centering operation; next, calculate the maximum and minimum values of each band in the hyperspectral remote sensing image; finally, subtract the minimum value of that band from the value of that band at each pixel point, and then divide by the difference between the maximum value and the minimum value of that band to complete the standardization operation;

[0052] For the principal component analysis for dimensionality reduction, first calculate the covariance matrix of the hyperspectral remote sensing image after mean-variance normalization, with a size of b×b, where b is the number of bands of the hyperspectral data; then perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and eigenvectors; next, sort the eigenvalues and select the eigenvectors corresponding to the top k eigenvalues to form a k×b matrix, which represents the direction of the data in the new coordinate system after dimensionality reduction; finally, project the hyperspectral data after mean-variance normalization into the new coordinate system composed of k×b to obtain the hyperspectral image after dimensionality reduction.

[0053] For the zero-value edge padding operation, zero-value elements are padded at the edges of the hyperspectral image after dimensionality reduction in the height and width dimensions, and the number is half of the neighborhood block size.

[0054] Step 5.2, In the hyperspectral image after data preprocessing, take a region with a fixed size of 21×21 centered on each pixel point as a neighborhood block; then allocate the generated 10,249 neighborhood blocks to the corresponding 16 category sets according to the category of their central pixel points; finally, randomly sample the neighborhood blocks in each category set at a training ratio of 10%, and the 1,024 sampled neighborhood blocks and their corresponding labels are used as the training set.

[0055] Step 6, Train the multi-scale dense connection and feature aggregation network.

[0056] Input the training set into the multi-scale dense connection and feature aggregation network, use the cross-entropy loss function to calculate the loss value between the predicted label vector and the true label vector, and use the gradient descent method to iteratively update the network parameters until the loss function of the network converges to obtain the trained multi-scale dense connection and feature aggregation network.

[0057] The specific steps of the network training are as follows:

[0058] In the first step, randomly select N training samples from the training set as a batch, input these samples into the model for forward propagation calculation, and obtain a set of predicted label vectors y n represents the predicted label vector corresponding to the nth training sample.

[0059] In the second step, use the following cross-entropy function to calculate the loss value between the predicted label vector and the true label vector:

[0060]

[0061] where N represents the total number of samples in the training set, n represents the serial number of the sample in the training set, y n represents the true label of the nth sample, In(·) represents the logarithmic operation with the natural constant e as the base, y′ n represents the predicted label of the nth sample, and L represents the loss value between the predicted label and the true label. The smaller the loss value, the closer the predicted result of the model is to the true result, and the better the performance of the model. During training, the cross-entropy function is usually used as the loss function, and the gradient is calculated through backpropagation and the parameters of the model are updated to minimize the loss function, thereby improving the accuracy and generalization ability of the model.

[0062] In the third step, adopt the gradient descent method to update the parameters of the model through the loss value:

[0063]

[0064] where ω represents the weight parameter of the model, η represents the learning rate with a value of 0.001, and ω′ represents the weight parameter of the model after ω is iteratively updated.

[0065] Step 7, classify the hyperspectral image.

[0066] Adopt the same method as in Step 5 to perform the same processing on the hyperspectral image to be classified; input the processed hyperspectral image into the trained multi-scale dense connection and feature aggregation network, and output the class labels of each pixel in the hyperspectral image to complete the hyperspectral image classification.

[0067] The following further illustrates the effect of the present invention in combination with simulation experiments:

[0068] 1. Simulation experiment conditions:

[0069] The hardware platform for the simulation experiment of the present invention is as follows: the processor is an Inter Xeon Silver 4114 CPU with a main frequency of 2.2 GHz, the memory is 128 G, and the graphics card is a Geforce RTX 2080Ti 12G.

[0070] The software platform for the simulation experiment of the present invention is as follows: the operating system is Windows 10, the programming language is python 3.7, the programming software is Pycharm 2022, and the deep learning framework is Pytorch.

[0071] The hyperspectral image dataset used in the simulation experiment of the present invention is Indian Pines.

[0072] The Indian Pines dataset is a hyperspectral remote sensing image with a size of 145×145 and a spatial resolution of 20 m, which is acquired by an airborne visible infrared imaging spectrometer. It contains 200 available spectral bands and 16 types of land cover, with a total of 10,249 labeled samples. The category and quantity of each type of ground object are shown in Table 1.

[0073] Table 1 Indian Pines sample categories and quantities

[0074]

[0075]

[0076] 2. Content and result analysis of the simulation experiment:

[0077] In the simulation experiment of the present invention, the present invention and three existing technologies (support vector machine SVM classification method, spectral-spatial residual network SSRN classification method, attention mechanism and weight sharing network MNWA classification method) are respectively used to classify the input Indian Pines hyperspectral image to obtain a classification result map.

[0078] In the simulation experiment, the three existing technologies adopted refer to:

[0079] The existing technology support vector machine SVM classification method refers to the hyperspectral image classification method proposed by Melgani et al. in "Classification of hyperspectral remote sensing images with support vector machines, IEEE Trans. Geosci. Remote Sens., vol. 42, no. 8, pp. 1778–1790, Aug. 2004", abbreviated as the support vector machine SVM classification method.

[0080] The prior art spectral-spatial residual network SSRN classification method refers to the hyperspectral image classification method proposed by Zilong Zhong et al. in "Spectral-Spatial Residual Network for Hyperspectral Image Classification: A 3-D Deep Learning Framework, IEEE Transactions on Geoscience and Remote Sensing, 2017: 847-858", which is abbreviated as the spectral-spatial residual network SSRN classification method.

[0081] The prior art attention mechanism and weight sharing network MNWA classification method refers to the hyperspectral image classification method proposed by Xidian University in its patent document "Hyperspectral Image Classification Method Based on Attention Mechanism and Weight Sharing" (Patent No.: ZL202110399194.7, Authorization Publication No.: CN113095409B), which is abbreviated as the attention mechanism and weight sharing network MNWA classification method.

[0082] To evaluate the simulation effect of the present invention, the classification results of four methods are evaluated using three evaluation indicators (overall accuracy OA, average accuracy AA, Kappa coefficient). The larger the value of each indicator, the better the classification effect.

[0083] The overall accuracy (OA) refers to the accuracy rate of classifying all samples, that is, the ratio of the number of correctly classified samples to the total number of samples.

[0084] The average accuracy (AA) refers to the average value of the accuracy rates of the classification model for classifying each type of sample, which can reflect the classification performance between different categories.

[0085] Kappa coefficient: The value range of the Kappa coefficient is [-1, 1]. Among them, the Kappa coefficient equal to 1 indicates that the classification of the classification model is completely correct, equal to 0 indicates that the classification of the classification model is equivalent to random classification, and less than 0 indicates that the classification effect of the classification model is worse than random classification.

[0086] The classification effects of the present invention and three prior arts on the Indian Pines hyperspectral dataset are compared using the above evaluation indicators, and the results are shown in Table 2.

[0087] Table 2 Comparison results of the present invention and three prior arts in classification accuracy

[0088]

[0089] As can be seen from Table 2, the overall accuracy OA of the present invention is 99.13%, the average accuracy AA is 98.57%, and the Kappa coefficient is 0.9901. These three indicators are all higher than those of the three existing technical methods, proving that the present invention can achieve higher hyperspectral image classification accuracy.

[0090] The above simulation experiments show that: the hyperspectral image classification method proposed by the present invention can effectively solve the problems of single feature scale and unutilized features at different levels during the training of convolutional neural networks. Specifically, this method constructs a spectral-spatial feature extraction module, a multi-scale feature extraction module, and a multi-level feature aggregation module, enabling the network to better extract and utilize the feature information of hyperspectral images. The advantages of these modules are that they can simultaneously process features in the spectral and spatial dimensions, extract feature information at different scales, and aggregate multi-level feature information, thereby obtaining better feature representations. In addition, this method can also better solve the problem of poor hyperspectral image classification performance. The experimental results show that this method can classify more accurately and obtain higher classification accuracy.

Claims

1. A hyperspectral image classification method based on a multi-scale dense connection and feature aggregation network, characterized in that Construct a spectral-spatial feature extraction module, a multi-scale feature extraction module, and a multi-level feature aggregation module respectively; the specific steps of this classification method are as follows: Step 1, construct a spectral-spatial feature extraction module: Build a spectral-spatial feature extraction module, whose structure includes: a first convolutional layer, a first normalization layer, a first activation function layer, a second convolutional layer, a second normalization layer, a second activation function layer, a third convolutional layer, a third normalization layer, a third activation function layer, a concatenate layer, and a fourth convolutional layer; among them, the first convolutional layer, the first normalization layer, the first activation function layer, the second convolutional layer, the second normalization layer, the second activation function layer, the third convolutional layer, the third normalization layer, the third activation function layer, the concatenate layer, and the fourth convolutional layer are cascaded; at the same time, the first activation function layer and the second activation function layer are also connected in parallel with the concatenate layer respectively; set the convolutional kernel sizes of the first to third convolutional layers to 3×3×3, and the numbers of convolutional kernels to 8, 16, 24 respectively, set the convolutional kernel size of the fourth convolutional layer to 1×1×1, and the number of convolutional kernels to 16, and set the dilation rates of the first to third convolutional layers to 1, 2, 3 respectively; the first to third normalization layers all use batch normalization functions, and the activation function layers all use rectified linear unit functions; Step 2, construct a multi-scale feature extraction module: Build a multi-scale feature extraction module, whose structure includes: a multi-scale branch, a residual branch, and a cross-attention layer; among them, the multi-scale branch and the residual branch are respectively connected to the cross-attention layer; the multi-scale branch includes: a first convolutional layer, a second convolutional layer, a second normalization layer, a second activation function layer, a first fusion layer, a third convolutional layer, a third normalization layer, a third activation function layer, a second fusion layer, a fourth convolutional layer, a fourth normalization layer, a fourth activation function layer, a third fusion layer, a fifth convolutional layer, a fifth normalization layer, a fifth activation function layer, a concatenate layer, and a sixth convolutional layer; among them, the first convolutional layer, the second convolutional layer, the second normalization layer, the second activation function layer, the first fusion layer, the third convolutional layer, the third normalization layer, the third activation function layer, the second fusion layer, the fourth convolutional layer, the fourth normalization layer, the fourth activation function layer, the third fusion layer, the fifth convolutional layer, the fifth normalization layer, the fifth activation function layer, the concatenate layer, and the sixth convolutional layer are cascaded; the first convolutional layer is connected in parallel with the first fusion layer, the second fusion layer, the third fusion layer, and the concatenate layer respectively, the second activation function layer is connected in parallel with the second fusion layer, the third fusion layer, and the concatenate layer respectively, the third activation function layer is connected in parallel with the third fusion layer and the concatenate layer respectively, and the fourth activation function layer is connected in parallel with the concatenate layer; set the convolutional kernel size of the first convolutional layer to 1×1, set the convolutional kernel sizes of the second to fifth convolutional layers to 3×3, set the number of convolutional kernels to 16 in sequence, and set the dilation rates to 1, 2, 3, 4 respectively, and set the convolutional kernel size of the sixth convolutional layer to 1×1; the second to fifth normalization layers all use the batch normalization function, and the activation function layers all use the rectified linear unit function; the residual branch is cascaded by a seventh convolutional layer, a seventh normalization layer, and a seventh activation function layer in sequence; set the convolutional kernel size of the seventh convolutional layer to 3×3; the seventh normalization layer uses the batch normalization function, and the activation function layer uses the rectified linear unit function; Step 3, construct a multi-level feature aggregation module: Build a multi-level feature aggregation module, whose structure includes: input feature 1, input feature 2, input feature 3, input feature 4, a first fusion layer, a first convolutional layer, a second fusion layer, a second convolutional layer, a third fusion layer, a third convolutional layer, aggregated feature 1, aggregated feature 2, aggregated feature 3, aggregated feature 4, a concatenate layer, and a fourth convolutional layer; among them, input feature 1, the third fusion layer, the third convolutional layer, and aggregated feature 1 are cascaded, input feature 2, the second fusion layer, the second convolutional layer, and aggregated feature 2 are cascaded, input feature 3, the first fusion layer, the first convolutional layer, and aggregated feature 3 are cascaded, input feature 4 and aggregated feature 4 are cascaded, the concatenate layer is respectively connected to aggregated feature 1, aggregated feature 2, aggregated feature 3, and aggregated feature 4, the first fusion layer is connected to input feature 4, the second fusion layer is connected to the first convolutional layer, and the third fusion layer is connected to the second convolutional layer; set the convolutional kernel sizes of the first to fourth convolutional layers to 3×3; Step 4, construct a multi-scale dense connection and feature aggregation network: Build a multi-scale dense connection and feature aggregation network, whose structure includes: a spectral-spatial feature extraction module, a first multi-scale feature extraction module, a second multi-scale feature extraction module, a third multi-scale feature extraction module, a multi-level feature aggregation module, and a linear layer; among them, the spectral-spatial feature extraction module, the first multi-scale feature extraction module, the second multi-scale feature extraction module, the third multi-scale feature extraction module, the multi-level feature aggregation module, and the linear layer are cascaded; the multi-level feature aggregation module is connected in parallel with the spectral-spatial feature extraction module, the first multi-scale feature extraction module, and the second multi-scale feature extraction module respectively; the linear layer is cascaded in turn by a global average pooling layer, a first fully connected layer, a first activation function layer, a second fully connected layer, and a second activation function layer; the first activation function layer uses a rectified linear unit function, and the second activation function uses a softmax function; Step 5, generate a training set: Select a hyperspectral remote sensing image with a size of 145×145×200 and a spatial resolution of 20m, which includes 200 available spectral bands, 16 types of land cover, and 10,249 labeled samples; perform preprocessing on the hyperspectral remote sensing image in sequence, including mean-variance normalization, principal component analysis for dimensionality reduction, and zero-value edge padding operations; generate neighborhood blocks for the preprocessed hyperspectral image, and randomly sample all neighborhood blocks; form a training set with all the sampled neighborhood blocks and their corresponding labels; Step 6, train the multi-scale dense connection and feature aggregation network: Input the training set into the multi-scale dense connection and feature aggregation network, use the cross-entropy loss function to calculate the loss value between the predicted label vector and the true label vector, and use the gradient descent method to iteratively update the network parameters until the loss function of the network converges, obtaining a trained multi-scale dense connection and feature aggregation network; Step 7, classify the hyperspectral image: Adopt the same method as in Step 5 to process the hyperspectral image to be classified; input the processed hyperspectral image into the trained multi-scale dense connection and feature aggregation network, and output the class labels of each pixel in the hyperspectral image to complete the classification of the hyperspectral image.

2. The hyperspectral image classification method based on multi-scale dense connection and feature aggregation according to claim 1, wherein The mean-variance normalization described in Step 5 refers to calculating the mean and standard deviation of each band in the hyperspectral remote sensing image; after subtracting the mean of each band from all pixel values of each band, divide by the standard deviation of that band to complete the centering operation; calculate the maximum and minimum values of each band in the hyperspectral remote sensing image; after subtracting the minimum value of that band from the value of that band at each pixel point, divide by the difference between the maximum value and the minimum value of that band to complete the standardization operation.

3. The hyperspectral image classification method based on multi-scale dense connection and feature aggregation according to claim 1, characterized in that, The principal component analysis dimensionality reduction described in step 5 refers to calculating the covariance matrix of the hyperspectral remote sensing image after mean-variance normalization, with a size of b×b, where b is the number of bands of the hyperspectral data; performing eigenvalue decomposition on the covariance matrix to obtain eigenvalues and eigenvectors; sorting the eigenvalues and selecting the eigenvectors corresponding to the top k eigenvalues to form a k×b matrix, which represents the direction of the data in the new coordinate system after dimensionality reduction; projecting the hyperspectral data after mean-variance normalization into the new coordinate system composed of k×b to obtain the hyperspectral image after dimensionality reduction.

4. The hyperspectral image classification method based on multi-scale dense connection and feature aggregation according to claim 1, characterized in that, The zero-value edge padding operation described in step 5 refers to padding zero-value elements at the edges of the hyperspectral image after dimensionality reduction in the height and width dimensions, and the number of such elements is half of the neighborhood block size.

5. The hyperspectral image classification method based on multi-scale dense connection and feature aggregation according to claim 1, wherein Generating neighborhood blocks for the preprocessed hyperspectral image described in step 5 refers to taking a fixed-size 21×21 area centered on each pixel point in the preprocessed image as the neighborhood block.

6. The hyperspectral image classification method based on multi-scale dense connection and feature aggregation according to claim 1, wherein Randomly sampling all neighborhood blocks described in step 5 refers to allocating all the generated neighborhood blocks to the corresponding 16 category sets according to the category of their central pixel points, randomly sampling the neighborhood blocks in each category set at a training ratio of 10%, and taking all the sampled neighborhood blocks and their corresponding labels as the training set; among them, the label of each neighborhood block is the label of its central pixel point.

7. The hyperspectral image classification method based on multi-scale dense connection and feature aggregation according to claim 1, characterized in that The cross-entropy loss function described in step 6 is as follows: Among them, N represents the total number of samples in the training set, n represents the serial number of the samples in the training set, and y n represents the true label of the nth sample, In(·) represents the logarithmic operation with the natural constant e as the base, and y n ' represents the predicted label of the nth sample, and L represents the loss value between the predicted label and the true label.

8. The hyperspectral image classification method based on multi-scale dense connection and feature aggregation according to claim 1, wherein The gradient descent method described in step 6 is as follows: Among them, ω represents the weight parameter of the model, η represents the learning rate with a value of 0.001, and ω′ represents the weight parameter of the model after ω is iteratively updated.

Citation Information

Patent Citations

  • Hyperspectral image classification method based on attention mechanism and weight sharing

    CN113095409B

  • Hyperspectral image classification method based on fusion of multi-scale and multi-dimensional spatial-spectral characteristics

    CN110321963A

  • Hyperspectral image classification method based on depth feature cross fusion

    CN111191736A