Small-Sample Hyperspectral Image Classification Method Based on 3D Deep Convolutional Neural Network
By constructing a 3D CNN model, combining multi-scale convolutional layer and combined null spectrum information processing, the problems of data redundancy and spatial structure information neglect in hyperspectral image classification are solved, and high-precision small-sample hyperspectral image classification is achieved.
Patent Information
- Application Number
- CN202210778135.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-29
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-06-29
AI Technical Summary
The prior art has high data redundancy and strong correlation between bands in hyperspectral image classification, resulting in low classification accuracy. Traditional methods ignore spatial structure information, making it difficult to effectively utilize the unique characteristics of hyperspectral data.
The small sample hyperspectral image classification method based on 3D deep convolutional neural network is adopted to build a 3D CNN model, combining multi-scale convolutional layers and combined null spectrum information processing, spatial and spectral features in hyperspectral data are extracted, and the adaptive learning rate optimization algorithm Adagrad is used for training.
The accuracy of hyperspectral image classification is improved, especially in the case of small samples, the identification ability of inconcentrated sample distribution or individual geographic categories is overcome, and the shortcomings of traditional methods are achieved, and higher classification accuracy and recognition ability are achieved.
Smart Images

Figure CN115147742B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image processing, and particularly relates to a small-sample hyperspectral image classification method based on a 3D deep convolutional neural network. Background Technique
[0002] Remote sensing plays an important role in providing rich information sources for various applications. Hyperspectral images are obtained through remote sensing satellites, which contain image data with high spectral resolution of continuous and narrow bands in the infrared and thermal infrared band ranges, with nanometer-level spectral resolution and up to hundreds of spectral bands. They can continuously provide various information, such as radiation, space, and characteristic spectra, and are widely used in military reconnaissance, environmental monitoring, vegetation surveys, geological exploration, medical applications, and deep space exploration. One of the key applications is the fine classification of ground objects in hyperspectral images with a small number of samples. However, due to the data characteristics of hyperspectral images, images of multiple bands are generated, with the characteristic of data redundancy, making the analysis of the fine classification problem somewhat challenging.
[0003] With the implementation of China's high-resolution earth observation system, especially the upcoming launch of the GF-5 satellite equipped with a nanometer-level hyperspectral camera, it can be seen that there are significant strategic application demands for hyperspectral remote sensing. One of the key applications is the fine classification of ground objects in hyperspectral images with a small number of samples. Although the spatial information and spectral information contained in hyperspectral images are very conducive to ground object classification, in the actual application process, due to the relatively high spectral dimension of hyperspectral image data, the correlation between bands is increased, making it difficult to distinguish between bands, and the information redundancy of hyperspectral data is relatively high. These problems cannot be solved by traditional image classification methods. In view of this, researchers have proposed pixel-based classification methods and proved through experiments that they have better classification effects than traditional methods and have been recognized by the industry for a long time. However, the pixel-based classification method only classifies based on the spectral information of the image, ignoring the spatial structure information of the image, resulting in waste of information and reduced classification accuracy. Therefore, it is very meaningful to study hyperspectral image classification methods based on the combination of spatial and spectral information.
[0004] With the rise of big data and artificial intelligence technologies, deep learning technologies have been widely applied in fields such as target detection and image recognition. Traditional hyperspectral image classification algorithms require manual annotation and expert-level derivation work. Using deep learning can better help workers with classification and speed up the training process. The 1D and 2D networks in deep learning make classification easier, but their accuracy needs to be improved. In addition, the processing speed is slow, wasting the unique data characteristics of hyperspectral images. 3D networks can make full use of the information in the data and simplify the model architecture.
[0005] Combined with the existing specific technologies and methods of hyperspectral classification, and aiming at the unique characteristics of hyperspectral data, a small-sample hyperspectral image classification method based on 3D deep convolutional neural network is proposed. This algorithm can not only improve the classification accuracy, but also play a key role in the further development of 3D CNN. Summary of the Invention
[0006] The object of the present invention is to provide a small-sample hyperspectral image classification method based on 3D deep convolutional neural network. By constructing a deep learning model of 3D CNN to learn the internal information structure of the data, the pixel points in the hyperspectral data can be classified.
[0007] The technical solution adopted by the present invention is a small-sample hyperspectral image classification method based on 3D deep convolutional neural network, which is specifically implemented according to the following steps:
[0008] Step 1: Input the hyperspectral data as a whole, and divide the data into a training set train and a test set test;
[0009] Step 2: Construct a deep learning network model of 3D CNN, determine the size of each layer's convolutional kernel, the number of convolutional kernels, the number of fully connected layers, and determine the loss function and its hyperparameters;
[0010] Step 3: Input the training set train of the hyperspectral data in Step 1 into the deep learning network model of 3D CNN constructed in Step 2 to achieve feature extraction;
[0011] Step 4: Input the extracted features into the classifier to learn the features, train the parameters, and determine the network model;
[0012] Step 5: Test the test samples to achieve the classification of each pixel point of the hyperspectral data.
[0013] The characteristics of the present invention also lie in that
[0014] Step 1 is specifically implemented according to the following steps:
[0015] Step 1.1: Process the hyperspectral data as 3D cube data, M represents that there are M pixel points in the hyperspectral data, M = m1×m2, where m1 and m2 respectively represent the length and width of the input data. Each pixel point is formed by the action of N spectral bands, N = {1, 2,..., n}. Therefore, the data set is also expressed as represents the hyperspectral image, with the length and width of X being m1 and m2, and a total of N bands. n represents the nth band;
[0016] Step 1.2: For each pixel M in the training set train, consider its 7*7 neighborhood centered on it, and use this neighborhood as the input for each pixel, and input it into the 3D CNN network model in Step 2.
[0017] Step 2: Construct a 3D CNN network structure model, specifically as follows:
[0018] The 3D CNN network model has a total of 6 layers, including three multi-scale convolutional layers, one convolutional layer for jointly processing spatial and spectral information, and two fully connected layers;
[0019] Step 2.1: For the first multi-scale convolutional layer, use small convolutional kernels in the spatial dimension and large convolutional kernels in the spectral dimension. The kernel sizes are 11*3*3, 7*2*2, and 5*1*1 respectively, and the corresponding convolutional strides stride are (1,1,1), (1,2,2), (1,1,1), and the convolutional padding blocks padding are (5,1,1), (3,3,3), (2,0,0);
[0020] Step 2.2: For the second multi-scale convolutional layer, only process the spatial information and use a 1*n1*n1 convolutional kernel, where n1 takes values of: 1, 2, 3, that is, the convolutional kernels are 1*3*3, 1*2*2, 1*1*1 respectively, and the corresponding convolutional strides stride are (1,1,1), (1,2,2), (1,1,1), and the convolutional padding blocks padding are (0,1,1), (0,3,3), (0,0,0);
[0021] Step 2.3: For the third multi-scale convolutional layer, process its spectral information and use an n2*1*1 convolutional kernel, where the three values of n2 are: 11, 7, 5; the convolutional kernels are 11*1*1, 7*1*1, 5*1*1 respectively, and the corresponding convolutional strides stride are (1,1,1), (1,1,1), (1,1,1), and the convolutional padding blocks padding are (5,0,0), (3,0,0), (5,0,0);
[0022] Step 2.4: Set up a convolutional layer for jointly processing spatial and spectral information. The size of the convolutional layer is: (3,2,2), the convolutional stride stride is (1,1,1), the convolutional padding block padding is (1,0,0), and then connect it to the pooling layer for dimensionality reduction. The pooling layer convolutional kernel is (3,2,2), the stride convolutional stride is (1,1,1), and the convolutional padding block padding is (0,0,0);
[0023] Step 2.5: Design two fully connected layers. The first fully connected layer FC1 first maps the feature S4 obtained in Step 2.4 to 1024 nodes to obtain feature S5. The second fully connected layer FC2 maps the feature S5 among the 1024 nodes to the category Label owned by the hyperspectral dataset in Step 1.
[0024] The specific structure of Step 3 is as follows:
[0025] Step 3.1: Input the training set train in Step 1 into the first multi-scale convolutional layer in Step 2.1 of the 3D CNN deep learning network model constructed in Step 2 to obtain features S11, S12, and S13. Add the features S11, S12, and S13 and then output through the RELU activation function to obtain feature S1, which is used as the input of the second multi-scale convolutional layer;
[0026] Step 3.2: After passing feature S1 through the second multi-scale convolutional layer in Step 2.2, obtain features S21, S22, and S23. Add the extracted features S21, S22, and S23 and then output through the RELU activation function to obtain feature S2, which is used as the input of the third multi-scale convolutional layer;
[0027] Step 3.3: After passing feature S2 through the third multi-scale convolutional layer in Step 2.3, obtain features S31, S32, and S33. Add the extracted features S31, S32, and S33 and then output through the RELU activation function to obtain feature S3;
[0028] Step 3.4: Input feature S3 into the convolutional layer for jointly processing spatial and spectral information in Step 2.4 and pass through the RELU activation function to obtain feature S41. Subsequently, perform a pooling operation on feature S41 to obtain feature S42, and then pass through the RELU activation function again to obtain feature S4.
[0029] The specific steps of Step 4 are as follows:
[0030] Input the feature S4 obtained in Step 3.4 into the first fully connected layer FC1 in Step 2.4. After the action of the activation function, input it into the second fully connected layer FC2 in Step 2.4. Finally, pull feature S5 into a one-dimensional vector corresponding to the number of categories and train in the 3D CNN deep learning network model. During the training process, use the adaptive learning rate optimization algorithm Adagrad algorithm to find the optimal parameters and save the weights of the model. The algorithm process is as follows:
[0031] First, it is necessary to set the global learning rate ε, the initial parameter θ, and the small constant δ;
[0032] Subsequently, initialize the gradient accumulation variable r = 0;
[0033] When the stopping criterion is not met:
[0034] Sample m samples {x (1) , …, x (m)} from the training samples, with the corresponding target being y (i) , representing the predicted output of the i-th sample;
[0035] Calculate the gradient g:
[0036] Accumulate the squared gradient r: r = r + g ⊙ g;
[0037] Calculate the updated parameter θ:
[0038] Apply the updated parameter θ: θ = θ + Δθ;
[0039] When updating the parameter, the learning rate becomes:
[0040]
[0041] ξ is for maintaining numerical stability to prevent the denominator from being zero,
[0042] Finally, a set of parameters corresponding to the 3D CNN deep learning model is obtained, and thus the 3D CNN deep learning network model is built.
[0043] The specific steps of Step 5 are as follows:
[0044] Step 3 and Step 4 are the processes of training the network, and this process is the testing process. Each pixel point of the test set test, considering the size of its surrounding 7*7 neighborhood, is input into the built 3D CNN deep learning network model, and finally the class label of the test sample is output, completing the classification of ground objects in the hyperspectral image.
[0045] The beneficial effects of the present invention are as follows. A small-sample hyperspectral image classification method based on a 3D deep convolutional neural network: 1) The present invention builds three different multi-scale convolutional layers for extracting different scale features in the hyperspectral image, overcoming the disadvantages of insufficient utilization of scale features and low classification accuracy by a single scale, and improving the classification accuracy of ground objects in the hyperspectral image. 2) The present invention designs a convolutional layer for jointly processing spatial and spectral information, performing convolution on both spatial and spectral information simultaneously, overcoming the deficiencies of insufficient fusion of spatial and spectral features in the prior art, and having poor classification effects for samples with non-concentrated distributions or very few samples of individual ground object classes, and improving the recognition ability for small-sample classes. 3) In the present invention, the non-linear activation function RELU is used after convolution, and strided convolution is used to replace the traditional convolution, so that the feature integration ability is more suitable for the non-linear characteristics of hyperspectral data than general linear means. Description of the Drawings
[0046] Figure 1 is the overall flowchart of the small-sample hyperspectral image classification method based on 3D deep convolutional neural network of the present invention;
[0047] Figure 2 is the structural diagram of the deep learning 3D CNN network of the present invention;
[0048] Figure 3 is the result diagram of comparison on the PaviaU dataset by selecting 5% training samples;
[0049] Figure 4 is the result diagram of comparison on the KSC dataset by selecting 5% training samples;
[0050] Figure 5 is the result diagram of comparison on the IndianPines dataset by selecting 20% training samples;
[0051] Figure 6 is the result diagram of comparison on the Salinas dataset by selecting 20% training samples. Detailed Implementation Manner
[0052] The present invention will be described in detail below in conjunction with the drawings and specific implementation manners.
[0053] The present invention proposes a small-sample hyperspectral image classification method based on 3D deep convolutional neural network: First, input the hyperspectral data. For each pixel point, consider its surrounding neighborhood and input each point as a 3D block. Subsequently, through the framework proposed by the algorithm of this article, extract joint spatial-spectral features, train the network model to reduce the loss error, and finally input the test samples to complete the prediction of the classification results, and output the Kappa coefficient and AA index of the prediction results.
[0054] For the small-sample hyperspectral image classification method based on 3D deep convolutional neural network, the overall network structure of the algorithm is as Figure 1 shown, and it is specifically implemented according to the following steps:
[0055] Step 1: Input the hyperspectral data as a whole without performing preprocessing of dimensionality reduction, and divide the data into a training set train and a test set test;
[0056] Step 1 is specifically implemented according to the following steps:
[0057] Step 1.1: Process the hyperspectral data as 3D cube data, Let \(M\) denote the total number of pixels in the hyperspectral data, where \(M = m_1\times m_2\), and \(m_1\) and \(m_2\) represent the length and width of the input data respectively. Each pixel is formed by the action of \(N\) spectral bands, where \(N=\{1,2,\cdots,n\}\). Thus, the dataset can also be represented as denotes the hyperspectral image, with length \(m_1\), width \(m_2\), and a total of \(N\) bands. \(n\) represents the \(n\)-th band;
[0058] Step 1.2: For each pixel \(M\) in the training set \(train\), consider its \(7\times7\) neighborhood centered around it, and use this neighborhood as the input for each pixel, and input it into the 3D CNN network model in Step 2.
[0059] Step 2: Construct a deep learning network model of 3D CNN, determine the size of the convolutional kernels, the number of convolutional kernels, the number of fully connected layers in each layer, and determine the loss function and its hyperparameters;
[0060] Step 2 constructs a 3D CNN network structure model, which is specifically as follows:
[0061] Determine the size of the convolutional kernels, the number of convolutional kernels, the number of fully connected layers in each layer, and determine the loss function and its hyperparameters. The structure diagram of the network is as Figure 2 shown.
[0062] The 3D CNN network model of the present invention has a total of 6 layers, including three multi-scale convolutional layers, one convolutional layer for jointly processing spatial and spectral information, and two fully connected layers;
[0063] Step 2.1: For the first multi-scale convolutional layer, use small convolutional kernels in the spatial dimension and large convolutional kernels in the spectral dimension. The kernel sizes are \(11\times3\times3\), \(7\times2\times2\), \(5\times1\times1\) respectively, and the corresponding convolutional strides \(stride\) are \((1,1,1)\), \((1,2,2)\), \((1,1,1)\) respectively, and the convolutional padding blocks \(padding\) are \((5,1,1)\), \((3,3,3)\), \((2,0,0)\);
[0064] Step 2.2: For the second multi-scale convolutional layer, only process the spatial information, and use convolutional kernels of \(1\times n_1\times n_1\), where \(n_1\) takes values of: 1, 2, 3, that is, the convolutional kernels are \(1\times3\times3\), \(1\times2\times2\), \(1\times1\times1\) respectively, and the corresponding convolutional strides \(stride\) are \((1,1,1)\), \((1,2,2)\), \((1,1,1)\) respectively, and the convolutional padding blocks \(padding\) are \((0,1,1)\), \((0,3,3)\), \((0,0,0)\);
[0065] Step 2.3: For the third multi-scale convolutional layer, process its spectral information using a convolutional kernel of n2*1*1, where the three values of n2 are: 11, 7, 5; the convolutional kernels are 11*1*1, 7*1*1, 5*1*1 respectively, and the corresponding convolutional strides stride are (1,1,1), (1,1,1), (1,1,1), and the convolutional padding blocks padding are (5,0,0), (3,0,0), (5,0,0);
[0066] Step 2.4: Set up a convolutional layer for jointly processing the spatial spectral information. The size of the convolutional layer is: (3,2,2), the convolutional stride stride is (1,1,1), and the convolutional padding block padding is (1,0,0). Then connect it to the pooling layer for dimensionality reduction. The convolutional kernel of the pooling layer is (3,2,2), the stride of the stride convolution is (1,1,1), and the convolutional padding block padding is (0,0,0);
[0067] Step 2.5: Design two fully connected layers: The first fully connected layer FC1 first maps the feature S4 obtained in Step 2.4 to 1024 nodes to obtain the feature S5. The second fully connected layer FC2 maps the feature S5 in the 1024 nodes to the category Label of the hyperspectral dataset in Step 1. For example, if the Pavia University (PaviaU) dataset is used in the experiment, this dataset is obtained by the ROSIS satellite, with a total of 610*340 pixels. After removing the bands affected by the atmosphere and water absorption, there are 103 spectral bands left, with a ground resolution of 1.3m, and it is divided into Asphalt, Meadows, Gravel, Trees, Painted metal sheets, Bare Soil, Bitumen, Self-Blocking Bricks, Shadows, a total of 9 categories. In this process, the feature S5 is mapped to 9 categories. The other datasets and their category Labels in the experiment can be seen in Table 1.
[0068] Step 3: Input the training set train of the hyperspectral data in Step 1 into the 3D CNN deep learning network model constructed in Step 2 to achieve feature extraction;
[0069] The specific structure of Step 3 is as follows:
[0070] Step 3.1: Input the training set train in Step 1 into the first multi-scale convolutional layer in Step 2.1 of the 3D CNN deep learning network model constructed in Step 2 to obtain features S11, S12, and S13. Add the features S11, S12, and S13 and then output through the RELU activation function to obtain feature S1, which serves as the input for the second multi-scale convolutional layer;
[0071] Step 3.2: After passing feature S1 through the second multi-scale convolutional layer in Step 2.2, obtain features S21, S22, and S23. Add the extracted features S21, S22, and S23 and then output through the RELU activation function to obtain feature S2, which serves as the input for the third multi-scale convolutional layer;
[0072] Step 3.3: After passing feature S2 through the third multi-scale convolutional layer in Step 2.3, obtain features S31, S32, and S33. Add the extracted features S31, S32, and S33 and then output through the RELU activation function to obtain feature S3;
[0073] Step 3.4: Input feature S3 into the convolutional layer for jointly processing spatial and spectral information in Step 2.4 and pass through the RELU activation function to obtain feature S41. Subsequently, perform a pooling operation on feature S41 to obtain feature S42, and then pass through the RELU activation function again to obtain feature S4.
[0074] Step 4: Input the extracted features into a classifier to learn the features, train the parameters, and determine the network model;
[0075] The specific steps of Step 4 are as follows:
[0076] Input feature S4 obtained in Step 3.4 into the first fully connected layer FC1 in Step 2.4. After the action of the activation function, input it into the second fully connected layer FC2 in Step 2.4. Finally, pull feature S5 into a one-dimensional vector corresponding to the number of categories and train in the 3D CNN deep learning network model. During the training process, use the adaptive learning rate optimization algorithm Adagrad algorithm to find the optimal parameters, save the weights of the model, and Adagrad performs parameter optimization, which is improved from SGD (Stochastic Gradient Descent). The algorithm process is as follows:
[0077] First, it is necessary to set the global learning rate ε, the initial parameter θ, and the small constant δ;
[0078] Subsequently, initialize the gradient accumulation variable r = 0;
[0079] When the stopping criterion is not reached:
[0080] Collect m samples {x (1) ,…,x (m)} in small batches, with the target being y (i) , representing the predicted output of the i-th sample;
[0081] Calculate the gradient g:
[0082] Accumulate the squared gradient r: r = r + g ⊙ g;
[0083] Calculate the updated parameter θ:
[0084] Apply the updated parameter θ: θ = θ + Δθ;
[0085] Briefly speaking, the Adagrad optimization algorithm is that when using a batch size of data for parameter update each time, the algorithm calculates the gradients of all parameters. Then the idea is that for each parameter, initialize a variable s to 0, and then each time accumulate the squared sum of the gradient of this parameter to this variable s. Then when updating the parameters, the learning rate becomes:
[0086]
[0087] ξ is to maintain numerical stability to prevent the denominator from being 0.
[0088] Compared with SGD, the difference is that Adagrad uses the accumulated squared gradient. After setting the global learning rate, the next learning rate is the square root of the sum of the squares of the historical gradients of the global learning rate for each parameter, making the learning rate of each parameter different, thus accelerating the training speed. Our initial learning rate is set to 0.01, and the L2 regularization parameter is used for parameter descent, with the small parameter δ being 0.0005. The adaptive learning rate can help the algorithm slow down the learning rate in the direction of parameters with large gradients and speed up the learning rate in the direction of parameters with small gradients.
[0089] Finally, a set of parameters corresponding to the 3D CNN deep learning model is obtained, and thus the 3D CNN deep learning network model of the present invention is built.
[0090] Step 5, test the test samples to achieve the classification of each pixel point of the hyperspectral data.
[0091] The specific steps of Step 5 are:
[0092] Step 3 and Step 4 are the processes of training the network. This process is the test process. Each pixel point of the test set test, considering the size of its surrounding 7*7 neighborhood, is input into the built 3D CNN deep learning network model, and finally the class label of the test sample is output to complete the classification of the ground objects in the hyperspectral image.
[0093] The classification results of the ground objects can be seenFigures 3 - 6 。
[0094] The following two evaluation metrics are used to evaluate the classification results.
[0095] 1) Kappa coefficient: An evaluation metric defined on the confusion matrix, which is a metric for measuring classification accuracy. It comprehensively considers the elements on the diagonal of the confusion matrix and the elements deviating from the diagonal, and more objectively reflects the classification performance of the algorithm. The value of Kappa is between -1 and 1, and the larger this value, the better the classification effect.
[0096] 2) Average accuracy (AA): Divide the number of correctly classified pixel points of each class on the test set by the total number of all pixels in that class to obtain the correct classification accuracy of that class. The average of the accuracies of all classes is called the average accuracy AA, and its value is between 0 and 100%. The larger this value, the better the classification effect.
[0097] The evaluation results can be seen in Tables 3 - 6.
[0098] Embodiment
[0099] The following further illustrates the effect of the present invention in combination with simulation experiments:
[0100] 1. Simulation experiment conditions:
[0101] 2. The simulation experiment environment platform of the present invention is: Intel(R)core(TM)i7 - 10700@2.90GHz, RAM 16.0GB, python3.8.5, Anaconda, torch.version 1.7.1.
[0102] 3. The hyperspectral dataset used in the simulation experiment of the present invention is shown in Table 1 below.
[0103] Table 1 Dataset Introduction
[0104]
[0105] Pavia University (PaviaU) dataset: It is obtained by the ROSIS satellite, with a total of 610 * 340 pixels. After removing the bands affected by the atmosphere and water absorption, there are 103 spectral bands left. The ground resolution is 1.3m, and it is divided into 9 classes.
[0106] IndianPines dataset: It is obtained by the AVIRIS satellite, with a total of 145 * 145 pixels. After removing 24 bands affected by the atmosphere, water absorption rate, etc., 200 bands are available for use. The ground resolution is 20m, and it is divided into 16 classes.
[0107] Salinas dataset: It was acquired by the AVIRIS satellite, with a total of 512 * 217 pixels. After removing 20 bands covering water absorption areas, 204 bands were finally selected for use. The spatial resolution is 3.7m, and it is divided into 16 classes.
[0108] KSC dataset: It was acquired by the AVIRIS satellite, with a total of 512 * 614 pixels. After removing the water absorption bands, 176 bands remained. Its spatial resolution is 18m, and it is divided into 13 classes.
[0109] The categories and corresponding labels in its dataset are shown in Table 2 below.
[0110] Table 2 Categories and Their Labels of Different Datasets
[0111]
[0112] 4. Parameter Settings
[0113] The classification of hyperspectral images is pixel - based classification. Previous scholars' experiments have verified that each pixel has a certain correlation with its surrounding pixels. For the central pixel, considering a 5 * 5 spatial area, the learning rate is set to 0.01, the Adagrad algorithm is used to find the optimal parameters, and the coefficient decay step size is set to 0.0005. The batchsize is 100, the epoch size is set to 200, and the cross - entropy function is used for classification.
[0114] Among them, the setting of the convolutional kernel size is obtained through comparative experiments and refers to the convolutional kernel size in HE. This experiment is an improvement on the HE architecture. Considering the full extraction of spatial and spectral details and adding a dropout layer to prevent overfitting. Experiments were conducted with dropout set to 0.5 and 0.6, and finally 0.6 was selected as the dropout size, which can make the classification accuracy of the experiment higher.
[0115] 5. Experimental Comparison
[0116] To verify the effectiveness of the method proposed in the present invention, under the same conditions, the classification results of the present invention are compared with those of four existing classification methods in the hyperspectral field. On the hyperspectral datasets PaviaU and Salinas, 5% training samples are selected, and the experimental results are the average of ten experiments. On the datasets IndianPines and KSC, 20% training samples are selected, and the experimental results are the average of five experiments. The comparison results are shown in Tables 3 - 6 below.
[0117] These four existing methods are respectively:
[0118] 1) Classical Support Vector Machine (SVM), which directly classifies spectral information through SVM.
[0119] 2) The method proposed by He et al. in their published paper "Multi-scale 3D deep convolutional neural network for hyperspectral image classification[J]. 2017 IEEE International Conference on Image Processing (ICIP), 2017, 3904 - 3908." that uses a one-layer multi-scale strategy to extract joint spatial-spectral information and a one-layer multi-scale convolutional kernel to extract spectral information.
[0120] 3) A method proposed by Li et al. in their published paper "Spectral-spatial classification of hyperspectral imagery with 3D convolutional neural network[J]. Remote Sensing, 2017, 9(1): 67." that uses two convolutional pooling layers to process data for feature extraction and finally classifies the features through a fully connected layer.
[0121] 4) A method proposed by Lee et al. in their published paper "Going deeper with contextual CNN for hyperspectral image classification[J]. in IEEE Transactions on Image Processing, 2017, 26(10): 4843 - 4855." that uses multi-scale convolution to extract features, splices the extracted features together, and then inputs the features into a residual network and uses a 1*1*n kernel for processing to prevent overfitting.
[0122] Table 3 Comparison results of the prior art and the present invention in classification accuracy on PaviaU
[0123]
[0124] It can be seen from the experimental results of HE and the algorithm in this paper that the structure of the multi-scale strategy using 5% for training is effective, and the average accuracy and Kappa of the algorithm in this paper exceed those of other comparison algorithms by 2%, indicating that the CNN network model designed in this paper can better fully exploit the information in the data under the condition of limited training samples.
[0125] Table 4 Comparison results of the prior art and the present invention in classification accuracy on Salinas
[0126]
[0127] Since there are many pixel points in the Salinas dataset, a small number of training samples are selected. It can be seen from the experimental results that the present invention is far superior to other comparative algorithms in terms of experimental results.
[0128] Table 5 Comparison results of the prior art and the present invention in classification accuracy on IndianPines
[0129]
[0130] For this dataset, the simpler the data model, the better the classification effect after feature extraction. When the training samples are 20%, he is the best and the present invention is the second best. It can be seen the effectiveness of the multi-scale strategy. In contrast, the large-scale kernel of the algorithm in this paper filters out some valid information in the spectral dimension. Secondly, the small-scale training kernel model of he is simple and the effect is better.
[0131] Table 6 Comparison results of the prior art and the present invention in classification accuracy on KSC
[0132]
[0133] Due to the sparsity of the KSC dataset, traditional parameter optimization algorithms are prone to fall into local extrema. The present invention uses the Adagrad algorithm to adaptively learn the rate, which reduces the occurrence of falling into local optima to a certain extent. It can be seen from the experimental results that the algorithm of the present invention is far superior to other comparative algorithms for the KSC dataset.
[0134] Based on the above analysis of the simulation experimental results, the method proposed by the present invention can effectively extract the joint spatial-spectral features of hyperspectral data, and has a certain adaptability to different datasets. The multi-scale strategy and multi-layer network structure adopted can improve the average accuracy AA, and the Kappa accuracy is also excellent.
Claims
1. A small-sample hyperspectral image classification method based on a 3D deep convolutional neural network, characterized in that, The implementation is specifically carried out according to the following steps: Step 1: Input the hyperspectral data as a whole, and divide the data into a training set train and a test set test; Step 2: Construct a deep learning network model of 3D CNN, determine the size of the convolutional kernels in each layer, the number of convolutional kernels, the number of fully connected layers, and determine the loss function and its hyperparameters; Construct a 3D CNN network structure model as follows: The 3D CNN network model has a total of 6 layers, including three multi-scale convolutional layers, one convolutional layer for jointly processing spatial and spectral information, and two fully connected layers; Step 2.1: For the first multi-scale convolutional layer, use small convolutional kernels in the spatial dimension and large convolutional kernels in the spectral dimension. The kernel sizes are 11*3*3, 7*2*2, 5*1*1 respectively, and the corresponding convolutional strides stride are (1,1,1), (1,2,2), (1,1,1), and the convolutional padding blocks padding are (5,1,1), (3,3,3), (2,0,0); Step 2.2: For the second multi-scale convolutional layer, only process the spatial information, and use convolutional kernels of 1*n1*n1, where n1 takes values of: 1, 2, 3, that is, the convolutional kernels are 1*3*3, 1*2*2, 1*1*1 respectively, and the corresponding convolutional strides stride are (1,1,1), (1,2,2), (1,1,1), and the convolutional padding blocks padding are (0,1,1), (0,3,3), (0,0,0); Step 2.3: For the third multi-scale convolutional layer, process its spectral information, and use convolutional kernels of n2*1*1, where the three values of n2 are: 11, 7, 5; the convolutional kernels are 11*1*1, 7*1*1, 5*1*1 respectively, and the corresponding convolutional strides stride are (1,1,1), (1,1,1), (1,1,1), and the convolutional padding blocks padding are (5,0,0), (3,0,0), (5,0,0); Step 2.4: Set a convolutional layer for jointly processing spatial and spectral information. The size of the convolutional layer is: (3,2,2), the convolutional stride stride is (1,1,1), and the convolutional padding block padding is (1,0,0). Then connect it to the pooling layer for dimensionality reduction. The convolutional kernel of the pooling layer is (3,2,2), the stride of the stride convolutional step is (1,1,1), and the convolutional padding block padding is (0,0,0); Step 2.5: Design two fully connected layers: The first fully connected layer FC1 first maps the feature S4 obtained in Step 2.4 to 1024 nodes to obtain the feature S5, and the second fully connected layer FC2 maps the feature S5 in the 1024 nodes to the category Label owned by the hyperspectral data set in Step 1; Step 3: Input the training set train of the hyperspectral data in Step 1 into the deep learning network model of 3D CNN constructed in Step 2 to achieve feature extraction; Step 4: Input the extracted features into the classifier to learn the features, train the parameters, and determine the network model; Step 5: Test the test samples to classify each pixel of the hyperspectral data.
2. The small-sample hyperspectral image classification method based on a 3D deep convolutional neural network according to claim 1, wherein, The step 1 is specifically implemented according to the following steps: Step 1.
1. Treat the hyperspectral data as 3D cube data, where M represents the total number of pixels in the hyperspectral data, M = m1 × m2, m1 and m2 represent the length and width of the input data respectively, and each pixel is formed by the action of N spectral bands, N = {1, 2, …, n}. Therefore represents the hyperspectral image, and the hyperspectral image has a length and width of m1 and m2 and a total of N bands, where n represents the nth band; Step 1.2: For each pixel in the training set train, consider the 7*7 area around it as the center, and use its area as the input of each pixel to the 3D CNN network model in step 2.
3. The small-sample hyperspectral image classification method based on a 3D deep convolutional neural network according to claim 2, wherein The specific structure of step 3 is: Step 3.1, input the training set train in step 1 into the first multi-scale convolutional layer in step 2.1 of the deep learning network model of the 3D CNN constructed in step 2, obtain features S11, S12, and S13, add features S11, S12, and S13, and then output feature S1 through a RELU activation function as the input of the second multi-scale convolutional layer; Step 3.2: After feature S1 passes through the second multi-scale convolution layer of step 2.2, features S21, S22, and S23 are obtained. The extracted features S21, S22, and S23 are added together and then output through the RELU activation function to obtain feature S2, which is used as the input of the third multi-scale convolution layer. Step 3.3, after passing feature S2 through the third multi-scale convolution layer of step 2.3, features S31, S32, and S33 are obtained. The extracted features S31, S32, and S33 are added together and then output through the RELU activation function to obtain feature S3; Step 3.4, input feature S3 into the convolution layer of step 2.4 for jointly processing spatial spectrum information and pass it through the RELU activation function to obtain feature S41, then perform pooling operation on feature S41 to obtain feature S42, and pass it through the RELU activation function again to obtain feature S4.
4. The small-sample hyperspectral image classification method based on a 3D deep convolutional neural network according to claim 3, characterized in that The specific steps of step 4 are: The feature S4 obtained in step 3.4 is input into the first fully connected layer FC1 in step 2.4, and then input into the second fully connected layer FC2 in step 2.4 after the activation function. Finally, the feature S5 is pulled into a one-dimensional vector corresponding to the number of categories and trained in the 3D CNN deep learning network model. During the training process, the adaptive learning rate optimization algorithm Adagrad algorithm is used to find the optimal parameters and save the weights of the model. The algorithm flow is as follows: First, you need to set the global learning rate ε, initial parameter θ, and small constant δ; Then, initialize the gradient accumulation variable r = 0; When the stopping criterion is not met: Collect a mini-batch containing m samples {x (1) , …, x (m)} from the training samples, with the corresponding target being y (i) , representing the predicted output of the i-th sample; Calculate the gradient g: Cumulative square gradient r: r = r + g⊙g; Calculate the updated parameter θ: Apply update parameter θ: θ = θ + Δθ; When updating the parameters, the learning rate becomes: ξ is to maintain numerical stability to prevent the denominator from being 0. Finally, a set of parameters corresponding to the 3D CNN deep learning model is obtained, and the 3D CNN deep learning network model is completed.
5. The small-sample hyperspectral image classification method based on a 3D deep convolutional neural network according to claim 4, characterized in that The specific steps of step 5 are: Steps 3 and 4 are the process of training the network. This process is the testing process. Each pixel point of the test set test, considering the size of the surrounding 7*7 area, is input into the built 3D CNN deep learning network model, and finally the category label of the test sample is output to complete the classification of the ground objects in the hyperspectral image.
Citation Information
Patent Citations
Hyperspectral image classification method based on depth feature cross fusion
CN111191736A