A hyperspectral image classification method based on small sample deep learning

By introducing small sample deep learning methods in hyperspectral image classification, the source domain data set is used to help the target domain data sets to classify, solving the problem of poor results in the training of few samples in the existing technology, achieving higher classification accuracy and deeper feature extraction.

CN116310510BActive Publication Date: 2025-05-09XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310095139.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-08
Publication Date
2025-05-09
Estimated Expiration
2043-02-08

AI Technical Summary

Technical Problem

The existing hyperspectral image classification methods are not effective in training with few samples, and it is difficult to make full use of key information in hyperspectral images, resulting in information loss or redundancy, and it is impossible to obtain more resolvable spectral semantic features.

Method used

A hyperspectral image classification method based on small sample deep learning is proposed. By learning the prior knowledge in the source domain hyperspectral dataset, the source domain dataset with a large number of label samples is used to help classify the target domain dataset with a small number of label samples. The method includes acquiring the hyperspectral image set of source and target domains, extracting image blocks, constructing a small sample spectral null feature extraction convolutional neural network, and optimizing the model through the combination of alternating training and multiple loss functions.

Benefits of technology

The accuracy of classification of land objects in hyperspectral images is improved, and the spatial information in hyperspectral images can be extracted more deeply, overcoming the problems of information loss and redundancy, and significantly improving the classification performance under small sample training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310510B_ABST
    Figure CN116310510B_ABST
Patent Text Reader

Abstract

The present invention relates to a hyperspectral image classification method based on small sample deep learning, comprising: obtaining a source domain hyperspectral image set and a target domain hyperspectral image set; extracting a plurality of first image blocks and a plurality of second image blocks with each pixel point of the source domain hyperspectral image and the target domain hyperspectral image after filling as the center; in each category, randomly selecting part of the first image blocks to form a source domain support set and part of the first image blocks to form a source domain query set, and randomly selecting part of the second image blocks to form a target domain support set and part of the first image blocks to form a target domain query set; training a small sample spectral space feature extraction convolutional neural network using the support set and the query set to obtain a trained network; inputting the hyperspectral image to be classified into the trained small sample spectral space feature extraction convolutional neural network to obtain a classification result. The present invention can improve the classification accuracy and can extract the spatial information contained in the hyperspectral image at a deeper level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing information processing and relates to a hyperspectral image classification method based on small sample deep learning. Background Art

[0002] Hyperspectral images record the continuous spectral characteristics of ground objects with their rich band information, and have the possibility of identifying more types of ground objects and classifying targets with higher accuracy. Unlike ordinary natural images, hyperspectral image data presents a three-dimensional structure, with very rich spectral information and relatively less spatial information. The key to hyperspectral image classification technology is to use the spatial and inter-spectral features of hyperspectral images to classify sample categories. However, hyperspectral image training samples are relatively small, and for classification methods that require a large number of parameters, overfitting problems are prone to occur. How to train an efficient classification model with a small number of samples is very important for hyperspectral image classification.

[0003] Kun Tan et al. proposed a novel semi-supervised HSI classification method in their paper “A novel semi-supervised hyperspectral image classification approach based on spatial neighborhood information and classifier combination” (ISPRS journal of photogrammetry and remote sensing, 2015), combining spatial neighborhood information with the classifier to enhance classification ability. Yue Wu et al. proposed a semi-supervised method in their paper “Semi-supervised hyperspectral image classification via spatially-regulated self-training” (Remote Sensing, 2020), which uses self-training to gradually assign highly confident pseudo-labels to unlabeled samples through clustering, and uses spatial constraints to regulate the self-training process. However, these methods assume that labeled and unlabeled samples come from the same dataset, which means that the classification performance is still limited by the number of labeled samples in the data to be classified (i.e., the target domain).

[0004] Bing Liu et al. proposed a small sample deep learning method to solve the small sample problem of HSI classification in their paper "Deep few-shot learning for hyperspectral image classification" (IEEE Transactions on Geoscience and Remote Sensing, 2018). This method helps classification better by learning the metric space from the training set. Kuiliang Gao et al. designed a new deep classification model based on relational network in their paper "Deep relation network for hyperspectral image few-shot classification" (Remote Sensing, 2020) and trained it with the idea of ​​meta-learning.

[0005] In addition to the hyperspectral image classification methods listed above, the current hyperspectral image classification methods based on deep convolutional neural networks are similar to the above methods. The commonality of these methods is that when extracting inter-spectral and spatial features, the information is lost due to insufficient utilization of the extracted features, or too much irrelevant information is retained to cause information redundancy, which cannot fully utilize the key information in the hyperspectral image bands and obtain more distinguishable spectral and spatial semantic features. In addition, a large number of hyperspectral samples are required to train the neural network during training, which leads to poor classification effects of these methods on hyperspectral images when training with few samples, and no deeper attention is paid to the differences in information between different spectra. In addition, considering the problem of fewer hyperspectral data sets, although there are methods that use small sample learning to solve the hyperspectral classification problem, no suitable method has been proposed to solve the domain adaptation problem caused by using different hyperspectral data sets. During classification, the loss function used by the above methods is too single and cannot meet the needs of high-precision classification. Summary of the invention

[0006] The purpose of the present invention is to address the deficiencies of the above-mentioned prior art and propose a hyperspectral image classification method based on small sample deep learning. By learning the prior knowledge in the source domain hyperspectral dataset, the source domain hyperspectral dataset with a large number of labeled samples is used to help the target domain hyperspectral dataset with a small number of labeled samples for classification, so as to improve the accuracy of ground object classification in the case of few sample training in the hyperspectral image. The technical problem to be solved by the present invention is achieved by the following technical solutions:

[0007] The embodiment of the present invention provides a hyperspectral image classification method based on small sample deep learning, and the hyperspectral image classification method includes:

[0008] Step 1: obtaining a source domain hyperspectral image set and a target domain hyperspectral image set, wherein the source domain hyperspectral image set includes a plurality of source domain hyperspectral images, and the target domain hyperspectral image set includes a plurality of target domain hyperspectral images;

[0009] Step 2, filling the edge parts of the source domain hyperspectral image and the target domain hyperspectral image respectively, and extracting a plurality of first image blocks and a plurality of second image blocks with each pixel point of the source domain hyperspectral image and the target domain hyperspectral image after the filling process as the center;

[0010] Step 3: In each category, randomly select part of the first image blocks to form a source domain support set and part of the first image blocks to form a source domain query set, and randomly select part of the second image blocks to form a target domain support set and part of the first image blocks to form a target domain query set;

[0011] Step 4: Based on the stochastic gradient descent method, the source domain support set, the source domain query set, the target domain support set and the target domain query set are used to alternately train the small sample spectral-spatial feature extraction convolutional neural network to obtain a trained small sample spectral-spatial feature extraction convolutional neural network, wherein the small sample spectral-spatial feature extraction convolutional neural network extracts features from both spectral and spatial aspects, and the total loss function of the small sample spectral-spatial feature extraction convolutional neural network consists of a cross entropy loss function, a correlation alignment loss function and a maximum mean difference loss function;

[0012] Step 5: input the hyperspectral image to be classified into the trained small sample spectral-spatial feature extraction convolutional neural network to obtain a classification result.

[0013] In one embodiment of the invention, step 2 comprises:

[0014] Step 2.1, filling pixels with a pixel value of 0 around the source domain hyperspectral image and the target domain hyperspectral image, respectively, to obtain a filled source domain hyperspectral image and a filled target domain hyperspectral image;

[0015] Step 2.2, taking each pixel point in the padded source domain hyperspectral image and the padded target domain hyperspectral image as the center, select the first image block and the second image block with a spatial size of (2t+1)×(2t+1) and a channel number of d, where t is an integer greater than 0.

[0016] In one embodiment of the invention, the structure of the small sample spectral-spatial feature extraction convolutional neural network includes a spectral branch network, a spatial branch network, a domain attention module, a first splicing layer, a fully connected layer and a softmax classifier. The spectral branch network and the spatial branch network are connected in parallel and then connected in series with the domain attention module in sequence. The domain attention module, the first splicing layer, the fully connected layer and the softmax classifier are connected in series. The spectral branch network includes 2 3D deformable convolution blocks and 2 first maximum pooling layers. The first 3D deformable convolution block, the first first maximum pooling layer, the second 3D deformable convolution block, and the second first maximum pooling layer are connected in series in sequence. The spatial branch network domain attention module includes a multi-scale spatial feature extraction module, a second splicing layer and a second maximum pooling layer connected in series in sequence.

[0017] In one embodiment of the invention, the 3D deformable convolution block includes three 3D deformable convolution layers and three first activation function layers. The first 3D deformable convolution layer, the first first activation function layer, the second 3D deformable convolution layer, the second first activation function layer, the third 3D deformable convolution layer, and the third first activation function layer are connected in series in sequence, and the output of the first first activation function layer is added to the output of the third 3D deformable convolution layer to form a residual structure.

[0018] In one embodiment of the invention, the multi-scale spatial feature extraction module includes 2 scale operation layers, 3 convolution layers, 3 normalization layers, and 3 second activation function layers;

[0019] The first convolution layer, the first normalization layer, and the first second activation function layer are connected in series in sequence;

[0020] The first scale operation layer, the second convolution layer, the second normalization layer, and the second second activation function layer are sequentially connected in series;

[0021] The second scale operation layer, the third convolution layer, the third normalization layer, and the third second activation function layer are sequentially connected in series;

[0022] The first second activation function layer, the second second activation function layer, and the third second activation function layer are connected in parallel and then connected in series with the second concatenation layer and the second maximum pooling layer in sequence.

[0023] In one embodiment of the invention, the total loss function of the small sample spectral space feature extraction convolutional neural network is:

[0024] L total =L fsl +L coral +L MMD

[0025] Among them, L total represents the total loss function of the convolutional neural network for small sample spectral feature extraction, L fsl represents the cross entropy loss, L coral represents the correlation alignment loss, L MMD represents the maximum mean difference loss;

[0026] Cross entropy loss L fsl It is expressed as:

[0027]

[0028]

[0029] Where d(·) represents the Euclidean distance, F ω (·) represents the feature extraction function with parameter ω, f l represents the feature of the lth category in the source domain support set or the target domain support set, C represents the number of categories, and x j represents a sample in the source domain query set or the target domain query set, y j Represents sample x j The label of , Q represents the source domain query set or the target domain query set;

[0030] Correlation alignment loss L coral It is expressed as:

[0031]

[0032] in, represents the Frobenius norm of the matrix, C S Represents the covariance matrix of the source domain features, C T Represents the covariance matrix of the target domain features, and d represents the dimension of the features;

[0033] Maximum mean difference loss L MMD It is expressed as:

[0034]

[0035] in, represents the spatial distance, φ(·) represents the mapping function, represents the source domain characteristics, represents the target domain features, X s represents the source domain dataset, X t represents the target domain dataset, n s Indicates the number of data in the source domain dataset, n t Indicates the number of data in the target domain dataset.

[0036] In one embodiment of the invention, step 4 comprises:

[0037] The initial learning rate of training is set to α, and the number of iterations is set to T. In odd-numbered iterations, the source domain support set and the source domain query set are sent to the small sample spectral-space feature extraction convolutional neural network for training, and the total loss function is used to calculate the loss value between the features of the source domain support set and the features of the source domain query set to update the parameters in the small sample spectral-space feature extraction convolutional neural network; in even-numbered iterations, the target domain support set and the target domain query set are sent to the small sample spectral-space feature extraction convolutional neural network for training, and the total loss function is used to calculate the loss value between the features of the target domain support set and the features of the target domain query set to update the parameters in the small sample spectral-space feature extraction convolutional neural network until the loss value of the small sample spectral-space feature extraction convolutional neural network no longer decreases and the current number of training rounds is less than the number of iterations T or the number of training rounds reaches the number of iterations T, then the training of the small sample spectral-space feature extraction convolutional neural network is stopped to obtain the trained small sample spectral-space feature extraction convolutional neural network.

[0038] In one embodiment of the invention, the weight vector W after the small sample spectral space feature extraction convolutional neural network is updated is new for:

[0039]

[0040] Among them, L total represents the total loss function of the small sample spectral space feature extraction convolutional neural network, W represents the weight vector of the small sample spectral space feature extraction convolutional neural network before updating, and R represents the learning rate.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] The present invention targets the rich spectral feature information and spatial feature information of hyperspectral images and classifies them using a small sample spectral-spatial feature extraction convolutional neural network. The small sample spectral-spatial feature extraction convolutional neural network extracts features from both spectral and spatial aspects, which can improve classification accuracy and extract the spatial information contained in hyperspectral images at a deeper level.

[0043] In the spectral branch, the present invention adopts 3D deformable convolution and residual structure, so that the convolution kernel can better extract deep information for irregular-shaped hyperspectral images and retain the information of shallow networks, thereby improving classification accuracy; the spatial branch adopts multi-scale operation. For the input hyperspectral sample, it is first copied twice, and then the edge pixels are discarded one by one, so that three input samples with different spatial resolution sizes are obtained, which can extract the spatial information contained in the hyperspectral image at a deeper level.

[0044] Other aspects and features of the present invention will become apparent from the following detailed description with reference to the accompanying drawings. It should be understood, however, that the drawings are designed for illustrative purposes only and are not intended to limit the scope of the present invention, as reference should be made to the appended claims. It should also be understood that, unless otherwise indicated, the drawings are not necessarily drawn to scale and are intended merely to conceptually illustrate the structures and processes described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is a flowchart of a hyperspectral image classification method based on small sample deep learning provided by an embodiment of the present invention;

[0046] Figure 2 It is a schematic diagram of a model structure of a convolutional neural network for extracting spectral and spatial features of small samples provided by an embodiment of the present invention;

[0047] Figure 3 It is a structural schematic diagram of a 3D deformable convolution module in a spectral branch provided by an embodiment of the present invention;

[0048] Figure 4 is a structural schematic diagram of a multi-scale spatial feature extraction module provided by an embodiment of the present invention;

[0049] Figure 5 It is a simulation diagram of the classification results of the present invention and the two existing networks on the University of Pavia data set;

[0050] Figure 6 It is a simulation diagram of the classification results of the present invention and the two existing networks on the Indian Pines data set. DETAILED DESCRIPTION

[0051] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0052] Embodiment 1

[0053] At present, all the hyperspectral image classification methods based on deep convolutional neural networks have some shortcomings. The common feature of these methods is that when extracting inter-spectral and spatial features, they suffer from information loss due to insufficient utilization of the extracted features, or retain too much irrelevant information, resulting in information redundancy. They are unable to fully utilize the key information in the hyperspectral image bands and obtain more distinguishable spectral-spatial semantic features. In addition, a large number of hyperspectral samples are required to train the neural network during training, which results in poor hyperspectral image classification results when training with a small number of samples, and they do not pay more attention to the differences in information between different spectra.

[0054] Based on this, the present invention proposes a hyperspectral image classification method based on small sample deep learning. Figure 1 , Figure 1 : is a flow chart of a hyperspectral image classification method based on small sample deep learning provided by an embodiment of the present invention. The hyperspectral image classification method based on small sample deep learning provided by an embodiment of the present invention may specifically include steps 1 to 4, wherein:

[0055] Step 1: obtain a source domain hyperspectral image set and a target domain hyperspectral image set, wherein the source domain hyperspectral image set includes multiple source domain hyperspectral images, and the target domain hyperspectral image set includes multiple target domain hyperspectral images.

[0056] Specifically, a hyperspectral image is a three-dimensional data S∈R h×w×c Each band in the hyperspectral image corresponds to a two-dimensional matrix S in the three-dimensional data i ∈R h×w , where ∈ represents the symbol, R represents the real number domain symbol, h represents the length of the hyperspectral image, w represents the width of the hyperspectral image, c represents the number of spectral bands of the hyperspectral image, and i represents the serial number of the spectral band in the hyperspectral image, i = 1, 2, …, c.

[0057] In this embodiment, the source domain hyperspectral image set adopts the Chikusei dataset, and the target domain hyperspectral image set adopts the UP dataset.

[0058] Step 2: Fill the edge parts of the source domain hyperspectral image and the target domain hyperspectral image respectively, and extract a plurality of first image blocks and a plurality of second image blocks with each pixel point of the filled source domain hyperspectral image and the target domain hyperspectral image as the center.

[0059] Specifically, this embodiment performs filling processing on the source domain hyperspectral image and the target domain hyperspectral image respectively, so that the first image block and the second image block at the edge can also be obtained, so that the first image block and the second image block at the edge contain the required information.

[0060] Step 2.1, fill the surrounding areas of the source domain hyperspectral image and the target domain hyperspectral image with pixels having a pixel value of 0 to obtain a filled source domain hyperspectral image and a filled target domain hyperspectral image.

[0061] Step 2.2, taking each pixel in the padded source domain hyperspectral image and the padded target domain hyperspectral image as the center, select the first image block and the second image block with a spatial size of (2t+1)×(2t+1) and a channel number of d, where t is an integer greater than 0, for example, the size of the image block is 9×9, so t=4.

[0062] Specifically, the first image block is a pixel block extracted with the pixel point in the filled source domain hyperspectral image as the center, and the second image block is a pixel block extracted with the pixel point in the filled target domain hyperspectral image as the center. The number of channels d is the same as the number of spectral bands of the hyperspectral image.

[0063] Step 3: In each category, randomly select part of the first image blocks to form a source domain support set and part of the first image blocks to form a source domain query set, and randomly select part of the second image blocks to form a target domain support set and part of the first image blocks to form a target domain query set.

[0064] Specifically, the first image block and the second image block are assigned to the set to which the category belongs according to the category of the central pixel point, and the category is, for example, water, glass, etc. Some of the first image blocks of each category are selected to form the source domain support set, and then some are selected to form the source domain query set; some of the second image blocks of each category are selected to form the target domain support set, and then some are selected to form the target domain query set.

[0065] For example, the source domain hyperspectral image set uses the Chikusei dataset, and 200 first image blocks are selected from each category to form the source domain dataset. Then, one image block is randomly selected from the source domain dataset for each category to form the source domain support set, and 19 first image blocks are randomly selected from each category to form the source domain query set; the target domain hyperspectral image set uses the UP dataset and other datasets, and 5 second image blocks are selected from each category to form the target domain dataset. After data enhancement (i.e., the target domain dataset is copied), one image block is randomly selected from the target domain dataset for each category to form the target domain support set, and 19 image blocks are randomly selected from each category to form the target domain query set.

[0066] Step 4. Based on the stochastic gradient descent method, the small sample spectral-space feature extraction convolutional neural network is alternately trained using the source domain support set, source domain query set, target domain support set and target domain query set to obtain a trained small sample spectral-space feature extraction convolutional neural network. The small sample spectral-space feature extraction convolutional neural network extracts features from both spectral and spatial aspects. The total loss function of the small sample spectral-space feature extraction convolutional neural network is composed of a cross entropy loss function, a correlation alignment loss function and a maximum mean difference loss function.

[0067] For details, see Figure 2 The structure of the convolutional neural network for small sample spectral-spatial feature extraction includes a spectral branch network, a spatial branch network, a domain attention module, a first splicing layer, a fully connected layer and a softmax classifier. The spectral branch network and the spatial branch network are connected in parallel and then connected in series with the domain attention module. The domain attention module, the first splicing layer, the fully connected layer and the softmax classifier are connected in series in sequence. The spectral branch network includes two 3D deformable convolution blocks and two first maximum pooling layers. The first 3D deformable convolution block, the first first maximum pooling layer, the second 3D deformable convolution block, and the second first maximum pooling layer are connected in series in sequence. The spatial branch network domain attention module includes a multi-scale spatial feature extraction module, a second splicing layer and a second maximum pooling layer connected in series in sequence.

[0068] The convolution kernel size of the first maximum pooling layer is set to 2*2*4, and the number of convolution kernels is set to 8. The convolution kernel size of the second first maximum pooling layer is set to 2*2*4, and the number of convolution kernels is set to 16.

[0069] See also Figure 3 The 3D deformable convolution block includes three 3D deformable convolution layers and three first activation function layers. The first 3D deformable convolution layer, the first first activation function layer, the second 3D deformable convolution layer, the second first activation function layer, the third 3D deformable convolution layer, and the third first activation function layer are connected in series in sequence. The output of the first first activation function layer is added to the output of the third 3D deformable convolution layer to form a residual structure. The 3D deformable convolution layer is a structure formed by adding a deformable convolution kernel to the 3D convolution.

[0070] The convolution kernel size of the 3D deformable convolution layer is set to 3*3*3; the activation function of each first activation function layer is set to the ReLU activation function, which is expressed as follows:

[0071] ReLU(x)=max(0,x)

[0072] Where x represents the input of the activation function.

[0073] See also Figure 4,The multi-scale spatial feature extraction module includes 2 scale operation layers, 3 convolution layers, 3 normalization layers, and 3 second activation function layers, where:

[0074] The first convolution layer, the first normalization layer, and the first second activation function layer are connected in series;

[0075] The first scale operation layer, the second convolution layer, the second normalization layer, and the second second activation function layer are connected in series in sequence;

[0076] The second scale operation layer, the third convolution layer, the third normalization layer, and the third second activation function layer are connected in series in sequence;

[0077] The first second activation function layer, the second second activation function layer, and the third second activation function layer are connected in parallel and then connected in series with the second concatenation layer and the second maximum pooling layer.

[0078] The first scale operation layer in the multi-scale spatial feature extraction module reduces one pixel on the edges of the selected image block, and the second scale operation layer reduces two pixels on the edges of the selected image block. The convolution kernel size of the first convolution layer is set to 5*5*4, the convolution kernel size of the second convolution layer is set to 3*3*4, and the convolution kernel size of the third convolution layer is set to 1*1*4. The number of convolution kernels is set to 16, and the activation function of each second activation function layer is set to the ReLU activation function.

[0079] The outputs of the three second activation function layers are all 16 features of size 5*5*25. After the second splicing layer, the splicing operation is performed to obtain 16 features of size 5*5*75, and then the second maximum pooling layer is used for pooling operation. The convolution kernel of the second maximum pooling layer is set to 2*2*8, and the number of convolution kernels is set to 16.

[0080] The domain attention module uses a 2D convolutional layer, which includes 1 spectral attention module and 2 spatial attention modules. The spectral attention module is 1 2D convolutional layer, and the 2 spatial attention modules are 2 2D convolutional layers. The three 2D convolutional layers are connected in series. The convolution kernel size of the spectral attention module is set to 9*9, and the number of convolution kernels is set to 1. The domain attention module contains spatial attention and inter-spectral attention.

[0081] In this embodiment, the spectral branch network and the spatial branch network are connected in parallel and then connected in series with the domain attention module, the first splicing layer, the fully connected layer and the softmax classifier to form a small sample spectral and spatial feature extraction convolutional neural network. The domain attention module is a convolutional layer. The domain attention module performs a convolution operation to obtain a weighted coefficient. The small sample spectral and spatial feature extraction convolutional neural network selects the cross entropy loss function, the correlation alignment loss function and the maximum mean difference loss function as the loss function of the small sample spectral and spatial feature extraction convolutional neural network.

[0082] In this embodiment, the total loss function of the small sample spectral space feature extraction convolutional neural network is:

[0083] L total =L fsl +L coral +L MMD

[0084] Among them, L total represents the total loss function of the convolutional neural network for small sample spectral feature extraction, L fsl represents the cross entropy loss function, L coral represents the correlation alignment loss function, L MMD represents the maximum mean difference loss function;

[0085] Cross entropy loss function L fsl It is expressed as:

[0086]

[0087]

[0088] Among them, L fsl represents the loss value between the predicted label vector and the true label vector, d(·) represents the Euclidean distance, and F ω (·) represents the feature extraction function with parameter ω, f l represents the feature of the lth category in the source domain support set or the target domain support set, C represents the number of categories, and x j represents a sample in the source domain query set or the target domain query set, y j Represents sample x j The label of , Q represents the source domain query set or the target domain query set.

[0089] Correlation alignment loss L coral It is expressed as:

[0090]

[0091] in, represents the Frobenius norm of the matrix, C SRepresents the covariance matrix of the source domain features, C T represents the covariance matrix of the target domain features. The source domain features are the features of the source domain support set and the source domain query set after the features are extracted by the spectral branch network and the spatial branch network, and then the features are spliced ​​after the domain attention module and the first splicing layer. The target domain features are the features of the target domain support set and the target domain query set after the features are extracted by the spectral branch network and the spatial branch network, and then the features are spliced ​​after the domain attention module and the first splicing layer. d represents the dimension of the feature.

[0092] Maximum mean difference loss L MMD It is expressed as:

[0093]

[0094] in, represents the spatial distance, which is measured by mapping the data into the Reproducing Hilbert Space (RKHS) by φ(·), where φ(·) represents the mapping function. represents the source domain characteristics, represents the target domain features, X s represents the source domain dataset, X t represents the target domain dataset, n s Indicates the number of data in the source domain dataset, n t Indicates the number of data in the target domain dataset.

[0095] Based on the small sample spectral space feature extraction convolutional neural network and its total loss function described above, the training method for the small sample spectral space feature extraction convolutional neural network is:

[0096] The initial learning rate of training is set to α, and the number of iterations is set to T. In odd-numbered iterations, the source domain support set and the source domain query set are sent to the small-sample spectral-space feature extraction convolutional neural network for training, and the total loss function is used to calculate the loss value between the features of the source domain support set and the features of the source domain query set to update the parameters in the small-sample spectral-space feature extraction convolutional neural network; in even-numbered iterations, the target domain support set and the target domain query set are sent to the small-sample spectral-space feature extraction convolutional neural network for training, and the total loss function is used to calculate the loss value between the features of the target domain support set and the features of the target domain query set to update the parameters in the small-sample spectral-space feature extraction convolutional neural network until the loss value of the small-sample spectral-space feature extraction convolutional neural network no longer decreases and the current number of training rounds is less than the number of iterations T or the number of training rounds reaches the number of iterations T, then the training of the small-sample spectral-space feature extraction convolutional neural network is stopped to obtain a trained small-sample spectral-space feature extraction convolutional neural network.

[0097] Specifically, in odd-numbered iterations, the source domain support set and source domain query set in the source domain dataset are sent to the small sample spectral-space feature extraction convolutional neural network for training, and the loss value between the features of the source domain support set and the features of the source domain query set is calculated; in even-numbered iterations, the target domain support set and target domain query set in the target domain dataset are sent to the small sample spectral-space feature extraction convolutional neural network for training, and the loss value between the support set features and the query set features is also calculated. Alternately, the source domain dataset and the target domain dataset are input respectively to train the small sample spectral-space feature extraction convolutional neural network, and the parameters in the network are updated by continuous iteration.

[0098] The learning rate R of each input hyperspectral image block is set to: R = α.

[0099] Perform T weight updates on the small sample spectral space feature extraction convolutional neural network to obtain the updated weight vector W new :

[0100]

[0101] Among them, L total Represents the total loss function, W represents the weight vector before updating of the small sample spectral space feature extraction convolutional neural network, and R represents the learning rate.

[0102] The next training sample set is input into the small sample spectral space feature extraction convolutional neural network, and the loss function value of the total loss function is updated so that the loss function value L total Keep decreasing until the loss function value L total If it no longer decreases and the current number of training rounds is less than the set number of iterations T, the training of the network is stopped to obtain a trained small sample spectral space feature extraction convolutional neural network; otherwise, when the number of training rounds reaches T, the training of the network is stopped to obtain a trained small sample spectral space feature extraction convolutional neural network.

[0103] Step 5: Input the hyperspectral image to be classified into the trained small sample spectral-spatial feature extraction convolutional neural network to obtain the classification result.

[0104] In this embodiment, in order to test the trained small sample spectral-spatial feature extraction convolutional neural network, the test sample can be input into the trained small sample spectral-spatial feature extraction convolutional neural network to obtain the category of the test sample and complete the classification of the hyperspectral image.

[0105] First, the spectral branch network constructed by the present invention can extract rich inter-spectral features through the 3D deformable convolution blocks therein. By paying attention to and screening these inter-spectral features through the 3D deformable convolution blocks therein, more discriminative inter-spectral features can be extracted, which overcomes the problem that the prior art cannot extract more useful information due to the fixed convolution kernel when extracting inter-spectral features, or retains too much irrelevant information causing information redundancy, thereby improving the classification accuracy of ground objects in hyperspectral images.

[0106] Secondly, the spatial branch network constructed by the present invention enables the small sample spectral-spatial feature extraction convolutional neural network to pay attention to spatial features of different scales through the multi-scale spatial feature extraction module therein, overcoming the shortcoming of the prior art that a single scale is used to extract spatial features of hyperspectral image blocks. Through the multi-way spatial attention mechanism module therein, these multi-scale spatial features can be paid attention to and screened to extract more discriminative spatial features, overcoming the information loss caused by insufficient utilization of the extracted features in the prior art during spatial feature extraction, or the information redundancy caused by retaining too much irrelevant information, thereby improving the classification ability of the convolutional neural network during training with a small number of samples.

[0107] Third, the small sample spectral-spatial feature extraction convolutional neural network of the present invention adopts a domain attention module, which is mainly aimed at the problem that the large number of spectral bands in hyperspectral images leads to excessive redundant information between bands. Useful inter-spectral features are extracted by spectral attention and useful spatial features are extracted by spatial attention, so that the neural network pays more attention to the useful information in the feature information. In order to reduce the domain transfer problem caused by training with different hyperspectral data sets, the present invention adopts the cross entropy loss function, the correlation alignment loss function and the maximum mean difference loss function as the loss function of the network. This makes the small sample spectral-spatial feature extraction convolutional neural network pay more attention to the categories of objects with unconcentrated sample distribution or small sample size.

[0108] The effect of the present invention is further described below in conjunction with simulation experiments.

[0109] Simulation experiment conditions:

[0110] The hardware platform of the simulation experiment of the present invention is: Inter core i7-6700, frequency is 3.4GHz, Nvidia GeForce RTX3090. The software of the simulation experiment of the present invention uses pytorch.

[0111] The simulation experiment of the present invention adopts the present invention and two existing RN-FSC and DCFSL methods to classify the ground objects in the University of Pavia and Indian Pines hyperspectral datasets respectively.

[0112] The RN-FSC method refers to a hyperspectral classification method proposed by Kuiliang Gao et al. in "Deep relation network for hyperspectral image few-shot classification" (Remote Sensing, 2020). Its feature learning module and relationship learning module can make full use of the spatial-spectral information in the hyperspectral image to achieve accurate classification of new hyperspectral images with a small number of labeled samples.

[0113] The DCFSL method refers to: a hyperspectral classification method of small-sample meta-learning that uses source class data to help classify target classes proposed by Rui Li et al. in “Deep crossdomain few-shot learning for hyperspectral image classification” (Remote Sensing, 2021).

[0114] The target domain datasets used in this invention are the University of Pavia and Indian Pines hyperspectral datasets, which are data collected by AVIRIS sensor at the University of Pavia in California and an Indian pine tree in Indiana, USA, respectively. Indian Pines is the earliest test data for hyperspectral image classification. In 1992, an Indian pine tree in Indiana, USA, was imaged by the Airborne Visible Infrared Imaging Spectrometer (AVIRIS), and then the image was cut into a size of 145×145 and annotated for hyperspectral image classification test purposes. Among them, the size of the University of Pavia hyperspectral dataset image is 610×340, with 103 bands, including 9 types of ground objects, and the category and quantity of each type of ground object are shown in Table 1.

[0115] Table 1. University of Pavia sample categories and quantities

[0116]

[0117]

[0118] The image size of the Indian Pines hyperspectral dataset is 145×145, with 200 bands and 16 types of ground objects. The category and number of each type of ground object are shown in Table 2.

[0119] Table 2 Indian Pines sample categories and quantities

[0120] Classification Feature Type quantity 1 Alfalfa 46 2 Corn-notill 1428 3 Corn-mintill 830 4 Corn 237 5 Grass-pasture 483 6 Grass-tree 730 7 Grass-pasture-mowed 28 8 Hay-windrowed 478 9 Oats 20 10 Soybean-notill 972 11 Soybean-mintill 2455 12 Soybean-clean 593 13 Wheat 205 14 Woods 1265 15 Buildings-Grass-Trees-Dribes 386 16 Stone-steel-Towers 93

[0121] The source domain dataset used in this paper is the Chikusei dataset, which is a hyperspectral image of Chikusei, Ibaraki, Japan, acquired by the Hyperspec-VNIR-CIRIS spectrometer. The ground sampling distance is 2.5m, the image size is 2517×2335 pixels, with a total of 512 bands, 128 bands in the spectral range of 363nm to 1018nm, and a total of 19 categories. The category and number of each type of ground feature are shown in Table 3.

[0122]

[0123]

[0124] In order to verify the high efficiency and good classification performance of the present invention, three evaluation indicators, namely, overall classification accuracy OA, average accuracy AA, and Kappa coefficient, are used.

[0125] The overall classification accuracy OA refers to the ratio of the number of correctly classified pixels in the test set to the total number of pixels, and its value is between 0 and 100%. The larger the value, the better the classification effect.

[0126] The average accuracy AA refers to dividing the number of correctly classified pixels of each category in the test set by the total number of all pixels in that category to obtain the correct classification accuracy of that category, and taking the average of the accuracy of all categories. Its value is between 0 and 100%. The larger the value, the better the classification effect.

[0127] The Kappa coefficient is an evaluation index defined on the confusion matrix. It comprehensively considers the elements on the diagonal of the confusion matrix and the elements deviating from the diagonal, and more objectively reflects the classification performance of the algorithm. The value of the Kappa coefficient is between -1 and 1. The larger the value, the better the classification effect.

[0128] 2. Simulation experiment content and result analysis:

[0129] Simulation 1: The present invention and two prior arts are tested on the University of Pavia hyperspectral dataset for classification. The results are shown in the figure below. Figure 5 As shown, where:

[0130] Figure 5 (a) Classification results of the existing RN-FSC method on the University of Pavia hyperspectral dataset;

[0131] Figure 5 (b) is the classification result of the existing DCFSL method on the University of Pavia hyperspectral dataset;

[0132] Figure 5 (c) is the classification result of the method of the present invention on the University of Pavia hyperspectral dataset.

[0133] from Figure 5 (c) It can be seen that the classification result of the present invention on the University of Pavia dataset is significantly better than Figure 5 (a) Figure 5 (b) Smoother and with sharper edges.

[0134] Simulation 2, the present invention and two prior arts are tested on Indian Pines hyperspectral dataset, and the simulation results are shown in the figure below: Figure 6 As shown, where:

[0135] Figure 6 (a) Classification results of the existing RN-FSC method on the Indian Pines hyperspectral dataset;

[0136] Figure 6 (b) is the classification result of the existing DCFSL method on the Indian Pines hyperspectral dataset;

[0137] Figure 6 (c) is the classification result of the method of the present invention on the Indian Pines hyperspectral dataset;

[0138] The classification accuracies of the present invention and the prior art in the above two simulations are compared on the University of Pavia hyperspectral dataset and the Indian Pines hyperspectral dataset, respectively, and the results are shown in Table 4.

[0139] Table 4 Comparison of classification accuracy of three networks under two different data sets

[0140]

[0141]

[0142] It can be seen from Table 4 that the method of the present invention achieves higher classification accuracy than the prior art RN-FSC method and DCFSL method under the University of Pavia and Indian Pines datasets, indicating that the present invention can more accurately predict the category of hyperspectral image samples.

[0143] The above simulation experiments show that the method of the present invention can more fully extract inter-spectral features by using the constructed inter-spectral 3D deformable convolution block, and the constructed multi-scale spatial feature extraction block can more fully extract spatial features. And the spatial features and inter-spectral features are spliced, and then more distinguishable spectral-spatial features can be obtained through the fully connected layer, and finally the hyperspectral image classification results are obtained through the softmax classifier. The present invention uses a total loss function composed of a cross entropy loss function, a correlation alignment loss function, and a maximum mean difference loss function to train the neural network, so that the small sample spectral-spatial feature extraction convolutional neural network pays more attention to the category of objects with unconcentrated sample distribution or a small sample size. It solves the problem of low classification accuracy in the case of few training samples due to insufficient utilization of the extracted features or information redundancy caused by retaining too much irrelevant information in the prior art during spatial feature extraction. It is a very practical hyperspectral image classification method for few training samples.

[0144] In the description of the invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the invention, "plurality" means two or more, unless otherwise clearly and specifically defined.

[0145] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristic data points described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristic data points described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification. The above content is a further detailed description of the present invention in conjunction with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be deemed to belong to the scope of protection of the present invention.

Claims

1. A hyperspectral image classification method based on small sample deep learning, characterized in that: The hyperspectral image classification method comprises: Step 1: obtaining a source domain hyperspectral image set and a target domain hyperspectral image set, wherein the source domain hyperspectral image set includes a plurality of source domain hyperspectral images, and the target domain hyperspectral image set includes a plurality of target domain hyperspectral images; Step 2, filling the edge parts of the source domain hyperspectral image and the target domain hyperspectral image respectively, and extracting a plurality of first image blocks and a plurality of second image blocks with each pixel point of the source domain hyperspectral image and the target domain hyperspectral image after the filling process as the center; Step 3: In each category, randomly select part of the first image blocks to form a source domain support set and part of the first image blocks to form a source domain query set, and randomly select part of the second image blocks to form a target domain support set and part of the first image blocks to form a target domain query set; Step 4: Based on the stochastic gradient descent method, the source domain support set, the source domain query set, the target domain support set and the target domain query set are used to alternately train the small sample spectral-spatial feature extraction convolutional neural network to obtain a trained small sample spectral-spatial feature extraction convolutional neural network, wherein the small sample spectral-spatial feature extraction convolutional neural network extracts features from both spectral and spatial aspects, and the total loss function of the small sample spectral-spatial feature extraction convolutional neural network consists of a cross entropy loss function, a correlation alignment loss function and a maximum mean difference loss function; Step 5: input the hyperspectral image to be classified into the trained small sample spectral and spatial feature extraction convolutional neural network to obtain a classification result; Among them, the structure of the small sample spectral-spatial feature extraction convolutional neural network includes a spectral branch network, a spatial branch network, a domain attention module, a first splicing layer, a fully connected layer and a softmax classifier. The spectral branch network and the spatial branch network are connected in parallel and then connected in series with the domain attention module in sequence. The domain attention module, the first splicing layer, the fully connected layer and the softmax classifier are connected in series. The spectral branch network includes 2 3D deformable convolution blocks and 2 first maximum pooling layers. The first 3D deformable convolution block, the first first maximum pooling layer, the second 3D deformable convolution block, and the second first maximum pooling layer are connected in series in sequence. The spatial branch network domain attention module includes a multi-scale spatial feature extraction module, a second splicing layer, and a second maximum pooling layer connected in series in sequence. The domain attention module adopts a 2D convolutional layer, which includes 1 spectral attention module and 2 spatial attention modules. The spectral attention module is 1 2D convolutional layer, and the 2 spatial attention modules are 2 2D convolutional layers. The three 2D convolutional layers are connected in series.

2. The hyperspectral image classification method based on small sample deep learning according to claim 1 is characterized in that: The step 2 comprises: Step 2.1, filling pixels with a pixel value of 0 around the source domain hyperspectral image and the target domain hyperspectral image, respectively, to obtain a filled source domain hyperspectral image and a filled target domain hyperspectral image; Step 2.2: Taking each pixel point in the padded source domain hyperspectral image and the padded target domain hyperspectral image as the center, select a space with a size of (2t+1)×(2t+1) and a number of channels. d The first image block and the second image block are as follows, wherein t is an integer greater than 0.

3. The hyperspectral image classification method based on small sample deep learning according to claim 1 is characterized in that: The 3D deformable convolution block includes three 3D deformable convolution layers and three first activation function layers. The first 3D deformable convolution layer, the first first activation function layer, the second 3D deformable convolution layer, the second first activation function layer, the third 3D deformable convolution layer, and the third first activation function layer are connected in series in sequence, and the output of the first first activation function layer is added to the output of the third 3D deformable convolution layer to form a residual structure.

4. The hyperspectral image classification method based on small sample deep learning according to claim 1 is characterized in that: The multi-scale spatial feature extraction module includes 2 scale operation layers, 3 convolution layers, 3 normalization layers, and 3 second activation function layers; The first convolution layer, the first normalization layer, and the first second activation function layer are connected in series in sequence; The first scale operation layer, the second convolution layer, the second normalization layer, and the second second activation function layer are sequentially connected in series; The second scale operation layer, the third convolution layer, the third normalization layer, and the third second activation function layer are sequentially connected in series; The first second activation function layer, the second second activation function layer, and the third second activation function layer are connected in parallel and then connected in series with the second concatenation layer and the second maximum pooling layer in sequence.

5. The hyperspectral image classification method based on small sample deep learning according to claim 1 is characterized in that: The total loss function of the small sample spectral space feature extraction convolutional neural network is: L total =L fsl +L coral +L MMD Among them, L total represents the total loss function of the convolutional neural network for small sample spectral feature extraction, L fsl represents the cross entropy loss, L coral represents the correlation alignment loss, L MMD represents the maximum mean difference loss; Cross entropy loss L fsl It is expressed as: Where d(·) represents the Euclidean distance, F ω (·) represents the feature extraction function with parameter ω, f l represents the feature of the lth category in the source domain support set or the target domain support set, C represents the number of categories, and x j represents a sample in the source domain query set or the target domain query set, y j Represents sample x j The label of , Q represents the source domain query set or the target domain query set; Correlation alignment loss L coral It is expressed as: in, represents the Frobenius norm of the matrix, C S Represents the covariance matrix of the source domain features, C T Represents the covariance matrix of the target domain features, and d represents the dimension of the features; Maximum mean difference loss L MMD It is expressed as: in, represents the spatial distance, φ(·) represents the mapping function, represents the source domain features, represents the target domain features, X s represents the source domain dataset, X t represents the target domain dataset, n s Indicates the number of data in the source domain dataset, n t Indicates the number of data in the target domain dataset.

6. The hyperspectral image classification method based on small sample deep learning according to claim 5 is characterized in that: The step 4 comprises: The initial learning rate of training is set to α, and the number of iterations is set to T. In odd-numbered iterations, the source domain support set and the source domain query set are sent to the small sample spectral-space feature extraction convolutional neural network for training, and the total loss function is used to calculate the loss value between the features of the source domain support set and the features of the source domain query set to update the parameters in the small sample spectral-space feature extraction convolutional neural network; in even-numbered iterations, the target domain support set and the target domain query set are sent to the small sample spectral-space feature extraction convolutional neural network for training, and the total loss function is used to calculate the loss value between the features of the target domain support set and the features of the target domain query set to update the parameters in the small sample spectral-space feature extraction convolutional neural network until the loss value of the small sample spectral-space feature extraction convolutional neural network no longer decreases and the current number of training rounds is less than the number of iterations T or the number of training rounds reaches the number of iterations T, then the training of the small sample spectral-space feature extraction convolutional neural network is stopped to obtain the trained small sample spectral-space feature extraction convolutional neural network.

7. The hyperspectral image classification method based on small sample deep learning according to claim 6 is characterized in that: The weight vector W after the small sample spectral feature extraction convolutional neural network is updated new for: Among them, L total represents the total loss function of the small sample spectral space feature extraction convolutional neural network, W represents the weight vector of the small sample spectral space feature extraction convolutional neural network before updating, and R represents the learning rate.

Citation Information

Patent Citations

  • Hyperspectral image classification method based on multi-scale spectral space convolutional neural network

    CN111639587A

  • Semantic convolution hyperspectral image classification method based on multi-path attention mechanism

    CN112052755A