Hyperspectral Image Classification Method Based on Pseudo-Label Aiding
By optimizing the alignment strategy in hyperspectral image classification and constructing the class-perceived maximum mean difference loss function, combined with pseudo-label assisted and reweighted weight pruning technology, the spectral displacement and portability problems in the cross-domain classification of hyperspectral images are solved, achieving efficient cross-domain classification and improved portability.
Patent Information
- Application Number
- CN202311004129.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-10
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-08-10
AI Technical Summary
The existing hyperspectral image classification methods are difficult to effectively solve the spectral displacement problem during cross-domain classification, resulting in a decrease in classification accuracy. In addition, traditional methods require that the last few layers of the deep network be linear layers, limiting the universality and portability of the method.
By optimizing the alignment strategy, a hyperspectral image classification method based on pseudo-label assisted is proposed, a class-perceived maximum mean difference loss function is constructed, the update process of deep neural networks and classifiers is optimized, and the target domain samples are pruned through the reweighted weight of the target samples, and high-quality pseudo-labels are screened to improve cross-domain transfer learning ability.
The cross-domain classification performance and portability of hyperspectral images are improved. By accurately aligning the use of subdomains and high-quality pseudo-labels, the feature extraction ability of deep neural networks and classifier discrimination ability are improved.
Smart Images

Figure CN117011714B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a hyperspectral image classification method, which can be used in intelligent agriculture and environmental monitoring. Background Art
[0002] Hyperspectral image (HSI) is an important type of remote sensing data, which is widely used in the field of earth observation, such as intelligent agriculture and environmental monitoring. HSI contains hundreds of bands, providing rich information for accurate classification of ground object categories. In recent years, many supervised learning methods for HSI classification have been proposed, including support vector machines, sparse representation, and convolutional neural networks. Supervised learning methods usually require a large number of labeled samples to maintain high classification accuracy. However, it is very difficult, time-consuming, and expensive to obtain sufficient labeled HSI. In addition, there are spectral shifts between HSIs at different times and different locations. Therefore, traditional classification models trained on one image and tested on another image with spectral shifts cannot achieve satisfactory results. To solve the problem of cross-domain remote sensing image classification, domain adaptation (DA) technology has been widely applied to cross-scene classification tasks.
[0003] With the rapid development of software and hardware levels, deep neural networks have been widely applied to DA due to their excellent ability to automatically extract hierarchical features and generate accurate feature representations, which can help improve the non-linear mapping ability of traditional domain adaptation methods for HSIs in different domains, thus better aligning the source domain and the target domain to achieve cross-scene classification. According to the processing method, the existing DA processing methods are mainly divided into feature adaptation-based methods and adversarial-based methods, among which:
[0004] A method based on feature adaptation, which mainly matches the marginal distribution or / and conditional distribution between the main matching domains, adds an adaptation layer to the original deep network architecture to achieve the adaptation of the source domain to the target domain, so as to achieve the representation for the task. Li et al. published an article "Adversarial Discriminative Active DeepLearning for Domain Adaptation in Hyperspectral Images Classification" in the journal "Remote Sensing". Aiming at the problem that deep learning methods require a large amount of labeled data in hyperspectral image classification, a two-stage deep domain adaptation method TDDA was proposed. It uses very few labeled samples in the target domain, minimizes the data shift between the two domains, and learns a more discriminative deep embedding space. In the first stage, the maximum mean discrepancy MMD criterion is used to minimize the distance between the source domain and the target domain, and a deep embedding space is learned; in the second stage, a spatial-spectral siamese network is used to reduce the data shift, and the pairwise loss is minimized to reduce the distance between samples of different domains but the same category, increase the distance between samples of different domains and categories, and learn a more discriminative deep embedding space. However, since this method directly uses the MMD loss to minimize the distribution difference between the two domains, on the one hand, it ignores the "same spectrum, different object" phenomenon of HSI images, resulting in samples with similar spectral characteristics in different categories of the source domain and the target domain being aligned into the same category, causing a decrease in classification accuracy; on the other hand, it is necessary to require the last few layers of the deep network to be linear layers, which is not conducive to the generality and portability of the method.
[0005] The adversarial-based method is to learn transferable and domain-invariant features through adversarial learning to minimize the cross-domain difference. Ma et al. published an article "Cross-Dataset Hyperspectral Image Classification Based on AdversarialDomain Adaptation" in the journal "IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING", proposing an adversarial domain adaptation method ADA-Net. Aiming at the problem that a large amount of labeled data is required in hyperspectral image classification, it uses the labeled information of other datasets and the pseudo-labels predicted by the network to minimize the data shift between the two domains, and learns a more discriminative deep embedding space. A discriminator is constructed using multiple classifiers, and a generator is composed of a variational autoencoder to drive the classification of target samples with the support of the source domain in an adversarial manner. However, since this method directly uses the pseudo-labels with a high noise rate predicted by the network as an auxiliary to minimize the data shift between the two domains, the cross-domain classification performance decreases. Summary of the Invention
[0006] The object of the present invention is to propose a hyperspectral image classification method based on pseudo-label assistance for the deficiencies of the above-mentioned existing technologies, so as to improve the cross-domain classification performance and portability of hyperspectral images.
[0007] The technical idea for achieving the object of the present invention is: by optimizing the alignment strategy, providing a sub-domain adaptation method that is easy to transplant, solving the situation where the last few layers of the deep network are required to be linear layers in the conventional sub-domain adaptation method, improving the portability of the method, and by evaluating the quality of the model pseudo-labels, that is, only selecting high-quality pseudo-labels to assist the cross-domain transfer learning ability of the model, improving the cross-domain classification performance.
[0008] According to the above idea, the technical solution of the present invention includes the following steps:
[0009] (1) Obtain two datasets with partially the same categories from the publicly available hyperspectral dataset, and use one of them as the source domain dataset SD with true labels, and the other as the target domain dataset TD without true labels;
[0010] (2) Input the samples of the source domain dataset SD and the target domain dataset TD into the deep neural network F respectively, and obtain the source domain output feature map f of the last layer of the network s , the target domain output feature map f t ;
[0011] (3) Input the source domain dataset SD output feature map f s and the target domain TD output feature map f t into the double classifiers C1 and C2 respectively, and obtain two prediction probabilities of the source domain SD samples two prediction probabilities of the target domain TD samples
[0012] (4) Input the two prediction probabilities of the source domain SD samples and the two prediction probabilities of the target domain TD samples into the classification loss function L cls , and then input the two prediction probabilities of the target domain TD samples into the classifier difference loss function L td , and update the deep neural network F and the double classifiers C1 and C2 by optimizing the two loss functions, so as to improve the feature extraction ability of the deep neural network F for samples and the discriminant classification ability of the classifiers C1 and C2, and obtain the first updated deep neural network F' and double classifiers C1' and C2';
[0013] (5) According to the source domain SD sample feature map f sand its corresponding true label y s and the feature map f of the target domain TD samples t Construct the class-aware maximum mean discrepancy loss function L c ,
[0014]
[0015] where δ(·) is the indicator function, k is the class serial number, n s represents the number of source domain samples, n t represents the number of source domain samples, if then otherwise is the layer parameter of the embedding space for converting the source domain and the target domain, is the feature map of the i-th source domain sample, is the true label of the i-th source domain sample, is the feature map of the j-th target domain sample, is the cosine similarity class of the target domain TD samples, C is the total number of classes, and H represents the Hilbert space;
[0016] (6) Minimize the class-aware maximum mean discrepancy loss function L c to update the deep neural network F′ and the binary classifiers C1′ and C2′ after the first update, and obtain the deep neural network F″ and the binary classifiers C1″ and C2″ after the second update;
[0017] (7) Update the target domain samples:
[0018] 7a) Maintain the feature maps f of all source domain SD samples s as a memory set M:
[0019] where is the feature map of the i-th source domain SD sample, is the true label of the i-th source domain SD sample;
[0020] 7b) Randomly select some feature maps and corresponding labels from M category by category as the projection o s (x s ) of the source domain samples SD, and take the feature map f of the target domain TD samples t as the projection o t (x t ) of the target domain TD;
[0021] 7c) Propagate the source domain projection o s (x s ) and its corresponding label to the target domain projection o t (x t) In it, obtain the label propagation probability Ψ of the target domain TD samples * ;
[0022] 7d) According to the two prediction probabilities of the target domain TD samples Calculate the reweighted weight r of the target samples:
[0023] r = max(Ψ * ·P t )
[0024] Where
[0025] 7e) Prune the target domain TD samples according to the reweighted weight r of the target samples, and obtain the pseudo-labels of the target domain TD samples
[0026]
[0027] Where I(·) is the indicator function, τ is the weight threshold, and N t is the total number of samples in the target domain dataset;
[0028] 7f) According to the pruned pseudo-labels Obtain the corresponding target domain credible samples And combine these two to form a new target domain dataset To replace the original target domain dataset TD and obtain a new target domain dataset TD′;
[0029] (8) Input the new target domain TD′ samples into the second updated deep neural network F″ and the two classifiers C1″ and C2″ in sequence, obtain their two prediction probabilities G1, G2, calculate the average probability G = (G1 + G2) / 2, and take the category corresponding to the maximum value in G as the classification result of the target domain samples.
[0030] Compared with the prior art, the present invention has the following advantages:
[0031] First, since the present invention constructs a class-aware maximum mean discrepancy loss function, it can, by virtue of the characteristics of the samples themselves, match the features of the target domain samples with the prototype features of all classes in the source domain, assign cosine-similar classes to the target domain samples, and then divide the target domain samples into the common sub-domains of the source domain to minimize the class-aware maximum mean discrepancy loss function, achieving precise alignment of the sub-domains and enhancing the feature extraction ability of the deep neural network for the target domain samples and the discriminant classification ability of the two classifiers; at the same time, since this loss function does not require the last few layers of the network to be fully connected layers, it has high portability.
[0032] Second, since the target domain dataset is updated in the present invention, that is, only high-quality target domain samples and their pseudo-labels are selected to assist the cross-domain transfer learning ability of the model. Pruning is performed on the target domain samples through the reweighted weights of the target samples to obtain screened target domain credible samples and their pseudo-labels, and a new target domain dataset is formed by them. Furthermore, the new target domain dataset can be used to accelerate cross-domain learning and improve the final classification performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is the implementation flowchart of the present invention;
[0034] Figure 2 are the false color map and the true value map of the Pavia dataset used in the experiment of the present invention;
[0035] Figure 3 are the false color map and the true value map of the Houston dataset used in the experiment of the present invention;
[0036] Figure 4 are the false color map and the true value map of the HyRANK dataset used in the experiment of the present invention;
[0037] Figure 5 is the classification result graph of the experiment of the present invention and five existing classification methods on the Pavia dataset;
[0038] Figure 6 is the classification result graph of the experiment of the present invention and five existing classification methods on the Houston dataset;
[0039] Figure 7 is the classification result graph of the experiment of the present invention and five existing classification methods on the HyRANK dataset. DETAILED DESCRIPTION OF THE INVENTION
[0040] The examples and effects of the present invention will be further described in detail below with reference to the drawings;
[0041] Refer to Figure 1 , the implementation steps of this example are as follows:
[0042] Step 1, construct a hyperspectral cross-domain dataset.
[0043] Obtain two hyperspectral datasets from a public website as a pair of datasets, and use one of the datasets as the source domain dataset SD and the other dataset as the target dataset TD.
[0044] Three pairs of public hyperspectral datasets are obtained in this example, where:
[0045] The first pair of datasets is the Pavia dataset, which includes two datasets, Pavia University and Pavia Center. These two datasets share 7 land cover classes. The Pavia University dataset is used as the source domain dataset, and the Pavia Center dataset is used as the target domain dataset. The detailed information is shown in Table 1, and the false color maps and ground truth labels are shown in Figure 2 , where Figure 2 (a) is the false color map of the Pavia University dataset, Figure 2 (b) is the ground truth label of the Pavia University dataset, Figure 2 (c) is the false color map of the Pavia Center dataset, Figure 2 (d) is the ground truth label of the Pavia Center dataset;
[0046] The second pair of datasets is the Houston dataset, which includes two datasets, Houston 2013 and Houston 2018. These two datasets share 7 land cover classes. The Houston 2013 dataset is used as the source domain dataset, and the Houston 2018 dataset is used as the target domain dataset. The detailed information is shown in Table 2, and the false color maps and ground truth labels are shown in Figure 3 , where Figure 3 (a) is the false color map of the Houston 2013 dataset, Figure 3 (b) is the ground truth label of the Houston 2013 dataset, Figure 3 (c) is the false color map of the Houston 2018 dataset, Figure 3 (d) is the ground truth label of the Houston 2018 dataset;
[0047] The third pair of datasets is the HyRANK dataset, which includes two datasets, Dioni and Loukia. These two datasets share 12 land cover classes. The Dioni dataset is used as the source domain dataset, and the Loukia dataset is used as the target domain dataset. The detailed information is shown in Table 3, and the false color maps and ground truth labels are shown in Figure 4 , where Figure 4 (a) is the false color map of the Dioni dataset, Figure 4 (b) is the ground truth label of the Dioni dataset, Figure 4 (c) is the false color map of the Loukia dataset, Figure 4 (d) is the ground truth label of the Loukia dataset;
[0048] Table 1 Number of source domain data samples and target domain data samples for each land cover class in the Pavia data
[0049] Category number Category name Number of source domain samples Number of target domain samples 1 Tree 3064 7598 2 Asphalt 6631 9248 3 Brick 3682 2685 4 Bitumen 1330 7287 5 Shadow 947 2863 6 Meadow 18649 3090 7 Bare soil 5029 6584
[0050] Table 2 Source domain data sample numbers and target domain data sample numbers for each land cover class of Houston data
[0051] Category number Category name Number of source domain samples Number of target domain samples 1 Grass healthy 345 1353 2 Grass stressed 365 4888 3 Trees 365 2766 4 Water 285 22 5 Residential buildings 319 5347 6 Non-residential buildings 408 32459 7 Road 443 6365
[0052] Table 3 Source domain data sample numbers and target domain data sample numbers for each land cover class of HyRANK data
[0053]
[0054]
[0055] All samples used in this example are randomly selected 180 samples from the source domain dataset SD and the target domain dataset TD, and
[0056] The selected source domain labeled samples are used to form the source domain set The selected target domain samples are used to form the target domain set Wherein:
[0057] Where n s is the total number of source domain samples extracted, and N t is the total number of target domain samples extracted, are the hyperspectral pixels with d bands in the source domain and the target domain respectively, is the label {1, 2, 3,..., C} corresponding to the source domain labeled samples, and C is the number of classes.
[0058] Step 2, select a deep neural network F and two classifiers C1 and C2 with the same structure.
[0059] Select a deep neural network F from existing deep neural networks, which includes one-dimensional convolutional neural network, two-dimensional convolutional neural network, three-dimensional convolutional neural network and graph neural network; select a classifier from existing classifiers, which includes softmax classifier, classifier composed of fully connected layers or other classifiers with updatable parameters.
[0060] In this example, but not limited to, a two-dimensional convolutional neural network F is selected, which includes a six-layer stacked structure, wherein:
[0061] The first layer is composed of a first convolutional layer, a first batch normalization layer and a ReLU activation function cascaded in sequence;
[0062] The second layer is composed of a second convolutional layer, a second batch normalization layer and a ReLU activation function cascaded in sequence;
[0063] The third layer is composed of a third convolutional layer, a third batch normalization layer, and a ReLU activation function cascaded in sequence;
[0064] The fourth layer is composed of a convolutional layer, a batch normalization layer, and a ReLU activation function cascaded in sequence;
[0065] The fifth layer is an average pooling layer with a pooling kernel size of 5×5 and a stride of 1;
[0066] The sixth layer is a fully connected layer with an input dimension of 200 and an output dimension of 200;
[0067] The convolutional kernel size of each convolutional layer is 1×1, the stride is 1, the padding is 0, and the input and output channel sizes are both 200.
[0068] In this example, two same-structured binary classifiers C1 and C2 composed of fully connected layers are selected but not limited to. Each classifier includes a three-layer stacked structure, specifically:
[0069] The first layer is composed of a first fully connected layer, a first batch normalization layer, and a ReLU activation function cascaded in sequence;
[0070] The second layer is composed of a second fully connected layer, a second batch normalization layer, and a ReLU activation function cascaded in sequence;
[0071] The third layer is the third fully connected layer;
[0072] The input dimension and output dimension of the first fully connected layer and the second fully connected layer are both 200. The input dimension of the third fully connected layer is 200, and the output dimension is C, where this output dimension represents the number of common classes in the source domain SD and the target domain TD.
[0073] Step 3: Extract the features and their respective two prediction probabilities of the source domain and the target domain respectively.
[0074] 3.1) Input the source domain dataset SD and the target domain dataset TD into the deep neural network F for feature extraction respectively to obtain the source domain feature map f s and the target domain feature map f t ;
[0075] 3.2) Input the source domain feature map f s , the target domain feature map f t into the binary classifiers C1 and C2 respectively for prediction to obtain the two prediction probabilities of the source domain and the two prediction probabilities of the target domain
[0076] Step 4: Update the deep neural network F and the binary classifiers C1 and C2 using the source domain and target domain prediction probabilities.
[0077] (4.1) Input the two predicted probabilities in the source domain and the two predicted probabilities in the target domain into the existing classification loss function and the classifier difference loss function simultaneously, to obtain the classification loss function \(L\) of the source domain and the target domain cls and the difference loss function \(L\) td , which are respectively expressed as follows:
[0078]
[0079]
[0080] Among them, \(p1\) represents the predicted probabilities of the source domain and target domain samples participating in the calculation output by the first classifier \(C1\), \(p2\) represents the predicted probabilities of the source domain and target domain samples participating in the calculation output by the second classifier \(C2\), \(y\) represents the true labels of the source domain and target domain samples participating in the calculation, \(N\) represents the total number of samples in the source domain and target domain participating in the calculation, \(y\) i represents the true label of the \(i\)-th sample, \(p\) 1,i represents the predicted probability of the \(i\)-th sample output by the first classifier \(C1\), \(p\) 2,i represents the predicted probability of the \(i\)-th sample output by the second classifier \(C2\);
[0081] \(m,n\in\{1,2,3,\cdots,K\}\), \(K\) is the number of categories, \(A\) is a probability correlation matrix with a diagonal of 0, of size \(K\times K\), used to represent the difference between classifiers, \(A\) mn represents the element in the \(m\)-th row and \(n\)-th column of matrix \(A\). \(N\) t represents the number of target domain samples, are the two predicted probabilities of the target domain TD samples;
[0082] (4.2) Optimize the above classification loss function \(L\) cls and the difference loss function \(L\) td to obtain the first updated deep neural network \(F'\) and the dual classifiers \(C1'\) and \(C2'\):
[0083] (4.2.1) Minimize the classification loss function \(L\) cls , that is, update the parameters \(\theta\) of \(F\) F , the parameters of classifier \(C1\) and the parameters of classifier \(C2\) through backpropagation, expressed as
[0084] (4.2.2) First maximize the difference loss function \(L\) td , that is, update the parameters \(\theta\) of classifier \(C1\) through backpropagationC1 and the parameters of classifier C2 are updated, denoted as Then minimize it, that is, the parameters θ of F are updated through backpropagation F are updated, denoted as
[0085] Finally, the updated deep neural network F′ and dual classifiers C1′ and C2′ after the first update are obtained.
[0086] Step 5, use the source domain and target domain feature maps to perform a second update on the once-updated deep neural network F′ and dual classifiers C1′ and C2′.
[0087] (5.1) According to the source domain SD sample feature map f s and its corresponding true label y s and the target domain TD sample feature map f t , construct a class-aware maximum mean discrepancy loss function L c :
[0088] (5.1.1) Calculate the mean of the feature map f of the source domain SD s to obtain the prototype vector Q of each class sample in the source domain SD k :
[0089]
[0090] where δ(·) is the indicator function, k represents the class serial number; if then otherwise Mean(·) represents calculating the mean, is the feature map of the i-th source domain sample, is the true label of the i-th source domain sample;
[0091] (5.1.2) Calculate the cosine similarity between the feature map f of each target domain TD sample t and the prototype vector of the source domain SD to obtain the cosine similarity label of this target domain sample:
[0092]
[0093] where Q is the prototype vector of all source domain classes, represents the cosine similarity label of the i-th target domain sample;
[0094] (5.1.3) According to the source domain SD sample feature map f s and its corresponding true label y s and the target domain TD sample feature map f t and its corresponding cosine similarity label qt , the following class-aware maximum mean discrepancy loss function \(L\) is constructed c :
[0095]
[0096] where \(C\) is the total number of classes, \(n\) s represents the number of source domain samples, \(n\) t represents the number of source domain samples, \(i\) and \(j\) represent the sample indices, is the feature map of the \(i\)-th source domain sample, is the true label of the \(i\)-th source domain sample, is the feature map of the \(j\)-th target domain sample, is the cosine similarity label of the \(j\)-th target domain sample, \(\varphi(\cdot)\) is the layer parameter of the embedding space that transforms the source domain and the target domain, \(H\) represents the Hilbert space, and \(k(\cdot,\cdot)\) is the Gaussian kernel function;
[0097] (5.2) Minimize the above class-aware maximum mean discrepancy loss function \(L\) c , that is, update the parameters \(\theta\) of \(F'\) through backpropagation F′ , the parameters of the classifier \(C1'\) and the parameters of the classifier \(C2'\) , which is expressed as to obtain the updated deep neural network \(F''\) and the dual classifiers \(C1''\) and \(C2''\) after the second update.
[0098] Step 6, update the target domain samples.
[0099] (6.1) Maintain the feature maps \(f\) of all source domain \(SD\) samples s into a memory set \(M\):
[0100] where is the feature map of the \(i\)-th source domain \(SD\) sample, is the true label of the \(i\)-th source domain \(SD\) sample;
[0101] (6.2) Randomly select some feature maps and corresponding labels from \(M\) for each class as the projection \(o\) of the source domain sample \(SD\) s (x s ), and use the feature map \(f\) of the target domain \(TD\) sample t as the projection \(o\) of the target domain \(TD\) t (x t );
[0102] (6.3) Propagate the source domain projection \(o\) s (x s ) and its corresponding label to the target domain projection \(o\) t (x t)In it, obtain the label propagation probability Ψ of the target domain TD samples * , and its label propagation process is as follows:
[0103] (6.3.1) Calculate the source domain projection o s (x s ) and the target domain projection o t (x t )'s cosine similarity to obtain the cosine similarity matrix U:
[0104] U = cos([o s (x s ), o t (x t )], [o s (x s ), o t (x t )])
[0105] where the size of U is (n + m) × (n + m), and U ij represents the element in the i-th row and j-th column of the matrix, n is the number of features in the source domain projection o s (x s ), and m is the number of features in the target domain projection o t (x t );
[0106] (6.3.2) Determine the symmetric adjacency matrix Ω with diagonal elements being 0 according to the cosine similarity matrix U:
[0107]
[0108] where, Ω ij represents the element in the i-th row and j-th column of the matrix, and λ is the decision threshold for whether to construct an adjacency relationship;
[0109] (6.3.3) Symmetrically normalize Ω to obtain the symmetrically normalized adjacency matrix
[0110]
[0111] where D is the degree matrix of Ω;
[0112] (6.3.4) According to the symmetrically normalized adjacency matrix calculate the closed set solution in the label propagation LPA to obtain the label propagation probability Ψ of the target domain TD samples * :
[0113]
[0114] where α ∈ (0, 1) represents the amount of information controlling label propagation, I is the identity matrix, and y s is the projection of the source domain sample o s (x s ) corresponding true label;
[0115] (6.4) Calculate the reweighted weight r of the target sample according to the two predicted probabilities of the target domain TD samples :
[0116] r = max(Ψ * ·P t )
[0117] where
[0118] (6.5) Prune the target domain TD samples according to the reweighted weight r of the target samples to obtain the pseudo-labels of the target domain TD samples
[0119]
[0120] where I(·) is the indicator function, τ is the weight threshold, and N t is the total number of samples in the target domain dataset;
[0121] (6.6) Obtain the corresponding target domain credible samples according to the pruned pseudo-labels and form a new target domain dataset with these two to replace the original target domain dataset TD to obtain a new target domain dataset TD′; Step 7, obtain the final classification result.
[0122] Input the new target domain TD′ samples into the second updated deep neural network F″ and the binary classifiers C1″ and C2″ in sequence to obtain their two predicted probabilities G1, G2;
[0123] Calculate the average probability G = (G1 + G2) / 2 of the two predicted probabilities G1, G2, and take the category corresponding to the maximum value in G as the classification result of the target domain samples.
[0124] The labels of the above steps are for a clearer description of the implementation scheme of the present invention, and their sequence numbers are not limited.
[0125] The following combines experiments to further illustrate the technical effects of the present invention.
[0126] 1. Experimental conditions
[0127]
[0128] The hardware platform for the experiment is as follows: the processor is Intel Xeon E5-2698 v3 (2.3 GHz), the memory is 128 GB, and the graphics card is NVIDIA GeForce RTX 3090 Ti 24 GB. The operating system is ubuntu 20.04. The software platform is python 3.8.11 and torch 1.10.0.
[0129] 2. Experimental parameters:
[0130] The datasets are the Pavia dataset, Houston dataset, and HyRANK dataset obtained in Step 1. The sample size in each dataset is set to 5×5, and the decision threshold λ for label propagation is set to 0.6.
[0131] For the Pavia dataset, τ is set to 0.85;
[0132] For the Houston dataset, τ is set to 0.60;
[0133] For the HyRANK dataset, τ is set to 0.55.
[0134] All experimental results are the average of 10 experiments.
[0135] 3. Experimental content and results:
[0136] Experiment 1: Classify the 7 categories of the Pavia dataset described in Table 1 using the present invention and five existing hyperspectral image classification methods, namely DAN, DeepCoral, DSAN, TSTnet, and CLDA. The results are as Figure 5 , and calculate the three evaluation metrics, namely the overall accuracy OA, average accuracy AA, and Kappa coefficient K, for each classification. The results are shown in Table 4.
[0137] Table 4 Comparison of classification accuracies of each classification method on the Pavia dataset (%)
[0138] Evaluation index DAN DeepCoral DSAN TSTnet CLDA The present invention OA 83.28 85.12 83.20 84.49 92.35 94.25 AA 83.06 84.61 82.53 83.30 90.49 93.04 K 79.88 82.12 79.86 81.35 90.76 93.06
[0139] As can be seen from Table 4, compared with the five existing methods, the present invention has the highest values of overall accuracy OA, average accuracy AA, and Kappa coefficient K, indicating that the high-quality target domain samples and their pseudo-labels selected by the present invention can improve the cross-domain transfer learning ability and ultimately enhance the classification performance of hyperspectral images.
[0140] By comparing each view in Figure 5 with the true labels of Pavia Center in Figure 2 (d), it can be seen that the present invention is closer to the true ground object distribution, and the area of misclassification is greatly reduced, further proving the effectiveness of the present invention in hyperspectral data classification.
[0141] Experiment 2: The seven categories of the Houston dataset described in Table 2 were classified using the present invention and five existing hyperspectral image classification methods, namely DAN, DeepCoral, DSAN, TSTnet, and CLDA. The results are as Figure 6 , and the three evaluation metrics, namely the overall accuracy OA, average accuracy AA, and Kappa coefficient K, were calculated for each classification, and the results are shown in Table 5.
[0142] Table 5 Comparison of classification accuracies of various classification methods on the Houston dataset (%)
[0143] Category DAN DeepCoral DSAN TSTnet CLDA PASDA OA 58.90 59.53 57.51 71.13 64.46 72.59 AA 69.62 71.43 66.57 53.35 75.13 74.62 K 43.63 44.83 42.74 57.01 51.47 59.79
[0144] As can be seen from Table 5, compared with the five existing methods, the present invention has the highest values of overall accuracy OA, average accuracy AA, and Kappa coefficient K, indicating that the high-quality target domain samples and their pseudo-labels selected by the present invention can improve the cross-domain transfer learning ability and ultimately enhance the classification performance of hyperspectral images.
[0145] By comparing each view in Figure 6 with the true labels of Houston 2018 in Figure 3 (d), it can be seen that the present invention is closer to the true ground object distribution, and the area of misclassification is greatly reduced, further proving the effectiveness of the present invention in hyperspectral data classification.
[0146] Experiment 3: The twelve categories of the HyRANK dataset described in Table 3 were classified using the present invention and five existing hyperspectral image classification methods, namely DAN, DeepCoral, DSAN, TSTnet, and CLDA. The results are as Figure 7 , and the three evaluation metrics, namely the overall accuracy OA, average accuracy AA, and Kappa coefficient K, were calculated for each classification, and the results are shown in Table 6.
[0147] Table 6 Comparison of classification accuracies of various classification methods on the HyRANK dataset (%)
[0148]
[0149] As can be seen from Table 6, compared with the five existing methods, the present invention has the highest values of overall accuracy OA, average accuracy AA, and Kappa coefficient K, indicating that the high-quality target domain samples and their pseudo-labels selected by the present invention can improve the cross-domain transfer learning ability and ultimately enhance the classification performance of hyperspectral images.
[0150] By comparing each view in Figure 7 with Figure 4From the comparison of the true labels of Loukia in (d), it can be seen that the present invention is closer to the actual ground object distribution, and the area of misclassification is greatly reduced, further proving the effectiveness of the present invention in hyperspectral data classification.
[0151] The sources of the above existing hyperspectral image classification methods of DAN, DeepCoral, DSAN, TSTnet, and CLDA are as follows:
[0152] DAN is a method proposed in the article "Learning transferable features with deep adaptation networks" in the 2015 conference International conference on machine learning;
[0153] DeepCoral is a method proposed in the article "Deep coral: Correlation alignment for deep domain adaptation" in the 2016 conference omputer Vision–ECCV;
[0154] DSAN is a method proposed in the article "Deep subdomain adaptation network for image classification" published in the journal "IEEE transactions on neural networks and learningsystems";
[0155] TSTnet is a method proposed in the article "Topological structure and semantic information transfer network for cross-scene hyperspectral image classification" published in the journal "IEEE transactions on neural networks and learningsystems";
[0156] CLDA is a method proposed in the article "Confident learning-based domain adaptation for hyperspectral image classification" published in the journal "IEEE Transactions on Geoscience and Remote Sensing".
[0157] The above description is only a specific example of the present invention and does not constitute any limitation to the present invention. Obviously, for professionals in the field, after understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle and structure of the present invention. However, these corrections and changes based on the idea of the present invention are still within the scope of protection of the claims of the present invention.
Claims
1. A hyperspectral image classification method based on pseudo-label assistance, characterized in that It includes the following steps: (1) Obtain two datasets with partially the same categories from the publicly available hyperspectral dataset, and use one of them as the source domain dataset SD with true labels, and the other as the target domain dataset TD without true labels; (2) Input the samples of the source domain dataset SD and the target domain dataset TD into the deep neural network F respectively to obtain the source domain output feature map f of the last layer of the network s , and the target domain output feature map f t ; (3) Output the feature map f of the source domain dataset SD s and the feature map f of the target domain TD t Input them into the binary classifiers C1 and C2 respectively to obtain two prediction probabilities of the source domain SD samples Two prediction probabilities of the target domain TD samples (4) Perform one update on the deep neural network F and the binary classifiers C1 and C2: 4a) Input the two predicted probabilities of the source domain SD samples and the two predicted probabilities of the target domain TD samples into the classification loss function, and then input the two predicted probabilities of the target domain TD samples into the classifier difference loss function to obtain the classification loss function L of the source domain and the target domain cls and the difference loss function L td ; 4b) Update the deep neural network F and the binary classifiers C1 and C2 by optimizing the two loss functions L cls and L td to obtain the updated deep neural network F′ and binary classifiers C1′ and C2′ after the first update; (5) According to the source domain SD sample feature map f s and its corresponding true label y s and the target domain TD sample feature map f t construct the class-aware maximum mean discrepancy loss function L c , where δ(·) is the indicator function, k is the class serial number, n s represents the number of source domain samples, n t represents the number of source domain samples, if then δ(y i s , k) = 1, otherwise φ(·) is the layer parameter of the embedding space that transforms the source domain and the target domain, is the feature map of the i-th source domain sample, is the true label of the i-th source domain sample, is the feature map of the j-th target domain sample, is the cosine similarity class of the target domain TD sample, C is the total number of classes, and H represents the Hilbert space; (6) Minimize the class-aware maximum mean discrepancy loss function \(L\). c Update the deep neural network \(F'\) and the binary classifiers \(C1'\) and \(C2'\) after the first update to obtain the deep neural network \(F''\) and the binary classifiers \(C1''\) and \(C2''\) after the second update. (7) Update the target domain samples: 7a) Maintain the feature maps f of all source domain SD samples as a memory set M: s where f i s , is the feature map of the i-th source domain SD sample, is the true label of the i-th source domain SD sample; 7b) Randomly extract some feature maps and corresponding labels from M category by category as the projection of the source domain sample SD o s (x s ), and take the feature map f of the target domain TD sample t as the projection of the target domain TD o t (x t ); 7c) Project the source domain o s (x s ) and its corresponding label are propagated to the target domain projection o t (x t ) to obtain the label propagation probability Ψ of the target domain TD samples * ; 7d) Calculate the reweighted weight r of the target sample based on the two predicted probabilities of the target domain TD samples : r = max(Ψ * ·P t ) Among them 7e) Prune the target domain TD samples according to the reweighted weight r of the target samples to obtain the pseudo-labels of the target domain TD samples where I(·) is the indicator function, τ is the weight threshold, and N t is the total number of samples in the target domain dataset; 7f) According to the pruned pseudo-labels obtain their corresponding target domain trustworthy samples and combine these two to form a new target domain dataset to replace the original target domain dataset TD and obtain a new target domain dataset TD'; (8) Input the new target domain TD′ samples into the deep neural network F″ and the binary classifiers C1″ and C2″ after the second update in sequence, obtain their two prediction probabilities G1 and G2, calculate the average probability G = (G1 + G2) / 2, and take the category corresponding to the maximum value in G as the classification result of the target domain samples.
2. The method according to claim 1, wherein In step (2), for the deep neural network F, it includes a six-layer stacked structure, where: The first layer is composed of a first convolutional layer, a first batch normalization layer, and a ReLU activation function cascaded in sequence; The second layer is composed of a second convolutional layer, a second batch normalization layer, and a ReLU activation function cascaded in sequence; The third layer is composed of a third convolutional layer, a third batch normalization layer, and a ReLU activation function cascaded in sequence; The fourth layer is composed of a convolutional layer, a batch normalization layer, and a ReLU activation function cascaded in sequence; The fifth layer is an average pooling layer with a pooling kernel size of 5×5 and a stride of 1; The sixth layer is a fully connected layer with an input dimension of 200 and an output dimension of 200; The convolutional kernel size of each convolutional layer is 1×1, the stride is 1, the padding is 0, and the input and output channel sizes are both 200.
3. The method according to claim 1, wherein In step (3), for the binary classifiers C1 and C2, they are two classifiers with the same structure, and each classifier includes a three-layer stacked structure, where: The first layer is composed of a first fully connected layer, a first batch normalization layer, and a ReLU activation function cascaded in sequence; The second layer is composed of a second fully connected layer, a second batch normalization layer, and a ReLU activation function cascaded in sequence; The third layer is a third fully connected layer; The input dimension and output dimension of the first two fully connected layers are both 200, and the input dimension of the third fully connected layer is 200, and the output dimension is C, which represents the number of common categories in the source domain SD and the target domain TD.
4. The method according to claim 1, wherein The classification loss function L of the source domain and the target domain obtained in step 4a) cls is expressed as follows: Among them, p1 represents the predicted probabilities of the source domain and target domain samples whose outputs of the first classifier C1 participate in the calculation, p2 represents the predicted probabilities of the source domain and target domain samples whose outputs of the second classifier C2 participate in the calculation, y represents the true labels of the source domain and target domain samples participating in the calculation, N represents the total number of source domain and target domain samples participating in the calculation, y i represents the true label of the i-th sample, p 1,i represents the predicted probability of the i-th sample output by the first classifier C1, p 2,i represents the predicted probability of the i-th sample output by the second classifier C2.
5. The method according to claim 1, wherein The double-classifier difference loss function L obtained in step 4a) td , is expressed as follows: Among them, N t represents the number of target domain samples, are the two predicted probabilities of the target domain TD samples; K is the number of classes, and A is a probability correlation matrix with a diagonal of 0, with a size of K×K, which is used to represent the differences between classifiers. A mn represents the element in the m-th row and n-th column of matrix A.
6. The method according to claim 1, characterized in that, In step 4b), by optimizing the classification loss function L cls and the difference loss function L td Updating the deep neural network F and the binary classifiers C1 and C2 with the two loss functions, the implementation steps include the following: 4b1) Minimize the classification loss function L cls by backpropagation to update the parameters θ F of F, the parameters of classifier C1, and the parameters of classifier C2, denoted as 4b2) Maximize the difference loss function L td first, that is, update the parameters of classifier C1 and the parameters of classifier C2 through backpropagation, denoted as Then minimize it, that is, update the parameter θ of F F through backpropagation, denoted as Finally, obtain the deep neural network F′ and the binary classifiers C1′ and C2′ after the first update.
7. The method according to claim 1, wherein In step (5), according to the source domain SD sample feature map f s and its corresponding true label y s and the target domain TD sample feature map f t construct a class-aware maximum mean discrepancy loss function L c , and the implementation steps are as follows: 7a) Calculate the mean of the feature map f s of the source domain SD to obtain the prototype vector Q of each class sample in the source domain SD k : where δ(·) is the indicator function and k represents the class serial number; if then otherwise Mean(·) represents taking the mean, is the feature map of the i-th source domain sample, is the true label of the i-th source domain sample; 7b) Calculate the feature map f of each target domain TD sample t and the cosine similarity with the source domain SD prototype vector to obtain the cosine similarity label of this target domain sample: Among them, Q is the prototype vector of all source domain categories, and f i s is the feature map of the i-th source domain sample, is expressed as the cosine similarity label of the i-th target domain sample; 7c) According to the source domain SD sample feature map f s and its corresponding true label y s and the target domain TD sample feature map f t and its corresponding cosine similarity label q t , construct the following class-aware maximum mean discrepancy loss function L c : where C is the total number of categories, n s represents the number of source domain samples, n t represents the number of source domain samples, i and j represent the sample serial numbers, f i s is the feature map of the i-th source domain sample, is the true label of the i-th source domain sample, is the feature map of the j-th target domain sample, is the cosine similarity label of the j-th target domain sample, φ(·) is the layer parameter of the embedding space that transforms the source domain and the target domain, H represents the Hilbert space, and k(·,·) is the Gaussian kernel function.
8. The method according to claim 1, wherein In step 7c), project the source domain o s (x s ) and its corresponding label are propagated to the target domain projection o t (x t ). The implementation steps are as follows: 8a) Calculate the source domain projection o s (x s ) and the target domain projection o t (x t ) to obtain the cosine similarity matrix U: U = cos([o s (x s ), o t (x t ), [o s (x s ), o t (x t )]) where the size of U is (n + m)×(n + m), and U ij represents the element in the i-th row and j-th column of the matrix, n is the number of features in the source domain projection o s (x s ), and m is the number of features in the target domain projection o t (x t ); 8b) Determine the symmetric adjacency matrix Ω with diagonal elements being 0 according to the cosine similarity matrix U: Among them, Ω ij represents the element in the i-th row and j-th column of the matrix, and λ is the decision threshold for whether to construct an adjacency relationship; 8c) Symmetrically normalize Ω to obtain a symmetrically normalized adjacency matrix where D is the degree matrix of Ω; 8d) According to the symmetrically normalized adjacency matrix Calculate the closed-set solution in label propagation LPA to obtain the label propagation probability Ψ of the target domain TD samples * : where α∈(0,1) represents the amount of information controlling label propagation, I is the identity matrix, and y s is the projection of the source domain samples o s (x s ) corresponding true label.
Citation Information
Patent Citations
Hyperspectral image classification method based on double-classifier adversarial enhancement network
CN114723994A
Hyperspectral image field adaptive method based on virtual classifier
CN115410088A