Hyperspectral Image Domain Adaptive Classification Method Based on Self-Training and Domain Adversarial

By constructing a debiased self-training domain adversarial adaptive model, combining multiple loss functions and gradient inversion layers, the pseudo-label noise and inter-domain alignment problems in hyperspectral image classification are solved, and classification accuracy and migration performance are improved.

CN117253074BActive Publication Date: 2025-07-29XIDIAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311017594.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-14
Publication Date
2025-07-29
Estimated Expiration
2043-08-14

AI Technical Summary

Technical Problem

The existing hyperspectral image classification algorithm has problems such as high pseudo-label noise and low classification accuracy in unsupervised domain adaptation, and the existing methods cannot effectively perform inter-domain alignment and knowledge migration.

Method used

A debiased self-training domain-adaptive adaptive model including feature extractor, three classifiers and gradient inversion layer was constructed. Combined with cross entropy loss, self-training loss and domain confusion loss, the joint distribution alignment of the source domain and the target domain through domain adversarial learning and debiased self-training methods is achieved, which alleviates pseudo-label noise and improves classification performance.

Benefits of technology

By debiased self-training domain against adaptive models, pseudo-label noise is reduced, the accuracy of hyperspectral image classification and migration generalization ability are improved, and higher classification accuracy and stability are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117253074B_ABST
    Figure CN117253074B_ABST
Patent Text Reader

Abstract

The present invention discloses a hyperspectral image domain adaptive classification method based on self-training and domain adversarial, which mainly solves the problems of large pseudo-label noise and low classification accuracy in the prior art. The implementation solution is as follows: obtain the source domain and target domain datasets; construct a debiased self-training domain adversarial adaptive model composed of a feature extractor, three classifiers and a gradient reversal layer, and define a loss function composed of cross-entropy loss, self-training loss, worst-case adversarial loss and domain confusion loss; calculate the loss value of the debiased self-training domain adversarial adaptive model, substitute it into the chain rule to calculate the gradients of each parameter of the model, and update the model parameters until the specified number of rounds is reached to complete the model training; input the target domain dataset into the trained model to obtain the classification result. The present invention reduces the noise of pseudo-labels and improves the hyperspectral domain adaptive classification accuracy, and can be used for the recognition of ground object targets in land planning, ecological protection and municipal construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and further relates to a hyperspectral image classification method, which can be used for identifying ground object targets in land planning, ecological protection, and municipal construction. Background Art

[0002] Hyperspectral images have high spatial and spectral resolutions. They contain dozens or even hundreds of bands, which can capture a large amount of spectral details, thus enabling more accurate ground object classification. This cannot be achieved in natural images or multispectral images. Therefore, hyperspectral images are widely used in fields such as agriculture and environmental protection. In recent years, many deep learning algorithms such as convolutional neural networks, recurrent neural networks, and attention mechanisms have been widely applied to hyperspectral image classification tasks. However, most deep learning classification algorithms require a large amount of labeled datasets, which requires time-consuming and laborious labeling work. In addition, due to factors such as collection locations, acquisition conditions (such as changes in lighting, seasons, and differences in sensors), there are band shifts between different hyperspectral datasets, which do not conform to the assumption of independent and identical distribution of the training set and the test set in traditional deep learning models. Therefore, the classification accuracy and robustness of the model are reduced, and the migration and generalization ability of the model is limited. Unsupervised domain adaptation aims to achieve the migration from a labeled source domain dataset to an unlabeled target domain dataset, and its core lies in reducing the distribution differences between datasets. Unsupervised domain adaptation algorithms based on statistical alignment, adversarial learning, and semi-supervised algorithms have all achieved good performance. However, none of the above domain adaptation methods attempt to combine various methods - using domain adversarial methods for domain alignment while using semi-supervised self-training to train the model with pseudo-labels step by step, resulting in insufficient domain alignment or noisy pseudo-labels, and causing poor classification performance.

[0003] China University of Mining and Technology disclosed "A Hyperspectral Image Domain Adaptation Method Based on Virtual Classifiers" in the patent document with the application number CN202211235431.7. Firstly, it clusters the samples in the source domain and the target domain into several clusters through a clustering algorithm; then, by aligning the cluster centers of the source domain and the target domain, it reduces the risk of aligning outlier source domain samples with target domain samples; next, it uses a similarity matrix to calculate the soft prototype contrast loss, thereby reducing the impact of noisy pseudo-labels. In addition, this method also constructs a virtual classifier to perform classification based on feature similarity measurement. By reducing the disagreement between the real and virtual classifiers, it encourages cross-domain samples with similar features to be assigned to the same class; finally, it reduces the overall distribution difference between domains through a domain adversarial strategy. The advantages of this method lie in aligning the cluster centers of the target domain samples and the source domain, reducing the risk of aligning outlier source domain samples with target domain samples; at the same time, it can assign corresponding confidence coefficients to each positive and negative sample, thereby reducing the impact of noisy pseudo-labels. However, since this method uses a similarity matrix to calculate the soft prototype contrast loss, it needs to calculate a large number of similarity matrices, resulting in a high computational complexity; at the same time, since this method only calculates pseudo-labels based on confidence, the pseudo-labels thus contain a large amount of noise, ultimately reducing the classification performance.

[0004] Chen et al. proposed a new semi-supervised algorithm in the paper "Debiased Self-Training for Semi-Supervised Learning" (《Advances in Neural Information Processing Systems》, 2022). Through debiased self-training, it aims to reduce the bias problem in the self-training process and designs three classifiers: a K classifier, a pseudo-label classifier, and a worst-case estimation classifier. On the one hand, it decouples the generation and utilization of pseudo-labels: uses the K classifier to generate pseudo-labels, and the pseudo-label classifier uses the pseudo-labels to update the model to achieve the purpose of training the model only with reliable samples, reducing the training bias; on the other hand, to reduce data bias, debiased self-training optimizes the representation by estimating the worst case of the training bias, thereby improving the quality of pseudo-labels. Although this algorithm can reduce the bias problem in the self-training process and improve the stability and performance balance of the model. However, since its debiased self-training does not reduce the domain shift, it cannot align the invariant features between domains and cannot achieve cross-dataset knowledge transfer. Therefore, this semi-supervised algorithm cannot be directly used for unsupervised domain adaptation. Summary of the Invention

[0005] The purpose of the present invention is to address the deficiencies of the above-mentioned existing technologies and propose a hyperspectral image domain adaptation classification method based on self-training and domain adversarial to alleviate the problem of noisy pseudo-labels, achieve domain alignment between the source domain and the target domain, and improve the hyperspectral domain adaptation classification accuracy.

[0006] The technical idea for achieving the object of the present invention is as follows: on the one hand, by debiased self-training, the training bias and data bias are minimized to alleviate the problem of noisy pseudo-labels. On the other hand, by combining the domain adversarial method with the self-training method, the joint distribution alignment of the source domain and the target domain is carried out, and the invariant features between the source domain and the target domain are utilized to perform effective knowledge transfer between domains.

[0007] According to the above idea, the technical solution of the present invention includes the following steps:

[0008] (1) Obtain the source domain and target domain data sets on a public website, and perform hyperspectral data preprocessing on their images to obtain the preprocessed source domain and target domain;

[0009] (2) Construct a debiased self-training domain adversarial adaptive model including a feature extractor, three classifiers, and a gradient reversal layer;

[0010] (3) Construct a debiased self-training domain adversarial adaptive loss function \(L\) total :

[0011]

[0012] where is the source domain cross-entropy loss of the first classifier \(C1\), is the confidence-based self-training loss of the second classifier \(C2\); \(L\) worst is the worst-case adversarial loss of the third classifier \(C3\); \(L\) domain is the domain confusion loss of the third classifier \(C3\); \(\lambda_1\) and \(\lambda_2\) are hyperparameters for balancing the loss values;

[0013] (4) Train the debiased self-training domain adversarial adaptive model;

[0014] (4a) Set the initial value of the iteration number to 0, input the source domain and target domain data into the debiased self-training domain adversarial adaptive model, and calculate its loss value using the domain adversarial adaptive loss function \(L\) total ;

[0015] (4b) Substitute the loss value into the chain rule to calculate the gradients of the parameters of the debiased self-training domain adversarial adaptive model, and update the parameters of the model;

[0016] (4c) Increment the iteration number by 1 and return to (4a);

[0017] (4c) Repeat (4b) and (4c) to continuously update the parameters of the model and reduce the loss value \(L\) total , until the iteration number reaches the specified number of rounds to obtain a trained debiased self-training domain adversarial adaptive model;

[0018] (5) Input the target domain dataset into the trained debiased self-training domain adversarial adaptation model to obtain the classification result of the target domain hyperspectral image.

[0019] Compared with the prior art, the present invention has the following advantages:

[0020] First, since the present invention constructs a debiased self-training domain adversarial adaptation model composed of a feature extractor, three classifiers, and a gradient reversal layer, the K classifier C1 can infer the source domain data and generate pseudo-labels for the target domain data, the pseudo-label classifier C2 can infer the target domain data, and the adversarial classifier C3 can perform domain classification and predict the worst classification situation. On the one hand, it alleviates the problem of noisy pseudo-labels, and on the other hand, it avoids the overfitting problem that appears in existing self-training, improving the classification performance of the debiased self-training domain adversarial adaptation model.

[0021] Second, the present invention constructs a debiased self-training domain adversarial adaptation loss function L total , which also includes the source domain cross-entropy loss self-training loss based on confidence worst-case adversarial loss L worst and domain confusion loss L domain . By combining domain adversarial learning and debiased self-training methods, and using multilinear mapping and gradient reversal layer to achieve the joint distribution alignment between the source domain and the target domain, it reduces the domain bias while stabilizing the self-training process, improving the accuracy of the debiased self-training domain adversarial adaptation model in migrating and generalizing from the source domain to the target domain. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is the implementation flowchart of the present invention

[0023] Figure 2 is the structural diagram of the debiased self-training domain adversarial adaptation model in the present invention

[0024] Figure 3 is the comparison chart of domain adaptation classification results between the present invention and the prior art on the HyRANK dataset;

[0025] Figure 4 is the comparison chart of domain adaptation classification results between the present invention and the prior art on the Shanghau-Hangzhou dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The following further illustrates the embodiments and effects of the present invention with reference to the accompanying drawings. This embodiment is only used to illustrate the present invention and does not constitute any limitation to the present invention.

[0027] Referring to Figure 1 , the implementation steps of this example include the following:

[0028] Step 1, construct the source domain dataset and the target domain dataset:

[0029] (1.1) Obtain the source domain dataset from public websites:

[0030] In this example, the obtained source domain datasets are Dioni in the HyRANK dataset and Shanghai in the ShanghaiHangzhou dataset respectively; the image sizes are 250×1376 and 1600×260 in sequence; the number of bands are 176 and 198 in sequence;

[0031] (1.2) Obtain the target domain dataset from public websites:

[0032] In this example, the obtained target domain datasets are Loukia in the HyRANK dataset and Hangzhou in the ShanghaiHangzhou dataset respectively; the image sizes are 249×945 and 590×230 in sequence; the number of bands are 176 and 198 in sequence;

[0033] (1.3) For each pixel of the source domain and target domain dataset images, extract a pixel block with a size of 27×27 pixels around it, and form a new dataset by combining the pixel blocks of the source domain and target domain with the labels corresponding to their central pixels;

[0034] (1.4) Remove the background class and redundant class data and their corresponding labels from the new dataset. After removal, the Dioni dataset and the Loukia dataset each have 12 categories, and the dataset sizes are 20024 and 12208 respectively; the Shanghai dataset and the Hangzhou dataset each have 3 categories, and the dataset sizes are 135700 and 368000 respectively;

[0035] (1.5) Perform Z-Score normalization on the new dataset after removing the labels to convert the data into a standard normal distribution:

[0036]

[0037] where \(X\) is the input dataset, \(X'\) is the output dataset after Z-Score normalization, \(\mu\) is the mean of \(X\), and \(\sigma\) is the standard deviation of \(X\).

[0038] Step 2, construct a debiased self-training domain adversarial adaptation model.

[0039] Refer to Figure 2 , and the implementation steps are as follows:

[0040] (2.1)Construct a feature extractor FE consisting of four convolutional layers, four batch normalization layers, and one one-dimensional flattening layer. Its structure is as follows: the 1st convolutional layer → the 1st batch normalization layer → the 2nd convolutional layer → the 2nd batch normalization layer → the 3rd convolutional layer → the 3rd batch normalization layer → the 4th convolutional layer → the 4th batch normalization layer → the one-dimensional flattening layer. The structural parameters of these four convolutional layers are as follows:

[0041] For the 1st convolutional layer, the number of convolutional kernels is 64, the size of the convolutional kernel is 3×3, the stride is 2×2, the pixel padding is 1×1, and the activation function is the ReLU function;

[0042] For the 2nd convolutional layer, the number of convolutional kernels is 128, the size of the convolutional kernel is 3×3, the stride is 2×2, the pixel padding is 1×1, and the activation function is the ReLU function;

[0043] For the 3rd convolutional layer, the number of convolutional kernels is 256, the size of the convolutional kernel is 3×3, the stride is 2×2, the pixel padding is 1×1, and the activation function is the ReLU function;

[0044] For the 4th convolutional layer, the number of convolutional kernels is 512, the size of the convolutional kernel is 3×3, the stride is 2×2, the pixel padding is 0×0, and the activation function is the ReLU function;

[0045] The parameters of all convolutional layers are initialized using Kaiming initialization;

[0046] The weights of all batch normalization layers are initialized to 1, and the biases are initialized to 0;

[0047] (2.2)Construct three classifiers, namely the K-classifier C1, the pseudo-label classifier C2, and the adversarial classifier C3. Each classifier consists of four cascaded fully connected layers. Their structural parameters are as follows:

[0048] For the 1st fully connected layer, the number of nodes is 256, and the activation function is the ReLU function;

[0049] For the 2nd fully connected layer, the number of nodes is 100, and the activation function is the ReLU function;

[0050] For the 3rd fully connected layer, the number of nodes is 100, and the activation function is the ReLU function;

[0051] The number of nodes in the 4th fully connected layer of the K-classifier C1 and the pseudo-label classifier C2 is the number of task-specific classes;

[0052] The number of nodes in the 4th fully connected layer of the adversarial classifier C3 has two classification branches for the number of task-specific classes and the number of domains;

[0053] (2.3) Construct a gradient reversal layer, which is used to multiply the gradient by a negative weight when the data features are passed to the subsequent layer to reversely update the features during the backpropagation process;

[0054] (2.4) Cascade the feature extractor FE with the K classifier C1 and the pseudo-label classifier C2 respectively, and cascade them with the adversarial classifier C3 through the gradient reversal layer to form a three-way parallel debiased self-training domain adversarial adaptation model.

[0055] Step 3, construct a debiased self-training domain adversarial adaptation loss function L total .

[0056] (3.1) Define the worst-case adversarial loss L of the adversarial classifier C3 worst (X s ,Y s ,X t ,Y t ):

[0057]

[0058] Among them, X s represents the data of the source domain dataset, Y s represents the label of the source domain dataset, X t represents the data of the target domain dataset, Y t represents the label of the target dataset, n s is the size of the source domain dataset, n t is the size of the target domain dataset, f is the feature extracted by FE, g is the output of the K classifier C1, y is the label of the dataset; is the label after being screened by the threshold α; is the multilinear mapping of the feature and the prediction, is the cross-entropy loss, C is the number of classes, is the probability that the K classifier C1 classifies the i-th sample as the c-th class;

[0059] (3.2) Define the domain confusion loss L of the adversarial classifier C3 domain (X s ,Y s ,X t ,Y t ):

[0060]

[0061] Among them, y domain is the domain label, and the data of the source domain is the domain label the domain label of the data of the target domain

[0062]

[0063] (3.3) Select the existing cross-entropy loss as the loss function of the K-classifier C1 and the confidence-based self-training loss as the loss function of the pseudo-label classifier C2;

[0064] (3.4) Define the debiased self-training domain adversarial adaptive loss function L total :

[0065]

[0066] where λ1 and λ2 are hyperparameters for balancing the loss values,

[0067]

[0068]

[0069] Step 4: Train the debiased self-training domain adversarial adaptive model.

[0070] (4.1) Set the initial value of the iteration number to 0, input the source domain and target domain data into the debiased self-training domain adversarial adaptive model, and use the domain adversarial adaptive loss function L total to calculate its loss value;

[0071] (4.2) Initialize the gradient of the loss function L total with respect to itself:

[0072]

[0073] (4.3) Calculate the gradient of the loss with respect to the linear output of the output layer:

[0074]

[0075] where a output is the output of the activation function of the output layer, and z output is the linear output of the output layer;

[0076] (4.4) Starting from the output layer, traverse each layer backward, and use the chain rule to calculate the gradients of each layer. For each hidden layer l, calculate the gradients of the activation function and the linear output

[0077]

[0078] where W l+1 is the weight matrix of the next layer;

[0079] (4.5) Calculate the weight gradients and the bias gradient

[0080]

[0081] where a l-1 is the output of the activation function of the (l - 1)-th hidden layer;

[0082] (4.6) Calculate the gradient passing through the gradient reversal layer and

[0083]

[0084] where is the gradient before passing through the gradient reversal layer, is the gradient after passing through the gradient reversal layer;

[0085] (4.7) Initialize the first-order moment momentum m t and the second-order moment velocity v t to 0, and update them iteratively according to the following formula:

[0086]

[0087]

[0088] where t is the time step, which is incremented by 1 at each iteration, m t-1 and v t-1 are the first-order moment momentum and the second-order moment velocity at time step t - 1, m t and v t are the first-order moment momentum and the second-order moment velocity at time step t; β1 and β2 are the momentum factor and the velocity factor respectively, and both of these factors are hyperparameters;

[0089] (4.8) Correct the scale biases of m t and v t :

[0090]

[0091] where and are m t and v t after correcting the scale biases respectively;

[0092] (4.9) Update the model parameters using the corrected momentum and velocity:

[0093]

[0094] where w are the debiased self-training domain adversarial adaptive model parameters, α is the hyperparameter learning rate, and ε is the constant 10-8 ;

[0095] (4.10)Increment the iteration count by 1 and return to (4.1);

[0096] (4.11)Repeat (4.2)-(4.10) to continuously update the model parameters and reduce the loss value L total until the iteration count reaches the specified number of rounds to obtain a trained debiased self-training domain adversarial adaptive model.

[0097] Step 5, input the target domain dataset into the trained debiased self-training domain adversarial adaptive model to obtain the classification result y of the target domain hyperspectral image pred :

[0098]

[0099] where x t is the target domain data, C1 is the K classifier, and FE is the feature extractor.

[0100] The labels of the above steps are only used for clear description of the embodiments of the present invention, and their sequence numbers are not limited.

[0101] The effects of the present invention can be further illustrated by the following simulation results.

[0102] I. Simulation experiment conditions

[0103] The hardware platform for the simulation experiment of the present invention is: the processor is an Intel i7 7820x CPU with a main frequency of 3.6 GHz and a memory of 64 GB. The software platform for the simulation experiment of the present invention is: Windows 10 operating system and python 3.8.

[0104] The source domain datasets used in the simulation experiment of the present invention are the Dioni dataset and the Shanghai dataset, with image sizes of 250×1376 and 1600×260 respectively; the number of bands are 176 and 198 respectively;

[0105] The target domain datasets used in the simulation experiment of the present invention are the Loukia dataset and the Hangzhou dataset, with image sizes of 249×945 and 590×230 respectively; the number of bands are 176 and 198 respectively;

[0106] The comparison algorithms used in the simulation include:

[0107] The domain adaptation image classification method proposed by Mingsheng Long et al. in "Learning Transferable Features with Deep Adaptation Networks", abbreviated as DAN;

[0108] The domain adaptation image classification method proposed by Mingsheng Long et al. in "Deep Transfer Learning with Joint Adaptation Networks", abbreviated as JAN;

[0109] The domain adaptation image classification method proposed by Mingsheng Long et al. in "Maximum Classifier Discrepancy for Unsupervised Domain Adaptation", abbreviated as MCD.

[0110] II. Simulation Content and Results

[0111] Simulation 1: Under the above conditions, the Loukia dataset is classified using the present invention and three existing comparison algorithms respectively. An exclusive RGB color is assigned to each category and overlaid on a black background image to form the classification result, as Figure 3 shown, where:

[0112] Figure 3 (a) is the result graph of classifying the Loukia dataset by the DAN method,

[0113] Figure 3 (b) is the result graph of classifying the Loukia dataset by the JAN method,

[0114] Figure 3 (c) is the result graph of classifying the Loukia dataset by the MCD method,

[0115] Figure 3 (d) is the result graph of classifying the Loukia dataset by the present invention.

[0116] From Figure 3 (a), Figure 3 (b), Figure 3 (c), it can be seen that there are many noise points in the classification results of the existing methods, and many classification results do not match the actual situation significantly.

[0117] From Figure 3 (d), it can be seen that the present invention has fewer noise points and good regional consistency, proving that the classification effect of the present invention is better than that of the prior art and the classification effect is relatively ideal.

[0118] The overall accuracy OA, average accuracy AA and Kappa coefficient of the classification results of the above four methods are calculated respectively for performance comparison, and the results are shown in Table 1:

[0119] Table 1 Quantitative analysis table of the present invention and existing comparison algorithms in Simulation 1

[0120] Algorithm OA(%) AA(%) Kappa(%) DAN 53.6 48.8 48.6 JAN 54.5 44.6 49.3 MCD 58.4 50.5 53.4 The present invention 60.1 49.9 55.2

[0121] The parameter calculations for each index in Table 1 are as follows:

[0122] Compare the classification result y of the target domain pred with the labels y of the target domain dataset to form a confusion matrix M:

[0123]

[0124] Among them, TP represents the number of samples that are actually positive and predicted as positive, FP represents the number of samples that are actually negative but predicted as positive, FN represents the number of samples that are actually positive but predicted as negative, and TN represents the number of samples that are actually negative and predicted as negative;

[0125] According to the confusion matrix, calculate the overall accuracy OA, average accuracy AA, and Kappa coefficient:

[0126]

[0127]

[0128]

[0129] Among them is the accuracy of random classification.

[0130] As can be seen from Table 1, the overall accuracy OA, average accuracy AA, and Kappa coefficient of the classification of the Loukia dataset by the present invention are all higher than those of the three existing comparison algorithms. Therefore, it can be obtained that the classification performance of the present invention is better than that of the three existing comparison algorithms.

[0131] Simulation 2: Under the above conditions, classify the Hangzhou dataset using the present invention and the three existing comparison algorithms respectively, assign a unique RGB color to each category, and overlay it on a black background image to form a classification result, as Figure 4 shown, where:

[0132] Figure 4 (a) is the result diagram of classifying the Hangzhou dataset by the DAN method,

[0133] Figure 4 (b) is the result diagram of classifying the Hangzhou dataset by the JAN method,

[0134] Figure 4 (c) is the result diagram of classifying the Hangzhou dataset by the MCD method,

[0135] Figure 4(d) is the result graph of the classification of the Hangzhou dataset by the present invention.

[0136] As can be seen from Figure 4 (a), there are obvious missed detections in the central island area of the river by the existing method DAN;

[0137] As can be seen from Figure 4 (b), there are a few missed detections in the tributary area of the river by the existing method JAN;

[0138] As can be seen from Figure 4 (c), there are obvious misclassifications in the central island area of the river and obvious missed detections in the tributary area of the river by the existing method MCD.

[0139] As can be seen from Figure 4 (d), the classification result of the present invention has a high consistency with the actual situation, can clearly and accurately classify different categories, and the classification effect is relatively ideal.

[0140] Using the above formula, calculate the overall accuracy OA, average accuracy AA and Kappa coefficient of the classification results of the four methods respectively, and draw the calculation results into Table 2:

[0141] Table 2 Quantitative analysis table of the present invention and the existing comparison algorithms in Simulation 2

[0142] Algorithm OA(%) AA(%) Kappa(%) DAN 91.3 91.0 87.3 JAN 91.3 91.2 87.1 MCD 92.0 92.3 88.1 The present invention 94.1 93.6 91.2

[0143] As can be seen from Table 2, the overall accuracy OA, average accuracy AA and Kappa coefficient of the classification results of the present invention for the Hangzhou dataset are all higher than those of the three existing comparison algorithms. It can be obtained therefrom that the classification performance of the present invention is better than that of the three existing comparison algorithms.

[0144] The above simulation results show that the debiased self-training domain adversarial adaptive model proposed by the present invention can effectively improve the classification performance of hyperspectral image domain adaptation.

[0145] The above description is only a specific example of the present invention and does not constitute any limitation to the present invention. Obviously, for professionals in the field, after understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle of the present invention. However, these corrections and changes based on the idea of the present invention are still within the protection scope of the claims of the present invention.

Claims

1. A hyperspectral image domain adaptation classification method based on self-training and domain adversarial, characterized in that, The following are included: (1) Obtain the source domain and target domain datasets on public websites, and perform hyperspectral data preprocessing on their images to obtain the preprocessed source domain and target domain; (2) Construct a debiased self-training domain adversarial adaptive model consisting of a feature extractor, three classifiers, and a gradient reversal layer; (3) Construct the debiased self-training domain adversarial adaptive loss function L total : where is the source domain cross - entropy loss of the first classifier C1, is the confidence - based self - training loss of the second classifier C2; L worst is the worst - case adversarial loss of the third classifier C3; L domain is the domain confusion loss of the third classifier C3; λ1 and λ2 are hyperparameters for balancing the loss values; (4) Train the debiased self-training domain adversarial adaptive model; (4a) Set the initial value of the iteration round to 0, input the source domain and target domain data into the debiased self-training domain adversarial adaptation model, and use the domain adversarial adaptation loss function L total to calculate its loss value; (4b) Substitute the loss value into the chain rule, calculate the gradients of each parameter of the debiased self-training domain adversarial adaptive model, and update the parameters of the model; (4c) Increment the number of iteration rounds by 1, and return to (4a); (4c) Repeat (4b) and (4c), continuously update the parameters of the model and reduce the loss value L total until the number of iterations reaches the specified number of rounds, and obtain the trained debiased self-training domain adversarial adaptive model; (5) Input the target domain dataset into the trained debiased self-training domain adversarial adaptive model to obtain the classification result of the target domain hyperspectral image.

2. The method according to claim 1, wherein In step (1), when performing hyperspectral data preprocessing on the obtained source domain and target domain dataset images, the implementation steps are as follows: (1a) For each pixel of the source domain and target domain dataset images, extract a pixel block with a size of 27×27 pixels around it, and form a new dataset by combining the pixel blocks of the source domain and target domain with the labels corresponding to their central pixels; (1b) Remove the background class and redundant class data and their corresponding labels from the new dataset; (1c) Perform Z-Score normalization on the new dataset after removing the labels, and convert the data into a standard normal distribution: where X is the input dataset, X' is the output dataset after Z-Score normalization, μ is the mean of X, and σ is the standard deviation of X.

3. The method according to claim 1, characterized in that, In step (2), when constructing a debiased self-training domain adversarial adaptive model consisting of a feature extractor, three classifiers, and a gradient reversal layer, the implementation steps are as follows: (2a) Construct a feature extractor FE including four convolutional layers, four batch normalization layers, and one one-dimensional flattening layer. Its structure is: the 1st convolutional layer → the 1st batch normalization layer → the 2nd convolutional layer → the 2nd batch normalization layer → the 3rd convolutional layer → the 3rd batch normalization layer → the 4th convolutional layer → the 4th batch normalization layer → one-dimensional flattening layer; All convolutional layer parameters are initialized with Kaiming; all batch normalization layer weights are initialized to 1, and the biases are 0; (2b) Construct three classifiers, namely the K classifier C1, the pseudo-label classifier C2, and the adversarial classifier C3. Each classifier includes four cascaded fully connected layers; (2c) Construct a gradient reversal layer, which is used to multiply the gradient by a negative weight when the data features are passed to the subsequent layers to inversely update the features during the backpropagation process; (2d) Cascade the feature extractor FE with the K classifier C1 and the pseudo-label classifier C2 respectively, and cascade with the adversarial classifier C3 through the gradient reversal layer to form a three-way parallel debiased self-training domain adversarial adaptive model.

4. The method according to claim 3, wherein For the four convolutional layers in step (2a), their parameter structures are as follows: The number of convolutional kernels in the 1st convolutional layer is 64, the convolutional kernel size is 3×3, the stride is 2×2, the pixel padding is 1×1, and the activation function is the ReLU function; The number of convolutional kernels in the 2nd convolutional layer is 128, the convolutional kernel size is 3×3, the stride is 2×2, the pixel padding is 1×1, and the activation function is the ReLU function; The number of convolution kernels in the third convolution layer is 256, the size of the convolution kernel is 3×3, the stride is 2×2, the pixel padding is 1×1, and the activation function is the ReLU function; The number of convolution kernels in the fourth convolution layer is 512, the size of the convolution kernel is 3×3, the stride is 2×2, the pixel padding is 0×0, and the activation function is the ReLU function.

5. The method according to claim 3, characterized in that, Each classifier in step (2b) includes four cascaded fully connected layers, and its structural parameters are as follows: The number of nodes in the first fully connected layer is 256, and the activation function is the ReLU function; The number of nodes in the second fully connected layer is 100, and the activation function is the ReLU function; The number of nodes in the third fully connected layer is 100, and the activation function is the ReLU function; The number of nodes in the fourth fully connected layer of the K-classifier C1 and the pseudo-label classifier C2 is the number of task-specific categories; The number of nodes in the fourth fully connected layer of the adversarial classifier C3 has two classification branches for the number of task-specific categories and the number of domains.

6. The method according to claim 1, characterized in that, In step (3), construct the debiased self-training domain adversarial adaptive loss function L total , and the implementation steps are as follows: (3a) Define the worst-case adversarial loss L for the adversarial classifier C3 worst : where n s is the size of the source domain dataset, n t is the size of the target domain dataset, f is the feature extracted by FE, g is the output of the K classifier C1, and y is the label of the dataset; is the label after being screened by the threshold α, MML is the multilinear mapping of the feature and the prediction, L ce is the cross-entropy loss, where C is the number of classes, is the probability that the K classifier C1 classifies the i-th sample as the c-th class; (3b) Define the domain confusion loss L of the adversarial classifier C3 domain : where y domain is a domain label, and the data of the source domain is the domain label the domain label of the data in the target domain (3c) Select the source domain cross-entropy loss of the existing K classifier C1 Confidence-based self-training loss of the pseudo-label classifier C2 (3e) Define the debiased self-training domain adversarial adaptive loss function \(L\) according to the above four losses total : Where λ1 and λ2 are hyperparameters for balancing the loss value, 7. The method according to claim 1, wherein In step (4b), calculating the gradients of the parameters of the debiased self-training domain adversarial adaptive model and updating the parameters of the model, the implementation steps include the following: (4b1) Initialize the loss function L total Gradient with respect to itself: (4b2) Calculate the gradient of the loss with respect to the linear output of the output layer: where a output is the output of the activation function of the output layer, and z output is the linear output of the output layer; (4b3) Starting from the output layer, traverse each layer in reverse and calculate the gradients of each layer using the chain rule. For each hidden layer l, calculate the gradients of the activation function and the linear output in sequence where W l+1 is the weight matrix of the next layer; (4b4) Calculate the weight gradient and the bias gradient respectively and the bias gradient where a l-1 is the output of the activation function of the (l-1)-th hidden layer; (4b5) Calculate the gradient passing through the gradient reversal layer and Among them is the gradient before passing through the gradient reversal layer, is the gradient after passing through the gradient reversal layer; (4b6) Initialize the first-order moment momentum m in the Adam optimizer t and the second-order moment velocity v t to 0, and update them iteratively according to the following formula: where t is the time step, which is incremented by 1 at each iteration, m t-1 and v t-1 are the first-order moment momentum and the second-order moment velocity at time t-1, m t and v t are the first-order moment momentum and the second-order moment velocity at time t; β1 and β2 are the momentum factor and the velocity factor respectively, and both of these factors are hyperparameters; (4b7) Correct m t and v t Scale deviation of: Among them and are m t and v t ; (4b8) Update the model parameters using the corrected momentum and velocity: where w are the parameters of the debiased self-training domain adversarial adaptation model, α is the hyperparameter learning rate, and ε is the constant 10 -8 .

Citation Information

Patent Citations

  • A Domain Adaptive Method for Hyperspectral Images Based on Virtual Classifiers

    CN115410088B