An Adaptive Threshold-Based Scene Classification Method for Pseudo-Labeled Remote Sensing Images

Through the pseudo-label method of adaptive thresholds, the problem of unstable threshold settings in remote sensing image scene classification is solved by using the main and auxiliary classifiers and adversarial learning, and the accuracy and stability of remote sensing image scene classification is improved, which is suitable for remote sensing image scene classification.

CN114549909BActive Publication Date: 2025-07-11ZHONGKE LIULIMA (GUANGDONG) INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210209902.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-03
Publication Date
2025-07-11
Estimated Expiration
2042-03-03

AI Technical Summary

Technical Problem

In the existing remote sensing image scene classification methods, manually setting thresholds leads to instability in the model and it is difficult to effectively form decision boundaries. Transfer learning is affected by the excessive gap in the characteristics of the source domain and the target domain, resulting in a decrease in classification accuracy.

Method used

Using the pseudo-label method of adaptive thresholds, by building the main classifier and auxiliary classifier, the adaptive threshold is calculated using variance regularization, and combining adversarial learning and Gaussian guidance, the source domain and target domain features are mapped to the same feature space to generate accurate pseudo-labels and reduce the impact of noise.

Benefits of technology

It improves the accuracy and stability of remote sensing image scene classification, reduces the impact of noise labels, and does not require manual settings to effectively form decision boundaries and improves the model's learning ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114549909B_ABST
    Figure CN114549909B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of remote sensing image recognition, and particularly relates to a method for classifying pseudo-labeled remote sensing image scenes based on an adaptive threshold, which includes first using a generative adversarial network to preliminarily expand the source domain dataset, then shallowly aligning the data through Gaussian guidance in the backbone network for the expanded dataset, and then connecting a pseudo-label module with two classifiers to generate pseudo-labels for the target domain; regularizing the two classification variances and using them as an adaptive threshold to correct the cross-entropy loss. When the prediction distance between the main classifier and the auxiliary classifier is large, it indicates that the label is incorrect, and for such uncertain samples, no penalty is imposed, that is, the incorrect label samples are discarded, and the samples with a smaller prediction distance are used as pseudo-labels; then the target domain and the source domain with pseudo-labels are sent into a domain discriminator, which better narrows the feature distance between the source domain and the target domain; the present invention can be better applied to remote sensing image scene classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing image processing and classification recognition, and particularly relates to a method for classifying remote sensing image scenes based on an adaptive threshold and pseudo-labels. Background Art

[0002] Remote sensing images are widely used in fields such as surveying and mapping, environmental monitoring, and urban planning. With the continuous increase in the acquisition volume of remote sensing images, scene classification and semantic segmentation have received extensive attention in the remote sensing community. Remote sensing scene classification is a practical task of inferring the correct scene category based on the image content. In the past few decades, traditional remote sensing image scene classification mainly relied on manually extracting features. However, the features extracted manually are low-level features, which can only represent the capabilities of the targets, meaning that using these features in the present invention will result in a higher classification error rate.

[0003] In recent years, convolutional neural networks have played a great role in promoting this field. It requires a large amount of data sets for training, which is a prerequisite for stronger model generalization capabilities. However, in practical tasks, there are not enough data sets, and manual labeling is also very time-consuming. Research shows that transfer learning can effectively solve this problem. However, in real task scenarios, RS images are affected by natural factors such as lighting, spatial resolution, rotation, and weather conditions. Therefore, transfer learning cannot be effectively applied to remote sensing scene classification, essentially because the feature gap between the source domain and the target domain is too large.

[0004] In order to make the features between the source domain and the target domain in remote sensing images closer, researchers have proposed domain adaptation methods to solve this problem. Domain adaptation methods can currently be mainly divided into traditional metric distance-based methods and deep learning-based methods. Deep learning-based methods often use adversarial learning to map the features of the source domain and the target domain into the same feature space. Domain adversarial learning consists of a domain discriminator and a generator. The main role of the generator is to confuse the features of the source domain and the target domain so that the domain discriminator cannot determine whether the features come from the source domain or the target domain. The role of the domain discriminator is the opposite of that of the generator, and its purpose is to correctly determine the source of the features as much as possible. In this way, an adversarial form is formed, and ultimately the source domain and the target domain will be confused together. However, this method cannot form an effective decision boundary, and the image categories of the two domains are completely confused together. Therefore, the method of adding pseudo-labels is often used to help form an effective decision boundary. However, the pseudo-label scheme sets the threshold manually, and it is often very difficult to determine the size of the threshold. Setting the threshold too high or too low will affect the learning of the model. Therefore, in order to more effectively set a reasonable threshold to form a decision boundary, an adaptive threshold strategy is needed to more effectively classify remote sensing image scenes. Summary of the Invention

[0005] For the above technical problems, the present invention provides a method for pseudo-label remote sensing image scene classification based on an adaptive threshold. The method can effectively solve the problem of model instability caused by manually setting the threshold, enabling noise labels to be filtered out by the threshold, and making the remote sensing image scene classification more accurate. The present invention not only effectively maps the source domain and target domain of remote sensing images to the same feature space, but also forms an effective decision boundary through the threshold, and this threshold does not need to be manually set and is automatically generated according to the characteristics of the dataset, improving the classification ability.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A method for pseudo-label remote sensing image scene classification based on an adaptive threshold, comprising the following steps:

[0008] S1. Divide the remote sensing image dataset into a source domain dataset and a target domain dataset, and input the source domain dataset with true labels into the adversarial generation network to expand the source domain dataset;

[0009] S2. Construct a first scene classification network including a feature extractor and a classifier, and construct a second scene classification network including a generator and a discriminator; and pre-train the first scene classification network using the expanded source domain dataset;

[0010] S3. Use the pre-trained feature extractor to extract data features of the target domain dataset at different levels, and align the data features of the shallow layer through Gaussian guidance;

[0011] S4. Copy the pre-trained classifier to form a main classifier and an auxiliary classifier, input the aligned data features of different levels in the target domain dataset into the main classifier and the auxiliary classifier, and output a first classification label and a second classification label;

[0012] S5. According to the first classification label and the second classification label of the target domain dataset, calculate the classification variance and cross-entropy loss between the two classifiers;

[0013] S6. Regularize the classification variance, use the variance regularization as an adaptive threshold to correct the cross-entropy, and use the corrected cross-entropy loss to update the model parameters of the first scene classification network, and determine the pseudo-labels of the target domain data;

[0014] S7. Input the target domain data with pseudo-labels and the source domain data with true labels into the second scene classification network, and use the adversarial training method to update the model parameters of the second scene classification network;

[0015] S8. Iteratively loop through steps S3 - S7 to update and optimize the parameters of each pre - trained scene classification network model until the classification target requirements are met, and output the remote sensing image scene classification results of the target domain dataset.

[0016] Advantages of the present invention:

[0017] The present invention is applicable to remote sensing image scene classification. Compared with existing methods, the advantages of the present invention are as follows:

[0018] (1) Using the idea of adversarial learning, the present invention uses a generator to extract data features from the source domain and the target domain, and a discriminator to determine whether the features come from the source domain or the target domain. Such a domain - adaptation method maps the data features of the source domain and the target domain of remote sensing images into the same feature space, reducing the problem of accuracy decline caused by domain transfer.

[0019] (2) The present invention proposes to construct a main classifier and an auxiliary classifier. By introducing variance regularization as an adaptive threshold, this threshold can filter out the noise in the dataset and more effectively learn the features in the remote sensing dataset. Description of the Drawings

[0020] Figure 1 is the flowchart of the method for pseudo - label remote sensing image scene classification based on an adaptive threshold according to an embodiment of the present invention;

[0021] Figure 2 is the network structure diagram of the present invention;

[0022] Figure 3 is a partial scene diagram of three datasets adopted by the present invention. Detailed Embodiments

[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0024] Figure 1 is a flowchart of a method for pseudo - label remote sensing image scene classification based on an adaptive threshold in an embodiment of the present invention. As Figure 1 shown, the method includes:

[0025] S1. Divide the remote sensing image dataset into a source domain dataset and a target domain dataset, and input the source domain dataset with true labels into the adversarial generation network to expand the source domain dataset;

[0026] In an embodiment of the present invention, the remote sensing image dataset can be directly retrieved from various public data platforms. For example, the UCAS AOD remote sensing image dataset can be collected for aircraft and vehicle detection. The aircraft dataset includes 600 images and 3,210 aircraft, while the vehicle dataset includes 310 images and 2,819 vehicles. All the images are carefully selected to ensure a uniform distribution of object directions in the dataset. Alternatively, the RSSCN7 Dataset remote sensing image dataset can be collected, which contains 2,800 remote sensing images from 7 typical scene categories, including grassland, forest, farmland, parking lot, residential area, industrial area, and river / lake. Each category contains 400 images, sampled based on 4 different scales. The pixel size of each image in this dataset is 400*400. The diversity of the scene images makes it challenging. These images are from different seasons and weather changes and are sampled at different scales. In addition, the present invention can obtain data from any public dataset, and no specific limitation is imposed in this embodiment.

[0027] In an embodiment of the present invention, assume that three remote sensing image datasets are selected, namely the AID, UC Merced, and RSI datasets. The resolutions and sizes of these three remote sensing image datasets are different. Three common scenes are selected as the datasets to be used. Then, six cross-dataset tasks will be established using these three datasets, which are and That is, the former is the source domain dataset, and the latter is the target domain dataset. The source domain datasets in these datasets are respectively input into the adversarial generation network to generate more data. Specifically, the generator in the adversarial generation network accepts a random noise vector, and the goal is to generate similar real samples through this noise vector. The discriminator in the adversarial generation network determines whether the sample is generated by its own generator or a real sample, and uses the loss function to optimize this adversarial generation network until the loss function of the adversarial generation network converges, outputting the augmented source domain dataset to enhance the dataset samples. Among them, the loss function can be expressed as: V(G,D) = E x~Pdata [logD(x)] + E x-pg [log(1 - D(x))].

[0028] Among them, V(G,D) represents the adversarial loss between the generator and the discriminator in the adversarial generation network, E[] represents taking the expectation, x ∼ Pdata represents the distribution of real images; x represents a real sample, and the probability that x is a real sample is judged through the discriminant network D(); z ∼ P z(z) represents the distribution of the generated image; z represents the noise input to the generation network, G(z) represents the noise sample z generated by the generation network G(), and D(G(z)) represents the probability of judging its true sample after the generated sample passes through the discriminant network.

[0029] S2. Construct a first scene classification network including a feature extractor and a classifier, and a second scene classification network including a generator and a discriminator; and pre-train the first scene classification network using the augmented source domain dataset;

[0030] In the embodiment of the present invention, a convolutional neural network needs to be constructed, as Figure 2 shown, mainly divided into three parts. The first part is the feature extractor. In the present invention, ResNet-101 is used as the backbone network of the model, and this network has been pre-trained on the ImageNet dataset, with a total of 5 groups of convolutions. The input size of the first group of convolutions is 224x224, and the output size of the fifth group of convolutions is 7x7, shrinking by 2 times each time, for a total of 5 times of shrinking, and each time the stride is 2 on the first layer of each group of convolutions. The second part is the pseudo-label module, which consists of two main classifiers and auxiliary classifiers connected to different layers. The structures of the two classifiers are the same, consisting of a convolutional layer, a BN layer, a ReLU activation function, and a max pooling layer. Connect this auxiliary classifier network to res4b22 in the backbone network, that is, the output of res4b22 is used as the input of this network. The main classifier adopts the same network structure and is connected to the res5c layer. Among them, during the pre-training process, the pseudo-label module has only one classifier for classifying the source domain dataset. The third part is the adversarial generation module. This adversarial generation network includes a generator, a discriminator, and a classifier. Here, the discriminator is composed of a domain discriminator, and the domain discriminator is composed of two convolutional layers, also using the ReLU activation function.

[0031] For the convenience of description, the present invention divides it into two scene classification networks. The purpose of the first scene classification network is to obtain the pseudo-labels of each target domain data in the target domain dataset, and the purpose of the second scene classification network is to determine the scene classification prediction results of each target domain data in the target domain dataset.

[0032] In the embodiment of the present invention, the first scene classification network is trained using the augmented source domain dataset to optimize the feature extractor and classifier in the first scene classification network; after the training is completed, the data features of the target domain dataset can be directly extracted using the feature extractor, and the classification results of the target domain dataset can also be obtained using the trained classifier as the corresponding pseudo-labels.

[0033] It can be understood that in the embodiments of the present invention, the specific structure and training process of the first scenario classification network of the present invention may not be specifically limited, and those skilled in the art can make adaptive changes according to existing classifier models.

[0034] S3. Use the pre-trained feature extractor to extract data features of the target domain dataset at different levels, and align the shallow data features through Gaussian guidance;

[0035] In the embodiments of the present invention, under the guidance of Gaussian prior, the data features of the source domain dataset and the data features of the target domain dataset are indirectly aligned. That is, the sample distributions of the two domains are constructed in a common feature space, which ultimately better promotes the alignment of features. To encourage the construction of a discriminative feature space of the source domain in the prior space, the labeled original samples are regularized using the softmax cross-entropy loss. At the same time, the latent features of the target domain should also be similar to the Gaussian prior. To effectively align the target latent distribution, the L1 distance is used to measure the gap between the features of the two, which more effectively plays a role of pre-guidance. The L1 distance is also called the Manhattan distance, and it is expressed as:

[0036] C = |x1 - x2| + |y1 - y2|

[0037] S4. Copy the pre-trained classifier to form a main classifier and an auxiliary classifier, input the aligned data features of the target domain dataset at different levels into the main classifier and the auxiliary classifier, and output the first classification label and the second classification label;

[0038] In the embodiments of the present invention, the classifier trained with the source domain dataset can effectively classify the target domain dataset, that is, generate pseudo-labels for the unlabeled target domain dataset. Among them, the traditional method can directly use the above-trained classifier to obtain the pseudo-labels of the target domain data, which is expressed as:

[0039] y t = argmaxF(x t |θ s )

[0040] Among them, y t represents the pseudo-label of the target domain data; the input is expressed as x t , F() represents the classifier; θ s represents the parameters learned from the remote sensing source domain dataset. However, due to the different data distributions of the source domain dataset and the target domain dataset, the pseudo-labels obtained by the traditional method are not accurate. It can be represented by the deviation:

[0041]

[0042] Among them, Bias(p t ) represents the bias of the pseudo-label; the first term is the difference between the predicted label F(x t |θ t ) and the pseudo-label , while the second term is the error between the true label p t and the pseudo-label . However, there is an inherent problem, that is, the labels inevitably contain noise. The mislabeling is passed from the original model to the final model. Therefore, the pseudo-labels of traditional techniques greatly reduce the training effect.

[0043] Based on this, the present invention uses two-level classifiers to separately obtain the classification results of the target domain dataset, and uses the classification variance between the two classification results to form an adaptive threshold control to generate the corresponding decision boundary.

[0044] Among them, the two-level classifiers in the embodiments of the present invention can directly adopt the classifiers trained with the source domain dataset, and the structures of the two-level classifiers are exactly the same; but they perform classification processing on the data features of different network layers in the feature extractor, that is, input the deep features in the aligned target domain dataset into the main classifier to obtain the first classification label of the target domain dataset, and input the shallow features in the aligned target domain dataset into the auxiliary classifier to obtain the second classification label of the target domain dataset. The one with a smaller distance can be selected from the two classification labels as the pseudo-label of the target domain data.

[0045] S5. Calculate the classification variance and cross-entropy loss between the two classifiers according to the first classification label and the second classification label of the target domain dataset;

[0046] In the embodiments of the present invention, first calculate the predicted variance:

[0047] V ar(pt) =E[(F(x t )-p t ) 2

[0048] Because p t represents the true label and is unknown in the target domain, so the present invention uses the pseudo-label to replace p t , so the corresponding formula changes to

[0049]

[0050] However, the present invention uses the feature extractor F(x) to obtain When optimizing the prediction bias, the variance will become smaller, so it cannot reflect the true prediction method during the training phase. Therefore, another similar judgment formula is adopted​

[0051] V ar(pt) ≈E[(F(x t ) - F aux (x t )) 2

[0052] where F aux (x t ) represents the output of the auxiliary classifier. Since the main classifier learns from deeper layers, the auxiliary classifier learns from relatively shallower layers, and the input activations between the two classifiers are different, which will lead to prediction differences. Secondly, the two classifiers have not been trained on the target domain dataset. Therefore, the two classifiers have different biases for the target domain dataset. In the present invention, the KL divergence is used to represent the classification variance between the two

[0053]

[0054] where D kl represents the KL divergence of the classification variance between the two classifiers; F(x t ) represents the first classification label output by the main classifier; F aux (x t ) represents the second classification label output by the auxiliary classifier; E[] represents taking the expectation.

[0055] If the two classifiers provide two different class predictions, the approximate variance will obtain a larger value. It reflects the uncertainty of the model for the prediction. The KL divergence is an alternative for variance calculation, and the distance formula can also be replaced to calculate the distance between the main classifier and the auxiliary classifier. The maximum mean discrepancy is a common metric distance formula, expressed as:

[0056]

[0057] where represents the mapping used to map the variable into the reproducing kernel Hilbert space, and the formula expresses the mean distance between the two data in this space.

[0058] In the embodiments of the present invention, since there is a certain gap between the predicted label and the pseudo-label, the cross-entropy function is used to minimize this bias. The cross-entropy loss is expressed as:

[0059]

[0060] where L ce represents the cross-entropy loss; represents the pseudo-label; F(x t |θ t ) represents the first classification label output by using the main classifier.​

[0061] S6. Compared with other methods of manually setting thresholds to filter out low-confidence ones, regularizing the classification variance is equivalent to forming a process of adaptive thresholding. The variance regularization is used as the adaptive threshold to correct the cross-entropy, and the corrected cross-entropy loss is used to update the model parameters of the first scene classification network and determine the pseudo-labels of the target domain data;

[0062] In the embodiments of the present invention, the classification variance is fixed and regularized to correct learning from incorrect labels. The corrected target loss function can be expressed as:

[0063]

[0064] Since the prediction variance is not minimized under all conditions. If the predicted variance gets a large value, the bias will not be penalized. At the same time, to prevent the model from always predicting a large variance, the present invention adds V ar(pt) to introduce regularization. In addition, since Var(p t ) may be zero, there may be a situation of division by zero. Therefore, to exclude this situation, exp(-Var) is used instead Therefore, the corrected loss function is rewritten as:

[0065] L rect = E[exp{-D kl}L ce + D kl

[0066] where L rect represents the corrected cross-entropy loss, D kl represents the KL divergence loss of the classification variance, E[] represents taking the expectation; L ce represents the cross-entropy loss.

[0067] S7. Input the target domain data with pseudo-labels and the source domain data with true labels into the second scene classification network, and use the method of adversarial training to update the model parameters of the second scene classification network;

[0068] ​In the embodiment of the present invention, through the above process, pseudo-labels can be assigned to all the target-domain data in the target-domain dataset. Therefore, all the source-domain data and target-domain data have labels. The target-domain data with pseudo-labels and the source data with true labels are sent into the domain discriminator together to determine the source of the features. The source domain and the target domain come from different distributions, so training on the source domain and testing on the target domain cannot obtain a good result. The main reason is that the feature distributions of the two are different, so what needs to be done is to narrow the feature distance between the source domain and the target domain. Through the feature extraction of the dataset in the previous steps, combined with the domain discriminator to determine the source of the features, an adversarial generation loop is formed, and finally the features of the two can be mapped into the same space.

[0069] S8. Iteratively loop through steps S3 - S7 to update and optimize the parameters of each pre-trained scene classification network model until the classification target requirements are met, and output the remote sensing image scene classification result of the target-domain dataset.

[0070] In the embodiment of the present invention, it is necessary to jointly optimize the first scene classification network and the second scene classification network until the classification target requirements are met, that is, the remote sensing image scene classification result of the target-domain dataset can be output.

[0071] It can be understood that the present invention first uses an adversarial generation network to preliminarily expand the source-domain dataset, and then performs shallow alignment of the data through Gaussian guidance in the backbone network for the expanded dataset. After that, a pseudo-label module is connected to generate pseudo-labels for the target domain. This pseudo-label module mainly consists of two classifiers, a main classifier and an auxiliary classifier, which are respectively connected behind different network layers to obtain temporary classification results. Due to the different network layers accessed and the different feature distributions of the labels, the classification results of the two classifiers will also be different. Therefore, cross-entropy loss is used for correction. When the prediction distance between the main classifier and the auxiliary classifier is large, it means that the label is incorrect, and such uncertain samples are not penalized, that is, the incorrect label samples are discarded, and the samples with a smaller prediction distance are used as pseudo-labels. Then, the target domain and the source domain with pseudo-labels are sent into the domain discriminator, which better narrows the feature distance between the source domain and the target domain. The present invention has fast operation, good effect, and is easy to deploy. First, the dataset is expanded through an adversarial generation network, and then a threshold is actively learned through the pseudo-label module in the case of uncertain differences between the source domain and the target domain, without worrying that the manually set threshold will regard valid labels as noise and affect the learning efficiency of the model, and it can be better applied to remote sensing image scene classification.

[0072] Figure 3 These are partial scene diagrams of the three datasets used in the embodiment of the present invention. The accuracies of these three datasets reach 94.6%, 94.1%, and 95.1% respectively, which can fully prove the effectiveness of the present invention.

[0073] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "coaxial", "bottom", "one end", "top", "middle", "the other end", "upper", "one side", "top", "inner", "outer", "front", "center", "both ends", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.

[0074] In the present invention, unless otherwise clearly defined and limited, the terms "installed", "set", "connected", "fixed", "rotated", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements or the interaction relationship between two elements. Unless otherwise clearly defined, for those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0075] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for pseudo-label remote sensing image scene classification based on an adaptive threshold, characterized in that Specifically, it includes the following steps: S1. Divide the remote sensing image dataset into a source domain dataset and a target domain dataset, and input the source domain dataset with true labels into the adversarial generation network to expand the source domain dataset; the generation network in the adversarial generation network accepts a random noise vector, and the goal is to generate real samples similar to the source domain dataset through this random noise vector. The discriminative network in the adversarial generation network judges whether the sample is generated by its generation network or a real source domain dataset sample until the loss function converges, and outputs the samples of the expanded source domain dataset; the loss function is expressed as: V(G,D)=E x~Pdata [logD(x)]+E z~Pz(z) [log(1 - D(G(z))], where V(G,D) represents the adversarial loss between the generator and the discriminator in the adversarial generation network, E[] represents taking the expectation, x~Pdata represents the distribution of real images; x represents a real sample, and the probability that x is a real sample is judged through the discriminative network D(); z~Pz(z) represents the distribution of generated images; z represents the noise input into the generation network, G(z) represents the noise sample z generated by the generation network G(), and D(G(z)) represents the probability of judging the generated sample as a real sample after passing through the discriminative network; S2. Construct a first scene classification network including a feature extractor and a classifier, and construct a second scene classification network including a generator and a discriminator; and pre-train the first scene classification network using the augmented source domain dataset; S3. Use the pre-trained feature extractor to extract data features of the target domain dataset at different levels, and align the shallow data features through Gaussian guidance; with the early access of Gaussian guidance, construct the sample distributions of the two domains in a common feature space, and finally better promote the alignment of features; S4. Duplicate the pre-trained classifier to form a main classifier and an auxiliary classifier, input the data features at different levels of the aligned target domain dataset into the main classifier and the auxiliary classifier, and output a first classification label and a second classification label; that is, input the deep features in the aligned target domain dataset into the main classifier, and the shallow features into the auxiliary classifier, and sequentially obtain the first classification label and the second classification label of the target domain data; S5. Calculate the classification variance and cross-entropy loss between the two classifiers according to the first classification label and the second classification label of the target domain dataset; The cross-entropy loss is expressed as: Among them, L ce represents the cross-entropy loss; represents the one-hot vector of the pseudo-label of the target domain data; F(x t |θ t ) represents the first classification label output by the main classifier, and E[] represents taking the expectation; S6. Regularize the classification variance, use the variance regularization as an adaptive threshold to correct the cross-entropy, and use the corrected cross-entropy loss to update the model parameters of the first scene classification network, and determine the pseudo-labels of the target domain data; The corrected cross-entropy loss is expressed as: L rect = E[exp{-D kl}L ce + D kl ​ Among them, L rect represents the corrected cross-entropy loss, D kl represents the KL divergence loss of the classification variance, and E[] represents the expectation; L ce represents the cross-entropy loss; S7. Input the target domain data with pseudo-labels and the source domain data with real labels into the second scene classification network, and use the adversarial training method to update the model parameters of the second scene classification network; use the generator in the second scene classification network to extract the data features of the target domain data and the source domain data respectively, use the domain classifier to determine whether the source of the data features is the target domain data or the source domain dataset, and through the adversarial training method, narrow the feature distance between the source domain dataset and the target domain dataset, and determine the adversarial loss until the adversarial loss converges, so as to optimize the model parameters of the second scene classification network; S8. Loop and iterate steps S3 - S7, update and optimize the model parameters of each pre-trained scene classification network until the classification target requirements are met, and output the remote sensing image scene classification result of the target domain dataset.

2. The method for pseudo-label remote sensing image scene classification based on an adaptive threshold according to claim 1, characterized in that, The classification variance between the two classifiers is expressed as: Among them, D kl represents the KL divergence of the classification variance between two classifiers; F(x t ) represents the first classification label output by the main classifier; F aux (x t ) represents the second classification label output by the auxiliary classifier; E[] represents taking the expectation.

Citation Information

Patent Citations

  • Unsupervised domain adaptation method based on adversarial learning loss function

    CN110837850A

  • Anti-counterfeit label identification method and device

    CN111414779A