Open set passive field adaptive method based on uncertain drive and unknown diffuser
Through the open-set passive domain adaptation method of uncertain drivers and unknown diffusers, the problem of difficult knowledge migration when the source data and target data cannot be accessed simultaneously is solved, and more effective knowledge migration and target domain adaptation effects are achieved.
Patent Information
- Application Number
- CN202510297722.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
Existing unsupervised domain adaptive approaches are difficult to effectively carry out knowledge migration when both source and target data are not accessible, especially when the target domain contains unknown categories, and may lead to feature space degradation and negative transfer problems.
The open-set passive domain adaptive method based on uncertain drivers and unknown diffusers is adopted to train by building an initial model, generating initial pseudo-label allocation, refining pseudo-label, estimating pseudo-label uncertainty, selecting uncertainty samples, and optimizing the model with negative learning loss and comparison loss.
Effectively carry out knowledge transfer, improve the adaptability of target areas, reduce the occurrence of feature space degradation and negative transfer problems, and improve the model's ability to deal with unknown categories.
Smart Images

Figure QLYQS_8 
Figure QLYQS_13 
Figure QLYQS_14
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of transfer learning methods, and particularly relates to an open-set source-free domain adaptation method based on uncertainty-driven and unknown diffusers. Background Art
[0002] Unsupervised domain adaptation (UDA) has shown great potential in reducing the cost of data annotation and dealing with practical problems such as domain distribution shift by transferring the knowledge of the source domain with sufficient labels to the unlabeled target domain. However, most UDA methods usually rely on two strong assumptions: First, both source data and target data are accessed simultaneously during the adaptation process, which may affect deployment in scenarios with limited data privacy or transmission bandwidth. Second, it is assumed that the source domain and the target domain share the same class space. In fact, the target domain often contains unknown classes not covered by the source domain, and the source data may not be available at any time. For this reason, open-set source-free domain adaptation (OS-SFDA) has emerged, which only relies on the source domain pre-trained model without accessing the source data, thereby achieving knowledge transfer in an open-set environment. Existing OS-SFDA methods mainly transfer the knowledge of known classes to the target domain through pseudo-supervised learning, while treating the private class samples in the target domain as outliers and aggregating them into an "unknown" class. However, this approach may lead to feature space degradation because these samples are semantically different. In addition, to distinguish unknown classes from known classes, previous OS-SFDA methods rely on statistical separation criteria, which may cause the problem of negative transfer, that is, the source domain knowledge actually degrades the performance during the target domain adaptation process.
[0003] Pseudo-label is a self-supervised method commonly used in the learning of unlabeled data. It generates "pseudo-labels" through the model's prediction of unlabeled data and uses these pseudo-labels for training in the target domain, thereby achieving knowledge transfer. To improve the reliability of pseudo-labels, uncertainty measures are combined to evaluate the confidence of the model's prediction of samples, so as to screen out more reliable pseudo-labels and effectively improve the effect of target domain adaptation. However, existing uncertainty measure methods usually highly rely on the model's prediction output, which may itself be biased or incorrect. This inaccurate screening may use incorrect pseudo-labels for training, thus affecting the model performance, especially when the class distribution in the target domain is significantly different from that in the source domain.
[0004] The basic idea of the diffuser is to transfer the knowledge of the source domain to the target domain through a certain mechanism, even if the target domain contains unknown classes that the source domain does not have. There are usually differences in the class distributions between the source domain and the target domain, and the target domain may have classes that have not been seen in the source domain. The role of the diffuser is to help the model effectively transfer the knowledge of the source domain to the target domain and handle these unknown classes. However, if the unknown classes in the target domain are similar to the existing classes in the source domain, the diffuser may misclassify them as known classes, resulting in a decrease in the adaptation accuracy of the model. In addition, if the distribution difference between the target domain and the source domain is too large, or if there are completely unknown classes in the target domain, the performance of the diffuser may drop significantly. At this time, it is difficult to effectively transfer the knowledge of the source domain to the target domain, and the model cannot adapt to the new data distribution, thus weakening its ability to handle new classes. Summary of the Invention
[0005] The purpose of the present invention is to provide an open-set source-free domain adaptation method based on uncertainty-driven and unknown diffuser, which solves the domain adaptation problem when the source data and the target data cannot be accessed simultaneously and the target domain contains unknown classes.
[0006] The technical solution adopted by the present invention is that the open-set source-free domain adaptation method based on uncertainty-driven and unknown diffuser includes the following steps: S1. Construct an initial model, obtain the source domain and the target domain datasets, pre-train the source domain dataset, retain the feature extractor , construct a classifier , and obtain a refined model; S2. Use the feature extractor to extract the feature representations of the target domain samples , and generate an initial pseudo-label assignment according to the refined model; S3. Refine the pseudo-labels by aggregating information from the nearest neighbor samples, optimize in the latest class space, and complete the exploration of the target class space; S4. Estimate the uncertainty of the pseudo-labels, select the samples with uncertainty as the training samples, and perform hierarchical class exploration on the target class space; S5. Through the negative learning classification loss, the contrastive loss NL-InfoNCELoss and the regularization term to prevent posterior collapse , and combine the optimization in the latest class space in S3 and the hierarchical class exploration in S4 to form a loss function; S6. Input the dataset into the network model for training, and optimize the model through the loss function.
[0007] The characteristics of the present invention also lie in: Use the feature extractor in S2 Extract the feature representation of the target domain samples The specific process is as follows: Use the feature extractor Extract features from the samples of the target domain dataset .
[0008] The specific process of generating the initial pseudo-label assignment according to the refinement model in S2 is as follows: For all the extracted features Perform unsupervised clustering to obtain K center points , where , , is the set of target samples assigned to the k-th cluster. By calculating the cosine similarity between the shared class prototype and the cluster center , align the shared class with some clusters, and match each column in with the cluster center with the highest similarity. The cluster centers not assigned to any shared class prototype will be regarded as the discovery prototypes of the target private classes and used to fill columns. After passing through the refinement model will be used to generate the initial pseudo-label assignment.
[0009] The specific process of refining the pseudo-labels by aggregating information from the nearest neighbor samples in S3 is as follows: Step 3.1: After clustering initialization, associate each target sample with a probability vector proportional to its similarity to the cluster center . Given the target feature and the cluster center , the probability vector is defined as the per-class similarity between and each cluster center . For each target sample, calculate the probability vector used to initialize the bank through the following formula : (1); where is the cosine distance and i is the class index; Step 3.2: Input the probability vector into the softmax function with the temperature parameter to obtain: (2); For all samples , the probability value of class c will be proportional to the similarity between the feature and the class center c; Step 3.3: Obtain the feature vector from the enhanced image , where is the target sample, is the weakly enhanced sample randomly drawn from the distribution . This vector is used to search for the neighbor of in the target feature space; Step 3.4: Aggregate the predictions of the neighbors through soft voting, and refine the pseudo-label as follows: (3); where I is the index set of the selected neighbors, is the softmax output, and c represents the class index; Step 3.5: To obtain the refined pseudo-label, select the class c with the highest probability, i.e.: (4); The refined pseudo-label is assigned to the target sample and used as a self-supervised signal.
[0010] The specific process of optimization in the latest class space in S3 is as follows: Step 3.6: The feature is the output of the feature extractor in the source pre-trained model, denoted as , where d is the dimension of the feature space, i represents the i-th sample, and the feature set Z consists of the features of n samples. The main network C in the unknown class diffuser maps the feature to the cluster assignment , where is the mapping operation of C, and the cluster assignment of the feature is denoted as ; Step 3.7: In the current class space K, obtain the discriminative cluster assignments by optimizing the cluster assignments, which are optimized according to the discriminative features Z, i.e.: obtain the pseudo-cluster assignment , and use the k-means algorithm to cluster the feature set Z on the current target class space K to generate the cluster assignments; Step 3.8: According to , generate the teacher cluster soft distribution based on the features, where (5); where, is the k-th element in , is the teacher cluster soft distribution of the feature , is the feature of the k-th center, is the number of features of the k-th main cluster, is the feature matrix The soft distribution of the teacher clusters for each feature is normalized; Step 3.9, Use the alignment loss to enforce the soft distribution of the teacher clusters to be consistent with the cluster assignment : (6); Step 3.10, By optimizing push the cluster assignment under the current target class space K to complete the exploration of the target class space.
[0011] The specific process of estimating the uncertainty of the pseudo-labels in S4 is as follows: Step 4.1, Uncertainty estimation by neighbor consensus method; Calculate the entropy of the pseudo-label as an estimator of the uncertainty of the pseudo-label. A smaller entropy value indicates low uncertainty, while a larger entropy value indicates high uncertainty. Define the uncertainty coefficient as follows: (7); where, represents the entropy of the pseudo-label, is the number of classes; Step 4.2, Uncertainty estimation by class separation method; Estimate the uncertainty of the pseudo-label by analyzing the distance between the sample and the prototypes of each class. For the target sample , select two classes i, j from whose values are the largest as the most likely pseudo-labels of the sample; Then, calculate the cosine distances between the target sample feature and the class prototypes W[i] and W[j]; Finally, obtain the uncertainty value as follows: (8).
[0012] The specific process of selecting samples with uncertainty as training samples in S4 is as follows: Step 4.3, Bernoulli sampling; Select training samples according to the uncertainty of the samples. Specifically, calculate the probability that the target sample has the correct pseudo-label and , respectively based on and : (9); (10); F is a monotonically decreasing function used to assign a higher probability to samples with low uncertainty and vice versa. Combine as follows: (11); Among them, is Bernoulli sampling with a success probability of 0 ≤ p ≤ 1, and ⊕ is a logical operator; finally, define as the set of samples selected based on uncertainty, and the samples in this set will be used for the calculation of classification loss.
[0013] The specific process of hierarchical category exploration of the target category space in S4 is as follows: Step 4.4, Obtain the discriminative cluster distribution in all splitting / merging cases based on the foregoing process, where the clusters before splitting / merging have been learned through the above method; Step 4.5, Use the sub-network in the unknown category diffuser to obtain the discriminative sub-cluster distribution. This sub-network contains K sub-clustering layers, and the output dimension of each layer is 2. Each sub-network generates sub-cluster assignments for each main cluster where is the operation of the k-th sub-network on the k-th main cluster; then, to generate the pseudo sub-cluster assignment, perform k-means clustering on the feature matrix of the k-th main cluster, and optimize the clusters through the following formula: (12); where is the element in the i-th row and j-th column of is the central feature of the j-th sub-cluster in the k-th main cluster; then, generate the teacher sub-cluster assignment for the k-th main cluster; Step 4.6, Use the alignment loss to optimize these sub-cluster assignments: (13); The training process of the sub-network is synchronized with the main network C to ensure the quality of sub-cluster assignments; Step 4.7, Generate a candidate set for splitting / merging according to the assignment distance between clusters; in the case of a large assignment distance between clusters, these cluster pairs or sub-cluster pairs are potential objects for splitting; while when the assignment distance between clusters is close, it indicates a high similarity between two main clusters and they need to be merged; based on this, the candidate set can be expressed as: (14); (15); where calculate the Euclidean distance of the cluster embedding center, and nif is the number of clusters to be merged or split; Step 4.8: By calculating the Hastings ratio between sub-clusters within the candidate set in the M-H framework and comparing it with 1, the target number of categories K is obtained, thereby exploring the target category space and refining the category structure of the target domain.
[0014] The specific process of S5 is as follows: To mitigate the impact of noisy pseudo-labels, in addition to adopting sample selection and exclusion strategies, negative learning loss is used as the classification loss, and positive loss is not used throughout the training process. Specifically, the classification loss used is as follows: (16); where is a complementary label, , which is randomly selected from the label set and does not include refined pseudo-labels; is the probability output of the strongly augmented image ; U is the set of samples with low uncertainty. During the training process, the refinement of pseudo-labels is gradually guided by negative learning loss and does not rely on positive loss; The contrastive loss NL-InfoNCELoss is adopted, introducing the negative learning principle into the standard InfoNCELoss. The formula of NL-InfoNCELoss is as follows: (17); (18); where is a negative sample randomly selected from the negative sample set, is the index set of samples that have never shared the same pseudo-label with the query sample in the past τ epochs. By optimizing NL-InfoNCELoss, the model is trained to increase the feature distance between the query sample and the randomly selected negative sample; The regularization term is adopted to prevent posterior collapse; the formula is defined as follows: (19); where , and σ is the softmax function; The total loss of the uncertainty estimation module is the weighted sum of the classification loss , the contrastive loss and the regularization term , specifically as follows: (20); where , and are hyperparameters; By combining the optimization in the latest category space and the hierarchical category exploration in steps S3 and S4, better results can ultimately be achieved in a broader category space. The total loss expression is as follows: (21).
[0015] The specific process of S6 is as follows: Input the dataset into the network model constructed in S2 to S5 for training, and optimize the model through the loss function in S5 to finally obtain the trained model and its performance evaluation results.
[0016] The beneficial effects of the present invention are as follows: The open-set unsupervised domain adaptation method based on uncertainty-driven and unknown diffuser provided by the present invention uses uncertainty-driven to assign pseudo-labels to the features of target samples, calculates the uncertainty through two methods for sample selection and pseudo-label refinement; introduces the NL-InfoNCE Loss contrast loss, which is a new loss function that integrates the negative learning paradigm into self-supervised contrast learning, helps to regularize the feature space and improve the robustness to pseudo-label noise; uses the unknown diffuser to explore a broader and more accurate target class space in the target domain, which is beneficial to the transfer of known knowledge and the generalization of unknown classes. By utilizing reliable known knowledge and clustering pseudo-labels in a broader class space, effective supervision is provided for the optimization process. Specific Embodiments
[0017] The present invention will be described in detail below in conjunction with specific embodiments.
[0018] Embodiment 1 The open-set unsupervised domain adaptation method based on uncertainty-driven and unknown diffuser proposed in this embodiment includes the following steps: S1. Construct an initial model, obtain the source domain and the target domain datasets, pre-train the source domain dataset, retain the feature extractor , construct a classifier , and obtain a refined model; S2. Use the feature extractor to extract the feature representations of the target domain samples, and generate an initial pseudo-label assignment according to the refined model; S3. Refine the pseudo-labels by aggregating information from the nearest neighbor samples, optimize in the latest category space, and complete the exploration of the target category space; S4. Estimate the uncertainty of the pseudo-labels, select the samples with uncertainty as the training samples, and conduct hierarchical category exploration of the target category space; S5. Through the negative learning classification loss, the contrastive loss NL-InfoNCE Loss, and the regularization term to prevent posterior collapse , and combined with the optimization in the latest class space in S3 and the hierarchical class exploration in S4 to form the loss function; S6. Input the dataset into the network model for training, and optimize the model through the loss function.
[0019] Embodiment 2 The open-set unsupervised domain adaptation method based on uncertainty-driven and unknown diffuser proposed in this embodiment includes the following steps: Based on Embodiment 1, in S2, use the feature extractor to extract the feature representation of the target domain samples The specific process is as follows: Use the feature extractor to extract features from the samples of the target domain dataset; ; The specific process of generating the initial pseudo-label assignment according to the refinement model in S2 is as follows: Perform unsupervised clustering on all the extracted features to obtain K center points , where , , is the set of target samples assigned to the k-th cluster. By calculating the cosine similarity between the shared class prototype and the cluster center , align the shared class with some clusters, and match each column in with the cluster center with the highest similarity. The cluster centers not assigned to any shared class prototype will be regarded as the discovery prototypes of the target private classes and used to fill 's columns. After passing through the refinement model will be used to generate the initial pseudo-label assignment.
[0020] Embodiment 3 The open-set unsupervised domain adaptation method based on uncertainty-driven and unknown diffuser proposed in this embodiment includes the following steps: Based on Embodiment 2, the specific process of refining the pseudo-labels by aggregating information from the nearest neighbor samples in S3 is as follows: Step 3.1. After clustering initialization, associate each target sample with a probability vector proportional to its similarity to the cluster center . Given the target feature and the cluster center , the probability vector is defined as with each cluster center For each type of similarity between, for each target sample, calculate the probability vector for initializing the bank through the following formula : (1); Among them, is the cosine distance, and i is the class index; Step 3.2. Input the probability vector into the softmax function with temperature parameter to obtain: (2); For all samples , the probability value of class c will be proportional to the similarity between the feature and the class center c; Step 3.3. Obtain the feature vector from the enhanced image, where is the target sample, is the weakly enhanced randomly drawn from the distribution , and this vector is used to search for neighbors of in the target feature space; Step 3.4. Aggregate the predictions of the neighbors through soft voting, and the pseudo-label is refined, and the formula is as follows: (3); Among them, I is the index set of the selected neighbors, is the softmax output, and c represents the class index; Step 3.5. To obtain the refined pseudo-label, select the class c with the highest probability, that is: (4); The refined pseudo-label is assigned to the target sample , and used as a self-supervised signal; The specific process of optimization in the latest category space in S3 is as follows: Step 3.6. The feature is the output of the feature extractor in the source pre-trained model, denoted as , where d is the dimension of the feature space, i represents the i-th sample, the feature set Z consists of the features of n samples, and the main network C in the unknown class diffuser maps the feature to the cluster assignment among them, is the mapping operation of C, and the cluster assignment of the feature is denoted as ; Step 3.7. Under the current class space K, obtain discriminative cluster assignments by optimizing the cluster assignments, which are optimized according to the discriminative features Z, that is: obtain the pseudo-cluster assignments , use the k-means algorithm to cluster the feature set Z on the current target class space K to generate cluster assignments; Step 3.8. According to , generate the teacher cluster soft distribution based on the features , where (5); Among them, is the k-th element in is the feature of the teacher cluster soft distribution, is the feature of the k-th center, is the number of features of the k-th main cluster, is the feature matrix the i-th element in; the teacher cluster soft distribution of each feature is normalized; Step 3.9. Use the alignment loss to enforce the consistency between the teacher cluster soft distribution and the cluster assignments : (6); Step 3.10. By optimizing , promote the cluster assignments under the current target class space K to complete the exploration of the target class space.
[0021] Example 4 The open-set unsupervised domain adaptation method based on uncertainty-driven and unknown diffuser proposed in this example includes the following steps: On the basis of Example 3, the specific process of estimating the uncertainty of the pseudo-labels in S4 is as follows: Step 4.1. Uncertainty estimation by neighbor consensus method; calculate the entropy of the pseudo-labels as the estimator of the pseudo-label uncertainty. A smaller entropy value indicates low uncertainty, and a larger entropy value indicates high uncertainty. Define the uncertainty coefficient as follows: (7); Among them, represents the entropy of the pseudo-labels, is the number of classes; Step 4.2. Uncertainty estimation by class separation method; estimate the uncertainty of the pseudo-labels by analyzing the distances between the samples and the prototypes of each class. For the target sample , from Select two classes i and j with the largest values from them as the most likely pseudo-labels of the sample; then, calculate the cosine distances between the target sample features and the class prototypes W[i] and W[j]; finally, obtain the uncertainty value as follows: (8); The specific process of selecting samples with uncertainty as training samples in S4 is as follows: Step 4.3, Bernoulli sampling; select training samples according to the uncertainty of the samples. Specifically, calculate the probability that the target sample has the correct pseudo-label and , respectively based on and : (9); (10); F is a monotonically decreasing function used to assign higher probabilities to samples with low uncertainty and vice versa. It is combined as follows: (11); where is Bernoulli sampling with a success probability of 0 ≤ p ≤ 1, and ⊕ is a logical operator; finally, define as the set of samples selected based on uncertainty, and the samples in this set will be used for the calculation of the classification loss; The specific process of hierarchical class exploration of the target class space in S4 is as follows: Step 4.4, obtain the discriminative cluster distribution in all split / merge cases based on the foregoing process, where the clusters before split / merge have been learned through the above method; Step 4.5, use the sub-network in the unknown class diffuser to obtain the discriminative sub-cluster distribution. This sub-network contains K sub-clustering layers, and the output dimension of each layer is 2. Each sub-network generates sub-cluster assignments for each main cluster , where is the operation of the k-th sub-network on the k-th main cluster; then, to generate the pseudo-sub-cluster assignment, perform k-means clustering on the feature matrix of the k-th main cluster, and optimize the clusters through the following formula: (12); where is the element in the i-th row and j-th column of , and is the central feature of the j-th sub-cluster in the k-th main cluster; then, generate the teacher sub-cluster assignment generated by the k-th main cluster ; Step 4.6. Optimize these sub-cluster assignments using the alignment loss: (13); The sub-network is trained synchronously with the main network C to ensure the quality of sub-cluster assignments; Step 4.7. Generate a candidate set for splitting / merging based on the assignment distances between clusters; in cases where the cluster assignment distances are far, these cluster pairs or sub-cluster pairs are potential objects for splitting; while when the cluster assignment distances are close, it indicates a high similarity between two main clusters and they need to be merged; based on this, the candidate set can be expressed as: (14); (15); where, calculate the Euclidean distance of the cluster embedding centers, and nif is the number of clusters to be merged or split; Step 4.8. By calculating the Hastings ratio between sub-cluster pairs within the candidate set in the M-H framework and comparing it with 1, obtain the target number of categories K, thereby exploring the target category space and refining the category structure of the target domain.
[0022] Example 5 The open-set passive domain adaptation method based on uncertainty-driven and unknown diffuser proposed in this example includes the following steps: Based on Example 4, The specific process of S5 is as follows: To reduce the impact of noisy pseudo-labels, in addition to using the sample selection and exclusion strategy, use the negative learning loss as the classification loss and do not use the positive loss throughout the training process. Specifically, the classification loss used is as follows: (16); where, is a complementary label, , is randomly selected from the label set and does not include refined pseudo-labels; is the probability output of the strongly augmented image ; U is a set of samples with low uncertainty. During the training process, the refinement of pseudo-labels is gradually guided by the negative learning loss and does not rely on the positive loss; Adopt the contrastive loss NL-InfoNCELoss, introduce the negative learning principle into the standard InfoNCELoss, and the formula of NL-InfoNCELoss is as follows: (17); (18); Among them, is a negative sample randomly selected from the negative sample set, is the index set of samples that have never shared the same pseudo-label with the query sample in the past τ cycles. By optimizing the NL-InfoNCE Loss, the model is trained to increase the feature distance between the query sample and the randomly selected negative sample; The regularization term is adopted to prevent posterior collapse; the formula is defined as follows: (19); Among them, , and σ is the softmax function; The total loss of the uncertainty estimation module is the classification loss , the contrastive loss , and the regularization term The weighted sum is as follows: (20); Among them, , and are hyperparameters; By combining the optimization in the latest category space and the hierarchical category exploration in steps S3 and S4, better results can be finally obtained in a wider category space. The total loss expression is: (21).
[0023] Embodiment 6 The open-set passive domain adaptation method based on uncertainty-driven and unknown diffuser proposed in this embodiment includes the following steps: S1. Construct an initial model, obtain the source domain and the target domain datasets, pre-train the source domain dataset, retain the feature extractor , construct a classifier , and obtain a refined model; In step 1, the first-stage training is first carried out, that is, the model is trained on a large amount of labeled source data. In the subsequent steps, the model will no longer contact the source samples but adapt to the unlabeled data in the target domain. During the pre-training process on the source dataset, the classifier will be trained to adapt to the categories in the shared class . The target domain not only contains the categories in the shared class , but also contains a set of private categories The samples of the target private class are divided into multiple "unknown" categories to make full use of the granularity of the target private class. The side effect of this choice is that the learned feature space will aggregate the samples from the private class semantically. Therefore, before the adaptation process, a new classifier needs to be constructed , whose weight matrix consists of two parts, namely . Among them, the weight set is initialized with the pre-trained weights obtained from the classifier , while is randomly initialized from a uniform distribution. This means that the pre-trained classifier is extended by adding columns that match the number of private classes that may exist in the target domain . It should be noted that is arbitrarily chosen because the true number of target private classes cannot be obtained in the open-set setting; S2. Use the feature extractor to extract the feature representation of the target domain samples , and generate the initial pseudo-label assignment according to the refinement model; The specific process of using the feature extractor to extract the feature representation of the target domain samples in S2 is as follows: Use the feature extractor to extract features from the samples of the target domain dataset; ; The specific process of generating the initial pseudo-label assignment according to the refinement model in S2 is as follows: Perform unsupervised clustering on all the extracted features to obtain K center points , where , , is the set of target samples assigned to the k-th cluster. By calculating the cosine similarity between the shared class prototype and the cluster center , align the shared class with some clusters, and match each column in with the cluster center with the highest similarity. This strategy solves the class misalignment problem by determining a suitable permutation, which rearranges the cluster centers according to the matching relationship between the shared class prototype and the cluster centers ; The cluster centers not assigned to any shared class prototype will be regarded as the discovery prototypes of the target private class and used to fill the columns. After passing through the refinement model will be used to generate the initial pseudo-label assignment; S3. Refine the pseudo-labels by aggregating the information from the samples in the nearest neighborhood, optimize in the latest class space, and complete the exploration of the target class space; The specific process of refining the pseudo-labels by aggregating information from the nearest neighbor samples in S3 is as follows: Step 3.1. After clustering initialization, each target sample is associated with a probability vector proportional to its similarity to the cluster center Given the target feature and the cluster center , the probability vector is defined as the per-class similarity between each cluster center For each target sample, the probability vector used to initialize the bank is calculated by the following formula : (1); where is the cosine distance and i is the class index; Step 3.2. The probability vector is input into the softmax function with the temperature parameter to obtain: (2); For all samples , the probability value of class c will be proportional to the similarity between the feature and the class center c; Step 3.3. Obtain the feature vector from the augmented image, where is the target sample and is the weakly augmented randomly drawn from the distribution . This vector is used to search for the neighbors of in the target feature space; Step 3.4. Aggregate the predictions of the neighbors by soft voting, and the pseudo-label is refined, and the formula is as follows: (3); where I is the set of indices of the selected neighbors, is the softmax output, and c represents the class index; Step 3.5. To obtain the refined pseudo-label, select the class c with the highest probability, that is: (4); The refined pseudo-label is assigned to the target sample and used as a self-supervised signal; The specific process of optimization in the latest category space in S3 is as follows: Step 3.6. The feature is the feature extractor in the source pre-trained model The output is denoted as , where d is the dimension of the feature space, i represents the i-th sample, the feature set Z consists of the features of n samples, and the main network C in the unknown class diffuser maps the features to cluster assignments in which is the mapping operation of C, and the feature cluster assignment is denoted as ; Step 3.7: In the current class space K, obtain discriminative cluster assignments by optimizing the cluster assignments, and these cluster assignments are optimized according to the discriminative features Z, that is: obtain the pseudo-cluster assignments , and use the k-means algorithm to cluster the feature set Z on the current target class space K to produce cluster assignments; Step 3.8: According to , generate the teacher cluster soft distribution based on the features , where (5); Among them, is the k-th element in is the feature teacher cluster soft distribution, is the feature of the k-th center, is the number of features of the k-th main cluster, is the feature matrix the i-th element in; the teacher cluster soft distribution of each feature is normalized; Step 3.9: Use the alignment loss to enforce the teacher cluster soft distribution to be consistent with the cluster assignment : (6); Step 3.10: By optimizing , promote the cluster assignment under the current target class space K to complete the exploration of the target class space; S4. Estimate the uncertainty of the pseudo-labels, select the samples with uncertainty as the training samples, and conduct hierarchical class exploration on the target class space; The specific process of estimating the uncertainty of the pseudo-labels in S4 is as follows: Step 4.1: Uncertainty estimation by the neighbor consensus method; calculate the entropy of the pseudo-labels as the estimator of the uncertainty of the pseudo-labels. A smaller entropy value indicates low uncertainty, and a larger entropy value indicates high uncertainty. Define the uncertainty coefficient as follows: (7); Among them, Denotes the entropy of the pseudo-label, is the number of classes; Step 4.2, Uncertainty estimation of the class separation method; Estimate the uncertainty of the pseudo-label by analyzing the distance between the sample and the prototypes of each class. For the target sample , select two classes i, j from whose values are the largest as the most likely pseudo-labels of the sample; Then, calculate the cosine distances between the target sample feature and the class prototypes W[i] and W[j]; Finally, obtain the uncertainty value as follows: (8); The specific process of selecting samples with uncertainty in S4 as training samples is: Step 4.3, Bernoulli sampling; Select training samples according to the uncertainty of the samples. Specifically, calculate the probability that the target sample has the correct pseudo-label and , respectively based on and : (9); (10); F is a monotonically decreasing function used to assign a higher probability to samples with low uncertainty and vice versa. It is combined as follows: (11); where, is Bernoulli sampling with a success probability of 0 ≤ p ≤ 1, and ⊕ is a logical operator; Finally, define as the set of samples selected based on uncertainty, and the samples in this set will be used for the calculation of the classification loss; The specific process of hierarchical class exploration of the target class space in S4 is: Step 4.4, Obtain the discriminative cluster distribution in all split / merge cases based on the foregoing process, where the clusters before split / merge have been learned by the above method; Step 4.5, Use the sub-network in the unknown class diffuser to obtain the discriminative sub-cluster distribution. This sub-network contains K sub-clustering layers, and the output dimension of each layer is 2. Each sub-network generates sub-cluster assignments for each main cluster , where is the operation of the k-th sub-network on the k-th main cluster; Then, to generate pseudo-sub-cluster assignments, perform k-means clustering on the feature matrix of the k-th main cluster, and optimize the clusters through the following formula: (12); Among them, is the element in the i-th row and j-th column of and ; Step 4.6. Optimize these sub-cluster assignments using the alignment loss: (13); The sub-network is trained synchronously with the main network C to ensure the quality of sub-cluster assignments; Step 4.7. Generate a candidate set for splitting / merging based on the assignment distance between clusters; in the case of a relatively large assignment distance between clusters, these cluster pairs or sub-cluster pairs are potential objects for splitting; while in the case of a relatively small assignment distance between clusters, it indicates a relatively high similarity between two main clusters and they need to be merged; based on this, the candidate set can be expressed as: (14); (15); Among them, Calculate the Euclidean distance of the cluster embedding center, and nif is the number of clusters to be merged or split; Step 4.8. By calculating the Hastings ratio between sub-cluster pairs within the candidate set in the M-H framework and comparing it with 1, obtain the target number of categories K, thereby exploring the target category space and refining the category structure of the target domain; S5. Through the negative learning classification loss, the contrastive loss NL-InfoNCELoss, and the regularization term to prevent posterior collapse, and combine the optimization in the latest category space in S3 and the hierarchical category exploration in S4 to form a loss function; The specific process of S5 is as follows: To reduce the impact of noisy pseudo-labels, in addition to adopting sample selection and exclusion strategies, use the negative learning loss as the classification loss and do not use the positive loss throughout the training process. Specifically, the classification loss used is as follows: (16); Among them, is a complementary label, , which is randomly selected from the label set and does not include refined pseudo-labels; is the probability output of the strongly augmented image ; U is a set of samples with low uncertainty. During the training process, the refinement of pseudo-labels is gradually guided by the negative learning loss and does not rely on the positive loss; The contrastive loss NL-InfoNCELoss is adopted to introduce the negative learning principle into the standard InfoNCELoss. The formula of NL-InfoNCELoss is as follows: (17); (18); where, is a negative sample randomly selected from the negative sample set, is the index set of samples that have never shared the same pseudo-label with the query sample in the past τ epochs. By optimizing NL-InfoNCELoss, the model is trained to increase the feature distance between the query sample and the randomly selected negative sample; The regularization term is adopted to prevent posterior collapse; the formula is defined as follows: (19); where, and σ is the softmax function; The total loss of the uncertainty estimation module is the weighted sum of the classification loss , the contrastive loss and the regularization term , specifically as follows: (20); where, , and are hyperparameters; By combining the optimization in the latest class space and the hierarchical class exploration in steps S3 and S4, better results can finally be obtained in a broader class space. The total loss expression is: (21); S6. Input the dataset into the network model for training and optimize the model through the loss function; The specific process of S6 is: Input the dataset into the network model constructed in S2 to S5 for training and optimize the model through the loss function in S5 to finally obtain the trained model and its performance evaluation results.
Claims
1. An open set passive field adaptive method based on uncertain drivers and unknown diffusers, characterized in that: The following steps are involved: S1. Build the initial model and obtain the source domain With the target domain Dataset, source domain The dataset is pre-trained and the feature extractor is retained , build a classifier , and obtain a refined model; S2. Using feature extractor Extract feature representation of target domain samples , generate initial pseudo-label assignments based on the refined model; S3, refine the pseudo-labels by aggregating information from the nearest neighbor samples, optimize them in the latest category space, and complete the exploration of the target category space; S4, estimate the uncertainty of pseudo labels, select samples with uncertainty as training samples, and perform hierarchical category exploration on the target category space; S5. Regularization term to prevent posterior collapse through negative learning classification loss, contrast loss NL-InfoNCELoss , and combines the optimization in the latest category space in S3 and the hierarchical category exploration in S4 to form a loss function; S6. Input the data set into the network model for training and optimize the model through the loss function.
2. The open set passive field adaptive method based on uncertain drive and unknown diffuser according to claim 1 is characterized in that: Using the feature extractor described in S2 Extract feature representation of target domain samples The specific process is: using feature extractor From the target domain Extract features from samples in the dataset .
3. The open set passive field adaptive method based on uncertain drive and unknown diffuser according to claim 1 is characterized in that: The specific process of generating the initial pseudo-label assignment according to the refined model described in S2 is: Perform unsupervised clustering to obtain K center points ,in , , is the target sample set assigned to the kth cluster, by calculating the shared class prototype With cluster center The cosine similarity between , aligns the shared classes with the partial clusters, and Each column in is matched with the cluster center with the highest similarity. The cluster centers that are not assigned to any shared class prototypes are considered as discovered prototypes of the target private class and are used to fill Columns, after the refined model Will be used to generate the initial pseudo-label assignments.
4. The open set passive field adaptive method based on uncertain drive and unknown diffuser according to claim 1 is characterized in that: The specific process of refining pseudo labels by aggregating information from the nearest neighbor samples described in S3 is: Step 3.1: After clustering is initialized, each target sample is associated with a cluster center. Similarity proportional probability vector associated, given the target feature and cluster centers , the probability vector Defined as With each cluster center For each target sample, the probability vector used to initialize the bank is calculated by the following formula: : (1); in, is the cosine distance, i is the class index; Step 3.2: Convert the probability vector Input to the temperature parameter In the softmax function, we get: (2); For all samples , the probability value of class c will be the same as the feature It is proportional to the similarity between the class centers c; Step 3.3: Obtain feature vector from enhanced image ,in is the target sample, From the distribution A weak boost randomly drawn from , which is used to search in the target feature space Neighbors; Step 3.4: Aggregate neighbors’ predictions through soft voting, pseudo labels After being refined, the formula is as follows: (3); Among them, I is the index set of the selected neighbors, is the softmax output, c represents the class index; Step 3.5: In order to obtain refined pseudo-labels, select the class c with the maximum probability, that is: (4); The refined pseudo-labels are assigned to the target samples , and used as a self-supervisory signal.
5. The open set passive field adaptive method based on uncertain drive and unknown diffuser according to claim 1 is characterized in that: The specific process of optimizing in the latest category space described in S3 is: Step 3.6: Feature Extractor in Source Pre-trained Model The output of , where d is the dimension of the feature space, i represents the i-th sample, the feature set Z consists of the features of n samples, and the main network C in the unknown class diffuser maps the features to cluster assignments Among them is a mapping operation of C, with the following characteristics: The cluster allocation is denoted as ; Step 3.7: Under the current category space K, obtain discriminative cluster allocations by optimizing cluster allocations. These cluster allocations are optimized according to the discriminative features Z, that is, obtain pseudo cluster allocations , use the k-means algorithm to cluster the feature set Z in the current target category space K to produce cluster assignments; Step 3.8: According to , Generate teacher cluster soft distribution based on features ,in (5); in, yes The kth element in It is a feature The teacher cluster soft distribution, is the characteristic of the kth center, is the number of features of the kth main cluster, is the feature matrix The i-th element in ; the teacher cluster soft distribution of each feature is normalized; Step 3.9: Use alignment loss to enforce soft distribution of teacher clusters With cluster allocation Consistent: (6); Step 3.10: Optimize and promote the cluster allocation under the current target category space K to complete the exploration of the target category space.
6. The open set passive field adaptive method based on uncertain drive and unknown diffuser according to claim 1 is characterized in that: The specific process of estimating the uncertainty of pseudo labels described in S4 is: Step 4.
1. Uncertainty estimation of the neighbor consensus method; calculating the entropy of the pseudo-label , as an estimator of pseudo-label uncertainty, a small entropy value indicates low uncertainty, and a large entropy value indicates high uncertainty. The uncertainty coefficient is defined as as follows: (7); in, represents the entropy of pseudo labels, is the number of categories; Step 4.2: Uncertainty estimation of class separation method: The uncertainty of pseudo labels is estimated by analyzing the distance between samples and various prototypes. ,from Select two classes i and j with the largest values as the most likely pseudo labels for the sample; then calculate the target sample features The cosine distance from the class prototypes W[i] and W[j]; finally, the uncertainty value is obtained as follows: (8)。 7. The open set passive field adaptive method based on uncertain drive and unknown diffuser according to claim 1 is characterized in that: The specific process of selecting uncertain samples as training samples in S4 is: Step 4.3, Bernoulli sampling: select training samples based on the uncertainty of the samples. Specifically, calculate the probability that the target sample has the correct pseudo-label and , based on and : (9); (10); F is a monotonically decreasing function, which is used to assign higher probabilities to samples with low uncertainty, and vice versa, and is combined as follows: (11); in, is Bernoulli sampling, with a success probability of 0≤p≤1, and ⊕ is a logical operator; finally, define To select a sample set based on uncertainty, the samples in this set will be used to calculate the classification loss.
8. The open set passive field adaptive method based on uncertain drive and unknown diffuser according to claim 1 is characterized in that: The specific process of performing hierarchical category exploration on the target category space described in S4 is: Step 4.4, based on the above process, obtain the discriminative cluster distribution of all split / merge cases, where the clusters before split / merge have been learned by the above method; Step 4.5: Use subnetwork in unknown class diffuser , obtain the discriminative sub-cluster distribution. The sub-network contains K sub-clustering layers, each with an output dimension of 2. Generate sub-cluster allocation for each main cluster ,in is the operation of the kth sub-network on the kth main cluster; then, to generate pseudo sub-cluster allocation, the feature matrix of the kth main cluster is Perform k-means clustering, and optimize the clusters using the following formula: (12); in, yes The element in row i and column j in is the central feature of the jth sub-cluster in the kth main cluster; then, the kth main cluster is generated to generate the teacher sub-cluster allocation ; Step 4.6: Use alignment loss to optimize these subcluster assignments: (13); Subnetwork The training process of is carried out synchronously with the main network C to ensure the quality of sub-cluster assignment; Step 4.7: Generate a candidate set for splitting / merging based on the distribution distance between clusters. When the cluster distribution distance is far, these cluster pairs or sub-cluster pairs are potential objects for splitting. When the cluster distribution distance is close, it means that the similarity between the two main clusters is high and needs to be merged. Based on this, the candidate set can be expressed as: (14); (15); in, Calculate the Euclidean distance of the cluster embedding center, nif is the number of clusters merged or split; Step 4.8: By calculating the Hasting ratio between sub-cluster pairs in the candidate set in the MH framework and comparing it with 1, the number of target categories K is obtained, thereby exploring the target category space and refining the category structure of the target domain.
9. The open set passive field adaptive method based on uncertain drive and unknown diffuser according to claim 1 is characterized in that: The specific process of S5 is as follows: In order to mitigate the impact of noisy pseudo-labels, in addition to adopting sample selection and exclusion strategies, negative learning loss is used as classification loss, and positive loss is not used throughout the training process. Specifically, the classification loss used is as follows: (16); in, is a complementary label. , is randomly selected from the label set and does not include the refined pseudo-labels; It is a strongly enhanced image The probability output of ; U is a set of samples with low uncertainty. During the training process, the refinement of pseudo labels is gradually guided by negative learning loss without relying on positive loss; The contrast loss NL-InfoNCELoss is used to introduce the negative learning principle into the standard InfoNCELoss. The formula of NL-InfoNCELoss is as follows: (17); (18); in, is a negative sample randomly selected from the negative sample set, is the set of indices of samples that have never shared the same pseudo-label with the query sample in the past τ cycles. By optimizing NL-InfoNCELoss, the model is trained to increase the feature distance between the query sample and the randomly selected negative samples; Using regularization term , to prevent posterior collapse; the formula is defined as follows: (19); in, , σ is the softmax function; The total loss of the uncertainty estimation module is the classification loss , contrast loss and the regularization term The weighted sum of is as follows: (20); in, , and is a hyperparameter; By combining the optimization in the latest category space and the hierarchical category exploration in steps S3 and S4, better results can be achieved in a wider category space. The total loss expression is: (21)。 10. The open set passive field adaptive method based on uncertain drive and unknown diffuser according to claim 1, characterized in that: The specific process of S6 is: inputting the data set into the network model constructed from S2 to S5 for training, and optimizing the model through the loss function in S5, and finally obtaining the trained model and its performance evaluation results.