Few-sample-domain adaptation method based on prototype guide diffusion alignment
By using the prototype-guided diffusion alignment method in the few-sample unsupervised domain adaptation scenario, the distribution alignment of the target domain samples is solved, and the classification accuracy is improved.
Patent Information
- Application Number
- CN202510344336.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-24
AI Technical Summary
In the scenario of few-sample unsupervised domain adaptation, the labeling of the source domain samples is sparse, making it difficult to conduct supervised training and accurate domain alignment. The existing methods have the problem of loss of discriminant semantic information.
Using a prototype-guided diffusion alignment method, by obtaining the category prototype of the source domain and the target domain, using the diffusion model to distribute and align the target domain samples, and classify them in the classifier trained on the source domain, accept classification results above the threshold, and re-align samples below the threshold.
The loss of discriminant semantic information in latent feature representation is avoided, and the accuracy of adaptation of unsupervised domains in small samples is improved.
Smart Images

Figure CN120198736A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the problem of few-shot unsupervised domain adaptation in the field of machine learning, specifically a few-shot domain adaptation method based on prototype-guided diffusion alignment. Background Art
[0002] In traditional machine learning, researchers usually directly apply the model (such as a classifier or a regressor) learned from the training set to the test set. This approach generally requires the training set and the test set to satisfy the independent and identically distributed assumption. However, in many real-world applications, due to differences in data sample sources, acquisition methods, feature expressions, etc., the training data and the test data have different sample distributions. To solve the above problems, unsupervised domain adaptation is proposed as a new machine learning paradigm and used to solve the transfer task when the source domain and the target domain are in different distributions but have the same task. Currently, unsupervised domain adaptation has achieved good performance in many scenarios such as object detection and cross-domain classification. However, in practical problems, such as in the medical field, due to the difficulty of collecting sample labels, the labels in the source domain are often sparse, that is, only partially labeled. Therefore, the few-shot unsupervised domain adaptation scenario is proposed to solve such problems.
[0003] The main challenge of few-shot unsupervised domain adaptation is that the labels of the source domain samples are sparse, so it is difficult to perform supervised training and accurate domain alignment. Due to the great challenges, there are few existing research methods and they are highly homogeneous. Specifically, the existing methods mainly construct auxiliary tasks through self-supervised training to learn a relatively reliable feature representation in the latent space, thereby implicitly achieving alignment. Although this approach has achieved certain results, since the self-supervised tasks rely on artificial construction and are essentially irrelevant to the downstream tasks, the learned latent feature representation will inevitably cause the loss of task-related discriminant information, resulting in unreliable distribution alignment. Therefore, a few-shot domain adaptation method based on prototype-guided diffusion alignment is proposed. Summary of the Invention
[0004] Object of the Invention: The object of the present invention is to aim at the deficiencies of the prior art and propose a few-shot domain adaptation method based on prototype-guided diffusion alignment, which performs domain alignment using comprehensive semantic information from the original space to solve the defect problems pointed out in the background art. By using the method disclosed in the present invention, the problem of losing discriminative semantic information in the latent feature representation can be avoided, thereby improving the accuracy of few-shot unsupervised domain adaptation classification.
[0005] Technical Solution: The few-shot domain adaptation method based on prototype-guided diffusion alignment specifically comprises the following steps:
[0006] S1. Obtain the source domain dataset and the target domain dataset. Input the images in the datasets into the representation network to train the feature extractor and the classifier. Input the output of the feature extractor into the clustering model, and use the labels obtained by the clustering model to retrain the network. Repeat the training until the trained feature extractor, classifier, and the class prototypes of the source domain and the target domain are finally obtained.
[0007] S2. Obtain the diffusion model, and use the source domain images and the class prototypes to train the trained diffusion model through the noise estimation loss.
[0008] S3. Domain adaptation: Input the target domain images and the class prototypes into the trained diffusion model obtained in step S2 to obtain the aligned target domain samples, and input them into the trained classifier obtained in step S1 to obtain their classification probabilities. Accept the classification results with classification probabilities higher than the threshold, and add more noise to the images with classification probabilities lower than the threshold and realign them with the trained diffusion model obtained in step S2.
[0009] Preferably, in S1, after normalizing and preprocessing the source domain and target domain images, input them into the representation network of S1.
[0010] Preferably, in step S1, the representation learning network includes a feature extractor E and a classifier C. The feature extractor E is used to extract the feature representation of the source domain image and the target domain image where i is the image index, N is the total number of source domain images, and N is the total number of source domain images; s is the total number of source domain images; T is the total number of source domain images;
[0011] The classifier C is used to classify the feature representation of the source domain image to obtain the pseudo label and classify the feature representation of the target domain image to obtain the pseudo label
[0012] In step S1, the clustering model P is used to extract the prototype from the feature representation of the source domain image and extract the prototype from the feature representation of the target domain image where k is the class index and K is the total number of classes. and extract the prototype from the feature representation of the target domain image where k is the class index and K is the total number of classes.
[0013] Preferably, the feature extractor E is trained by the semantic invariance loss and the spatial proximity loss as follows:
[0014]
[0015] Among them, represents the semantic invariance loss, represents the spatial proximity loss, f i represents the feature representation of the i-th image extracted by the feature extractor E, D ls , D us , D t , respectively represent the labeled source domain, the unlabeled source domain, and the target domain. P() represents obtaining the prototype of the features in the parentheses, and P*() represents obtaining the prototype of the features in the parentheses in the corresponding opposite domain, u j represents the feature prototype of the j-th class, represents the feature prototype of the j-th class in the domain opposite to the current domain, and τ is the temperature parameter;
[0016] The classifier C is trained by the cross-entropy loss as follows:
[0017]
[0018] Among them, represents the classification loss,, and CE() represents the cross-entropy loss, represents the pseudo-label of the i-th sample, y i represents the true label of the i-th sample;
[0019] The clustering model P obtains the clustering prototype as follows:
[0020]
[0021] Among them, represents the feature prototype of the j-th class in the target domain at the r-th round of update, represents the feature representation of the i-th image in the target domain extracted by the feature extractor E, represents the set of all samples in the j-th class in the source domain, and φ is the temperature parameter.
[0022] Preferably, in the step S2, the diffusion model is a U-Net network, and the training of the diffusion model is as follows:
[0023]
[0024] Among them, x t represents the image obtained by adding t steps of noise to the image x0, ∈ represents random Gaussian noise, ε θ represents the noise predicted by the diffusion model, P(x0) represents the class prototype corresponding to x0, and t is the noise step size.
[0025] Preferably, in the step S3, add T i steps of noise to the target domain image x, Ti ∈ [100, 200, ..., 1000], and then perform alignment sampling as follows:
[0026]
[0027] where β i ∈ (0, 1) represents the noise scale, represents the image after adding noise with -1 steps, x represents the original image without noise, represents the noise predicted by the biased diffusion network, calculated as:
[0028]
[0029] where w represents the hyperparameter.
[0030] Preferably, in step S3, first, perform T0 times of alignment sampling on the target domain image x to obtain the aligned target domain image Input it into the trained classifier C obtained in step S1 to obtain the maximum classification probability g. When g is greater than or equal to the given threshold h, accept the category of the target domain image x as the category corresponding to the maximum classification probability g. Otherwise, add T1 steps of noise to the target domain image x and perform alignment sampling and classification discrimination again. When T i = 1000 and the classification result of the target domain image x is still not accepted, then select the category corresponding to the maximum classification probability among the previous classification results as the category of the target domain image x.
[0031] Beneficial effects: The few-shot domain adaptation method based on prototype-guided diffusion alignment of the present invention learns the class prototypes of the source domain and the target domain to guide the alignment of the target domain sample distribution towards the source domain distribution in the original space, and then uses the classifier on the source domain to classify the aligned samples, accepts the classification results greater than the threshold, and realigns and classifies the samples with classification probabilities less than the threshold, thereby detecting unknown class samples. Compared with the existing few-shot unsupervised domain adaptation methods, this method performs distribution alignment in the original space, avoiding the loss of discriminative semantic information, and thus improving the accuracy of few-shot unsupervised domain adaptation classification. Brief Description of the Drawings
[0032] Figure 1 is a flowchart of an open-set recognition method based on representation learning and diffusion models;
[0033] Figure 2 is a model structure diagram of an open-set recognition method based on representation learning and diffusion models; Detailed Embodiments
[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0035] Embodiment
[0036] Please refer to Figure 1-2 , the present invention provides a technical solution: a few-shot domain adaptation method based on prototype-guided diffusion alignment, including the following steps:
[0037] S1. Obtain the source domain dataset and the target domain dataset, input the images in the dataset into the representation network to train the feature extractor and the classifier, input the output of the feature extractor into the clustering model, retrain the network with the labels obtained by the clustering model, and repeat the training to finally obtain the trained feature extractor, classifier, and the class prototypes of the source domain and the target domain;
[0038] S2. Obtain the diffusion model, and train the obtained diffusion model through the noise estimation loss using the source domain images and the class prototypes;
[0039] S3. Domain adaptation: Input the target domain images and the class prototypes into the trained diffusion model obtained in step S2 to obtain the aligned target domain samples, input them into the trained classifier obtained in step S1 to obtain their classification probabilities, accept the classification results with classification probabilities higher than the threshold, and add more noise to the images with classification probabilities lower than the threshold and realign them using the trained diffusion model obtained in step S2.
[0040] In this embodiment, specifically: in the S1, after normalizing and preprocessing the source domain and target domain images, input them into the representation network in S1.
[0041] In this embodiment, specifically: in the step S1, the representation learning network includes a feature extractor E and a classifier C. The feature extractor E is used to extract the feature representation of the source domain image and the target domain image Extract the feature representation of the source domain image and the feature representation of the target domain image i is the image index, N s is the total number of source domain images, N T is the total number of source domain images;
[0042] The classifier C is used to classify the feature representation of the source domain image to obtain the pseudo label and the feature representation of the target domain image Perform classification to obtain pseudo-labels
[0043] In the step S1, the clustering model P is used to extract prototypes from the feature representations of source domain images Extract prototypes And extract prototypes from the feature representations of target domain images Extract prototypes k is the class index, and K is the total number of classes
[0044] In this embodiment, specifically: the feature extractor E is trained by semantic invariance loss and spatial proximity loss as follows
[0045]
[0046]
[0047] Among them, Represents the semantic invariance loss Represents the spatial proximity loss, f i Represents the feature representation of the i-th image extracted by the feature extractor E, D ls , D us , D t , respectively represent the labeled source domain, the unlabeled source domain, and the target domain. P() represents obtaining the prototype of the feature in the parentheses, P * () represents obtaining the prototype of the feature in the parentheses in the corresponding opposite domain, u j Represents the feature prototype of the j-th class Represents the feature prototype of the j-th class in the domain opposite to the current domain, and τ is the temperature parameter
[0048] The classifier C is trained by cross-entropy loss as follows
[0049]
[0050] Among them, Represents the classification loss, and CE() represents the cross-entropy loss Represents the pseudo-label of the i-th sample, y i Represents the true label of the i-th sample
[0051] The clustering model P obtains clustering prototypes as follows
[0052]
[0053] Among them, Represents the feature prototype of the j-th class in the target domain during the r-th update Represents the feature representation of the i-th image in the target domain extracted by the feature extractor E denotes the set of all samples in the j-th class in the source domain, and φ is the temperature parameter.
[0054] In this embodiment, specifically: in the step S2, the diffusion model is a U-Net network, and the training of the diffusion model is as follows:
[0055]
[0056] where x t denotes the image obtained by adding t steps of noise to the image x0, ∈ denotes random Gaussian noise, and ε θ denotes the noise predicted by the diffusion model, P(x0) denotes the class prototype corresponding to x0, and t is the noise step size.
[0057] In this embodiment, specifically: in the step S3, T i steps of noise are added to the target domain image x, and T i ∈ [100, 200,..., 1000], and then the alignment sampling is performed as follows:
[0058]
[0059] where β i ∈ (0, 1) denotes the noise scale, denotes the image after adding t - 1 steps of noise, x denotes the original image without noise, denotes the noise predicted by the biased diffusion network, and is calculated as:
[0060]
[0061] where w represents the hyperparameter.
[0062] In this embodiment, specifically: in the step S3, first, after performing T0 times of alignment sampling on the target domain image x, the aligned target domain image is input into the trained classifier C obtained in the step S1 to obtain the maximum classification probability g. When g is greater than or equal to the given threshold h, the class of the target domain image x is accepted as the class corresponding to the maximum classification probability g, otherwise T1 steps of noise are added to the target domain image x to re-perform alignment sampling and classification discrimination. When T i = 1000 and the classification result of the target domain image x is still not accepted, then the class corresponding to the maximum classification probability among the previous classification results is selected as the class of the target domain image x.
[0063] Working principle or structural principle. When in use, the present invention proposes a few-shot domain adaptation method based on prototype-guided diffusion alignment, which uses prototype learning to obtain the prototypes and pseudo-labels of the source domain and the target domain, and then trains a diffusion network to align the target domain distribution to the source domain distribution. After the samples are aligned, the classifier on the source domain is used to predict the classification probability and determine whether the sample accepts the classification result according to a threshold. If not, more significant noise is added to realign the sample.
[0064] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A few-shot domain adaptation method based on prototype-guided diffusion alignment, characterized by: The following steps are involved: S1. Obtain source domain data sets and target domain data sets, input images in the data sets into the representation network to train feature extractors and classifiers, input the output of the feature extractor into the clustering model, retrain the network using the labels obtained by the clustering model, and repeat the training to finally obtain the trained feature extractors, classifiers, and category prototypes of the source domain and target domain; S2, obtain the diffusion model, use the source domain image and the category prototype to obtain the trained diffusion model through noise estimation loss training; S3, domain adaptation: input the target domain image and category prototype into the trained diffusion model obtained in step S2 to obtain the aligned target domain sample, input it into the trained classifier obtained in step S1 to obtain its classification probability, receive the classification results with classification probability higher than the threshold, increase the noise of the image with classification probability lower than the threshold and re-align it with the trained diffusion model obtained in step S2.
2. The method for few-shot domain adaptation based on prototype-guided diffusion alignment according to claim 1, characterized in that: In S1, the source domain and target domain images are normalized and preprocessed before being input into the representation network of S1.
3. The method for few-shot domain adaptation based on prototype-guided diffusion alignment according to claim 1, characterized in that: In step S1, the representation learning network includes a feature extractor E and a classifier C. The feature extractor E is used to extract the feature from the input source domain image. and the target domain image Extract feature representation of source domain image and the feature representation of the target domain image i is the image index, N s is the total number of source domain images, N T is the total number of source domain images; The classifier C is used to represent the features of the source domain image. Classify and obtain pseudo labels And the feature representation of the target domain image Classify and obtain pseudo labels In step S1, the clustering model P is used to represent the feature of the source domain image Extract the prototype And the feature representation of the target domain image Extract the prototype k is the category index and K is the total number of categories.
4. The method for few-sample domain adaptation based on prototype-guided diffusion alignment according to claim 3, characterized in that: The feature extractor E is trained by semantic invariance loss and spatial proximity loss as follows: in, represents the semantic invariance loss, represents the spatial proximity loss, f i represents the feature representation of the i-th image extracted by the feature extractor E, D ls , D us , D t , respectively represent the labeled source domain, the unlabeled source domain and the target domain, P() represents the prototype for obtaining the features in the brackets, P * () indicates obtaining the prototype of the features in the brackets in the corresponding opposite domain, u j represents the feature prototype of the jth class, represents the characteristic prototype of the jth class in the domain opposite to the current domain, and τ is the temperature parameter; The classifier C is trained by the cross entropy loss as follows: in, represents classification loss, CE() represents cross entropy loss, represents the pseudo label of the i-th sample, y i Represents the true label of the i-th sample; The clustering model P obtains the clustering prototype as follows: in, represents the feature prototype of the jth class in the target domain during the rth round of update, represents the feature representation of the i-th image in the target domain extracted by the feature extractor E, represents the set of all samples in the jth class in the source domain, and φ is the temperature parameter.
5. The method for few-shot domain adaptation based on prototype-guided diffusion alignment according to claim 1, characterized in that: In step S2, the diffusion model is a U-Net network, and the training of the diffusion model is as follows: Among them, x t represents the image obtained by adding t steps of noise to the image x0, ε represents random Gaussian noise, ε θ represents the noise predicted by the diffusion model, P(x0) represents the class prototype corresponding to x0, and t is the noise step size.
6. The method for few-shot domain adaptation based on prototype-guided diffusion alignment according to claim 1, characterized in that: In step S3, T is added to the target domain image x. i Step noise, T i ∈[100,200,...,1000], and then perform aligned sampling as follows: in represents the noise scale, represents the image after adding t-1 steps of noise, and x represents the original image without noise. represents the noise predicted by the biased diffusion network and is calculated as: Where w represents a hyperparameter.
7. The method for few-sample domain adaptation based on prototype-guided diffusion alignment according to claim 6, characterized in that: In step S3, first, the target domain image x is aligned and sampled T0 times to obtain an aligned target domain image It is input into the trained classifier C obtained in step S1 to obtain the maximum classification probability g. When g is greater than or equal to the given threshold h, the category of the target domain image x is accepted as the category corresponding to the maximum classification probability g. Otherwise, T1-step noise is added to the target domain image x to re-align and sample and perform classification. i = 1000, if the classification result of the target domain image x is still not accepted, the category corresponding to the largest classification probability in all previous classification results is selected as the category of the target domain image x.