Confrontation fine tuning and generation method and device for generalization of single source domain
By constructing a domain generalization model, using a progressive adversarial learning strategy with category and domain style prompts, the problem of insufficient diversity of generated image samples in the existing technology is solved, and the effect of generating diversified images when the target domain feature abstraction is achieved.
Patent Information
- Application Number
- CN202510340723.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-21
AI Technical Summary
It is difficult for the prior art to generate diverse image samples, especially when the target domain features are too abstract, the existing text-generating image models are difficult to accurately capture and generate image samples of the corresponding domain, and the generated image samples are insufficient in diversity.
Build a domain generalization model, including text encoder, autoencoder and conditional denoising module, train the model through an incremental adversarial learning strategy, and use category and domain style tips to generate diversified images to adapt to the generalization needs of single-source domain data sets.
The ability to generate diverse image samples when abstracting the feature of the target domain is achieved, improving the model's performance in multiple unknown and distributed offset target domains.
Smart Images

Figure CN120298542A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of domain generalization image classification technology, and particularly to an adversarial fine-tuning and generation method and device for single-source domain generalization. Background Art
[0002] In recent years, machine learning models (such as deep neural networks) have demonstrated excellent performance in tasks such as computer vision and natural language processing. However, their success depends on the key assumption that the training data and test data satisfy independent and identically distributed. However, in real-world scenarios, domain shift is prevalent. For example, the test data may come from an unknown domain (i.e., out-of-distribution data) with a significant distribution difference from the training data, and this distribution difference will lead to a significant decline in model performance. To address the domain shift problem, traditional methods mainly focus on domain adaptation and domain generalization. However, domain adaptation and domain generalization rely on the availability of target domain data and the fact that only a single labeled source domain data can be obtained in practical applications, which is difficult to apply in scenarios where data privacy is restricted or the annotation cost is high. Therefore, single-domain generalization has become a more challenging and practically valuable research direction, whose goal is to train a model with a single source domain data so that it can perform well in multiple unknown and distribution-shifted target domains. Current single-domain generalization methods mainly expand the data distribution of the source domain through data augmentation techniques or adopt adaptive data normalization methods. However, the sample diversity of these two methods is limited by the inherent distribution characteristics of the source domain and it is difficult to cover the complex and diverse domain shift patterns in real-world scenarios; and there is a lack of theoretical constraints on the distribution difference between the generated samples and the potential target domains, resulting in a deviation between the distribution expansion of the generated samples and the requirements of the real target domains.
[0003] Currently, the breakthrough of the text-to-image foundation model provides a new idea for single-domain generalization. The text-to-image foundation model has a powerful ability to generate images from text descriptions. Since the text itself contains rich semantic information, the text-to-image foundation model can generate diverse images across semantic domains based on text prompts. In theory, it can generate data covering multi-domain features by designing diverse text prompts, thereby expanding the source domain distribution. However, in practical applications, when the features of the target domain are too abstract to be accurately described in language, the existing text-to-image foundation models are difficult to accurately capture and generate corresponding domain image samples; the existing text-to-image models rely on manually designed prompt words and lack an automated generation mechanism, resulting in a limited domain coverage of the generated data and leading to the generated image samples being single-domain and lacking diversity. Summary of the Invention
[0004] The objective of the embodiments of the present invention is to provide an adversarial fine-tuning and generation method and device for single-source domain generalization, so as to solve the problems that when the features of the target domain are too abstract to be accurately described in language, it is difficult for the prior art to generate image samples corresponding to the domain, and the image samples generated by the prior art are single-domain and lack diversity.
[0005] To solve the above technical problems, the embodiments of the present invention provide the following technical solutions:
[0006] The first aspect of the present invention provides an adversarial fine-tuning and generation method for single-source domain generalization, including:
[0007] Obtain a dataset for the domain generalization task, where the dataset includes multiple sub-datasets, and the domain styles of each sub-dataset are different and the number of categories is the same;
[0008] Construct a domain generalization model, which includes a text encoder, an autoencoder, and a conditional denoising module connected in sequence. The text encoder is used to add category hints and domain style hints to multiple sub-datasets;
[0009] Input multiple sub-datasets into the domain generalization model to train the domain generalization model using a progressive adversarial learning strategy to obtain a trained domain generalization model;
[0010] Input the single-source domain dataset into the trained domain generalization model to enable the trained domain generalization model to perform single-source domain generalization and output target images, where the target images are images with different types and different domain styles.
[0011] The second aspect of the present invention provides an adversarial fine-tuning and generation device for single-source domain generalization, including:
[0012] An acquisition module, configured to obtain a dataset for the domain generalization task, where the dataset includes multiple sub-datasets, and the domain styles of each sub-dataset are different and the number of categories is the same;
[0013] A construction module, configured to construct a domain generalization model, which includes a text encoder, an autoencoder, and a conditional denoising module connected in sequence. The text encoder is used to add category hints and domain style hints to multiple sub-datasets;
[0014] A training module, configured to input multiple sub-datasets into the domain generalization model to train the domain generalization model using a progressive adversarial learning strategy to obtain a trained domain generalization model;
[0015] A single-source domain generalization module, configured to input the single-source domain dataset into the trained domain generalization model to enable the trained domain generalization model to perform single-source domain generalization and output target images, where the target images are images with different types and different domain styles.
[0016] Compared with the prior art, the adversarial fine-tuning and generation method and device for single-source domain generalization provided by the present invention obtain a dataset for the domain generalization task. The dataset includes multiple sub-datasets, and the domain styles of each sub-dataset are different and the number of categories is the same. A domain generalization model is constructed. The domain generalization model includes a text encoder, an autoencoder, and a conditional denoising module connected in sequence. The text encoder is used to add category hints and domain style hints to multiple sub-datasets. The multiple sub-datasets are input into the domain generalization model to train the domain generalization model by using a progressive adversarial learning strategy, and a trained domain generalization model is obtained. The single-source domain dataset is input into the trained domain generalization model to enable the trained domain generalization model to perform single-source domain generalization and output a target image, where the target image is an image with different types and different domain styles. In this way, through the text encoder in the domain generalization model, by learning the added category hints and domain style hints, abstract hints can be learned, so that when the target domain features are too abstract to be accurately described in language, image samples of the corresponding domain can be generated. By the added category hints and domain style hints, multiple categories and domain styles can be learned, so that the generation can perform single-source domain generalization on the single-source domain dataset and generate diverse images of multiple domain styles and multiple categories. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown by way of illustration and not limitation, and the same or corresponding reference numerals indicate the same or corresponding parts, wherein:
[0018] Figure 1 Schematically shows a flowchart of the adversarial fine-tuning and generation method for single-source domain generalization;
[0019] Figure 2 Schematically shows a flowchart of obtaining a trained domain generalization model;
[0020] Figure 3 Schematically shows a structural diagram of the adversarial fine-tuning and generation device for single-source domain generalization. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.
[0022] It should be noted that: unless otherwise specified, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those skilled in the art to which the present invention pertains.
[0023] The method in the embodiments of the present invention will be described in detail below.
[0024] Figure 1 A flowchart of the adversarial fine-tuning and generation method for single-source domain generalization in the embodiments of the present invention is schematically shown. Refer to Figure 1 As shown, the adversarial fine-tuning and generation method for single-source domain generalization may include:
[0025] S101. Obtain a data set for the domain generalization task.
[0026] Among them, the data set includes multiple sub-data sets, and the domain styles of each sub-data set are different and the number of categories is the same. Each of the multiple sub-data sets includes multiple preset texts and multiple preset images corresponding to the multiple preset texts, and the multiple preset texts are used to express category names.
[0027] The samples in each sub-data set do not intersect. One sub-data set has only one domain style, and the domain styles of the multiple sub-data sets are different from each other.
[0028] The multiple sub-data sets can be expressed as Among them, is the first sub-data set, is the second sub-data set, is the third sub-data set, is the (m + 1)-th sub-data set, is the -th sub-data set. The first sub-data set is used as the source domain, and is used as the multiple target domains That is, the multiple target domains Among them, the second sub-data set is used as the first target domain, the third sub-data set is used as the second target domain, the (m + 1)-th sub-data set is used as the m-th target domain, and the -th sub-data set is used as the -th target domain, is the total number of the multiple target domains. Each sub-data set has a label, that is, the source domain and the multiple target domains all have labels, and the labels are multiple preset texts used to express category names. Let be expressed as Among them, x i′is the i'-th preset image, i.e., the sample, of the first sub-dataset, y i′ is the preset text, i.e., the label, corresponding to the i'-th preset image of the first sub-dataset, N s is the number of multiple preset images of the first sub-dataset.
[0029] The number of categories in each sub-dataset is K. Denote as where x i″ is the x i″ -th preset image of the (m + 1)-th sub-dataset, y i″ is the preset text corresponding to the x i″ -th preset image of the (m + 1)-th sub-dataset, is the number of multiple preset images of the (m + 1)-th sub-dataset. The goal of the domain generalization task is to solve the domain distribution shift problem by generating diverse domain-style images, so as to build a model that can perform well on all unseen target domains, i.e., the domain generalization model in the following step S102.
[0030] S102. Build a domain generalization model.
[0031] Among them, the domain generalization model includes a text encoder, an autoencoder, and a conditional denoising module connected in sequence. The text encoder is used to add class hints and domain-style hints to multiple sub-datasets.
[0032] Specifically, the text encoder is an encoder based on the Contrastive Language-Image Pre-training (CLIP) model. The autoencoder includes an encoder and a decoder. The encoder is used to map the input multiple preset images into the latent space to convert the input multiple preset images into latent representations, and the decoder is used to reconstruct images from the latent representations. The conditional denoising module is a module based on the U-Net network with cross-attention layers. The U-Net network with cross-attention layers introduces cross-attention layers into the U-Net network. The conditional denoising module is used to add noise to the latent representations converted by the encoder, and is also used to predict noise and denoise.
[0033] S103. Input multiple sub-datasets into the domain generalization model to train the domain generalization model using the progressive adversarial learning strategy, and obtain the trained domain generalization model.
[0034] Specifically, Figure 2 schematically shows the flowchart of obtaining the trained domain generalization model. See Figure 2As shown in the figure, multiple sub-datasets are input into the domain generalization model to train the domain generalization model by using the progressive adversarial learning strategy, and the trained domain generalization model is obtained, including:
[0035] Step S1031: Input multiple preset texts into the text encoder to add class hints to the multiple preset texts by using the text encoder, so that the text encoder outputs the first text embedding representation feature.
[0036] Specifically, step S1031 includes:
[0037] Step A1: Tokenize the multiple preset texts by using the first input string to obtain multiple first tokens.
[0038] The expression of the input string is:
[0039]
[0040] Among them, is the first input string, * is the first placeholder, k is the kth class, and K is the number of classes in each sub-dataset. The first placeholder is used to indicate the position of the class hint in the first input string.
[0041] Step A2: Convert each first token into a first embedding vector.
[0042] Among them, the first embedding vector includes the first placeholder.
[0043] Assume M c is the token length corresponding to the first input string except the first placeholder, that is, the first token length, then the corresponding first embedding vector is defined as: Among them, is the first embedding vector, is the first embedding vector corresponding to the first token length of the 1st of the kth class, is the first embedding vector corresponding to the jth first token length of the kth class, is the first embedding vector corresponding to the M c th first token length of the kth class.
[0044] Step A3: Add class hints at the first placeholder according to the embedding function corresponding to the multiple first tokens to obtain the intermediate feature with class hints.
[0045] The class hint added at the first placeholder is a learnable class hint.
[0046] Specifically, the expression of the intermediate feature with class hints is:
[0047]
[0048] Among them, is the intermediate feature with category hint for the k-th category, ε(·) is the embedding function corresponding to multiple first tokens, is the first token corresponding to the length of the first token of the 1st of the k-th category, is the first token corresponding to the length of the j-th first token of the k-th category, is the first token corresponding to the length of the M c -th first token of the k-th category, is the first embedding vector corresponding to the length of the first token of the 1st of the k-th category, is the first embedding vector corresponding to the length of the j-th first token of the k-th category, is the first embedding vector corresponding to the length of the M c -th first token of the k-th category, is the embedding function of the first embedding vector corresponding to the length of the first token of the 1st of the k-th category, is the embedding function of the first embedding vector corresponding to the length of the j-th first token of the k-th category, is the embedding function of the first embedding vector corresponding to the length of the M c -th first token of the k-th category, is the category hint for the k-th category.
[0049] Step A4: Input the intermediate feature with category hint into the feature extractor so that the feature extractor outputs the first text embedding representation feature.
[0050] The feature extractor is the extractor in the text encoder based on the CLIP model. Represent the feature extractor as Φ(·), then the first text embedding representation feature is represented as
[0051] Step S1032: Input multiple preset images into the autoencoder so that the autoencoder outputs the first latent representation.
[0052] The expression of the first latent representation is: z0 = f E (X), where z0 is the first latent representation, X is multiple preset images, and f E (·) is the encoder in the autoencoder.
[0053] Step S1033: Input the first text embedding representation feature and the first latent representation into the conditional denoising module so that the conditional denoising module outputs the predicted noise.
[0054] Specifically, step S1033 includes:
[0055] Step B1: Add Gaussian noise to the first latent representation to obtain the noise-added first latent representation.
[0056] The amount of added Gaussian noise varies according to the training epoch t and follows a variance schedule where T′ is the total number of training epochs, and β t is the variance schedule for the t-th training epoch. Add Gaussian noise ∈ to the first latent representation z0, and the expression for the noise-added first latent representation is:
[0057]
[0058] where z t is the noise-added first latent representation, α t = 1 - β t , α t is the intermediate parameter for the t-th training epoch, is the average value of the total intermediate parameters corresponding to the total number of training epochs, z0 is the first latent representation, ∈ is Gaussian noise, and N(0, 1) follows a normal distribution with a mean of 0 and a standard deviation of 1.
[0059] Step B2: Determine the predicted noise according to the noise-added first latent representation, the first text embedding representation feature, and the training epoch.
[0060] Pass the first text embedding representation feature to the cross-attention layer ∈ θ (·) in the conditional denoising module, and predict the predicted noise based on the noise-added first latent representation and the t-th training epoch. The expression for the predicted noise is:
[0061]
[0062] where, is the predicted noise, ∈ θ (·) is the cross-attention layer, z t is the noise-added first latent representation, t is the t-th training epoch, is the first text embedding representation feature, is the intermediate feature with class prompt for the k-th class, and Φ(·) is the feature extractor.
[0063] Step S1034: Based on the predicted noise and the Gaussian noise, use the backpropagation algorithm to determine the updated class prompt.
[0064] Specifically, Step S1034 includes:
[0065] Step C1: Determine the latent space diffusion model loss according to the predicted noise and the Gaussian noise.
[0066] The expression for the loss of the Latent Diffusion Model (LDM) is as follows:
[0067]
[0068] Among them, is the loss of the latent diffusion model, E ∈,t is the expectation under the Gaussian noise in the t-th training round, ‖·‖2 is the L2 norm, is the predicted noise, and ∈ is the Gaussian noise.
[0069] Step C2: Using the backpropagation algorithm, backpropagate the loss of the latent diffusion model to obtain the updated class prompt.
[0070] After a total of T' training rounds, K class prompts can be learned. The K class prompts include the updated class prompts, and each class prompt corresponds to a class.
[0071] The above steps S1031 to S1034 can be referred to as class prompt tuning, which can capture class information with invariant domain style from the sub-dataset corresponding to the source domain. In actual scenarios, some concepts or classes are difficult to be understood by existing diffusion models through a single class label or name. Therefore, the present invention uses a domain generalization model to train class prompts to represent each class, and uses the class prompts as the abstract representation of the corresponding class.
[0072] Step S1035: Input the updated class prompt into the text encoder to use the text encoder to add domain style prompts in multiple preset texts, so that the text encoder outputs the second text embedding representation feature.
[0073] Specifically, step S1035 includes:
[0074] Step D1: Tokenize the updated class prompt using the second input string to obtain multiple second tokens.
[0075] The expression for the second input string is:
[0076]
[0077] Among them, is the second input string, ◇ is the second placeholder, i is the i-th domain style, and N' is the total number of domain styles.
[0078] Step D2: Convert each second token into a second embedding vector.
[0079] Among them, the second embedding vector includes the second placeholder.
[0080] Assume Md is the token length corresponding to the second input string except for the second placeholder, that is, the second token length. Accordingly, the corresponding second embedding vector is defined as: where is the second embedding vector, is the second embedding vector corresponding to the first second token length of the i-th domain style, is the second embedding vector corresponding to the j'-th second token length of the i-th domain style, is the second embedding vector corresponding to the M d -th second token length of the i-th domain style.
[0081] Step D3: Add a domain style hint at the second placeholder according to the embedding functions corresponding to multiple second tokens to obtain an intermediate feature with a domain style hint.
[0082] where the domain style hint is the hint initialized in the memory bank.
[0083] Specifically, the expression of the intermediate feature with a domain style hint is:
[0084]
[0085] where is the intermediate feature with a domain style hint of the i-th domain style, ε(·) is the embedding function corresponding to multiple second tokens, is the second token corresponding to the first second token length of the i-th domain style, is the second token corresponding to the j'-th second token length of the i-th domain style, is the second token corresponding to the M d -th second token length of the i-th domain style, is the second embedding vector corresponding to the first second token length of the i-th domain style, is the second embedding vector corresponding to the j'-th second token length of the i-th domain style, is the second embedding vector corresponding to the M d -th second token length of the i-th domain style, is the embedding function of the second embedding vector corresponding to the first second token length of the i-th domain style, is the embedding function of the second embedding vector corresponding to the j'-th second token length of the i-th domain style, is the embedding function of the second embedding vector corresponding to the M d -th second token length of the i-th domain style, is the domain style hint of the i-th domain style.
[0086] Step D4: Input the intermediate feature with domain-style prompt into the feature extractor so that the feature extractor outputs the second text embedding representation feature.
[0087] The feature extractor is the extractor in the text encoder based on the CLIP model. Represent the feature extractor as Φ(·), then the second text embedding representation feature is represented as
[0088] Step S1036: Input the second text embedding representation feature into the conditional denoising module to use the conditional denoising module to denoise the randomly generated noise and output the second latent representation.
[0089] Since random noise will be generated in the process of determining the second text embedding representation feature, it is necessary to remove the generated random noise. The second text embedding representation feature can be passed to the cross-attention layer ∈ θ (·) in the conditional denoising module, and based on the noise-added first latent representation and the t-th training round, make the second latent representation as t identical as possible to the noise-added first latent representation z
[0090] The expression of the second latent representation is:
[0091]
[0092] where z0′ is the second latent representation, z t is the noise-added first latent representation, α t = 1 - β t , α t is the intermediate parameter of the t-th training round, β t is the variance schedule of the t-th training round, is the average value of the total intermediate parameters corresponding to the total number of training rounds, T′ is the total number of training rounds, ∈ θ (·) is the cross-attention layer, is the second text embedding representation feature, is the intermediate feature with domain-style prompt of the i-th domain style.
[0093] Step S1037: Based on the second latent representation and the second text embedding representation feature, use the backpropagation algorithm to determine the updated domain-style prompt.
[0094] Specifically, Step S1037 includes:
[0095] Step E1: Determine the adversarial domain prompt tuning loss according to the second latent representation, the second text embedding representation feature, and the hyperparameters.
[0096] First, determine the probability that the second latent representation belongs to the i-th domain style according to the second latent representation, the second text embedding representation feature, and the hyperparameters; then, determine the adversarial domain loss according to the probability that the second latent representation belongs to the i-th domain style; then, determine the class consistency loss according to the second latent representation and the first text embedding representation feature; finally, determine the adversarial domain prompt tuning loss according to the adversarial domain loss and the class consistency loss.
[0097] The expression for the probability that the second latent representation belongs to the i-th domain style is:
[0098]
[0099] where, is the probability that the second latent representation belongs to the i-th domain style, z0′ is the second latent representation, is the second text embedding representation feature, τ is the hyperparameter, which is used to control the sharpness of the output, <·,·> is the dot product operation, is the intermediate feature with domain style prompts of the i-th domain style, i is the i-th domain style, and N′ is the total number of domain styles.
[0100] The adversarial domain loss is used to minimize the similarity between the second latent representation z0′ and all previous domain styles in the memory bank The expression for the adversarial domain loss is:
[0101]
[0102] where, is the adversarial domain loss, i is the i-th domain style, N′ is the total number of domain styles, is the probability that the second latent representation belongs to the i-th domain style, z0′ is the second latent representation, is the second text embedding representation feature.
[0103] The class consistency loss is used to maximize the semantic consistency between the second latent representation z0′ and the intermediate feature with class prompts of the k-th class, ensuring that after optimizing the adversarial domain loss, the second latent representation z0′ can maintain the class semantics. The expression for the class consistency loss is:
[0104]
[0105] where, is the class consistency loss, is the probability that the second latent representation belongs to the k-th class, and z0' is the second latent representation. is the feature of the first text embedding representation. is the intermediate feature with class hint for the k-th class.
[0106] The expression of the adversarial domain hint tuning loss is:
[0107]
[0108] where is the adversarial domain hint tuning loss. is the class consistency loss. is the adversarial domain loss, and λ is the target hyperparameter, which is used to maintain the trade-off between the updated domain style hint and its corresponding updated class hint.
[0109] Step E2: Using the backpropagation algorithm, backpropagate the adversarial domain hint tuning loss to obtain the updated domain style hint.
[0110] The above steps S1035 to S1037 can be called domain style hint tuning. The entire steps of domain style hint tuning implement learning the domain style hint using the progressive adversarial learning strategy to ensure that the updated domain style hint is different from the domain style stored in the memory bank. stored in the memory bank.
[0111] Step S1038: Store the updated domain style hint into the memory bank to obtain the trained domain generalization model.
[0112] By storing the updated domain style hint into the memory bank in the memory bank is updated. The domain style hint will be obtained in each training round until the preset number of training rounds is reached, and the iteration will stop. At this time, the updated memory bank stores the domain style hints obtained in each round. That is to say, the updated memory bank includes the domain style hint of the previous round and the updated domain style hint of the new round. The preset number of training rounds can be 50 times.
[0113] The expression of the domain style hint finally stored in the memory bank is: where is the updated memory bank. is the previous memory bank, which is used to store the domain style hint of the previous round. is the updated domain style hint of the new round.
[0114] When determining the updated domain style hint, the updated domain style hint can be increased. The domain style hint added previously, i.e., the domain style hint of the i-th domain style The distance in the memory bank. In this way, the updated domain style hint is different from the previously added domain style hint in terms of distribution, and a more challenging and abstract domain style can be learned.
[0115] Store the updated domain style hint in the memory bank until a preset number of training rounds is reached, and a trained domain generalization model can be obtained.
[0116] S104: Input the single-source domain dataset into the trained domain generalization model so that the trained domain generalization model performs single-source domain generalization and outputs the target image.
[0117] Among them, the target image is an image with different types and different domain styles.
[0118] The simulation experiment of the present invention is trained for 50 training rounds with a batch size of 64 and a learning rate of 0.001. The Stochastic Gradient Descent (SGD) optimizer is adopted, and the training is carried out on a single NVIDIA RTX 3090 GPU. The total number N' of the domain styles of the present invention is set to 5, and the target hyperparameter λ is set to 10. To ensure the reliability of the adversarial fine-tuning and generation method for single-source domain generalization of the present invention, all simulation experiments are independently repeated five times, and the average result of the five simulation experiment results is taken. The Visual Learning Challenge (VLCS) dataset commonly used in the domain generalization task is adopted for evaluation. The VLCS dataset includes 4 domain styles, 10,729 samples, 5 categories. The 4 domain styles are natural pictures, daily items, animals, and transportation tools respectively, and the 5 categories are birds, cars, chairs, dogs, and humans respectively. Table 1 shows the average results of the adversarial fine-tuning and generation method for single-source domain generalization of the present invention and the existing methods. The existing methods include the Augmix method, the Error Rate Minimization (ERM) method, the Parametric Adaptive Instance Normalization (pAdaIn) method, the MixStyle method, the Energy-based Feature Distribution Mixing (MixStyle) method, the Database Storage Unit (DSU) method, the Attention-based Cross-view Figure 1Attention-based Cross-View Consistency (ACVC) method and Mean Absolute Deviation (MAD) method. The first column in Table 1 is the method, and the second column is the venue. The venues include the International Conference on Learning Representations 2020 (ICLR'20), International Conference on Learning Representations 2021 (ICLR'21), IEEE Conference on Computer Vision and Pattern Recognition 2021 (CVPR'21), IEEE Conference on Computer Vision and Pattern Recognition 2022 (CVPR'22), International Conference on Learning Representations 2022 (ICLR'22), Computer Vision and Pattern Recognition 2022 Workshops (CVPRW*22), and IEEE Conference on Computer Vision and Pattern Recognition 2023 (CVPR'23). From the average result (Avg) of the test data classification accuracy rate in the 5 simulation experiments in the seventh column of Table 1, it can be seen that the average result of the adversarial fine-tuning and generation method for single-source domain generalization of the present invention is the best, and better domain generalization can be achieved.
[0119] Average Results of Test Data Classification Accuracy Rate in Simulation Experiment of Table 1
[0120] Method Venue Natural picture Household items Animal Vehicle Avg(%) Augmix ICLR'20 75.25 59.52 45.90 57.43 59.53 ERM ICLR'21 76.72 58.86 44.95 57.71 59.56 pAdaIn CVPR'21 76.03 65.21 43.17 57.94 60.59 Mixstyle ICLR'21 75.73 61.29 44.66 56.57 59.56 EFDMix CVPR'22 72.35 61.41 52.34 63.28 62.33 DSU ICLR'22 76.93 69.20 46.54 58.36 62.76 ACVC CVPRW*22 76.15 61.23 47.43 60.18 61.25 MAD CVPR'23 76.15 69.36 48.04 61.74 63.82 The present invention 73.16 74.69 69.66 75.96 73.37
[0121] Based on the above Figure 1From the implementation method, it can be seen that the adversarial fine-tuning and generation method for single-source domain generalization in the embodiments of the present invention obtains a dataset for the domain generalization task. The dataset includes multiple sub-datasets, and the domain styles of each sub-dataset are different and the number of categories is the same; a domain generalization model is constructed. The domain generalization model includes a text encoder, an autoencoder, and a conditional denoising module connected in sequence. The text encoder is used to add category hints and domain style hints to the multiple sub-datasets; the multiple sub-datasets are input into the domain generalization model to train the domain generalization model using a progressive adversarial learning strategy to obtain a trained domain generalization model; the single-source domain dataset is input into the trained domain generalization model to enable the trained domain generalization model to perform single-source domain generalization and output a target image. The target image is an image with different types and different domain styles. In this way, through the text encoder in the domain generalization model, learning the added category hints and domain style hints, abstract hints can be learned. When the target domain features are too abstract to be accurately described in language, image samples corresponding to the domain can be generated. By the added category hints and domain style hints, multiple categories and domain styles can be learned, enabling the generation to perform single-source domain generalization on the single-source domain dataset and generate diverse images with multiple domain styles and multiple categories.
[0122] Based on the same inventive concept, as an implementation of the above-mentioned adversarial fine-tuning and generation method for single-source domain generalization, the embodiments of the present invention also provide an adversarial fine-tuning and generation device for single-source domain generalization. Figure 3 For the structural diagram of the adversarial fine-tuning and generation device for single-source domain generalization in the embodiments of the present invention, see Figure 3 As shown, the adversarial fine-tuning and generation device for single-source domain generalization may include:
[0123] An acquisition module 301, configured to acquire a dataset for the domain generalization task. The dataset includes multiple sub-datasets, and the domain styles of each sub-dataset are different and the number of categories is the same;
[0124] A construction module 302, configured to construct a domain generalization model. The domain generalization model includes a text encoder, an autoencoder, and a conditional denoising module connected in sequence. The text encoder is used to add category hints and domain style hints to the multiple sub-datasets;
[0125] A training module 303, configured to input the multiple sub-datasets into the domain generalization model to train the domain generalization model using a progressive adversarial learning strategy to obtain a trained domain generalization model;
[0126] The single-source domain generalization module 304 is configured to input a single-source domain dataset into the trained domain generalization model, so that the trained domain generalization model performs single-source domain generalization and outputs a target image, where the target image is an image with different types and different domain styles.
[0127] It should be noted here that the above description of the embodiments of the adversarial fine-tuning and generation device for single-source domain generalization is similar to the description of the embodiments of the adversarial fine-tuning and generation method for single-source domain generalization, and has beneficial effects similar to those of the embodiments of the adversarial fine-tuning and generation method for single-source domain generalization. For technical details not disclosed in the embodiments of the adversarial fine-tuning and generation device for single-source domain generalization of the present invention, please refer to the description of the embodiments of the adversarial fine-tuning and generation method for single-source domain generalization of the present invention for understanding.
[0128] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An adversarial fine-tuning and generation method for single-source domain generalization, characterized in that Including: Obtain a dataset for the domain generalization task, where the dataset includes multiple sub-datasets, and each sub-dataset has a different domain style and the same number of categories; Construct a domain generalization model, which includes a text encoder, an autoencoder, and a conditional denoising module connected in sequence. The text encoder is used to add category hints and domain style hints to the multiple sub-datasets; Input the multiple sub-datasets into the domain generalization model to train the domain generalization model using a progressive adversarial learning strategy to obtain a trained domain generalization model; Input a single-source domain dataset into the trained domain generalization model to enable the trained domain generalization model to perform single-source domain generalization and output a target image, where the target image is an image with different types and different domain styles.
2. The adversarial fine-tuning and generation method for single-source domain generalization according to claim 1, wherein Each of the multiple sub-datasets includes multiple preset texts and multiple preset images corresponding to the multiple preset texts. The multiple preset texts are used to express category names. The step of inputting the multiple sub-datasets into the domain generalization model to train the domain generalization model using a progressive adversarial learning strategy to obtain a trained domain generalization model includes: Input the multiple preset texts into the text encoder to use the text encoder to add the category hint in the multiple preset texts, so that the text encoder outputs a first text embedding representation feature; Input the multiple preset images into the autoencoder to enable the autoencoder to output a first latent representation; Input the first text embedding representation feature and the first latent representation into the conditional denoising module to enable the conditional denoising module to output a predicted noise; Based on the predicted noise and Gaussian noise, use the backpropagation algorithm to determine an updated category hint; Input the updated category hint into the text encoder to use the text encoder to add the domain style hint in the multiple preset texts, so that the text encoder outputs a second text embedding representation feature; Input the second text embedding representation feature into the conditional denoising module to use the conditional denoising module to denoise randomly generated noise and output a second latent representation; Based on the second latent representation and the second text embedding representation feature, use the backpropagation algorithm to determine an updated domain style hint; Store the updated domain style hint in a memory bank to obtain the trained domain generalization model.
3. The adversarial fine-tuning and generation method for single-source domain generalization according to claim 2, wherein The step of inputting the multiple preset texts into the text encoder to use the text encoder to add a category hint in the multiple preset texts, so that the text encoder outputs a first text embedding representation feature includes: Tokenize the multiple preset texts using a first input string to obtain multiple first tokens; Convert each first token into a first embedding vector, where the first embedding vector includes a first placeholder; Add the category hint at the first placeholder according to the embedding function corresponding to the multiple first tokens to obtain an intermediate feature with a category hint; Input the intermediate feature with category hint into the feature extractor, so that the feature extractor outputs the first text embedding representation feature.
4. The adversarial fine-tuning and generation method for single-source domain generalization according to claim 2, wherein The step of inputting the first text embedding representation feature and the first latent representation into the conditional denoising module, so that the conditional denoising module outputs predicted noise, includes: Adding the Gaussian noise to the first latent representation to obtain the first latent representation with added noise; Determining the predicted noise according to the first latent representation with added noise, the first text embedding representation feature, and the training round.
5. The adversarial fine-tuning and generation method for single-source domain generalization according to claim 2, wherein Based on the predicted noise and the Gaussian noise, using the backpropagation algorithm to determine the updated category hint, includes: Determining the latent space diffusion model loss according to the predicted noise and the Gaussian noise; Using the backpropagation algorithm to backpropagate the latent space diffusion model loss to obtain the updated category hint.
6. The adversarial fine-tuning and generation method for single-source domain generalization according to claim 2, wherein The step of inputting the updated category hint into the text encoder to use the text encoder to add the domain style hint to the multiple preset texts, so that the text encoder outputs the second text embedding representation feature, includes: Segmenting the updated category hint using a second input string to obtain multiple second tokens; Converting each second token into a second embedding vector, where the second embedding vector includes a second placeholder; Adding the domain style hint at the second placeholder according to the embedding function corresponding to the multiple second tokens to obtain an intermediate feature with domain style hint, where the domain style hint is the hint initialized in the memory bank; Inputting the intermediate feature with domain style hint into the feature extractor, so that the feature extractor outputs the second text embedding representation feature.
7. The adversarial fine-tuning and generation method for single-source domain generalization according to claim 2, wherein Based on the second latent representation and the second text embedding representation feature, using the backpropagation algorithm to determine the updated domain style hint, includes: Determining the adversarial domain hint tuning loss according to the second latent representation, the second text embedding representation feature, and the hyperparameters; Using the backpropagation algorithm to backpropagate the adversarial domain hint tuning loss to obtain the updated domain style hint.
8. The adversarial fine-tuning and generation method for single-source domain generalization according to claim 3, wherein The expression of the intermediate feature with category hint is: Among them, is the intermediate feature with category hint for the k-th category, ε(·) is the embedding function corresponding to the multiple first tokens, is the first token corresponding to the length of the first token of the 1st of the k-th category, is the first token corresponding to the length of the j-th first token of the k-th category, is the first token corresponding to the length of the Mc-th first token of the k-th category, is the first embedding vector corresponding to the length of the first token of the 1st of the k-th category, is the first embedding vector corresponding to the length of the j-th first token of the k-th category, is the first embedding vector corresponding to the length of the Mc-th first token of the k-th category, is the embedding function of the first embedding vector corresponding to the length of the first token of the 1st of the k-th category, is the embedding function of the first embedding vector corresponding to the length of the j-th first token of the k-th category, is the embedding function of the first embedding vector corresponding to the length of the Mc-th first token of the k-th category, is the category hint for the k-th category.
9. The adversarial fine-tuning and generation method for single-source domain generalization according to claim 6, characterized in that, The expression of the intermediate feature with domain style hint is: Among them, is the intermediate feature with domain style hint of the i-th domain style, ε(·) is the embedding function corresponding to the multiple second tokens, is the second token corresponding to the length of the first second token of the i-th domain style, is the second token corresponding to the length of the j'-th second token of the i-th domain style, is the second token corresponding to the length of the Md-th second token of the i-th domain style, is the second embedding vector corresponding to the length of the first second token of the i-th domain style, is the second embedding vector corresponding to the length of the j'-th second token of the i-th domain style, is the second embedding vector corresponding to the length of the Md-th second token of the i-th domain style, is the embedding function of the second embedding vector corresponding to the length of the first second token of the i-th domain style, is the embedding function of the second embedding vector corresponding to the length of the j'-th second token of the i-th domain style, is the embedding function of the second embedding vector corresponding to the length of the Md-th second token of the i-th domain style, is the domain style hint of the i-th domain style.
10. An adversarial fine-tuning and generation device for single-source domain generalization, characterized in that, Including: An acquisition module for acquiring a dataset for domain generalization tasks, where the dataset includes multiple sub-datasets, and the domain styles of each sub-dataset are different and the number of categories is the same; A construction module for constructing a domain generalization model, where the domain generalization model includes a text encoder, an autoencoder, and a conditional denoising module connected in sequence, and the text encoder is used to add category hints and domain style hints to the multiple sub-datasets; A training module for inputting the multiple sub-datasets into the domain generalization model to train the domain generalization model using a progressive adversarial learning strategy to obtain a trained domain generalization model; A single-source domain generalization module for inputting a single-source domain dataset into the trained domain generalization model, enabling the trained domain generalization model to perform single-source domain generalization and output a target image, where the target image is an image with different types and different domain styles.
Citation Information
Patent Citations
Attribute face image generate method, device, system and readable storage medium
CN109147010A
Cross-domain diversity image generation method and system based on generative adversarial network
CN111292384A
Method for diversifying style words through entropy maximization to realize domain generalization
CN118587723A
Domain-invariant feature-based meta-knowledge fine-tuning method and platform
WO2022151553A1
AU2020103905A4