Single-sample unsupervised domain adaptive processing method based on LoRA training

The Stable Diffusion model is fine-tuned by a LoRA training method, and a high-quality target domain data set is generated by combining the BLIP model and weighted prompt words. This solves the problems of uncontrollable content generation and high resource consumption in the existing single-sample unsupervised domain adaptation processing method, and achieves better domain adaptation effects.

CN120047774APending Publication Date: 2025-05-27HUNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510212615.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing single-sample unsupervised domain adaptation processing methods have problems such as style transfer methods that cannot control image content generation, diffusion model-based methods are prone to overfitting, uncontrollable model weights, and high training resource consumption.

Method used

The Stable Diffusion model is fine-tuned by using LoRA training method, the BLIP model is used to generate image tags, and optimized weighted prompt words are designed. Combined with the image quality evaluation method ARNIQA, a high-quality target domain data set is generated.

Benefits of technology

Overcoming the problem of overfitting and uncontrollable model weights of Dreambooth training method, improving the controllability and quality of image content, and achieving better domain adaptation effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047774A_ABST
    Figure CN120047774A_ABST
Patent Text Reader

Abstract

The invention discloses a single-sample unsupervised domain adaptive processing method based on LoRA training, and relates to the field of artificial intelligence and semantic segmentation, and the method comprises the steps: firstly, randomly selecting a picture from a target domain data set as a label-free single-sample target domain image, and obtaining 35 randomly cut images; secondly, using a BLIP model to carry out image label generation on the randomly cut image; a large model training fine tuning method LoRA is adopted to carry out training fine tuning on the Stable Diffusion model; inputting a series of weighted cue words which are designed and optimized according to experience into a Stable Diffusion model so as to generate an image which is similar to a target domain in style and controllable in image content; an image quality evaluation method ARNIQA is adopted to remove low-quality images and regenerate the low-quality images, and an expanded high-quality target domain data set is obtained; a single-sample unsupervised domain adaptation problem is converted into an unsupervised domain adaptation problem, and finally, an existing mature UDA method is directly used, so that the segmentation model achieves an excellent domain adaptation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of single-sample unsupervised domain adaptation, and particularly to a single-sample unsupervised domain adaptation processing method based on LoRA training. Background Art

[0002] Semantic segmentation, as one of the core tasks in the field of computer vision, is crucial for numerous application scenarios such as autonomous driving. However, due to the distribution differences between the training data and the test data, semantic segmentation models often struggle to achieve optimal performance or even perform poorly when facing unknown data domains. In this context, adapting a segmentation model trained on a labeled source domain well to an unlabeled target domain has become a key research topic, namely the unsupervised domain adaptation problem. In particular, when there is only one available unlabeled image in the target domain, this problem becomes even more challenging, known as the single-sample unsupervised domain adaptation problem. Currently, the methods for dealing with the single-sample unsupervised domain adaptation problem are mainly divided into two categories, namely the method based on style transfer and the method based on diffusion model. By the method of style transfer, the source domain and the target domain are aligned in the pixel space. The method based on style transfer has good effects when the target data is scarce. However, the existing methods based on style transfer can only transfer the style information of the image to the source image and cannot control the generation of the image content. Different from the method based on style transfer, researchers have proposed a solution to the single-sample unsupervised domain adaptation problem based on the diffusion model: DATUM. Different from the methods that usually only transfer the target texture information, DATUM uses a text-to-image diffusion model to generate a synthetic target dataset, which has a similar style to the target domain dataset and has novel and rich image content. The text-to-image interface in DATUM enables guiding the generation of the image towards the desired semantic concept while respecting the style characteristics of a single target domain training image. However, in the target domain dataset generated by DATUM, there are situations where some images have low quality and the content of the prompt words cannot be generated. Moreover, it selects a large model training and fine-tuning method: Dreambooth. Once this training method is completed, the action weights of the model cannot be changed, which is not convenient for adjusting the image generation effect.

[0003] In summary, the problems of the current methods for single-sample unsupervised domain adaptation processing include: (1) The method based on style transfer can only transfer the style information of the image to the source image and cannot control the generation of the image content; (2) The method based on the diffusion model uses the Dreambooth training method, which has problems such as being prone to overfitting, uncontrollable model action weights, and large consumption of training resources; (3) Sometimes, the prompts of the method based on the diffusion model cannot well guide the image generation towards the desired semantic concepts; (4) The method based on the diffusion model sometimes generates low-quality target images, affecting the domain adaptation effect. Summary of the Invention

[0004] To overcome and improve the defects in the above-mentioned prior art that limit the further improvement of the domain adaptation effect due to the selection of the training method, the design of the text-to-image prompts, and the influence of low-quality images, the present invention provides a single-sample unsupervised domain adaptation processing method based on LoRA training.

[0005] To solve the above technical problems, the technical solution provided by the present invention includes the following steps: Step 1: Randomly select an image from the complete target domain dataset as an unlabeled single-sample target domain image, and randomly crop the image to obtain 35 randomly cropped images. Step 2: Use the BLIP model to generate image labels for the 35 randomly cropped images generated in Step 1. After the label generation, adopt the large model training fine-tuning method LoRA to fine-tune the Stable Diffusion model. Step 3: Input a series of weighted prompts designed and optimized according to experience into the Stable Diffusion model fine-tuned in Step 2 to generate images similar to the target domain style and with controllable image content. Subsequently, adopt the image quality assessment method ARNIQA to remove low-quality images and regenerate them to obtain an expanded high-quality target domain dataset. Step 4: According to the expanded target domain dataset obtained in Step 3, successfully transform the single-sample unsupervised domain adaptation problem into an unsupervised domain adaptation (UDA) problem. Finally, directly use the existing mature UDA method to enable the segmentation model to adapt to the unlabeled target domain dataset.

[0006] Further, the use of the BLIP model to generate image labels for the 35 randomly cropped images generated in Step 1 in Step 2 includes: Based on the BLIP text-image multi-modal network, generate corresponding label text files for each of the 35 images randomly cropped in Step 1. The content of the label text file is some words or phrases separated by commas. After generating the image labels with the BLIP model, add a specific recognition prompt in front of each label text file so that images with a style similar to the target domain can be generated through the specific recognition prompt in Step 3.

[0007] Furthermore, the training and fine-tuning of Stable Diffusion using the large model training and fine-tuning method LoRA in step 2 includes: Based on the text-to-image function of the Stable Diffusion model, using the model weight file of Stable Diffusion v1.4 version, with the cropped image dataset and the label file generated using the BLIP model as the training dataset, and using the large model training method LoRA (Low-Rank Adaptation of Large Language Models) to train and fine-tune the Stable Diffusion model, so that it can learn the style features of the target domain images, thus having the ability to generate images similar to the style of the target domain images and being able to controllably generate the required image content.

[0008] Furthermore, the training parameters in the training and fine-tuning stage of Stable Diffusion described in step 2 include: the number of iterations steps is set to 7000, the number of training epochs epoch is set to 20, the batch size batch_size is set to 1, the optimizer type optimizer_type is set to AdaFactor, the mixed precision mixed_precision is set to fp16, the network dimension network_dim is set to 128, and the network dimension limit network_alpha is set to 128.

[0009] Furthermore, inputting a series of weighted prompt words optimized according to experience into the Stable Diffusion model after training and fine-tuning in step 3 includes: using the prompt words optimized according to experience, including a series of image content nouns, and weighting the image content that is expected to be generated with emphasis in the image to be generated, so that the Stable Diffusion model can generate the target image dataset with rich semantic content while retaining the style of the target domain images.

[0010] Furthermore, using the image quality assessment method ARNIQA to remove low-quality images and regenerate in step 3 includes: During the process of circularly expanding the target domain image dataset, every time the Stable Diffusion model generates a target domain image, the image quality assessment method ARNIQA is used to evaluate the quality of the generated image without referring to a reference image. Images that do not meet the image quality requirements are discarded, regenerated, and evaluated again, and so on.

[0011] Furthermore, the use of existing mature UDA methods in step 4 to enable the segmentation model to adapt to the unlabeled target domain dataset includes: Since the target domain dataset has been expanded, the one-shot unsupervised domain adaptation problem (OSUDA) has been successfully converted into an unsupervised domain adaptation problem (UDA). Therefore, various existing mature UDA methods can be directly used for domain adaptation, enabling the segmentation model to perform well on the target domain dataset.

[0012] Compared with the prior art, the beneficial effects of the present invention include: The present invention proposes a one-shot unsupervised domain adaptation processing method based on LoRA training. The LoRA, a large model training fine-tuning method, is used to enable the StableDiffusion model to learn the style features of a single sample image in the target domain, so as to overcome the limitations of the Dreambooth training method, such as easy overfitting, uncontrollable model action weights, and large consumption of training resources. The LoRA training method is more suitable for learning the painting styles of small datasets; by designing optimized weighted prompt words, the present invention can better guide the image generation towards the desired semantic concepts. By inputting the prompt words into the Stable Diffusion model and using the text-to-image generation ability of the Stable Diffusion model, the expanded target domain dataset can be generated; the ARNIQA image quality assessment method is used to screen the image quality, further improving the quality of the generated images; finally, the existing mature UDA method can be directly used to complete the domain adaptation task, achieving a good domain adaptation effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a flowchart of the present invention.

[0014] Figure 2 is a schematic diagram of randomly cropping a single sample image in the target domain.

[0015] Figure 3 is a schematic diagram of generating labels for the cropped image using the BLIP model.

[0016] Figure 4 is a schematic diagram of expanding the target domain dataset through optimized weighted prompt words.

[0017] Figure 5 is a flowchart of evaluating the image quality using the ARNIQA image quality assessment method and regenerating the image. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the objectives, technical solutions, and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention; moreover, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other. The present invention will be described in more detail below with reference to the accompanying drawings. In each of the drawings, the same elements are denoted by similar reference numerals. For the sake of clarity, the various parts in the drawings are not drawn to scale.

[0019] The present invention is a single-sample unsupervised domain adaptation processing method based on LoRA (Low-Rank Adaptation) training, which is used to train and fine-tune a diffusion model in the field of semantic segmentation to complete a single-sample unsupervised domain adaptation task.

[0020] The flowchart of the present invention is as Figure 1 shown. First, a picture is randomly selected from the complete target domain dataset as an unlabeled single-sample target domain image, and this image is randomly cropped to obtain 35 randomly cropped images. The cropping effect is as Figure 2 shown; secondly, the BLIP (Bootstrapping Language-Image Pre-training) model is used to generate image labels for the above-mentioned 35 randomly cropped images. The schematic diagram of the label generation process is as Figure 3 shown. The content of this label text file is some words or phrases, separated by commas; after the label is generated, the LoRA large model training and fine-tuning method is used to train and fine-tune the Stable Diffusion model; thirdly, a series of weighted prompt words designed and optimized according to experience are input into the trained and fine-tuned Stable Diffusion model, so as to generate an image similar to the target domain style and with controllable image content, as Figure 4 shown; further, the ARNIQA image quality assessment method is used to remove low-quality images and regenerate them to obtain an expanded high-quality target domain dataset, as Figure 5 shown; finally, according to the obtained expanded target domain dataset, the single-sample unsupervised domain adaptation problem is successfully transformed into an unsupervised domain adaptation (UDA) problem, and finally the existing mature UDA method is directly used, so that the segmentation model can adapt to the unlabeled target domain dataset.

[0021] First, as Figure 2As shown, a picture is randomly selected from the complete target domain dataset as an unlabeled single-sample target domain image. First, use the torchvision.transforms.Resize() method to reset the resolution of the original high-resolution image. Using the bilinear interpolation method, the resolution is reset to 1024*1024. Then use the random cropping method RandomCrop() of torchvision.transform to randomly crop the image to obtain 35 randomly cropped images with a resolution of 512*512, which will be used as the training dataset to fine-tune the Stable Diffusion model in the follow-up.

[0022] Furthermore, as Figure 3 shown, use the BLIP model to generate image labels for the above 35 randomly cropped images: Based on the BLIP image-text multi-modal network, generate corresponding label text files for each of the 35 randomly cropped images one by one. The content of the label text file is some words or phrases, separated by commas. After generating the image labels with the BLIP model, add a specific recognition prompt word in front of each label text file, such as "CityA,urban scene", so that images with a style similar to the target domain can be generated through this specific recognition prompt word in the subsequent steps.

[0023] Furthermore, adopt the large model training fine-tuning method LoRA to fine-tune Stable Diffusion: Based on the text-to-image function of the Stable Diffusion model, use the model weight file of Stable Diffusion v1.4 version. Using the cropped image dataset and the label file generated by the BLIP model as the training dataset, fine-tune the Stable Diffusion model through the low-rank adaptation training method LoRA (Low-Rank Adaptation of Large Language Models) of the large model, so that it can learn the style features of the target domain images, thus having the ability to generate images with a style similar to the target domain images and being able to controllably generate the required image content. The training parameters in the fine-tuning stage of Stable Diffusion in this step include: the number of iterations steps is set to 7000, the number of training epochs epoch is set to 20, the batch size batch_size is set to 1, the optimizer type optimizer_type is set to AdaFactor, the mixed precision mixed_precision is set to fp16, the network dimension network_dim is set to 128, and the network dimension limit network_alpha is set to 128.

[0024] Furthermore, as Figure 4 shown, a series of weighted prompts optimized according to experience are input into the Stable Diffusion model that has completed training and fine-tuning: Use prompts optimized according to experience, including a series of image content nouns, and weight the image content that is desired to be generated with emphasis in the image to be generated, in the form of "chengshiA, urbanscene, (a train: 1.3), (train track: 0.9), road, tree", so that the Stable Diffusion model can generate a target image dataset with rich semantic content while retaining the image style of the target domain.

[0025] Furthermore, as Figure 5 shown, the image quality assessment method ARNIQA is used to remove low-quality images and regenerate them: During the process of circularly expanding the target domain image dataset, every time the Stable Diffusion model generates a target domain image, the image quality assessment method ARNIQA is used to evaluate the quality of the generated image without referring to a reference image. Images that do not meet the image quality requirements are discarded, regenerated, and evaluated again, and so on.

[0026] Furthermore, by using existing mature UDA methods, the segmentation model can be made to adapt to the unlabeled target domain dataset: Since the target domain dataset has been expanded, the one-shot unsupervised domain adaptation problem (OSUDA) has now been successfully converted into an unsupervised domain adaptation problem (UDA). Therefore, various existing mature UDA methods can be directly used for domain adaptation, enabling the segmentation model to perform well on the target domain dataset.

[0027] Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

Claims

1. A single-sample unsupervised domain adaptation processing method based on LoRA training, characterized in that: The steps include: Step 1: Randomly select a picture from the complete target domain dataset as an unlabeled single-sample target domain image, and randomly crop the image to obtain 35 randomly cropped images; Step 2: Use the pre-trained model of unified visual language understanding and generation, referred to as BLIP, to generate image labels for the 35 randomly cropped images generated in step 1. After the labels are generated, use the low-rank adaptation training method of large models, referred to as LoRA, to train and fine-tune the stable diffusion model, referred to as Stable Diffusion; Step 3: Input a series of weighted prompt words designed and optimized based on experience into the Stable Diffusion model trained and fine-tuned in Step 2 to generate images with similar style to the target domain and controllable image content. Then, an image quality assessment method is used: learning the distortion manifold for image quality assessment, referred to as ARNIQA, to remove low-quality images and regenerate them to obtain an expanded high-quality target domain dataset. Step 4: Based on the expanded target domain dataset obtained in step 3, the single-sample unsupervised domain adaptation problem is successfully transformed into an unsupervised domain adaptation problem, referred to as UDA. Finally, the existing mature UDA method is directly used to enable the segmentation model to adapt to the unlabeled target domain dataset.

2. According to a single-sample unsupervised domain adaptation processing method based on LoRA training according to claim 1, it is characterized in that: The step 2 uses the BLIP model to generate image labels for the 35 randomly cropped images generated in step 1, including: based on the BLIP image-text multimodal network, generating corresponding label text files for the 35 images randomly cropped in step 1 one by one, and after generating the image labels using the BLIP model, adding a specific recognition prompt word in front of each label text file, so as to generate an image with an image style similar to the target domain through the specific recognition prompt word in step 3.

3. According to a single-sample unsupervised domain adaptation processing method based on LoRA training according to claim 1, it is characterized in that: The step 2 adopts the large model training fine-tuning method LoRA to train and fine-tune the Stable Diffusion, including: based on the text image function of the Stable Diffusion model, using the Stable Diffusion model weight file, using the trimmed image data set and the label file generated based on the BLIP image-text multimodal network as the training data set, and training and fine-tuning the Stable Diffusion model through the large model training fine-tuning method LoRA, so that it can learn the style features of the target domain image, thereby having the ability to generate images with a similar style to the target domain image, and can generate the required image content in a controllable manner.

4. According to the single-sample unsupervised domain adaptation processing method based on LoRA training according to claim 3, it is characterized in that: The training parameters of the training fine-tuning phase for Stable Diffusion in step 2 include: the number of iterations steps is set to 7000, the number of training rounds epoch is set to 20, the batch size batch_size is set to 1, the optimizer type optimizer_type is set to AdaFactor, the mixed precision mixed_precision is set to fp16, the network dimension network_dim is set to 128, and the network dimension limit network_alpha is set to 128.

5. According to a single-sample unsupervised domain adaptation processing method based on LoRA training according to claim 1, it is characterized in that: The step 3 includes: using image content nouns that are optimized based on experience, and weighting the image content that is expected to be generated in the image to be generated, so that the Stable Diffusion model can generate a target image dataset containing rich semantic content while retaining the style of the target domain image.

6. According to a single-sample unsupervised domain adaptation processing method based on LoRA training according to claim 1, it is characterized in that: The step 3 includes: in the process of cyclically expanding the target domain image data set, each time the Stable Diffusion model generates a target domain image, the image quality assessment method ARNIQA is used; without the need for a reference image, the quality of the generated image can be evaluated, and images that do not meet the image quality requirements are discarded, regenerated and evaluated again.