A Diffusion Model LoRA Fine-tuning Optimization Method and System Based on CLIP Loss and Perceptual Loss

By introducing CLIP loss and perceived loss in the diffusion model LoRA and dynamically adjusting the loss weight, the problem of high computing resources and model parameters of existing fine-tuning technology is solved, high-quality image generation is achieved and training costs are reduced.

CN119478587BActive Publication Date: 2025-06-10NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510027124.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-06-10
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

The high demand for existing fine-tuning technologies in terms of computing resources and model parameters has hindered its popularity, especially for ordinary users and small teams.

Method used

The LoRA fine-tuning optimization method based on CLIP loss and perceived loss is adopted to dynamically adjust the loss weight, optimize the performance of semantic consistency and visual details, reduce the computing resource requirements and reduce the amount of model parameters.

Benefits of technology

It significantly improves the overall quality of generated images, reduces the number of times and number of pictures required for fine-tuning training, reduces the computational overhead, and improves the training efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119478587B_ABST
    Figure CN119478587B_ABST
Patent Text Reader

Abstract

The present invention proposes a LoRA fine-tuning optimization method and system for a diffusion model based on CLIP loss and perceptual loss. The method includes: Step 1, during the LoRA fine-tuning process, dynamically adjust the weights of the CLIP loss and the perceptual loss by combining the CLIP loss and the perceptual loss; Step 2, use the CLIP model to calculate the semantic similarity between the denoised intermediate image and the target text, and optimize the noise prediction ability of the diffusion model according to the similarity difference; Step 3, adopt the perceptual loss to calculate the difference between the intermediate image and the target image in the feature space, and optimize the noise prediction ability of the diffusion model to improve the visual quality and detail fidelity of the generated image; Step 4, adjust whether to enable the CLIP loss and the perceptual loss according to the training progress. By introducing the CLIP loss, the model can better align the image with the text during the fine-tuning training process, making the generated image more conform to the description of the text prompt.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep neural networks, and particularly relates to a LoRA fine-tuning optimization method and system for a diffusion model based on CLIP loss and perceptual loss. Background Art

[0002] Diffusion models are generative models based on probability distributions. The basic idea is to gradually add noise to real data until it finally approaches the standard Gaussian distribution, which is called forward diffusion; then train a model to learn the reverse process, gradually restoring real data from random noise to achieve data generation. In the forward diffusion process, the model injects noise into the data in fixed steps, gradually destroying its structure until the data completely becomes unstructured Gaussian noise. The goal of reverse generation is to gradually remove the noise and restore a realistic data distribution. At present, the requirements of a large number of users and industrial demands for AIGC (Artificial Intelligence Generated Content) have become more refined. Content generated in different fields and by individual users needs to more precisely meet specific scenario requirements. This trend has promoted the birth and development of fine-tuning technology. Fine-tuning technology enables the model to adapt to specific tasks or user needs while maintaining its original capabilities by efficiently adjusting a pre-trained model, thereby achieving more personalized and customized generation effects.

[0003] Fine-tuning technology is a very important research direction in the field of diffusion models. The core idea of fine-tuning technology is to build on a pre-trained large model and, through a small amount of new task data and targeted training, make the model focus on a specific field or task. Traditionally, training deep learning models usually requires a large amount of data and computing resources, but fine-tuning technology can utilize the broad adaptability of pre-trained models to significantly reduce the training cost of new tasks and avoid building models from scratch. Common fine-tuning methods include full-parameter fine-tuning and partial-parameter fine-tuning, such as freezing some parameters and only optimizing specific layers. However, most fine-tuning technologies still face problems such as high computational resource requirements and large model parameter quantities after training. Since fine-tuning technology usually relies on large-scale pre-trained models, and the training and optimization of these models often require powerful computing resources, such as high-performance GPU clusters and large-capacity storage devices, this makes it difficult for most ordinary users and small teams to bear the corresponding costs and technical complexities. Even though fine-tuning has significantly reduced resource requirements compared to training a model from scratch, for many users, they still need to be familiar with deep learning frameworks, adjust hyperparameters, and prepare high-quality domain data, and these thresholds have hindered the popularization of fine-tuning technology.

[0004] To address this issue, in recent years, the LoRA fine-tuning technique has emerged. LoRA (Low-Rank Adaptation) is a lightweight model fine-tuning technique aimed at reducing the computational cost and storage requirements for fine-tuning large pre-trained models while maintaining the flexibility and effectiveness of fine-tuning. The core idea of LoRA is to utilize low-rank matrix factorization to efficiently update model parameters instead of adjusting the weights of the entire pre-trained model. During the fine-tuning process, LoRA freezes all the original parameters of the pre-trained model and inserts trainable low-rank matrices into the specified network layers. These low-rank matrices are used to capture the specific task information newly added during the fine-tuning process, thereby achieving efficient fine-tuning while reducing the number of parameter updates. Summary of the Invention

[0005] Object of the Invention: The present invention proposes a fine-tuning optimization method for a diffusion model LoRA (Low-Rank Adaptation) based on CLIP loss and perceptual loss. This method significantly improves the overall quality of generated images by dynamically optimizing the performance of semantic consistency and visual details. Specifically, it includes the following steps:

[0006] Step 1, during the LoRA fine-tuning process, combine the CLIP loss and the perceptual loss, and ensure the overall performance of the generated images in terms of semantic consistency and visual details by dynamically adjusting the weights of the CLIP loss and the perceptual loss;

[0007] Step 2, randomly sample the time step t, calculate the semantic similarity between the denoised intermediate image and the target text using the CLIP model, and optimize the noise prediction ability of the diffusion model according to the similarity difference;

[0008] Step 3, calculate the difference between the intermediate image and the target image in the feature space using the perceptual loss, and optimize the noise prediction ability of the diffusion model to improve the visual quality and detail fidelity of the generated images;

[0009] Step 4, adjust whether to enable the CLIP loss and the perceptual loss according to the training progress;

[0010] In Step 1, the StableDiffusion model is used as the pre-trained diffusion model. To accelerate the training time, the StableDiffusion model uses the VAE (Variational Autoencoder) model to compress the image into the latent space for calculation during the training phase. The latent space means reducing the image size but increasing the number of channels, aiming to reduce the computational amount during per-pixel point calculation; the VAE model includes an encoder and a decoder. The encoder is used to compress the image into the latent space, and the decoder is used to restore the image size.

[0011] During the LoRA fine-tuning process, in the initial fine-tuning training stage, first let the StableDiffusion model focus on predicting noise. At this time, since the model's noise prediction is not yet stable, adding the CLIP loss term too early is likely to cause the model's noise prediction effect to deteriorate. The standard deviation of the regular loss is calculated using the following formula :

[0012]

[0013] where k is the sliding window size, which is default set to 10; a threshold θ is set, usually set to 0.07; is the average loss value within the sliding window; when the standard deviation of the regular StableDiffusion model loss is less than or equal to the threshold θ, it is determined that the regular loss tends to be stable, and at this time, the CLIP loss is enabled; represents the current time step, t represents all the time steps that can be sampled from to in the window; represents the loss of the regular StableDiffusion model calculated at the current sampling step;

[0014] The loss of the regular StableDiffusion model is defined as:

[0015]

[0016] where represents the true noise, represents the noise predicted by the StableDiffusion model; N represents the number of samples in a batch;

[0017] In the present invention, the loss calculation formulas for the CLIP loss and the perceptual loss are respectively improved. Under this loss formula, the CLIP loss and the perceptual loss can adaptively adjust the weights. In the CLIP loss, this improvement can help prevent the CLIP loss from causing instability in model training due to the excessive difference between the semantics and the noise-added image when the number of noise-added steps sampled is too large when intervening in the total loss. Under the perceptual loss, this improvement helps the perceptual loss to automatically adjust the weights of each layer and better adjust the optimization parameters.

[0018] In step 2, after enabling the CLIP loss term, use the pre-trained CLIP model to calculate the semantic similarity between the image and the target text. In the present invention, considering that the CLIP model is applied in the training stage, if the time step t sampled in the training stage is too large, there is too much noise in the noisy image x t At this time, performing CLIP calculation is likely to cause unstable loss. Therefore, in the present invention, a loss function is designed for the CLIP model to adapt to the noisy images with different degrees in the StableDiffusion model. The loss function formula is defined as:

[0019]

[0020]

[0021] Where T represents the maximum time step in the Stable Diffusion model, with a default value of 1000; Represents the vector value after CLIP encoding of the image; Represents the vector value after text encoding; Represents the CLIP loss term; Represents the weight value, calculated according to the time step t;

[0022] Step 2.1, in order to use the perceptual loss, the noise addition process of the Stable Diffusion model is modified. First, obtain the target fine-tuning image X from the target fine-tuning dataset 0 , and then use the encoder in the VAE model to compress the fine-tuning image X 0 to obtain the compressed target fine-tuning image , and then add noise to the compressed target fine-tuning image for t - 1 steps to obtain the noisy image x t-1 , this step is to provide the real x t-1 image for the calculation of the perceptual loss in the subsequent steps. Then add one more step of noise to the current noisy image x t-1 to obtain the noisy image x t , and at the same time retain the Gaussian noise randomly sampled at the t-th step for calculating the conventional loss pixel by pixel using MSE (Mean Squared Error Loss);

[0023] Perform a single-step denoising of the current noisy image x t using the conventional Stable Diffusion model, specifically including: encoding the corresponding text text of the compressed target fine-tuning image into a text vector T text , and then inputting T text and the noisy image x t into the Stable Diffusion model. The Stable Diffusion model will combine the internal parameters of LoRA to predict the noise and obtain the predicted noise , and then use to remove the noise added at the t-th step and obtain the predicted noisy image ;

[0024] Step 2.2, perform CLIP image encoding on the predicted noisy image to obtain the image vector I;

[0025] Step 2.3, calculate the text vector T textThe semantic similarity with the image vector I. As described in step 2.1, during the process of predicting noise, LoRA participates in the calculation of the Stable Diffusion model. Therefore, this part of the calculation will be passed to LoRA during backpropagation. During backpropagation, the gradient descent algorithm will be calculated based on the conventional loss term and the CLIP loss term in the present invention, and finally used to adjust the parameters in LoRA.

[0026] Step 2.1 includes:

[0027] Step 2.1.1, randomly sample a step of Gaussian noise , the Gaussian noise is used for denoising at the (t - 1)-th step, and then the Stable Diffusion model's one-step denoising formula x t-1 is used to obtain the noisy target image x that has been denoised for t-1 steps;

[0028] Among them, in the Stable Diffusion model, a set of parameters , refers to the t-th value in a set of parameters , refers to ;

[0029] Step 2.1.2, obtain random Gaussian noise and perform another single-step denoising on the current noisy target image x t-1 that has been denoised for t - 1 steps. The denoising formula is:

[0030] ,

[0031] to obtain the denoised image , and at the same time retain the Gaussian noise as the true noise; refers to the t-th data in a set of parameters predefined by the linear data sampling scheduling method in the Stable Diffusion model;

[0032] Step 2.1.3, use the CLIP model to perform word embedding operations on the text text. First, split the text text into independent word or sub-word units, and then perform embedding processing on the units through the text encoder of the CLIP model to generate a text vector T corresponding to the semantics of the text text ; The text vector T text can capture the semantic information of the text and be used as a reference for image generation during the multi-modal alignment process;

[0033] Step 2.1.4, the text vector T text, the noisy image and the current time step t are input into the StableDiffusion model combined with LoRA, allowing the model to predict the amount of noise added at the t-th step , for the image Remove the noise at the step, and use the following formula to obtain the noisy image:

[0034] .

[0035] Step 2.2 includes:

[0036] Step 2.2.1, considering the application limitations of the CLIP model, after obtaining the noisy image , before calculating the CLIP loss, first restore the size of the noisy image through the decoder of the VAE model to obtain the noisy original-size image , the purpose of this step is to match the input requirements of the CLIP model;

[0037] Step 2.2.2, use the image encoder of CLIP to extract the feature vector from the noisy original-size image , convert the noisy original-size image into the image vector I. The image vector I captures the key semantic information of the image and can be used to measure the similarity between the image and the text description, so as to achieve semantic alignment or optimization in multimodal tasks.

[0038] Step 2.3 includes:

[0039] Step 2.3.1, use the loss formula to calculate the loss between the text vector T text and the image vector I, where is replaced by T, is replaced by I, and the obtained is used as the semantic similarity between the text vector and the image vector I;

[0040] Step 2.3.2, judge the weight setting at this time according to the current time step t , considering that when t is large, the calculation of the CLIP loss is meaningless. At this time, if a large weight is set, it is often easy to cause the model to lose control of the loss prediction. Through the formula , adjust the weight value, so that a large weight value can be retained when t is large, and the weight is adjusted to an appropriate value when t is small;

[0041] Step 2.3.3, since in the present invention, when calculating the CLIP loss, the noisy original-size image is used, and the LoRA bypass parameters in the denoising process directly participate in the denoising For the prediction, the CLIP loss is introduced into the gradient calculation graph, and the calculated CLIP loss gradient will be backpropagated along the optimization path to adjust the parameters in LoRA.

[0042] Step 3 includes:

[0043] In step 3.1, considering the disadvantages of the traditional perceptual loss, which cannot perceive the importance of the losses at different layers, in the traditional perceptual loss, first, a deep convolutional network is used to extract features from the image, features are extracted at different levels respectively, and finally, simple summation is used as the total loss. In the present invention, a new perceptual loss function is adopted to optimize the network result, and the formula is:

[0044]

[0045]

[0046] Where represents the total loss value of the perceptual loss, represents the feature extraction of the image by each layer of the convolutional network, represents the dynamic weight term, which will adjust the composition of the final total perceptual loss according to the loss values of different layers; exp is the natural exponential function; represents the feature extraction network of the k-th layer;

[0047] In step 3.2, considering the design of the perceptual loss, the method of step 2.2.1 is used to reduce the size of the noisy image and then calculate using the perceptual loss to ensure the integrity of the subsequent input features;

[0048] In step 3.3, the perceptual loss is used to calculate the perceptual loss for the noisy original-size image and the target fine-tuning image X 0 By calculating the difference values between the noisy original-size image and the target fine-tuning image X 0 at different levels in the perceptual model, it can better guide the adjustment of the internal parameters of LoRA during backpropagation. Because of the adaptive loss function of the present invention, the perceptual loss will adjust the parameters in LoRA according to the importance of each layer, enabling the model to more carefully capture the details and semantic differences beyond the pixel level, so as to improve the visual quality and detail fidelity of the generated image.

[0049] Step 3.3 specifically includes:

[0050] Step 3.3.1: Layer the feature extraction network VGG16, and select the convolutional outputs of the 0th, 1st, 2nd, 3rd, 4th, and 5th layers as the five feature vectors of the image; add a linear layer to the output features for linear transformation. Here, neurons are killed with a probability of 20% to prevent overfitting; the input channels of the feature extraction network VGG16 are adjusted to 64, 128, 256, 512, 512 to better capture the feature information of the image under multiple channels.

[0051] Step 3.3.2: Input the original-sized noisy target image and the target fine-tuned image X 0 together into the feature extraction network VGG16. During the calculation process, the feature extraction network VGG16 will save the calculation results in the 0th, 1st, 2nd, 3rd, 4th, and 5th layers for calculating the loss.

[0052] Step 3.3.3: Calculate the second-order norm loss of the original-sized noisy target image 0 and the target fine-tuned image X after passing through different feature extraction layers, and record the loss value. This step is used to calculate the size of the weighting value, and the formula is:

[0053] Calculate and save the size of the weighting value for each layer, and finally use the formula to sum the total loss:

[0054] .

[0055] Among them represents the perceptual loss term.

[0056] Step 4 includes: Define the following final loss function:

[0057]

[0058] Among them, refers to the total loss value, is the CLIP loss, is the perceptual loss. When the CLIP loss enabling condition is not met, the parameter is set to 0; after the CLIP loss is enabled, is set to 1.

[0059] Through backpropagation, the total loss signal is transmitted into the LoRA model to optimize its parameter update. This process can effectively guide the model to capture detailed and semantic features, further improving the overall quality of the generated image and the text alignment effect.

[0060] The present invention also provides a diffusion model LoRA fine-tuning optimization system based on CLIP loss and perceptual loss, including:

[0061] A CLIP loss calculation module, configured to: encode the text text corresponding to the target fine-tuning image to obtain a text vector, and then compress the target fine-tuning image After adding noise for t steps, use the stable diffusion model merged with LoRA to predict the noise at the t-th step, and then subtract the predicted noise to obtain a noisy image The purpose is to introduce the CLIP loss into the gradient calculation graph for backpropagation to adjust LoRA; then Use the decoder of the VAE model to restore the size, then use the CLIP model of the present invention to encode the resized image into a vector, and finally use the adaptive weight loss calculation formula in the present invention to calculate the difference between the text vector and the image vector;

[0062] A perceptual loss calculation module, configured to: first, input the target fine-tuning image X 0 and the noisy image from which the noise at the t-th step has been removed into the feature extraction network VGG16. This model is specified to extract features on the 0, 1, 2, 3, 4, and 5 convolutional layers in VGG16. Finally, in order to enable the model to adaptively adjust the different importance of each layer of the perceptual loss, use the adaptive perceptual loss in the present invention to calculate the features extracted above.

[0063] Finally, after combining the CLIP loss and the perceptual loss with the basic loss, perform backpropagation to guide the update of the parameters inside LoRA.

[0064] Beneficial effects: The present invention integrates the CLIP loss and the perceptual loss into the LoRA fine-tuning, thereby enhancing the effect of LoRA fine-tuning, and at the same time reducing the number of fine-tuning training times and the number of pictures required. By introducing the CLIP loss, the model can better align the image with the text during the fine-tuning training process, making the generated image more in line with the description of the text prompt. The introduction of the perceptual loss enables the model to learn to align the generated image with the target image during the training process, thereby improving the quality and visual effect of the generated image. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0066] Figure 1 It is a flowchart of adding noise to the model in the present invention.

[0067] Figure 2It is a flowchart of noise prediction by combining the model with LoRA in the present invention.

[0068] Figure 3 It is a calculation process of two additional loss terms proposed in the present invention.

[0069] Figure 4 It shows a comparison chart of the effects of the present invention, the original LoRA, and SD-2.1V. Detailed implementation manner

[0070] The embodiment of the present invention provides a LoRA fine-tuning optimization method for a diffusion model based on CLIP loss and perceptual loss. Figure 1 It shows the noise addition process of the model, which starts from obtaining the data picture X from the target fine-tuning image. First, the target picture X is input into the encoder of the variational autoencoder (VAE) for processing. In the present invention, the original image X is a picture with a size of 512×512. When scaling the image, the experiment selects the strategy of enlarging it to 8 times the original size. The experimental results show that when the scaling ratio is between 8 and 16 times, not only can the image details be well retained, avoiding a large amount of detail loss during decoding, but also the consumption of computational resources is effectively reduced.

[0071] After being processed by the encoder, the image is compressed into a latent feature matrix Z, which is a highly condensed latent representation for subsequent noise addition and loss calculation. Then, the latent feature matrix Z is gradually added with noise, first added to the t−1 step. In this process, the noise components of the t−1 step are gradually mixed into the latent feature matrix Z, and in this state, it is used to calculate the perceptual loss to ensure that the addition of noise does not significantly affect the feature representation ability.

[0072] Subsequently, the Z that already contains the noise of the t−1 step is added with noise in a single step again to make it added with noise to the currently randomly selected time step t. In this noise addition process, the noise term ε is additionally retained for the calculation of subsequent conventional losses (such as L2 loss). In this way, while the model adds noise step by step, it can effectively balance the relationship between the retention of image details and the computational complexity, and lay a foundation for high-quality generation and reconstruction.

[0073] Figure 2 It shows the overall process of the model predicting noise, and details how the latent feature vector Z and the text information T efficiently predict noise through the U-Net network. In this process, first, the latent feature vector Z and the text information T obtained from the previous stage are jointly input into the U-Net network for processing. The architecture of the U-Net network includes an encoder and a decoder, and multiple cross-attention modules are embedded inside, which are specifically used to effectively fuse text and image information.

[0074] The first step of the process is to feed the input into the encoder module. In the encoder, each layer contains a cross-attention mechanism that completes the deep fusion of the text and image features by calculating the attention weights between them. During this process, the LoRA (Low-Rank Adaptation) parameters and the original parameters in the cross-attention module are combined to complete the calculation. However, the original parameters of the cross-attention module are frozen at this stage, meaning that only the LoRA parameters will be optimized during backpropagation. This design not only reduces the number of parameters that need to be adjusted but also significantly improves the training efficiency of the model while ensuring the stability of the core attention mechanism.

[0075] After being processed by the encoder, a compressed latent feature matrix that combines text information and image information is generated. This matrix not only retains the spatial representation of the image features but also contains the semantic information of the text description, providing rich context for subsequent generation.

[0076] Next, the compressed latent feature matrix is passed to the decoder module. In the decoder, the network restores the low-resolution feature matrix to an output that matches the resolution of the original image through successive upsampling operations. During this process, the decoder not only employs a cross-attention module but also introduces low-rank adaptation (LoRA) parameters to achieve the deep fusion of text and image information through these components. This design ensures that each step of the generation process makes full use of the text guidance information, thereby enhancing the relevance of the generated result to the text description while retaining the detailed performance and sense of hierarchy of the image, providing a twofold guarantee for the quality of the final generated image.

[0077] Figure 3 Shows the calculation process of two additional loss terms proposed in the present invention. Specifically, first, the predicted latent feature vector Z containing noise at step t - 1 is processed by the decoder to enlarge its size back to the resolution of the original image. At the same time, the original image with noise at step t - 1 is also restored to the original size through the decoder to meet the input requirements of the subsequent CLIP model and VGG16 network. In the specific process of calculating the loss, the CLIP loss is calculated first. The text information of the current original image is extracted through the CLIP model and calculated with the latent feature Z whose size has been restored, obtaining the CLIP loss term to measure the matching degree between the generated content and the text description.

[0078] Next, the restored predicted image X and the latent feature Z are input into the VGG16 network to calculate the multi-layer perception loss. The multi-layer perception loss can more comprehensively capture the differences in details and styles of the generated images by comparing and analyzing the perception results of different feature layers, further improving the generation quality. Finally, the average value of the multi-layer perception loss is taken to return a scalar loss term for guiding model optimization.

[0079] After obtaining the CLIP loss, the multi-layer perception loss, and the conventional noise loss term, the weighted loss calculation formula is used to combine these loss terms to generate the final overall loss value. Specifically, by reasonably setting the weights of each loss term, it is ensured that their contributions to model training during the optimization process are balanced and effective. Subsequently, parameter optimization is performed through the backpropagation algorithm. During the gradient descent process, the model automatically updates the internal parameters of LoRA, while the weights of the diffusion model are frozen and do not participate in the update calculation. This design significantly reduces the computational overhead required for training and effectively utilizes the efficient characteristics of LoRA.

[0080] In the present invention, the total number of parameters of the model reaches 872,550,340, but the actual parameters participating in the optimization calculation are only 6,639,616, accounting for less than 1%. This significant reduction in the number of parameters greatly reduces the computational burden and makes the training process more efficient. Taking an ordinary machine equipped with an RTX 4070 SUPER graphics card as an example, running 5000 training steps only takes about half an hour to complete. This high training speed greatly improves the practicality of the model and also provides the possibility for large-scale model training in resource-constrained hardware environments.

[0081] In addition, by freezing the weights of the diffusion model, it is ensured that its original performance is not affected, while the optimization of LoRA parameters can fully explore the potential of text-image matching. This targeted optimization strategy not only significantly improves the quality of the generated images but also reduces the dependence on hardware performance, providing a reliable basis for the implementation and further improvement of the model.

[0082] Figure 4 The comparative experimental results of the LoRA fine-tuning method proposed in the present invention with the traditional LoRA fine-tuning method and the conventional diffusion model are shown. From the leftmost column to the rightmost column are: the method proposed in the present invention, the traditional LoRA fine-tuning method, and the original Stable Diffusion model. The experiment used the rich woman dataset, which is characterized by distinct human figures and includes diverse features such as wearing bright clothes and accessories like sunglasses. Under the same generation seed and prompt text conditions, the optimized LoRA fine-tuning method of the present invention shows significant advantages, not only being able to more accurately understand and follow the prompt text but also effectively improving the quality and detail fidelity in the generated images.

[0083] The present invention provides a LoRA fine-tuning optimization method and system for a diffusion model based on CLIP loss and perceptual loss. There are many methods and ways to specifically implement this technical solution. The above is only the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented using existing technologies.

Claims

1. A LoRA fine-tuning optimization method for a diffusion model based on CLIP loss and perceptual loss, characterized in that: The steps include: Step 1: In the LoRA fine-tuning process, the perceptual loss is combined and the CLIP loss is introduced according to the magnitude of the loss change. The weights of the CLIP loss and the perceptual loss are dynamically adjusted during training. In step 1, a stable diffusion model is used as a pre-trained diffusion model, and a VAE model is used to compress the image into a latent space for calculation during the training phase; the VAE model includes an encoder and a decoder, the encoder is used to compress the image into the latent space, and the decoder is used to restore the image size; During the LoRA fine-tuning process, the following formula is used to calculate the standard deviation of the conventional loss: Where k is the sliding window size; is the average loss value in the sliding window; when the loss of the conventional stable diffusion model is within k steps, the standard deviation When it is less than or equal to the threshold θ, the conventional loss is judged to be stable, and the CLIP loss is enabled at this time; represents the current time step, t represents the time from arrive All time steps that can be sampled in the window; L t It represents the loss of the conventional stable diffusion model calculated under the current sampling step; The conventional stable diffusion model loss is defined as: where y i represents the real noise, represents the prediction noise of the stable diffusion model; N represents the number of samples in a batch; Step 2: randomly sample time step t, use the CLIP model to calculate the semantic similarity between the denoised intermediate image and the target text, and optimize the noise prediction ability of the diffusion model based on the similarity difference; Step 3: Use perceptual loss to calculate the difference between the intermediate image and the target image in the feature space and optimize the noise prediction ability of the diffusion model; Step 4: Dynamically adjust the weighted loss function based on whether CLIP loss is enabled during training; Step 4 includes: defining the following final loss function: L total =L t +αL clip +L perc , Among them, L total Refers to the total loss value, L clip is the CLIP loss, L perc For perceptual loss, when the CLIP loss enabling condition is not met, the parameter α is set to 0; after the CLIP loss is enabled, α is set to 1.

2. The method according to claim 1, characterized in that In step 2, after enabling the CLIP loss term, the pre-trained CLIP model is used to calculate the semantic similarity between the image and the target text. A loss function is designed for the CLIP model to adapt to different degrees of noisy images in the stable diffusion model. The loss function formula is defined as: Where T represents the maximum time step in the stable diffusion model; f img Represents the vector value after CLIP encoding of the image; f text Represents the vector value after text encoding; L CliP represents the CLIP loss term; Represents the weight value; Step 2.1: In order to use perceptual loss, the stable diffusion model denoising process is modified. First, the target fine-tuning image X0 is obtained from the target fine-tuning dataset. Then, the fine-tuning image X0 is compressed using the encoder in the VAE model to obtain the compressed target fine-tuning image x0. Then, the compressed target fine-tuning image x0 is denoised by t-1 steps to obtain the noise image x0. t-1 , and then for the current noise image x t-1 Add one more step of noise to get the noisy image x t , while retaining the Gaussian noise ε randomly sampled in the tth step t Used for MSE to calculate conventional loss pixel by pixel; For the current noise image x t Perform single-step denoising of the conventional stable diffusion model, specifically including: converting the text text corresponding to the compressed target fine-tuning image x0 into a text vector T through CLIP encoding text , then T text and the noisy image x t Input into the stable diffusion model, the stable diffusion model will combine the internal parameters of LoRA to predict the noise and obtain the predicted noise Then use Remove the noise added in step t to obtain the predicted noisy image Step 2.2, for the predicted noisy image Perform CLIP image encoding to obtain image vector I; Step 2.3, calculate the text vector T text The semantic similarity between the image vector I and the image vector I is calculated by the gradient descent algorithm according to the conventional loss term and the CLIP loss term during the back propagation process.

3. The method according to claim 2, characterized in that Step 2.1 includes: Step 2.1.1, randomly sample one step of Gaussian noise ε t-1 , Gaussian noise ε t-1 Used for the t-1th step of noise addition, and then use the stable diffusion model one-step noise addition formula Get the noisy target image x with t-1 steps of noise added t-1 ; Among them, a set of parameters a, α are predefined in the stable diffusion model. t refers to the tth value in a set of parameters a, refers to Step 2.1.2, obtain random Gaussian noise ε t For the current noisy target image x that has been noisy for t-1 steps t-1 Perform another single-step noise addition, the noise addition formula is: Get the noisy image x t , while retaining the Gaussian noise ε t As the real noise; β t Refers to the tth data in a set of predefined parameters β in the stable diffusion model; Step 2.1.3: Use the CLIP model to perform word embedding on the text. First, split the text into independent words or subword units, and then embed the units through the text encoder of the CLIP model to generate a text vector T corresponding to the text semantics. text ; Step 2.1.4, transform the text vector T text , the image after adding noise x t and the current time step t are passed into the stable diffusion model combined with LoRA, allowing the model to predict the amount of noise added at step t For image x t Remove the noise in step t and use the following formula to obtain the noisy image:

4. The method according to claim 3, characterized in that: Step 2.2 includes: Step 2.2.1, after obtaining the noisy image After that, the noisy image is first converted to The decoder of the VAE model restores the size to obtain the original size image with noise Step 2.2.2: Use CLIP’s image encoder to decode the noisy original size image. Extract feature vectors and convert the original noisy image into Convert to image vector I.

5. The method according to claim 4, characterized in that Step 2.3 includes: Step 2.3.1, use the loss formula For the text vector T text And the image vector I is used for loss calculation, where f text By T text Replace, f img Replaced by I, the obtained L CliP As the semantic similarity between the text vector and the image vector I; Step 2.3.2, determine the weight setting at this time according to the current time step t By formula Adjust the weight value; In step 2.3.3, the CLIP loss is introduced into the gradient calculation graph, and the calculated CLIP loss gradient is passed back along the optimization path to adjust the parameters in LoRA.

6. The method according to claim 5, characterized in that Step 3 includes: Step 3.1, use a perceptual loss function to optimize the network results, the formula is: L perC =∑ l w l ·|F l (x t )-F l (x0)| 2 , Where L perC represents the total loss value of the perceptual loss, F l (·) represents the feature extraction of each layer of convolutional network on the image, w l represents the dynamic weight term; exp is the natural exponential function; F k represents the feature extraction network of the kth layer; Step 3.2, for noisy images After resizing, use perceptual loss for calculation; Step 3.3, use perceptual loss L perc For the original noisy image The perceptual loss is calculated by calculating the original size image with noise The difference between the target fine-tuning image X0 and the layers in the perception model is used to guide the adjustment of LoRA internal parameters during back propagation.

7. The method according to claim 6, characterized in that Step 3.3 specifically includes: Step 3.3.1, the feature extraction network VGG16 is layered, and the outputs of the 0th, 1st, 2nd, 3rd, 4th, and 5th layers of convolution are selected as the five feature vectors of the image; a linear layer is added to the output features, and a linear change is performed to kill neurons with a set probability; the input channels of the feature extraction network VGG16 are adjusted to 64, 128, 256, 512, 512; Step 3.3.2: Convert the original size noisy target image The target fine-tuned image X0 is input into the feature extraction network VGG16. During the calculation process, the feature extraction network VGG16 saves the calculation results in the 0th, 1st, 2nd, 3rd, 4th, and 5th layers for calculating the loss. Step 3.3.3, calculate the original size noisy target image and the second-order norm loss of the target fine-tuned image X0 after passing through different feature extraction layers |F l (x t )-F l (x0)| 2 , and record the loss value, the formula is: Calculate the weighted value of each layer and save it, and finally use the formula to sum the total loss: L perC =∑ l w l ·|F l (x t )-F l (x0)| 2 , Where L perC represents the perceptual loss term.

8. A LoRA fine-tuning optimization system for a stable diffusion model based on CLIP loss and perception loss implemented according to the method described in any one of claims 1 to 7, characterized in that: include: The CLIP loss calculation module is used to: encode the text text corresponding to the target fine-tuning image to obtain the text vector, then add noise to the compressed target fine-tuning image x0 for t steps, use the stable diffusion model combined with LoRA to predict the noise of the tth step, and then subtract the predicted noise to obtain the noisy image Then the noisy image Use the decoder of the VAE model to restore the size, then use the CLIP model to encode the restored image into a vector, and finally use the adaptive weight loss calculation formula to calculate the difference between the text vector and the image vector; The perceptual loss calculation module is used to: firstly calculate the target fine-tuned image X0 and the noisy image from which the noise of the tth step has been removed Input the feature extraction network VGG16, specify to extract features on the 0th, 1st, 2nd, 3rd, 4th, and 5th convolution layers in VGG16, and use the adaptive perceptual loss to calculate the features extracted above; finally, after merging the CLIP loss and the perceptual loss with the basic loss, back propagate to guide the parameter update inside LoRA.

Citation Information

Patent Citations

  • Training method and device of image generation model

    CN117541459A

  • Method and system for optimizing text generation texture chartlet based on diffusion model

    CN118230090A