A training method for optimizing the degree of compliance of generated content with prompts through staggered diffusion
By interleaving different diffusion steps and prompt word optimization during the training process of the Stable Diffusion model, the deviation problem of the model in compliance with text prompts is solved, and the clarity and realism of the generated image is significantly improved.
Patent Information
- Application Number
- CN202411838570.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-12-13
AI Technical Summary
The Stable Diffusion model has room for improvement in complying with text prompts, and the generated images sometimes deviate from the input description, resulting in some ambiguity and uncertainty in the results.
A training method for the degree of compliance of the generated content for prompt words is proposed. By interleaved diffusion optimization, different diffusion steps and prompt words optimization are used in the training process, the prompt word compliance of the model is improved and the blurring and uncertainty of the generated images are reduced.
It significantly improves the model's compliance with prompt words, effectively reduces the ambiguity and uncertainty of the generated images, and enhances the practicality and reliability of the model in the field of precise image generation.
Smart Images

Figure CN119295605B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of optimizing the Stable-Diffusion training process, and particularly to a training method for optimizing the compliance degree of generated content with respect to prompts by interleaved diffusion. Background Art
[0002] The Stable Diffusion model is an advanced text-to-image generation technology that uses a deep neural network to convert an input text prompt into a detailed image. Although this technology has made significant progress in the field of image generation, there is still room for improvement in its compliance with text prompts. The generated images sometimes deviate from the input description, resulting in certain ambiguity and uncertainty in the results. Summary of the Invention
[0003] In view of this, the object of the present invention is to propose a training method for optimizing the compliance degree of generated content with respect to prompts by interleaved diffusion, which can strengthen the training process of Stable Diffusion, so that the generated image results are closer to the prompt. By interleaving different diffusion steps and prompt optimizations during the training process, the prompt compliance of the model is improved, and the ambiguity and uncertainty of the generated images are reduced, thereby enhancing the practicality and reliability in precise image generation applications.
[0004] According to one aspect of the present invention, there is provided a training method for optimizing the compliance degree of generated content with respect to prompts by interleaved diffusion, the method comprising the following steps:
[0005] Obtain picture data from a dataset and the picture data as well as the text data corresponding to the picture data and the text data ;
[0006] Fuse the picture data with a first noise to obtain picture data , and input the fused picture data and the text data into a diffusion generation network model to obtain a first predicted noise; recover the picture data based on the first predicted noise to obtain picture data ; analyze the picture data and the picture data to obtain a target noise;
[0007] Fuse the picture data with a first noise Obtain image data and input the image data along with the text data into the diffusion generation network model again to obtain the second predicted noise;
[0008] Train the neural network model based on the target noise and the second predicted noise.
[0009] In the above technical solution, the interleaved diffusion training method of this case aims to enhance the training efficiency of the Stable Diffusion model to generate images that better meet the requirements of the prompt. This method alternately uses different diffusion steps and prompt optimization strategies during the training phase, significantly improving the model's compliance with the prompt, effectively reducing the blurriness and uncertainty of the generated images, and thus enhancing the practicality and reliability of the model in the field of accurate image generation. Specifically, this method injects noise into the image data and trains the model to recover the original image from this noise, thereby enhancing the model's robustness to various noises and improving its generalization ability in diverse practical application scenarios. In addition, inputting the image data together with the corresponding text data enables the model to learn and master the internal connection between the image content and the text description, and thus achieve better performance in image generation or text-to-image conversion tasks. Introducing noise during the training process not only enables the model to make accurate predictions under conditions of interference but also is of great significance for dealing with various interferences and non-ideal conditions that may be encountered in practical applications. By comparing the predicted noise with the target noise, the performance of the model can be more accurately evaluated, thereby enabling more targeted optimization of the model training process. This method has good adaptability and can be applied to different data sets and task types. It does not depend on a specific data distribution but rather learns how to extract and recover information from noise, thereby comprehensively improving the performance of the model. This flexibility and efficiency make the interleaved diffusion training method proposed in this patent have broad application prospects in the field of image generation.
[0010] In some embodiments, the noise addition method adopts the following steps:
[0011] At time , the image data after fusing the noise is obtained by multiplying the original image data by a decay factor and adding it to the first noise multiplied by another factor ;
[0012] wherein, represents the decay coefficient, which gradually decreases as time passes.
[0013] In the above technical solution, compared with traditional noise addition methods, this method gradually incorporates noise into the data by introducing a time-dependent attenuation factor. This process simulates the natural accumulation of noise in the physical world and greatly promotes the model's in-depth learning of the underlying data structure and noise patterns. Through the attenuation coefficient, the noise level can be adjusted at different time points, providing the model with all-round adaptive training from a slightly noisy to a severely noisy environment. Compared with methods that only rely on simple forward and reverse processes, the step-by-step denoising strategy of this method can generate higher-quality images, ensuring the clarity and realism of the generated results. In addition, through this progressive denoising method, the model can process the data distribution more smoothly, effectively reducing the risk of discriminator overfitting, and thus improving the model's inference speed and computational efficiency.
[0014] In some embodiments, the picture data and the picture data satisfy the following conditional formula:
[0015]
[0016]
[0017] Wherein, represents the attenuation coefficient, represents the target noise.
[0018] In the above technical solution, the reconstruction target is still , that is, the above relational expression must hold. The relational expression ensures the consistency between the reconstruction target and the original image data . If the relational expression does not hold, then the reconstructed image may be significantly different from the original image, resulting in reconstruction failure.
[0019] According to another aspect of the present invention, there is provided a training device for optimizing the degree of compliance of generated content with prompts by staggered diffusion, including:
[0020] An acquisition module: used to acquire picture data and picture data as well as the corresponding text data and text data from a data set;
[0021] A first noise processing module: used to fuse the first noise with the picture data to obtain picture data , and to combine the fused picture data with the text data Input the diffusion generation network model to obtain the first predicted noise; restore the image data based on the first predicted noise , get the image data ; Analyze the image data With picture data , obtain the target noise;
[0022] The second noise processing module is used to process the image data Fusion First Noise Get image data , and the image data With text data Input the diffusion generation network model again to obtain the second prediction noise;
[0023] Training module: used to train the neural network model based on the target noise and the second predicted noise.
[0024] In the above technical solution, in order to better use the above method, the present application proposes a training device for the compliance degree of the generated content with the prompt words by staggered diffusion optimization. Each module corresponds to each step of the above method. The specific principle has been described above and will not be repeated here.
[0025] According to another aspect of the present invention, a stable diffusion model is provided, which is trained based on the above training method.
[0026] In the above technical solution, the model relies on the training method, and the model obtained based on the training method improves the prompt word compliance of the model, reduces the ambiguity and uncertainty of the generated image, and thus improves the practicality and reliability in the application of accurate image generation. It should be noted that the principle and effect of each step have been described above and will not be described in detail here.
[0027] According to another aspect of the present invention, there is provided a training device for the compliance of content generated by staggered diffusion optimization with prompt words, comprising: at least one processor and a memory in communication connection with the at least one processor;
[0028] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the above method.
[0029] In the above technical solution, in order to better run and process the method, the above method is stored in a memory, and the stored method is executed by a processor. It should be noted that the principle and effect of each step have been described above and will not be described in detail here.
[0030] According to the last aspect of the present invention, there is provided a computer-readable storage medium storing a computer program, which when executed by a processor implements the above-mentioned method.
[0031] In the above technical solution, in order to better run and use the method, the above method is stored in a computer-readable storage medium and the above method is implemented by using a processor. It should be noted that the principle and effect of each step have been described above and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0033] Figure 1 is a schematic flowchart of an embodiment of a training method for improving the compliance of generated content with prompts by staggered diffusion according to the present invention;
[0034] Figure 2 is a schematic diagram of a training framework of an embodiment of a training method for improving the compliance of generated content with prompts by staggered diffusion according to the present invention;
[0035] Figure 3 is a schematic structural diagram of an embodiment of a training device for improving the compliance of generated content with prompts by staggered diffusion according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The present invention will be further described in detail below with reference to the drawings and embodiments. It should be specifically noted that the following embodiments are only used to illustrate the present invention, but do not limit the scope of the present invention. Similarly, the following embodiments are only partial embodiments of the present invention rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0037] The present invention provides a training method for improving the compliance of generated content with prompts by staggered diffusion, which can strengthen the training process of Stable Diffusion, so that the generated image results are closer to the prompt. By alternately using different diffusion steps and prompt optimization during the training process, the prompt compliance of the model is improved, and the ambiguity and uncertainty of the generated images are reduced, thereby enhancing the practicality and reliability in accurate image generation applications.
[0038] Embodiment 1
[0039] Please refer to Figure 1 , a training method for optimizing the degree of compliance of generated content with prompts by staggered diffusion, the method comprising the following steps:
[0040] S1. Obtain image data from a dataset and the image data as well as the corresponding text data and the text data ;
[0041] S2. Fuse the first noise (the first noise is a randomly generated noise) to the image data
[0042] to obtain image data , and input the fused image data and the text data into a diffusion generation network model (text + noise, generating an image is the standard process of diffusion, the text is used to constrain and specify the generated content to simulate deviation in order to enable the training process to learn an "ability to get back on track", which is a setting solution not available in the standard training method), to obtain a first predicted noise (the first predicted noise is the noise predicted by the network when input); recover the image data based on the first predicted noise to obtain image data ; analyze the image data and the image data
[0043] to obtain a target noise (the target noise is the final learning target calculated); in this step, during the standard process, deviation is simulated and the new learning target required to correct the deviation is precisely defined, enabling the network to master the ability of automatic deviation correction S3. Fuse the first noise to the image data to obtain image data , and input the image data and the text data
[0044] into the diffusion generation network model again to obtain a second predicted noise (the second predicted noise is the noise predicted by the network when At * is input);
[0045] S4. Train a neural network model based on the target noise and the second predicted noise.
[0045] In this embodiment, the interleaved diffusion training method of this case aims to enhance the training efficiency of the Stable Diffusion model to generate images that better meet the requirements of the prompt. This method alternately uses different diffusion steps and prompt optimization strategies during the training phase, significantly improving the model's compliance with the prompt, effectively reducing the blurriness and uncertainty of the generated images, and thus enhancing the practicality and reliability of the model in the field of precise image generation. Specifically, this method injects noise into the image data and trains the model to recover the original image from this noise, thereby enhancing the model's robustness to various types of noise and improving its generalization ability in diverse practical application scenarios. In addition, the image data and the corresponding text data are input into the model together, enabling the model to learn and master the internal connection between the image content and the text description, and thus achieving better performance in image generation or text-to-image conversion tasks. Introducing noise during the training process not only enables the model to make accurate predictions under interfering conditions but also is of great significance for dealing with various interferences and non-ideal conditions that may be encountered in practical applications. By comparing the predicted noise with the target noise, the performance of the model can be more accurately evaluated, thereby enabling more targeted optimization of the model training process. This method has good adaptability and can be applied to different datasets and task types. It does not depend on a specific data distribution but comprehensively improves the model's performance by learning how to extract and recover information from noise. This flexibility and efficiency make the interleaved diffusion training method proposed in this patent have broad application prospects in the field of image generation.
[0046] In this embodiment, the noise addition method adopts the following steps:
[0047] At time , the image data after fusing the noise is obtained by multiplying the original image data by a decay factor and the first noise multiplied by another factor and then adding them together;
[0048] Among them, represents the decay coefficient, which gradually decreases as time passes.
[0049] In this embodiment, the standard noise addition method is to fuse a random noise with a graph as the input, and the network is required to predict this "random noise added to the photo". The method in this case redefines the input, which is to simulate the situation of deviating from the right track. Further, this case redefines the prediction target, no longer predicting the "random noise added to the picture", but changing to predicting the specific noise that can correct the deviation. The solution formula for the specific noise is derived in detail below and will not be elaborated here. Compared with the traditional noise addition method, this method gradually integrates the noise into the data by introducing a time-dependent attenuation factor. This process simulates the natural accumulation of noise in the physical world, greatly promoting the model's in-depth learning of the underlying structure of the data and the noise pattern. Through the attenuation coefficient, the noise level can be regulated at different time points, providing the model with all-round adaptive training from a slightly noisy to a severely noisy environment. Compared with those methods that only rely on simple forward and reverse processes, the step-by-step denoising strategy of this method can generate higher-quality images, ensuring the clarity and realism of the generated results. In addition, through this progressive denoising method, the model can process the data distribution more smoothly, effectively reducing the risk of discriminator overfitting, and thus improving the inference speed and computational efficiency of the model.
[0050] In this embodiment, the picture data and the picture data satisfy the following conditional formula:
[0051]
[0052]
[0053] where represents the attenuation coefficient, represents the target noise.
[0054] In this embodiment, the reconstruction target is still , that is, the above relational expression must hold. The relational expression ensures the consistency between the reconstruction target and the original image data . If the relational expression does not hold, then the reconstructed image may be significantly different from the original image, resulting in reconstruction failure.
[0055] To better explain the present invention, the following will be further elaborated. Please refer to Figure 2 , the steps are as follows:
[0056] 1. Extract unequal data A and B from the data set Z, and are pictures, , are texts describing A0 and .
[0057]
[0058]
[0059] 2. By adding noise, from obtain
[0060]
[0061] 3. The neural network under the text condition from recover , simulating the situation when deviating from the correct path
[0062]
[0063] 4. Then from add noise to obtain
[0064]
[0065] 5. And the reconstruction target is still , that is, it must hold
[0066]
[0067]
[0068] Among them, represents the attenuation coefficient, represents the target noise.
[0069] 6. This changes the training target of the neural network net:
[0070]
[0071]
[0072]
[0073]
[0074] Among them, is the supervision loss for the learnable parameter ; is the MSE loss calculated between the network prediction and the convergence target; is the diffusion noise of At; is the target noise; is the number of diffusion time steps; is a learnable parameter; After constructing the distribution deviation of the new learning objective (when using as the reconstruction guidance), the neural network net must follow the path of correction, and the retrained model will be significantly more compliant with the prompt guidance than the baseline.
[0075] Embodiment 2
[0076] Please refer to Figure 3 , a training device for optimizing the compliance degree of generated content with respect to prompts by interleaved diffusion, including:
[0077] An acquisition module: used to acquire image data from a dataset and image data as well as the corresponding text data of the image data and text data ;
[0078] A first noise processing module: used to fuse the first noise with the image data to obtain image data , and input the fused image data and the text data into the diffusion generation network model to obtain a first predicted noise; recover the image data based on the first predicted noise to obtain image data ; analyze the image data and the image data
[0079] to obtain the target noise; A second noise processing module: used to fuse the first noise with the image data to obtain image data , and input the image data and the text data
[0080] into the diffusion generation network model again to obtain a second predicted noise;
[0081] A training module: used to train the neural network model based on the target noise and the second predicted noise.
[0082] In this embodiment, in order to better use the above method, the present application proposes a training device for optimizing the compliance degree of generated content with respect to prompts by interleaved diffusion. Each module corresponds to each step of the above method, and its specific principle has been described above and will not be elaborated here.
[0082] Embodiment 3
[0083] A Stable Diffusion model trained based on the training method described in one of the embodiments.
[0084] In the above technical solution, the model depends on the training method, and the prompt compliance of the model obtained based on this training method is improved, reducing the ambiguity and uncertainty of the generated images, thereby enhancing the practicality and reliability in precise image generation applications. It should be noted that the principle and effect of each step have been described above and will not be elaborated here.
[0085] Embodiment 4
[0086] A training device for optimizing the compliance degree of generated content with prompts by staggered diffusion, including: at least one processor and a memory communicatively connected to the at least one processor;
[0087] Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in one of the embodiments.
[0088] In this embodiment, in order to better run and process the method described in one of the embodiments, the above method is stored in the memory, and the processor is used to execute the stored method. It should be noted that the principle and effect of each step have been described above and will not be elaborated here.
[0089] Embodiment 5
[0090] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above method is implemented.
[0091] In this embodiment, in order to better run and use the method described in one of the embodiments, the above method is stored in the computer-readable storage medium, and the processor is used to implement the above method. It should be noted that the principle and effect of each step have been described above and will not be elaborated here.
[0092] The above are only some embodiments of the present invention, and thus do not limit the protection scope of the present invention. Any equivalent device or equivalent process transformation made using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A training method for the compliance of generated content with prompt words by staggered diffusion optimization, characterized in that: The method comprises the following steps: Get image data from a dataset and image data And the text data corresponding to the image data and text data ; The image data Fusion First Noise Get image data , and the fused image data With text data Input the diffusion generation network model to obtain the first predicted noise; restore the image data based on the first predicted noise , get the image data ; Analyze the image data With picture data , obtain the target noise; The image data Fusion First Noise Get image data , and the image data With text data Input the diffusion generation network model again to obtain the second prediction noise; The diffusion generation network model is trained based on the target noise and the second prediction noise.
2. A training method for the compliance of content to prompt words by staggered diffusion optimization as claimed in claim 1, characterized in that: The noise adding method adopts the following steps: In time When the noise-fused image data is multiplied by an attenuation factor, the original image data and the first noise Multiply by another factor Add together to get; in, represents the attenuation coefficient, as time goes by As time goes by, it gradually decreases.
3. A training method for the compliance of content to prompt words by staggered diffusion optimization as claimed in claim 1, characterized in that: The picture data With picture data The following conditions are met: in, represents the attenuation coefficient, Represents the target noise.
4. A training device for the degree of compliance of generated content with prompt words by staggered diffusion optimization, characterized in that: include: Acquisition module: used to obtain image data from a data set and image data And the text data corresponding to the image data and text data ; The first noise processing module is used to process the image data Fusion First Noise Get image data , and the fused image data With text data Input the diffusion generation network model to obtain the first predicted noise; restore the image data based on the first predicted noise , get the image data ; Analyze the image data With picture data , obtain the target noise; The second noise processing module is used to process the image data Fusion First Noise Get image data , and the image data With text data Input the diffusion generation network model again to obtain the second prediction noise; Training module: used to train the diffusion generation network model based on the target noise and the second predicted noise.
5. A training device for the compliance of content generated by staggered diffusion optimization to prompt words as claimed in claim 4, characterized in that: The noise adding method adopts the following steps: In time When the noise-fused image data is multiplied by an attenuation factor, the original image data and the first noise Multiply by another factor Add together to get; in, represents the attenuation coefficient, as time goes by As time goes by, it gradually decreases.
6. A training device for the compliance of content generated by staggered diffusion optimization to prompt words as claimed in claim 4, characterized in that: The picture data With picture data The following conditions are met: in, represents the attenuation coefficient, Represents the target noise.
7. A training device for the compliance of generated content with prompt words by staggered diffusion optimization, characterized in that: include: at least one processor and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as claimed in any one of claims 1 to 3.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
Image generation content suppression method and system based on text graph diffusion model
CN117251589A
SAR image generation method based on de-noising diffusion probability model
CN118230191A