Diffusion model step-by-step reward learning and optimization method, system, device and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-08-11
AI Technical Summary
但扩散模型的中间潜表示高度噪声化、语义可解释性弱,导致难以直接为这些潜变量构建可靠的奖励标签
[0011]由上述本发明提供的技术方案可以看出,通过对偏好与非偏好噪声潜变量各自对应的奖励之间的差值进行约束,构建标准成对偏好损失可以使噪声潜变量的生成更加可靠,并且,通过构建的一致性损失可以使提升分步奖励模型的平滑度,显著减少偏好翻转;此外,通过分阶段训练避免分步奖励模型退化,从而兼顾判别能力与一致性。最终,本发明在扩散模型的偏好优化中表现出显著优势,使生成图像在审美质量、语义一致性、多属性组合、空间关系等多个指标上超过现有技术,适用性强,可直接用于稳定扩散模型等主流框架的偏好对齐训练。
Smart Images

Figure CN122366535B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of diffusion model optimization technology, and in particular to a method, system, device and medium for step-by-step reward learning and optimization of diffusion models. Background Technology
[0002] Current text-to-image generation techniques are primarily based on diffusion models. Diffusion models generate images through a multi-step denoising process, with each step operating within a latent space at different noise levels. To make the generated images more consistent with human preferences and semantics, a class of training methods based on human preferences has emerged in recent years, such as human feedback reinforcement learning and direct preference optimization. However, these methods mostly operate on the final output of the model, lacking fine-grained supervision of intermediate noise latent variables.
[0003] Previous studies have attempted to introduce step-by-step reward models into each denoising step of diffusion models to evaluate the latent variables at each step. However, the intermediate latent representations of diffusion models are highly noisy and have weak semantic interpretability, making it difficult to directly construct reliable reward labels for these latent variables. Some methods assume that "adding the same noise to preference pairs does not change the original preference relationship," but this assumption lacks theoretical support. When noise is amplified, samples that were originally preferred are prone to preference reversal, leading to the accumulation of step-by-step reward errors and affecting the optimization process.
[0004] Furthermore, existing step-by-step reward models often rely on simple noise addition methods to obtain intermediate noise latent variables, resulting in large forward mapping errors. The step-by-step reward models themselves may also lack sufficient smoothness and be highly sensitive to noise perturbations. Therefore, how to obtain a stable and reliable preference signal in the noise latent space, and thus assist the diffusion model to improve the quality of the generated image, is a pressing problem to be solved in the current step-by-step preference optimization of diffusion models.
[0005] In view of this, the present invention is hereby proposed. Summary of the Invention
[0006] The purpose of this invention is to provide a method, system, device, and medium for step-by-step reward learning and optimization of a diffusion model, which can improve the preference consistency of the diffusion model throughout the denoising and generation process, reduce the preference reversal of noise latent variables, construct a reliable step-by-step reward model, and enhance the preference alignment effect of the diffusion model, ultimately improving the preference consistency, semantic understanding ability, and aesthetic quality of the diffusion model in image generation.
[0007] The objective of this invention is achieved through the following technical solution: A step-by-step reward learning and optimization method for a diffusion model includes: Construct a training set consisting of clean latent space sample pairs of preferences and non-preferences under specified text conditions; Through the forward mapping process of the diffusion model, clean latent space samples of preferences and non-preferences under specified text conditions are obtained, along with the corresponding noisy latent variables of preferences and non-preferences at each time step. The step-by-step reward model is constructed and trained, including: predicting the rewards corresponding to the preferred and non-preferred noisy latent variables at each time step using the step-by-step reward model, and constraining the difference between the rewards to construct a standard pairwise preference loss; predicting the rewards corresponding to the preferred and non-preferred clean latent space samples using the step-by-step reward model, and jointly constraining the difference between the reward of the preferred noisy latent variable and the reward of the preferred clean latent space sample at each time step, as well as the difference between the reward of the non-preferred noisy latent variable and the reward of the non-preferred clean latent space sample at each time step to construct a consistency loss; training the step-by-step reward model using the standard pairwise preference loss within the training warm-up step range, and then training the step-by-step reward model by combining the standard pairwise preference loss and the consistency loss. The trained reward model guides the preference alignment training process of the diffusion model, and the trained diffusion model outputs a denoised and preference-aligned image.
[0008] A diffusion model step-by-step reward learning and optimization system includes: Training set construction unit, used to construct a training set consisting of clean latent space sample pairs of preferences and non-preferences under specified text conditions; The diffusion model-based mapping unit is used to obtain clean latent space samples of preferences and non-preferences under specified text conditions through the forward mapping process of the diffusion model, and the corresponding preference and non-preference noise latent variables at each time step. The step-by-step reward model construction and training unit is used to construct and train the step-by-step reward model, including: predicting the rewards corresponding to the preferred and non-preferred noise latent variables at each time step using the step-by-step reward model, and constraining the difference between the rewards to construct a standard pairwise preference loss; predicting the rewards corresponding to the preferred and non-preferred clean latent space samples using the step-by-step reward model, and jointly constraining the difference between the reward of the preferred noise latent variable and the reward of the preferred clean latent space sample at each time step, as well as the difference between the reward of the non-preferred noise latent variable and the reward of the non-preferred clean latent space sample at each time step to construct a consistency loss; training the step-by-step reward model using the standard pairwise preference loss within the training warm-up step range, and then training the step-by-step reward model by combining the standard pairwise preference loss and the consistency loss. The image generation unit is used to guide the preference alignment training process of the diffusion model using the trained reward model, and outputs a denoised and preference-aligned image from the trained diffusion model.
[0009] A processing device includes: one or more processors; and a memory for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.
[0010] A readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method.
[0011] As can be seen from the technical solution provided by this invention, by constraining the difference between the rewards corresponding to the preferred and non-preferred noise latent variables, constructing a standard pairwise preference loss can make the generation of noise latent variables more reliable. Furthermore, the constructed consistency loss can improve the smoothness of the step-by-step reward model and significantly reduce preference flipping. In addition, staged training avoids the degradation of the step-by-step reward model, thus balancing discriminative ability and consistency. Ultimately, this invention demonstrates significant advantages in preference optimization of diffusion models, enabling the generated images to surpass existing technologies in multiple metrics such as aesthetic quality, semantic consistency, multi-attribute combination, and spatial relationships. It has strong applicability and can be directly used for preference alignment training in mainstream frameworks such as stable diffusion models. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart of a diffusion model step-by-step reward learning and optimization method provided in an embodiment of the present invention.
[0014] Figure 2 This is a schematic diagram of a diffusion model step-by-step reward learning and optimization system provided in an embodiment of the present invention.
[0015] Figure 3 This is a schematic diagram of a processing device provided in an embodiment of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0017] First, the following explanations are provided for the terms that may be used in this article: The terms "comprising," "including," "containing," "having," or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.) should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements that are not expressly listed and are well-known in the art.
[0018] The following provides a detailed description of a diffusion model stepwise reward learning and optimization method, system, device, and medium provided by this invention. Contents not described in detail in the embodiments of this invention are prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of this invention, they are performed according to conventional conditions in the art or conditions recommended by the manufacturer. Instruments used in the embodiments of this invention, unless otherwise specified by the manufacturer, are all commercially available conventional products.
[0019] Example 1 like Figure 1 The diagram shown is a flowchart of a step-by-step reward learning and optimization method for a diffusion model provided in an embodiment of the present invention, which mainly includes the following steps: Step 1: Construct the training set.
[0020] In this embodiment of the invention, a training set consisting of clean latent space sample pairs of preferences and non-preferences under a specified text condition is constructed. Specifically, the clean latent space sample pairs of preferences and non-preferences under text condition c can be defined as follows: , w refers to preference. Refers to non-preference.
[0021] Step 2: Forward or backward mapping based on the diffusion model.
[0022] In this embodiment of the invention, the clean latent space samples of preferences and non-preferences under specified text conditions are obtained through the forward mapping process of the diffusion model, and the corresponding preference and non-preference noise latent variables at each time step.
[0023] Specifically, for time step t, the corresponding latent variables of preference and non-preference noise are obtained by the following formula: ; in, Let be the latent noise variable at time step t. For clean latent space samples, when hour, For the preference noise latent variable at time step t, To favor clean latent space samples, when hour, For time step t, the unbiased noise latent variable, This refers to clean latent space samples that are not preferred. This refers to the forward noise addition process in the diffusion model. It is Gaussian noise. The mean is 0 and the variance is Gaussian distribution; These are predefined scalar coefficients that vary with time step t. This is the inversion mapping process of the diffusion model; both the forward noise addition and the inversion mapping belong to the forward mapping.
[0024] Step 3: Build and train the step-by-step reward model.
[0025] In this embodiment of the invention, a step-by-step reward model is used to predict the rewards corresponding to the preferred and non-preferred noise latent variables at each time step, and the difference between the rewards is constrained to construct a standard pairwise preference loss. The step-by-step reward model is also used to predict the rewards corresponding to the preferred and non-preferred clean latent space samples, and the difference between the rewards of the preferred noise latent variables and the preferred clean latent space samples at each time step, as well as the difference between the rewards of the non-preferred noise latent variables and the non-preferred clean latent space samples at each time step, are jointly constrained to construct a consistency loss. Within the training warm-up step range, the standard pairwise preference loss is used to train the step-by-step reward model, and then the standard pairwise preference loss and the consistency loss are combined to train the step-by-step reward model.
[0026] In this embodiment of the invention, the standard pairwise preference loss is expressed as: ; Where L is the standard pairwise preference loss. For mathematical expectation, It is a Sigmoid function. This indicates that time step t is uniformly sampled from the interval [0,T], where T is the maximum time step; This represents a clean latent space sample representing the preference under textual condition c. Clean latent space samples with non-preference The sample pairs are formed, where D is the training set and log is the logarithmic function. This is a step-by-step reward model used for reward prediction at time step t. For the parameters of the step-by-step reward model, for Based on the preference noise latent variable at time step t Predicted rewards for Based on the non-preference noise latent variable at time step t Predicted rewards.
[0027] In this embodiment of the invention, the consistency loss is expressed as: ; in, For consistency loss, For mathematical expectation, This indicates that time step t is uniformly sampled from the interval [0,T], where T is the maximum time step; This represents a clean latent space sample representing the preference under textual condition c. Clean latent space samples with non-preference The sample pairs are formed, and D is the training set; and Both belong to the step-by-step reward model, used to predict rewards for corresponding time steps t and 0. For the parameters of the step-by-step reward model, for Based on the preference noise latent variable at time step t Predicted rewards for Based on preference for clean dive space samples Predicted rewards for Based on the non-preference noise latent variable at time step t Predicted rewards for Based on non-preferred clean latent space samples Predicted rewards; It is the absolute value symbol.
[0028] Step 4: Image generation.
[0029] In this embodiment of the invention, the trained reward model is used to guide the preference alignment training process of the diffusion model, and the trained diffusion model outputs a denoised and preference-aligned image.
[0030] In this embodiment of the invention, a reward value is provided for the noise latent variable (e.g., the noise latent variable at time step t) at each time step of the back-mapping process of the diffusion model based on the trained stepwise reward model. At each step, multiple latent variables are sampled and the reward value is calculated. Then, sample pairs with a reward difference greater than a certain threshold are taken as positive and negative sample pairs. The parameters of the diffusion model are updated using a stepwise preference optimization algorithm (preference alignment training). Finally, the diffusion model with updated parameters is used to perform the image generation task.
[0031] The above-described solution provided in this invention is a step-by-step reward learning framework with preference invariance. It can be used to train a diffusion model for text-to-image generation with preference alignment, thereby improving the diffusion model's preference consistency, semantic understanding, and aesthetic quality in image generation. In practical deployments, this invention can be applied to AIGC (Artificial Intelligence Generated Content) creation platforms, image generation tools, content moderation systems, and other scenarios, making the generated results more in line with human preferences, improving visual quality, reducing illegal elements and sensitive content, and enhancing the semantic understanding and controllability of the diffusion model.
[0032] To more clearly demonstrate the technical solution and its effects provided by the present invention, the method provided by the embodiments of the present invention will be described in detail below with reference to specific examples.
[0033] 1. Analysis and explanation of the preference invariance of the noise latent space vector.
[0034] In preference optimization based on diffusion models, reward models Training is typically performed on preference pairs of the following form: ,in and Let represent the clean latent space samples (clean latent variables) of preference and non-preference under textual condition c, respectively. This is a preference symbol.
[0035] Those skilled in the art will understand that the latent space is the internal space of a generative model (e.g., a diffusion model), which is a compressed representation space of data. For a diffusion model, the reverse mapping process of the diffusion model is to iterate a noisy latent variable (a vector in the latent space) from a noisy state along the time axis in the latent space until the time step t=0 is reached, thereby obtaining the latent variable in the clear state (i.e., the clean latent variable), which can then be converted into a clean image.
[0036] To extend reward modeling to the intermediate steps of the diffusion process, a step-by-step reward model is defined. Used to estimate the latent noise variable at time step t The reward. Noise latent variable. From clean latent variables, forward mapping can be used For example, forward mapping can be achieved using forward noise addition or inversion mapping.
[0037] set up Indicates forward mapping; This represents a reverse mapping, for example, implemented using the DDIM (Denoising Diffusion Implicit Model) denoising operator. Due to obtaining... Only involves The process assumes that the reward at time step t can be obtained by... Inverse mapping to clean latent variables To conduct an assessment: , where the symbol The representation is defined as follows: Indicates to Reverse mapping is performed to obtain , In response to The reward.
[0038] Under the above definition, we can analyze whether the preference relations remain consistent when transitioning from clean latent variables to noisy latent variables. The following theorem provides both strict and approximate conditions for preference invariance.
[0039] (1) Strict preference invariance.
[0040] If the forward mapping process and the reverse mapping process are exact inverse operations of each other, that is... Then for any satisfying For sample pairs, the step-by-step reward predictions still strictly maintain the preference relationship: ;in, , The corresponding representation of the reward model against , The reward In response to The forward mapping, the output is , In response to The forward mapping, the output is .
[0041] This demonstrates that, in the case of complete reversibility, noisy latent variables completely retain the preference relationships of their corresponding clean latent variables.
[0042] (2) Approximate preference invariance.
[0043] In practice, and The mapping is often not completely invertible. Suppose the reconstruction error satisfies... ,in, To reconstruct the upper bound of the error, and the reward model For L R -Lipschitz continuity, the meaning of the above reward model formula is: reward model It is a clean latent variable To the set of real numbers The mapping function. Let This represents the reward gap between clean latent space samples. Then, when the following conditions are met... At that time, the preference relationship will not change.
[0044] In this embodiment of the invention, the reward model This can be understood as the distributed reward model at time step t=0 in this invention. ,For example, .
[0045] Based on this, the following two pieces of information can be determined:
[0046] (2.1) When the upper bound of the reconstruction error Reward gap between samples relative to clean latent space When the noise level is sufficiently low, the preference relationships between noisy latent variables are more easily preserved. Therefore, in Section 2, a plug-and-play approach to reduce reconstruction error is adopted to achieve more accurate preference-guided step-by-step reward learning.
[0047] (2.2) When the reward model The Lipschitz constant. The noise level is not too high; in other words, when the reward model is sufficiently smooth, the noisy latent variables can better maintain preference invariance. Therefore, in Sections 3 and 4, we will further analyze how to improve the smoothness of the step-by-step reward model.
[0048] 2. Plug-and-play step-by-step reward model training scheme.
[0049] The input to the step-by-step reward model consists of quadruples. The composition, this information, is obtained only through the forward mapping process. Therefore, reducing the upper bound of the reconstruction error... Essentially, this corresponds to mitigating the errors introduced during the process. Therefore, employing more accurate or reversible inversion methods is crucial for obtaining reliable noisy latent space image pairs.
[0050] To this end, this invention proposes a plug-and-play step-by-step reward model training framework, enabling seamless integration of the diffusion-based forward noise addition process or any existing inversion mapping algorithm into the training workflow, thereby achieving flexible and efficient step-by-step reward modeling enhancement. Its training objectives are as follows: ; .
[0051] Since the meanings of each symbol have already been explained one by one in the preceding text, they will not be repeated here.
[0052] 3. Probability analysis of preference reversal.
[0053] Preference reversal refers to the loss of the original preference relationship. The smaller the probability of preference reversal, the stronger the preference invariance, which forms the starting point for this section's discussion.
[0054] (1) Necessary conditions for preference reversal.
[0055] make If a preference reversal occurs, then: ;in, The reward model predicts the reward based on information x. For reward models According to information The reward For the forward mapping of information x, Regarding the reward bias of information x, .
[0056] The proof is as follows: To simplify the expression, an intermediate parameter is introduced. , , , , , , : , , , , The actual deviation is expressed as a sign perturbation: ,in Preference reversal occurs if and only if Substituting, we get: And because Therefore, if a flip occurs, then the following condition must be met. .
[0057] (2) Sufficient conditions for preference reversal.
[0058] like and Then when At that time, a reversal of preferences is inevitable.
[0059] Based on the analysis of the above two conditions, the following conclusion can be drawn: when the expected consistency error... As the value decreases, the probability of preference reversal also decreases; among which, the expected consistency error... For about Expected value ,when hour, Indicates about The expected value, when hour, Indicates about The expected value.
[0060] The proof of the above conclusion is as follows: When hour, Where Pr represents probability; according to Markov's inequality: we have , Established, and thus .
[0061] 4. Training objectives with consistent regularization terms.
[0062] Based on the analysis in Section 3 above, this invention designs training objectives to explicitly reduce preference flipping while maintaining the model's discriminative ability.
[0063] Specifically, a consistency regularization term (consistency loss) is introduced to penalize the deviation between the prediction of noisy latent variables and the reward of the corresponding clean latent variables: .
[0064] Since the meanings of each symbol have already been explained one by one in the preceding text, they will not be repeated here.
[0065] 5. Phased training strategy.
[0066] Considering that training using only a consistency term would cause the model to degenerate into a constant function—meaning that although the consistency error is zero, effective preference ranking is impossible—this invention employs a warm-up strategy to avoid this collapse: In the initial stage, a stable ranking benchmark is established using standard pairwise preference loss; after warming up, a consistency regularization term is added to obtain the final training objective. : ;
[0067] in, The regularization coefficient is... The number of warm-up steps. Standard pairwise preference loss L guarantees relative preference relationships and avoids trivial solutions; consistency loss. Improving the robustness of the model in the noisy latent space effectively reduces the probability of preference reversal, which is consistent with the conclusion in Section 3 above.
[0068] Since the specific training process can be implemented by referring to conventional techniques, it will not be elaborated upon.
[0069] In the embodiments of this invention described above, the invariance of the preference of the noise latent variable is theoretically guaranteed, so that the step-by-step reward model can maintain stable performance under different noise scales. Forward error is reduced through a plug-and-play inversion structure (i.e., the plug-and-play step-by-step reward model training scheme mentioned above), making the generation of noise latent variables more reliable. The smoothness of the step-by-step reward model is improved through consistency loss, significantly reducing preference flipping. Degradation of the step-by-step reward model is avoided through staged training, balancing discriminative ability and consistency. Ultimately, this invention demonstrates significant advantages in the preference optimization of the diffusion model, enabling the generated image to surpass existing technologies in multiple indicators such as aesthetic quality, semantic consistency, multi-attribute combination, and spatial relationships. It has strong applicability and can be directly used for preference alignment training in mainstream frameworks such as the Stable Diffusion model (SD1.5, SDXL). Here, SD1.5 (Stable Diffusion v1.5) refers to version 1.5 of the Stable Diffusion model, and SDXL (Stable Diffusion XL) refers to the XL version (i.e., the advanced upgrade version) of the Stable Diffusion model.
[0070] Through the above description of the embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software, or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.), including several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0071] Example 2 This invention also provides a diffusion model step-by-step reward learning and optimization system, which is mainly used to implement the methods provided in the foregoing embodiments, such as... Figure 2 As shown, the system mainly includes: Training set construction unit, used to construct a training set consisting of clean latent space sample pairs of preferences and non-preferences under specified text conditions; The diffusion model-based mapping unit is used to obtain clean latent space samples of preferences and non-preferences under specified text conditions through the forward mapping process of the diffusion model, and the corresponding preference and non-preference noise latent variables at each time step. The step-by-step reward model construction and training unit is used to construct and train the step-by-step reward model, including: predicting the rewards corresponding to the preferred and non-preferred noise latent variables at each time step using the step-by-step reward model, and constraining the difference between the rewards to construct a standard pairwise preference loss; predicting the rewards corresponding to the preferred and non-preferred clean latent space samples using the step-by-step reward model, and jointly constraining the difference between the reward of the preferred noise latent variable and the reward of the preferred clean latent space sample at each time step, as well as the difference between the reward of the non-preferred noise latent variable and the reward of the non-preferred clean latent space sample at each time step to construct a consistency loss; training the step-by-step reward model using the standard pairwise preference loss within the training warm-up step range, and then training the step-by-step reward model by combining the standard pairwise preference loss and the consistency loss. The image generation unit is used to guide the preference alignment training process of the diffusion model using the trained reward model, and outputs a denoised and preference-aligned image from the trained diffusion model.
[0072] In this embodiment of the invention, the process of obtaining clean latent space samples of preferences and non-preferences under specified text conditions through forward or backward mapping of the diffusion model includes the following preferences and non-preference noise latent variables at each time step: Define the clean latent space samples corresponding to preferences and non-preferences under textual condition c as follows: , w refers to preference. Refers to non-preference; For time step t, the corresponding latent variables of preference and non-preference noise are obtained by the following formula: ; in, Let be the latent noise variable at time step t. For clean latent space samples, when hour, For the preference noise latent variable at time step t, To favor clean latent space samples, when hour, For time step t, the unbiased noise latent variable, This refers to clean latent space samples that are not preferred. This refers to the forward noise addition process in the diffusion model. It is Gaussian noise. The mean is 0 and the variance is Gaussian distribution; These are predefined scalar coefficients that vary with time step t. This is the inversion mapping process of the diffusion model; both the forward noise addition and the inversion mapping belong to the forward mapping.
[0073] In this embodiment of the invention, the standard pairwise preference loss is expressed as: ; Where L is the standard pairwise preference loss. For mathematical expectation, It is a sigmoid function. This indicates that time step t is uniformly sampled from the interval [0,T], where T is the maximum time step; This represents a clean latent space sample representing the preference under textual condition c. Clean latent space samples with non-preference The sample pairs are formed, where D is the training set and log is the logarithmic function. This is a step-by-step reward model used for reward prediction at time step t. For the parameters of the step-by-step reward model, for Based on the preference noise latent variable at time step t Predicted rewards for Based on the non-preference noise latent variable at time step t Predicted rewards.
[0074] In this embodiment of the invention, the consistency loss is expressed as: ; in, For consistency loss, For mathematical expectation, This indicates that time step t is uniformly sampled from the interval [0,T], where T is the maximum time step; This represents a clean latent space sample representing the preference under textual condition c. Clean latent space samples with non-preference The sample pairs are formed, and D is the training set; and Both belong to the step-by-step reward model, and are used to predict rewards for time step t and time step 0, respectively. For the parameters of the step-by-step reward model, for Based on the preference noise latent variable at time step t Predicted rewards for Based on preference for clean dive space samples Predicted rewards for Based on the non-preference noise latent variable at time step t Predicted rewards for Based on non-preferred clean latent space samples Predicted rewards; It is the absolute value symbol.
[0075] Since the main technical details of each part of the above system have been described in detail in the previous embodiments, they will not be repeated here.
[0076] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.
[0077] Example 3 The present invention also provides a processing device, such as Figure 3 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiments.
[0078] Furthermore, the processing device also includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.
[0079] In this embodiment of the invention, the specific types of the memory, input device, and output device are not limited; for example: Input devices can be touchscreens, image acquisition devices, physical buttons, or mice, etc. The output device can be a display terminal; The memory can be random access memory (RAM) or non-volatile memory, such as disk storage.
[0080] Example 4 The present invention also provides a readable storage medium storing a computer program that, when executed by a processor, implements the method provided in the foregoing embodiments.
[0081] In this embodiment of the invention, the readable storage medium is a computer-readable storage medium and can be disposed in the aforementioned processing device, for example, as a memory in the processing device. Furthermore, the readable storage medium can also be any medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0082] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. The information disclosed in the background section is intended only to enhance the understanding of the overall background technology of the present invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art.
Claims
1. A step-by-step reward learning and optimization method for a diffusion model, characterized in that, include: Construct a training set consisting of clean latent space sample pairs of preferences and non-preferences under specified text conditions; By using the forward mapping process of the diffusion model, the preference and non-preference noise latent variables corresponding to the clean latent space samples of preference and non-preference under specified text conditions are obtained at each time step. Constructing and training a step-by-step reward model includes: constructing the step-by-step reward model Used to estimate the latent noise variable at time step t The reward is calculated as follows: The reward is predicted for each time step using a step-by-step reward model, specifically for the preferred and non-preferred noise latent variables. A standard pairwise preference loss is constructed by constraining the difference between these rewards. Similarly, the reward is predicted for each clean latent space sample using the same step-by-step reward model. A consistency loss is constructed by jointly constraining the difference between the reward for the preferred noise latent variable and the reward for the clean latent space sample, as well as the difference between the reward for the non-preferred noise latent variable and the reward for the clean latent space sample. Within the training warm-up step range, the step-by-step reward model is trained using the standard pairwise preference loss. Then, the step-by-step reward model is trained again by combining the standard pairwise preference loss and the consistency loss. The trained reward model guides the preference alignment training process of the diffusion model, and the trained diffusion model outputs a denoised and preference-aligned image.
2. The diffusion model step-by-step reward learning and optimization method according to claim 1, characterized in that, The forward mapping process using the diffusion model to obtain the preference and non-preference clean latent space samples at each time step under specified text conditions includes the following: Define the clean latent space samples corresponding to preferences and non-preferences under textual condition c as follows: , w refers to preference. Refers to non-preference; For time step t, the corresponding latent variables of preference and non-preference noise are obtained by the following formula: ; in, Let be the latent noise variable at time step t. For clean latent space samples, when hour, For the preference noise latent variable at time step t, To favor clean latent space samples, when hour, For time step t, the unbiased noise latent variable, This refers to clean latent space samples that are not preferred. This refers to the forward noise addition process in the diffusion model. It is Gaussian noise. The mean is 0 and the variance is Gaussian distribution; These are predefined scalar coefficients that vary with time step t. This is the inversion mapping process of the diffusion model; both the forward noise addition and the inversion mapping belong to the forward mapping.
3. The diffusion model step-by-step reward learning and optimization method according to claim 2, characterized in that, The standard pairwise preference loss is expressed as: ; Where L is the standard pairwise preference loss. For mathematical expectation, For the Sigmoid function, This indicates that time step t is uniformly sampled from the interval [0,T], where T is the maximum time step; This represents a clean latent space sample representing the preference under textual condition c. Clean latent space samples with non-preference The sample pairs are formed, where D is the training set and log is the logarithmic function. This is a step-by-step reward model used for reward prediction at time step t. For the parameters of the step-by-step reward model, for Based on the preference noise latent variable at time step t Predicted rewards for Based on the non-preference noise latent variable at time step t Predicted rewards.
4. The diffusion model step-by-step reward learning and optimization method according to claim 2, characterized in that, The consistency loss is expressed as: ; in, For consistency loss, For mathematical expectation, This indicates that time step t is uniformly sampled from the interval [0,T], where T is the maximum time step; This represents a clean latent space sample representing the preference under textual condition c. Clean latent space samples with non-preference The sample pairs are formed, and D is the training set; and Both belong to the step-by-step reward model, used to predict rewards for corresponding time steps t and 0. For the parameters of the step-by-step reward model, for Based on the preference noise latent variable at time step t Predicted rewards for Based on preference for clean dive space samples Predicted rewards for Based on the non-preference noise latent variable at time step t Predicted rewards for Based on non-preferred clean latent space samples Predicted rewards; It is the absolute value symbol.
5. A diffusion model step-by-step reward learning and optimization system, characterized in that, include: Training set construction unit, used to construct a training set consisting of clean latent space sample pairs of preferences and non-preferences under specified text conditions; The diffusion model-based mapping unit is used to obtain the preference and non-preference noise latent variables corresponding to the clean latent space samples of preferences and non-preferences under specified text conditions at each time step through the forward mapping process of the diffusion model. The step-by-step reward model construction and training unit is used to construct and train the step-by-step reward model, including: constructing the step-by-step reward model. Used to estimate the latent noise variable at time step t The reward is calculated as follows: The reward is predicted for each time step using a step-by-step reward model, specifically for the preferred and non-preferred noise latent variables. A standard pairwise preference loss is constructed by constraining the difference between these rewards. Similarly, the reward is predicted for each clean latent space sample using the same step-by-step reward model. A consistency loss is constructed by jointly constraining the difference between the reward for the preferred noise latent variable and the reward for the clean latent space sample, as well as the difference between the reward for the non-preferred noise latent variable and the reward for the clean latent space sample. Within the training warm-up step range, the step-by-step reward model is trained using the standard pairwise preference loss. Then, the step-by-step reward model is trained again by combining the standard pairwise preference loss and the consistency loss. The image generation unit is used to guide the preference alignment training process of the diffusion model using the trained reward model, and outputs a denoised and preference-aligned image from the trained diffusion model.
6. The diffusion model step-by-step reward learning and optimization system according to claim 5, characterized in that, The forward mapping process using the diffusion model to obtain the preference and non-preference clean latent space samples at each time step under specified text conditions includes the following: Define the clean latent space samples corresponding to preferences and non-preferences under textual condition c as follows: , w refers to preference. Refers to non-preference; For time step t, the corresponding latent variables of preference and non-preference noise are obtained by the following formula: ; in, Let be the latent noise variable at time step t. For clean latent space samples, when hour, For the preference noise latent variable at time step t, To favor clean latent space samples, when hour, For time step t, the unbiased noise latent variable, This refers to clean latent space samples that are not preferred. This refers to the forward noise addition process in the diffusion model. It is Gaussian noise. The mean is 0 and the variance is Gaussian distribution; These are predefined scalar coefficients that vary with time step t. This is the inversion mapping process of the diffusion model; both the forward noise addition and the inversion mapping belong to the forward mapping.
7. The diffusion model step-by-step reward learning and optimization system according to claim 5, characterized in that, The standard pairwise preference loss is expressed as: ; Where L is the standard pairwise preference loss. For mathematical expectation, For the Sigmoid function, This indicates that time step t is uniformly sampled from the interval [0,T], where T is the maximum time step; This represents a clean latent space sample representing the preference under textual condition c. Clean latent space samples with non-preference The sample pairs are formed, where D is the training set and log is the logarithmic function. This is a step-by-step reward model used for reward prediction at time step t. For the parameters of the step-by-step reward model, for Based on the preference noise latent variable at time step t Predicted rewards for Based on the non-preference noise latent variable at time step t Predicted rewards.
8. The diffusion model step-by-step reward learning and optimization system according to claim 5, characterized in that, The consistency loss is expressed as: ; in, For consistency loss, For mathematical expectation, This indicates that time step t is uniformly sampled from the interval [0,T], where T is the maximum time step; This represents a clean latent space sample representing the preference under textual condition c. Clean latent space samples with non-preference The sample pairs are formed, and D is the training set; and Both belong to the step-by-step reward model, and are used to predict rewards for time step t and time step 0, respectively. For the parameters of the step-by-step reward model, for Based on the preference noise latent variable at time step t Predicted rewards for Based on preference for clean dive space samples Predicted rewards for Based on the non-preference noise latent variable at time step t Predicted rewards for Based on non-preferred clean latent space samples Predicted rewards; It is the absolute value symbol.
9. A processing device, characterized in that, include: One or more processors; Memory, used to store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method as described in any one of claims 1 to 4.
10. A readable storage medium storing a computer program, characterized in that, When a computer program is executed by a processor, it implements the method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Diffusion model preference optimization method and system based on adaptive gradient adjustment
CN120597717A
Personalized design generation method, system and equipment based on user portrait, and medium
CN121999086A