Image generation model training method based on reinforcement learning

Through the multi-dimensional preference optimization and reinforcement learning training framework, the diffusion walk length and the denoising process are adjusted, which solves the problems of slow generation speed, high cost and limited diversity of diffusion model, and achieves high-quality and low-cost image generation.

CN120495805APending Publication Date: 2025-08-15GIANT MOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510573009.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing diffusion model is slow to generate, high computational cost, and the generation mode is prone to collapse, diversity is limited, and traditional loss functions are difficult to optimize image quality and stability.

Method used

A multi-dimensional preference optimization method and reinforcement learning training framework are adopted to adjust the diffusion walk length through multi-dimensional preference optimization, introduce dynamic gradient changes, optimize the denoising process, and combine the reference model to prevent deviation from the data distribution, and use multi-frame compatibility enhancement.

Benefits of technology

It improves the quality and diversity of image generation, reduces computational costs, enhances training stability, avoids pattern crashes, and produces more realistic and rich in details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495805A_ABST
    Figure CN120495805A_ABST
Patent Text Reader

Abstract

The invention relates to an image generation model training method based on reinforcement learning, and the method comprises the following steps: S1, employing a multi-dimensional preference optimization method to guarantee the balance optimization on all key dimensions, and S2, training by adopting a training framework to generate the model, and taking gradient change of multi-dimensional preference optimization as dynamic change. According to the method, the denoising process can be optimized, and the diffusion step length can be adjusted, so that the calculation cost is greatly reduced and the training stability is improved while high-quality generation of the model is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of model training technology, and in particular to an image generation model training method based on reinforcement learning. Background Art

[0002] The prior art has the following deficiencies:

[0003] (1) High computational cost and slow generation speed:

[0004] Traditional diffusion models (such as DDPM and DDIM) require multiple denoising sampling steps during the generation process, resulting in slow generation and high computational cost. Furthermore, the denoising step of existing diffusion models typically uses a fixed denoising network, failing to adaptively adjust the step size or optimize the denoising strategy based on the input, resulting in limited image quality and sampling efficiency.

[0005] (2) Generative model collapse and limited diversity:

[0006] During long training periods, the model may tend to favor certain patterns, resulting in a lack of diversity in generated images, a phenomenon known as mode collapse. Traditional diffusion model training relies on mean squared error (MSE) loss, but MSE is a relatively crude measure of image quality, making it difficult to directly optimize both perceptual quality and generation stability.

[0007] Therefore, it is necessary to provide an image generation model training method based on reinforcement learning to optimize the denoising process and adjust the diffusion step size so that the model can significantly reduce the computational cost and improve training stability while ensuring high-quality generation. Summary of the Invention

[0008] The purpose of the present invention is to provide an image generation model training method based on reinforcement learning to optimize the denoising process and adjust the diffusion step size so that the model can significantly reduce the computational cost and improve the training stability while ensuring high-quality generation.

[0009] In order to solve the problems existing in the prior art, the present invention provides an image generation model training method based on reinforcement learning, comprising the following steps:

[0010] S1: A multi-dimensional preference optimization method is used to ensure balanced optimization across all key dimensions. The multi-dimensional preference optimization is as follows:

[0011] S11: Set represents n preference dimensions;

[0012] S12: For two images x i and x j , whose preference factor vectors are and The dominant preference condition is defined as follows:

[0013]

[0014] in, Denotes a natural number k between 1 and n, and the overall formula is expressed as an image is considered a more preferable choice only if it is superior to another image in all relevant preference dimensions;

[0015] Modify the loss function of DPO as follows:

[0016]

[0017] Among them, DPO is Direct Preference Optimization, represents the expectation, y represents all samples, x w and x l represents the winning samples and losing samples in the data set, D is the data distribution of the training data, and λ k is the preference dimension d k The weight of σ() represents the sigmoid function to ensure smooth optimization, β k To control the optimization intensity of each preference dimension, π θ and π ref Represent the current strategy and reference strategy respectively;

[0018] S2: The model is generated by training using a training framework, and the gradient change of multi-dimensional preference optimization is used as a dynamic change. The formula for gradient change is as follows:

[0019]

[0020] Among them, L MDPO Represents the training loss value of MPO, is the gradient about the parameter θ, that is, the direction used to update the parameter during the optimization process, α represents a constant scaling factor, β represents a scaling factor, σ() represents a sigmoid function to ensure smooth optimization, and z represents the logarithmic ratio of the preference scores. represents the gradient of the log-ratio of preference scores with respect to the parameter θ.

[0021] Optionally, in the image generation model training method based on reinforcement learning,

[0022] The multi-dimensional preference optimization method is Multi-Aspect Dimensional Preference Optimization;

[0023] The dominant preference condition is the Dominant Preference Condition;

[0024] Multi-Aspect Dimensional Preference Optimization.

[0025] Optionally, in the reinforcement learning-based image generation model training method, the preference dimensions include text-image alignment, fidelity, and aesthetics.

[0026] Compared with the prior art, the present invention has the following advantages:

[0027] (1) This paper proposes an image generation model training method based on reinforcement learning (RL) to improve the image generation quality, generation efficiency and diversity.

[0028] (2) Compared with the existing diffusion model, the present invention introduces reinforcement learning to optimize the denoising process, reward signal design, and adjusts the diffusion step size, so that the model can significantly reduce the computational cost and improve the training stability while ensuring high-quality generation. This training method can effectively solve the problems of mode collapse and unstable generation quality in the existing generation model, and provide more expressive and detailed image generation results.

[0029] (3) The present invention's reinforcement learning (RL)-based image generation model training method optimizes multi-dimensional human preference mechanisms, and reinforcement learning can guide the diffusion model to generate clearer and more detailed images. This avoids the over-smoothing problem caused by traditional mean square error (MSE) optimization, making the generated results more realistic. In addition, the technology of the present invention can be efficiently trained on a variety of image generation models, demonstrating its wide applicability and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 A flowchart of the model training steps provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0031] The following is a more detailed description of the specific embodiments of the present invention with reference to schematic diagrams. The advantages and features of the present invention will become more apparent from the following description. It should be noted that the drawings are greatly simplified and not to exact scale, and are only used for the purpose of conveniently and clearly illustrating the embodiments of the present invention.

[0032] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.

[0033] The existing technologies have the following deficiencies: (1) High computational cost and slow generation speed: Traditional diffusion models (such as DDPM and DDIM) require multiple steps of denoising sampling during the generation process, resulting in slow generation speed and high computational cost. At the same time, the denoising step of the existing diffusion model usually adopts a fixed denoising network, and fails to adaptively adjust the step size or optimize the denoising strategy according to the input, resulting in limited image quality and sampling efficiency. (2) Generation mode collapse and limited diversity: During long-term training, the model may tend to certain specific modes, resulting in a lack of diversity in the generated images, which is the mode collapse phenomenon. Traditional diffusion model training relies on mean square error (MSE) loss, but MSE is relatively rough in measuring image quality, making it difficult to directly optimize perceptual quality and generation stability.

[0034] In order to solve the problems existing in the prior art, the present invention provides an image generation model training method based on reinforcement learning, such as Figure 1 Said method comprises the following steps:

[0035] In order to stabilize the training process, the present invention introduces a reference model to prevent the optimized model from deviating too far from the original data distribution.

[0036] Although Direct Preference Optimization (DPO) excels at aligning text-image models with human preferences, it can over-optimize on certain dimensions (such as aesthetics), compromising other key factors (such as security and fidelity). To alleviate these trade-offs, we propose a Multi-Aspect Dimensional Preference Optimization (MDPO) approach to ensure balanced optimization across all key dimensions.

[0037] Specifically, S1: adopts a multi-dimensional preference optimization method to ensure balanced optimization in all key dimensions, wherein the multi-dimensional preference optimization is as follows:

[0038] S11: Set Represents n preference dimensions, including text-image alignment, fidelity, and aesthetics;

[0039] S12: For two images x i and x j , whose preference factor vectors are and The dominant preference condition is defined as follows:

[0040]

[0041] in, Denotes a natural number k between 1 and n, and the overall formula is expressed as an image is considered a more preferable choice only if it is superior to another image in all relevant preference dimensions;

[0042] Modify the loss function of DPO as follows:

[0043]

[0044] Among them, DPO is Direct Preference Optimization, represents the expectation, y represents all samples, x w and x l represents the winning samples and losing samples in the data set, D is the data distribution of the training data, and λ k is the preference dimension d k The weight of σ() represents the sigmoid function to ensure smooth optimization, β k To control the optimization intensity of each preference dimension, π θ and π ref Represent the current strategy and reference strategy respectively;

[0045] S2: The model is generated by training using a training framework, and the gradient change of multi-dimensional preference optimization is used as a dynamic change. The formula for gradient change is as follows:

[0046]

[0047] Among them, L MDPO Represents the training loss value of MPO, is the gradient about the parameter θ, that is, the direction used to update the parameter during the optimization process, α represents a constant scaling factor that is independent of the derivative, β represents a scaling factor, σ() represents a sigmoid function that ensures smooth optimization, and z represents the logarithmic ratio of the preference scores. represents the gradient of the log-ratio of preference scores with respect to the parameter θ.

[0048] The present invention proposes a diffusion model training scheme combined with reinforcement learning, which introduces multi-dimensional preference optimization (MDPO), multi-frame compatibility enhancement, and dynamic step size adjustment to improve the efficiency and stability of training image generation models.

[0049] Compared with the prior art, the present invention has the following advantages:

[0050] (1) This paper proposes an image generation model training method based on reinforcement learning (RL) to improve the image generation quality, generation efficiency and diversity.

[0051] (2) Compared with the existing diffusion model, the present invention introduces reinforcement learning to optimize the denoising process, reward signal design, and adjusts the diffusion step size, so that the model can significantly reduce the computational cost and improve the training stability while ensuring high-quality generation. This training method can effectively solve the problems of mode collapse and unstable generation quality in the existing generation model, and provide more expressive and detailed image generation results.

[0052] (3) The present invention's reinforcement learning (RL)-based image generation model training method optimizes multi-dimensional human preference mechanisms, and reinforcement learning can guide the diffusion model to generate clearer and more detailed images. This avoids the over-smoothing problem caused by traditional mean square error (MSE) optimization, making the generated results more realistic. In addition, the technology of the present invention can be efficiently trained on a variety of image generation models, demonstrating its wide applicability and efficiency.

[0053] The above description is merely a preferred embodiment of the present invention and does not limit the present invention in any way. Any person skilled in the art who, without departing from the scope of the present invention, makes any equivalent substitution, modification, or other changes to the technical solution and technical content disclosed in the present invention shall be deemed to be within the scope of the present invention and still fall within the scope of protection of the present invention.

Claims

1. A method for training an image generation model based on reinforcement learning, characterized in that: The following steps are involved: S1: A multi-dimensional preference optimization method is used to ensure balanced optimization across all key dimensions. The multi-dimensional preference optimization is as follows: S11: Set represents n preference dimensions; S12: For two images x i and x j , whose preference factor vectors are and The dominant preference condition is defined as follows: in, Denotes a natural number k between 1 and n, and the overall formula is expressed as an image is considered a more preferable choice only if it is superior to another image in all relevant preference dimensions; Modify the loss function of DPO as follows: Among them, DPO is Direct Preference Optimization, represents the expectation, y represents all samples, x w and x l represents the winning samples and losing samples in the data set, D is the data distribution of the training data, and λ k is the preference dimension d k The weight of σ() represents the sigmoid function to ensure smooth optimization, β k To control the optimization intensity of each preference dimension, π θ and π ref Represent the current strategy and reference strategy respectively; S2: The model is generated by training using a training framework, and the gradient change of multi-dimensional preference optimization is used as a dynamic change. The formula for gradient change is as follows: Among them, L MDPO Represents the training loss value of MPO, is the gradient about the parameter θ, that is, the direction used to update the parameter during the optimization process, α represents a constant scaling factor, β represents a scaling factor, σ() represents a sigmoid function to ensure smooth optimization, and z represents the logarithmic ratio of the preference scores. represents the gradient of the log-ratio of preference scores with respect to the parameter θ.

2. The image generation model training method based on reinforcement learning according to claim 1, characterized in that: The multi-dimensional preference optimization method is Multi-Aspect Dimensional Preference Optimization; The dominant preference condition is the Dominant Preference Condition; Multi-Aspect Dimensional Preference Optimization.

3. The image generation model training method based on reinforcement learning according to claim 1, characterized in that: Preference dimensions include text-image alignment, fidelity, and aesthetics.