Image generation method and apparatus, device, storage medium and program product

By adjusting the parameters of the image generation model through the target reward model, the fine-grained optimization problem of the Wensheng graph diffusion model when generating images is solved, achieving better image generation effects and diversity, and being suitable for a variety of application scenarios.

WO2025200899A1PCT designated stage Publication Date: 2025-10-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/078809
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-27
Filing Date
2025-02-24
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing cultural image diffusion models have difficulty achieving fine-grained dimensionality optimization when generating images, resulting in poor image generation and difficulty in adapting to different application scenarios.

Method used

The target reward model is used to adjust the parameters of the preset image generation model. The target reward model predicts the feature information of the image to be optimized, and optimizes the image generation model at multiple fine-grained levels, including adjustments in dimensions such as style consistency, content consistency, color, texture, atmosphere and layout.

Benefits of technology

The image generation model has improved its image generation effect in different dimensions, can be applied to a variety of application scenarios, and has improved the granularity and diversity of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025078809_02102025_PF_FP_ABST
    Figure CN2025078809_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of image generation. Disclosed are an image generation method and apparatus, a device, a storage medium and a program product. The method comprises: acquiring target prompt information; and, on the basis of the target prompt information and a target image generation model, generating a target image for the target prompt information, the target image generation model being obtained by adjusting parameters of a preset image generation model on the basis of a target reward model, inputs of the target reward model comprising first sample prompt information and an image to be optimized obtained by the preset image generation model on the basis of the first sample prompt information, and the target reward model being used for predicting image feature information to be optimized of the image to be optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Image generation method, device, equipment, storage medium and program product

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application No. 202410362124.8, filed on March 27, 2024, entitled “Image Generation Method, Device, Equipment, Storage Medium and Program Product”. The entire contents of that application are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to the field of image generation technology, and in particular to an image generation method, apparatus, device, storage medium, and program product. Background Art

[0004] Currently, the primary approach to generating diverse images given a given prompt text is through the use of the text-based diffusion model. In related art, feedback learning for the text-based diffusion model often involves pre-training a global reward model, which is then used to fine-tune the model to optimize its image generation performance. Summary of the Invention

[0005] In view of this, the present disclosure provides an image generation method, apparatus, device, storage medium, and program product to solve the problem of poor image generation effect.

[0006] In a first aspect, the present disclosure provides an image generation method, the method comprising:

[0007] Get target prompt information;

[0008] Based on the target prompt information and the target image generation model, a target image of the target prompt information is generated. The target image generation model is obtained by adjusting the parameters of the preset image generation model based on the target reward model. The input of the target reward model includes first sample prompt information and the image to be optimized obtained by the preset image generation model based on the first sample prompt information. The target reward model is used to predict the image feature information to be optimized in the image to be optimized.

[0009] In a second aspect, the present disclosure provides an image generating device, the device comprising:

[0010] Prompt information acquisition module, used to obtain target prompt information;

[0011] A target image generation module is used to generate a target image of the target prompt information based on the target prompt information and a target image generation model. The target image generation model is obtained by adjusting the parameters of a preset image generation model based on a target reward model. The input of the target reward model includes first sample prompt information and an image to be optimized obtained by the preset image generation model based on the first sample prompt information. The target reward model is used to predict the image feature information to be optimized in the image to be optimized.

[0012] In a third aspect, the present disclosure provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the image generation method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0013] In a fourth aspect, the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the image generation method of the first aspect or any corresponding embodiment thereof.

[0014] In a fifth aspect, the present disclosure provides a computer program product, comprising computer instructions for causing a computer to execute the image generation method of the first aspect or any corresponding embodiment thereof.

[0015] The image generation method provided in the disclosed embodiments uses first sample prompt information and an image to be optimized, obtained by a preset image generation model based on the first sample prompt information, as input to a target reward model. The target reward model is then used to predict image feature information to be optimized in the image to be optimized, thereby determining the dimension to be optimized by the preset image generation model when generating the image. The target reward model is then used to adjust the parameters of the preset image generation model so that the preset image generation model optimizes toward the dimension to be optimized. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] FIG1 is a schematic flow chart of an image generation method according to an embodiment of the present disclosure;

[0018] FIG2 is a flow chart of a method for determining a target image generation model according to an embodiment of the present disclosure;

[0019] FIG3 is a flow chart of another method for determining a target image generation model according to an embodiment of the present disclosure;

[0020] FIG4 is a flow chart of a method for determining a target reward model according to an embodiment of the present disclosure;

[0021] FIG5 is a flow chart of another method for determining a target reward model according to an embodiment of the present disclosure;

[0022] FIG6 is a structural block diagram of an image generating apparatus according to an embodiment of the present disclosure;

[0023] FIG7 is a structural block diagram of a computer device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0024] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present disclosure.

[0025] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0026] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0027] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0028] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0029] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0030] Currently, the main method for generating diverse images given a given prompt text is through the use of the VEG diffusion model. In related technologies, feedback learning for the VEG diffusion model often involves pre-training a global reward model, which is then used to fine-tune the VEG diffusion model to optimize its image generation performance. However, in this fine-tuning approach, the global reward model provides a global reward to the VEG diffusion model, typically fine-tuning the VEG diffusion model to optimize its diffusion along a specific dimension, such as the overall aesthetics of the image. This makes it difficult to achieve fine-grained dimensional optimization, making the VEG diffusion model difficult to adapt to different application scenarios and resulting in poor image generation performance.

[0031] In view of this, according to an embodiment of the present disclosure, an image generation method embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0032] In this embodiment, an image generation method is provided, which can be used in the above-mentioned computer devices, such as mobile phones, tablet computers, etc. FIG1 is a flow chart of an image generation method according to an embodiment of the present disclosure. As shown in FIG1 , the flow chart includes the following steps:

[0033] Step S101: Obtain target prompt information.

[0034] It should be noted that the target prompt information may be a prompt text or a prompt image, and the type of the target prompt information is not limited here.

[0035] Step S102: Generate a target image of the target prompt information based on the target prompt information and the target image generation model. The target image generation model is obtained by adjusting the parameters of the preset image generation model based on the target reward model. The input of the target reward model includes the first sample prompt information and the image to be optimized obtained by the preset image generation model based on the first sample prompt information. The target reward model is used to predict the image feature information to be optimized in the image to be optimized.

[0036] It should be noted that the target image generation model can employ a large pre-trained model such as a text-to-image generative adversarial network model, a diffusion model, or a large language model. The image feature information corresponding to the target image generation model includes at least one of style consistency, content consistency, color, texture, atmosphere, and layout. Furthermore, the image feature information may include other feature information such as outlines or extension relationships.

[0037] The image generation method provided in this embodiment uses the first sample prompt information and the image to be optimized obtained by the preset image generation model based on the first sample prompt information as the input of the target reward model, and predicts the image feature information to be optimized in the image to be optimized based on the target reward model to determine the dimension to be optimized when the preset image generation model generates the image. Then, the target reward model is used to adjust the parameters of the preset image generation model so that the preset image generation model is optimized towards the dimension to be optimized. Therefore, it is possible to improve the image generation effect of the preset image generation model in the dimension to be optimized, so as to improve the image generation effect of the target image generation model in different dimensions, thereby being suitable for different application scenarios.

[0038] In some optional embodiments, the target reward model includes at least two target reward sub-models, wherein the target reward sub-model includes multiple local dimension nodes to be optimized with a hierarchical affiliation, the target reward sub-model corresponds to the image features of the global dimension to be optimized, the global dimension to be optimized includes at least one local dimension to be optimized, and the leaf nodes in the local dimension nodes to be optimized are used to represent the image features of the local dimension to be optimized.

[0039] Specifically, the target reward model of the present invention can be regarded as a reward model with a tree structure of multiple levels, and the input end of the target reward model can be regarded as a classifier for predicting the classification probability of the current input data corresponding to each target reward sub-model from the global dimension. The local dimension nodes to be optimized at the first level of the target reward sub-model correspond to the image features of the global dimension to be optimized. For example, assuming that at least two global dimensions to be optimized include the image-text consistency dimension and the aesthetic dimension. Then, the target reward model includes a target reward sub-model for the image-text consistency dimension and a target reward sub-model for the aesthetic dimension. The local dimension nodes to be optimized at the first level of the target reward sub-model for the image-text consistency dimension correspond to the image features of the image-text consistency dimension, and the local dimension nodes to be optimized at the first level of the target reward sub-model for the aesthetic dimension correspond to the image features of the aesthetic dimension. The local dimension nodes to be optimized at the second level of the target reward sub-model correspond to the image features of the local dimension to be optimized. For example, assuming that the local dimensions to be optimized under the image-text consistency dimension include the style consistency dimension and the content consistency dimension, and the local dimensions to be optimized under the aesthetic dimension include the color dimension, the texture dimension, the atmosphere dimension, and the layout dimension. Then, the target reward sub-model of the image-text consistency dimension has two local dimension nodes to be optimized at the second level, which correspond to the image features of the style consistency dimension and the image features of the content consistency dimension respectively. The target reward sub-model of the aesthetic dimension has four local dimension nodes to be optimized at the second level, which correspond to the image features of the color dimension, the image features of the texture dimension, the image features of the atmosphere dimension and the image features of the layout dimension respectively. In addition, according to actual conditions, the local dimension nodes to be optimized at the third, fourth or other levels can be added under the global dimension nodes to be optimized at the second level, and the local dimension nodes to be optimized at the last level can be used as leaf nodes. For example, if the reward model has only two levels, the local dimension nodes to be optimized corresponding to the above-mentioned style consistency dimension and content consistency dimension can be regarded as leaf nodes of the target reward sub-model of the image-text consistency dimension, and the local dimension nodes to be optimized corresponding to the above-mentioned color dimension, texture dimension, atmosphere dimension and layout dimension can be regarded as leaf nodes of the target reward sub-model of the aesthetic dimension.

[0040] The image generation method provided in this embodiment includes a target reward model that includes at least two target reward sub-models, each of which includes multiple local dimension nodes to be optimized with a hierarchical relationship. The target reward sub-model corresponds to the image features of the global dimension to be optimized, the global dimension to be optimized includes at least one local dimension to be optimized, and the leaf nodes in each local dimension node to be optimized are used to represent the image features of the local dimension to be optimized. Therefore, the target reward model can be used to evaluate images generated by a preset image generation model from multiple fine-grained hierarchical dimensions to accurately determine the dimensions that the preset image generation model lacks when generating images.

[0041] In some optional implementations, as shown in FIG2 , a method for determining the target image generation model includes:

[0042] Step S201: Obtain first sample prompt information.

[0043] Step S202: input the first sample prompt information into a preset image generation model to obtain an image to be optimized.

[0044] In step S203 , the image to be optimized and the first sample prompt information are input into the target reward model to obtain image loss, where the image loss is used to characterize the image feature information to be optimized in the image to be optimized.

[0045] Step S204: updating the parameters of the preset image generation model based on the image loss to obtain a target image generation model.

[0046] The image generation method provided in this embodiment inputs first sample prompt information and the image to be optimized, obtained by a preset image generation model based on the first sample prompt information, into a target reward model. Thus, the target reward model can be used to determine the image feature information to be optimized, i.e., the image loss, in the image currently generated by the preset image generation model. The image loss is then used to update the parameters of the preset image generation model, thereby optimizing the image generation performance of the preset image generation model in terms of the image feature information to be optimized.

[0047] In some optional embodiments, the preset image generation model is constructed based on a diffusion model, and the step S202 of inputting the first sample prompt information into the preset image generation model to obtain the image to be optimized includes: inputting the first sample prompt information into the preset image generation model for noise addition and denoising; extracting the denoised image obtained within the preset denoising step range to obtain the image to be optimized.

[0048] Optionally, the preset denoising step range is the last 5 to 10 steps of the denoising process of the preset image generation model. It should be noted that the preset denoising step range should be selected so that the selected denoised image includes a relatively large amount of non-noise image information. In actual operation, the preset denoising step range can be adjusted based on implementation circumstances.

[0049] The image generation method provided in this embodiment inputs first sample prompt information into a preset image generation model constructed based on a diffusion model for noise addition and denoising. The denoised image obtained within a preset denoising step range is then extracted as the image to be optimized for the target reward model. Because the denoised image obtained within the preset denoising step range includes a significant amount of information that can reflect the image generation effect of the preset image generation model, using the denoised image within the preset denoising step range as the image to be optimized for the target reward model enables the target reward model to accurately determine the dimensions missing from the image currently generated by the preset image generation model, further improving the fine-tuning accuracy of the preset image generation model.

[0050] In some optional implementations, the step S203 of inputting the image to be optimized and the first sample prompt information into the target reward model to obtain the image loss includes:

[0051] In step a1, the image to be optimized and the first sample prompt information are input into the target reward model, and at least two target reward sub-models are used to predict the global dimension to be optimized to obtain a first prediction probability corresponding to the target reward sub-model.

[0052] It can be understood that the target reward model and its various nodes can be regarded as a classifier. The target reward sub-model of the global dimension to be optimized can determine the image generation effect of the image to be optimized in each global dimension to be optimized based on the image to be optimized and the first sample prompt information, so as to judge in which global dimension to be optimized the image to be optimized is deficient, so as to obtain the first prediction probability corresponding to the target reward sub-model.

[0053] For example, let's assume that the worse the image generation performance of the image to be optimized in a particular global dimension, the higher the corresponding first prediction probability. As shown in Figure 3, if the first preset probability of the target reward sub-model predicting the image to be optimized in the image-text consistency dimension is 0.2, and the first preset probability of the target reward sub-model predicting the image to be optimized in the aesthetic dimension is 0.8, this indicates that the image generation performance of the image to be optimized in the aesthetic dimension is poor. Therefore, in the subsequent process, it is necessary to fine-tune the preset image generation model based on the aesthetic dimension.

[0054] Step a2: Based on the prediction results of the global dimension to be optimized, the second prediction probability and reward value corresponding to the leaf node are predicted respectively.

[0055] Specifically, after obtaining the prediction results of the global dimension to be optimized, the local dimension node to be optimized under the target reward sub-model will further determine the image generation effect of the predicted image to be optimized in each local dimension to be optimized, so as to judge which local dimension to be optimized the predicted image to be optimized has deficiencies in, so as to obtain the second prediction probability and reward value corresponding to the leaf node.

[0056] For example, it is assumed that the worse the image generation effect of the image to be optimized in a certain local dimension to be optimized, the smaller the reward value of the corresponding leaf node. As shown in Figure 3, if the reward value output by the local node to be optimized in the style consistency dimension is 0.6, the reward value output by the local node to be optimized in the content consistency dimension is 2.0, the reward value output by the leaf node in the color dimension is 1.4, the reward value output by the leaf node in the texture dimension is 0.9, the reward value output by the leaf node in the atmosphere dimension is -2.1, and the reward value output by the leaf node in the layout dimension is -2.7, then it indicates that the image generation effect of the image to be optimized in the atmosphere dimension and the layout dimension is not good. Therefore, in the bar chart shown in Figure 3, the bars corresponding to the leaf nodes in the atmosphere dimension and the layout dimension are relatively high, indicating that in the subsequent fine-tuning process of the image generation model, it is necessary to focus on the image generation effect of the model in the atmosphere dimension and the layout dimension.

[0057] In step a3, for each target reward sub-model, a first loss is obtained based on the fusion of the second predicted probability and the reward value.

[0058] Specifically, the second predicted probability can be used as the first weight of the corresponding reward value, and all reward values ​​of the same target reward sub-model are weightedly calculated based on the first weight to obtain the first loss of each target reward sub-model.

[0059] In step a4, the first losses corresponding to all target reward sub-models are fused based on the first predicted probability to obtain the image loss.

[0060] Specifically, the first prediction probability is used as the second weight of the corresponding first loss, and all first losses are weightedly calculated based on the second weight to obtain the image loss.

[0061] The image generation method provided in this embodiment utilizes target reward sub-models for at least two global dimensions to be optimized to determine a first predicted probability for the image to be optimized corresponding to each target reward sub-model. Based on the prediction results for the global dimensions to be optimized, predictions are then made for the local dimensions to be optimized to determine a second predicted probability and reward value for the image to be optimized corresponding to each local dimension to be optimized. Then, based on the first and second predicted probabilities and reward values, an image loss is determined. This allows a preset image generation model to determine the current optimization direction based on the image loss, thereby improving the image generation effect through fine-tuning.

[0062] In some optional implementations, as shown in FIG4 , the target reward model may be determined by:

[0063] Step S301: Obtain second sample prompt information, an image pair corresponding to the second sample prompt information, and a dimension label to be optimized, where the image pair includes a positive sample image and a negative sample image corresponding to the dimension label to be optimized.

[0064] Step S302: Determine the leaf node corresponding to the dimension label to be optimized in the preset reward model to determine the reward sub-model to which it belongs.

[0065] Step S303: Using the corresponding reward sub-model, predict the dimensions to be optimized for the second sample prompt information and the image pair. The dimensions to be optimized include the global dimensions to be optimized and the local dimensions to be optimized.

[0066] Step S304: Determine the second loss based on the predicted result and the dimension label to be optimized.

[0067] Step S305: Update the parameters of the corresponding reward sub-model based on the second loss to determine the target reward model.

[0068] The image generation method provided in this embodiment determines the leaf node corresponding to the label of the dimension to be optimized in the preset reward model to determine the corresponding reward sub-model. Then, the corresponding reward sub-model is used to predict the dimension to be optimized for the second sample prompt information and the image pair, so as to measure the difference between the output of the current reward model and the expected output based on the predicted result and the label of the dimension to be optimized, and obtain the second loss. In this way, the parameters of the corresponding reward sub-model can be updated based on the second loss, so that the updated reward sub-model can accurately judge the image generation effect of the input image in the corresponding dimension to be optimized, thereby improving the accuracy of the target reward model in judging the image generation effect of the input image in multiple dimensions.

[0069] Exemplarily, as shown in FIG5 , preference data of multiple fine-grained dimensions can be collected in advance to obtain training samples of multiple dimensions. Each training sample includes the second sample prompt information, the image pair corresponding to the second sample prompt information, and the dimension label to be optimized. Assuming that the dimension label to be optimized in the current training sample is style consistency, it can be determined that the reward sub-model to which it belongs is the reward sub-model of the image-text consistency dimension. The local dimension nodes to be optimized corresponding to the image-text consistency dimension and the style consistency dimension in the reward sub-model of the image-text consistency dimension can be used to predict the global dimension to be optimized and the local dimension to be optimized for the second sample prompt information and the image pair to obtain the predicted result.

[0070] In some optional embodiments, the above-mentioned step S303 uses the reward sub-model to predict the dimension to be optimized for the second sample prompt information and the image pair, including: determining the first classification probability corresponding to the global dimension to be optimized in the reward sub-model based on the second sample prompt information and the image pair; determining the second classification probability corresponding to the leaf node in the reward sub-model based on the second sample prompt information and the image pair, and the predicted result includes the first classification probability and the second classification probability.

[0071] It should be noted that the reward sub-model can be regarded as a classifier. When the second sample prompt information and the image pair are input into the reward sub-model, the reward sub-model will classify the input data to obtain the first classification probability of the input data corresponding to the global dimension to be optimized and the second classification probability corresponding to the leaf node.

[0072] In some optional embodiments, determining the second loss based on the predicted results and the dimension labels to be optimized in the above-mentioned step S304 includes: determining the reward value based on the second classification probability and the dimension labels to be optimized; and determining the second loss based on the reward value and the fusion result of the first classification probability.

[0073] It should be noted that in order for the corresponding reward sub-model to accurately determine the image generation effect of the input image in the dimension to be optimized, it is necessary to use the second classification probability and the label of the dimension to be optimized to determine the reward value with the goal of increasing the difference between the positive sample image and the negative sample image, so that the reward value can reflect the difference between the positive sample image and the negative sample image corresponding to the dimension to be optimized. Then, based on the fusion result of the reward value and the first classification probability, a second loss is determined, and the parameters of the preset reward model are updated using the second loss. This ensures that the reward value output by the leaf node of the target reward model after training can reflect the image generation effect of the input image in the corresponding dimension.

[0074] It is worth noting that the image generation method disclosed herein constructs a tree-structured reward model according to multiple fine-grained preference dimensions, and trains the reward model using training samples of different preference dimensions to obtain a target reward model that can be used to judge the image generation effect of an input image in different preference dimensions. Therefore, when the image generation model undergoes feedback learning, the target reward model can be used to score the image generation effect of the image generation model in multiple preference dimensions, so as to adaptively predict the preference dimensions that the image generation model lacks when generating images, and then optimize the image generation model towards the missing preference dimensions to improve the image generation effect of the image generation model in multiple preference dimensions.

[0075] This embodiment also provides an image generation device for implementing the above-mentioned embodiments and preferred implementations. Details already described will not be repeated. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0076] This embodiment provides an image generation device, as shown in FIG6 , including:

[0077] Prompt information acquisition module 401, used to obtain target prompt information;

[0078] The target image generation module 402 is used to generate a target image of the target prompt information based on the target prompt information and the target image generation model. The target image generation model is obtained by adjusting the parameters of the preset image generation model based on the target reward model. The input of the target reward model includes the first sample prompt information and the image to be optimized obtained by the preset image generation model based on the first sample prompt information. The target reward model is used to predict the image feature information to be optimized in the image to be optimized.

[0079] In some optional embodiments, the target reward model in the target image generation module 402 includes at least two target reward sub-models, wherein the target reward sub-model includes multiple local dimension nodes to be optimized with a hierarchical affiliation, the target reward sub-model corresponds to the image features of the global dimension to be optimized, the global dimension to be optimized includes at least one local dimension to be optimized, and the leaf nodes in the local dimension nodes to be optimized are used to represent the image features of the local dimension to be optimized.

[0080] In some optional embodiments, the image generation device further includes an image model determination module for determining a target image generation model. The image model determination module includes:

[0081] A first information acquisition module, used to acquire first sample prompt information;

[0082] A first image acquisition module, configured to input the first sample prompt information into a preset image generation model to obtain an image to be optimized;

[0083] An image loss calculation module is used to input the image to be optimized and the first sample prompt information into the target reward model to obtain image loss, where the image loss is used to characterize the image feature information to be optimized in the image to be optimized;

[0084] The image model updating module is used to update the parameters of the preset image generation model based on the image loss to obtain the target image generation model.

[0085] In some optional embodiments, the preset image generation model is constructed based on a diffusion model, and the first image acquisition module includes:

[0086] An image processing unit, configured to input the first sample prompt information into a preset image generation model for noise addition and denoising;

[0087] The image extraction unit is used to extract the denoised image obtained within a preset denoising step number range to obtain the image to be optimized.

[0088] In some optional implementations, the image loss calculation module includes:

[0089] a global dimension prediction unit, configured to input the image to be optimized and the first sample prompt information into the target reward model, and use at least two target reward sub-models to predict the global dimension to be optimized, thereby obtaining a first prediction probability corresponding to the target reward sub-model;

[0090] A local dimension prediction unit, configured to predict the second prediction probability and reward value corresponding to each leaf node based on the prediction result of the global dimension to be optimized;

[0091] A first loss calculation unit is configured to obtain, for each target reward sub-model, a first loss based on a fusion of the second predicted probability and the reward value;

[0092] The model loss fusion unit is used to fuse the first losses corresponding to all target reward sub-models based on the first prediction probability to obtain the image loss.

[0093] In some optional embodiments, the image generation device further includes a reward model determination module for determining a target reward model. The reward model determination module includes:

[0094] A second information acquisition module is configured to acquire second sample prompt information, an image pair corresponding to the second sample prompt information, and a dimension label to be optimized, where the image pair includes a positive sample image and a negative sample image corresponding to the dimension label to be optimized;

[0095] The optimization path determination module is used to determine the leaf node corresponding to the dimension label to be optimized in the preset reward model to determine the reward sub-model to which it belongs;

[0096] An optimization dimension prediction module, configured to use the corresponding reward sub-model to predict the dimensions to be optimized for the second sample prompt information and the image pair, where the dimensions to be optimized include the global dimensions to be optimized and the local dimensions to be optimized;

[0097] A prediction loss calculation module is used to determine the second loss based on the prediction result and the dimension label to be optimized;

[0098] The reward model updating module is used to update the parameters of the corresponding reward sub-model based on the second loss to determine the target reward model.

[0099] In some optional implementations, the optimization dimension prediction module includes:

[0100] A first classification unit is configured to determine a first classification probability corresponding to a global dimension to be optimized in the corresponding reward sub-model based on the second sample prompt information and the image pair;

[0101] The second classification unit is used to determine the second classification probability corresponding to the leaf node in the reward sub-model based on the second sample prompt information and the image pair. The predicted result includes the first classification probability and the second classification probability.

[0102] In some optional implementations, the predicted loss calculation module includes:

[0103] A reward value determining unit, configured to determine a reward value based on the second classification probability and the dimension label to be optimized;

[0104] The second loss calculation unit is used to determine a second loss based on the reward value and the fusion result of the first classification probability.

[0105] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0106] The image generating device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0107] The embodiment of the present disclosure further provides a computer device having the image generating device shown in FIG6 .

[0108] Please refer to Figure 7, which is a block diagram of a computer device provided by an optional embodiment of the present disclosure. As shown in Figure 7, the computer device includes: one or more processors 501, a memory 502, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 7 takes a processor 501 as an example.

[0109] Processor 501 may be a central processing unit, a network processor, or a combination thereof. Processor 501 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.

[0110] The memory 502 stores instructions that can be executed by at least one processor 501, so as to enable the at least one processor 501 to execute the method shown in the above embodiment.

[0111] The memory 502 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 502 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 502 may optionally include a memory remotely located relative to the processor 501, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0112] The memory 502 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid state drive; the memory 502 may also include a combination of the above types of memory.

[0113] The computer device further includes an input device 503 and an output device 504. The processor 501, the memory 502, the input device 503 and the output device 504 may be connected via a bus or other means, and FIG7 shows a bus connection as an example.

[0114] The input device 503 can receive input digital or character information and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touch pad, an indicator stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 504 can include a display device, an auxiliary lighting device (e.g., an LED), and a tactile feedback device (e.g., a vibration motor). The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display, and a plasma display. In some optional embodiments, the display device can be a touch screen.

[0115] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.

[0116] A portion of the present disclosure may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present disclosure through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes but is not limited to a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium that can be accessed by the computer.

[0117] Although the embodiments of the present disclosure have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for generating an image, comprising: Get target prompt information; Based on the target prompt information and the target image generation model, a target image of the target prompt information is generated. The target image generation model is obtained by adjusting the parameters of the preset image generation model based on the target reward model. The input of the target reward model includes first sample prompt information and the image to be optimized obtained by the preset image generation model based on the first sample prompt information. The target reward model is used to predict the image feature information to be optimized in the image to be optimized.

2. The image generation method according to claim 1, wherein the target reward model comprises at least two target reward sub-models, wherein: The target reward sub-model includes multiple local dimension nodes to be optimized with a hierarchical affiliation. The target reward sub-model corresponds to the image features of the global dimension to be optimized. The global dimension to be optimized includes at least one local dimension to be optimized. The leaf nodes in the local dimension nodes to be optimized are used to represent the image features of the local dimension to be optimized.

3. The image generation method according to claim 2, wherein the target image generation model is determined by: Obtaining the first sample prompt information; Inputting the first sample prompt information into the preset image generation model to obtain the image to be optimized; Inputting the image to be optimized and the first sample prompt information into the target reward model to obtain an image loss, where the image loss is used to characterize image feature information to be optimized in the image to be optimized; The parameters of the preset image generation model are updated based on the image loss to obtain the target image generation model.

4. The image generation method according to claim 3, wherein the preset image generation model is constructed based on a diffusion model, and inputting the first sample prompt information into the preset image generation model to obtain the image to be optimized comprises: Inputting the first sample prompt information into the preset image generation model for noise addition and denoising; The denoised image obtained within a preset denoising step number range is extracted to obtain the image to be optimized.

5. The image generation method according to claim 4, wherein inputting the image to be optimized and the first sample prompt information into the target reward model to obtain the image loss comprises: Inputting the image to be optimized and the first sample prompt information into the target reward model, and using the at least two target reward sub-models to predict the global dimension to be optimized, to obtain a first prediction probability corresponding to the target reward sub-model; Based on the prediction results of the global dimension to be optimized, respectively predicting the second prediction probability and the reward value corresponding to the leaf node; For each target reward sub-model, obtaining a first loss based on a fusion of the second predicted probability and the reward value; The first losses corresponding to all the target reward sub-models are fused based on the first predicted probability to obtain the image loss.

6. The image generation method according to claim 2, wherein the target reward model is determined by: Obtaining second sample prompt information, an image pair corresponding to the second sample prompt information, and a dimension label to be optimized, where the image pair includes a positive sample image and a negative sample image corresponding to the dimension label to be optimized; Determine the leaf node corresponding to the dimension label to be optimized in the preset reward model to determine the reward sub-model to which it belongs; Using the reward sub-model, predict the dimensions to be optimized for the second sample prompt information and the image pair, where the dimensions to be optimized include the global dimension to be optimized and the local dimension to be optimized; Determining a second loss based on the predicted result and the dimension label to be optimized; The parameters of the corresponding reward sub-model are updated based on the second loss to determine the target reward model.

7. The image generation method according to claim 6, wherein the step of using the corresponding reward sub-model to predict the dimension to be optimized for the second sample prompt information and the image pair comprises: Determining, based on the second sample prompt information and the image pair, a first classification probability corresponding to the global dimension to be optimized in the corresponding reward sub-model; Based on the second sample prompt information and the image pair, a second classification probability corresponding to the leaf node in the corresponding reward sub-model is determined, and the predicted result includes the first classification probability and the second classification probability.

8. The image generation method according to claim 7, wherein determining the second loss based on the predicted result and the dimension label to be optimized comprises: Determine a reward value based on the second classification probability and the dimension label to be optimized; The second loss is determined based on a fusion result of the reward value and the first classification probability.

9. An image generating device, comprising: Prompt information acquisition module, used to obtain target prompt information; A target image generation module is used to generate a target image of the target prompt information based on the target prompt information and a target image generation model. The target image generation model is obtained by adjusting the parameters of a preset image generation model based on a target reward model. The input of the target reward model includes first sample prompt information and an image to be optimized obtained by the preset image generation model based on the first sample prompt information. The target reward model is used to predict the image feature information to be optimized in the image to be optimized.

10. A computer device comprising: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the image generation method according to any one of claims 1 to 8 by executing the computer instructions. 11 . A computer-readable storage medium having computer instructions stored thereon, wherein the computer instructions are used to enable a computer to execute the image generation method according to claim 1 . 12 . A computer program product comprising computer instructions, wherein the computer instructions are configured to cause a computer to execute the image generation method according to claim 1 .

Citation Information

Patent Citations

  • Text-to-image generation method based on fine-grained semantic reward

    CN116883530A

  • Text-image generation method, system and device and storage medium

    CN117095083A

  • Data processing method, text and graph generation method and related devices

    CN117671055A

  • Reinforced generation: reinforcement learning for text and knowledge graph bi-directional generation using pretrained language models

    US20240070404A1