Image generation method and device, equipment and storage medium

By training the low-resolution image generation model and fine-tuning the reward model, the high-resolution image generation model is optimized, solving the problem of matching the generated high-resolution images with human intentions, improving the performance of the super-resolution task and reducing the computational and storage requirements.

CN121661432APending Publication Date: 2026-03-13BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-09
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing image generation models struggle to match human intent when generating high-resolution images and perform poorly in super-resolution tasks, exhibiting issues with initial model selection, signal-to-noise ratio (SNR) choices, and sampling strategies.

Method used

The performance of the high-resolution image generation model is optimized by training a low-resolution image generation model, fine-tuning it using a second training data, and then fine-tuning it using a reward model.

Benefits of technology

It improves the performance of high-resolution image generation models in super-resolution tasks, ensuring that the generated images better meet user expectations, and reduces computation and storage requirements through memory optimization strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661432A_ABST
    Figure CN121661432A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image generation method and device, equipment and a storage medium. The method includes obtaining a trained first image generation model, the first image generation model configured to generate an image having a first resolution. A second image generation model is obtained by training the first image generation model using second training data including an image having a second resolution, the second image generation model being configured to generate an image having the second resolution, and the second resolution being higher than the first resolution. And training a second image generation model by using the second reward model. Therefore, the image generation model subjected to fine adjustment of the reward model can be obtained, so that the performance of the high-resolution image generation model is improved, and the high-resolution image generation model is enabled to better meet the expectation of a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatus, devices, and computer-readable storage media for image generation. Background Technology

[0002] With the increasing maturity of machine learning technology, image generation models based on machine learning have been widely used in generative applications. These models can generate a wide variety of images, greatly satisfying the diverse image generation needs of users across various industries. Among these applications, the generation of high-resolution images has become a key focus. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for image generation is provided. The method includes: acquiring a trained first image generation model, the first image generation model being configured to generate images having a first resolution; training the first image generation model using second training data, the second training data including images having a second resolution, the second image generation model being configured to generate images having a second resolution higher than the first resolution; and training the second image generation model using a second reward model.

[0004] In a second aspect of this disclosure, a method for image generation is provided. The method includes: obtaining descriptive text for an image generation target; generating a first image based on the descriptive text using a first image generation model, the first image having the first resolution; and generating a second image based on the first image using a second image generation model having a second resolution greater than the first resolution, the second image generation model being trained according to the method of the first aspect.

[0005] In a third aspect of this disclosure, an apparatus for image generation is provided. The apparatus includes: an acquisition module configured to acquire a trained first image generation model, the first image generation model being configured to generate an image having a first resolution; a first training module configured to train the first image generation model using second training data to obtain a second image generation model, the second training data including an image having a second resolution, the second image generation model being configured to generate an image having a second resolution higher than the first resolution; and a second training module configured to train the second image generation model using a second reward model.

[0006] In a fourth aspect of this disclosure, an apparatus for image generation is provided. The apparatus includes: an acquisition module configured to acquire descriptive text for an image generation target; a first image generation module configured to generate a first image based on the descriptive text using a first image generation model, the first image having a first resolution; and a second image generation module configured to generate a second image based on the first image using a second image generation model having a second resolution greater than the first resolution, the second image generation model being trained according to the apparatus of the third aspect.

[0007] In a fifth aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0008] In a sixth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;

[0012] Figure 2 An architecture diagram of an example of a training system for a second image generation model according to some embodiments of the present disclosure is shown;

[0013] Figure 3 An architecture diagram of an example of a training system for a first image generation model according to some embodiments of the present disclosure is shown;

[0014] Figure 4 An architectural diagram of an example model for image generation according to some embodiments of the present disclosure is shown;

[0015] Figure 5A schematic diagram of an example of a memory optimization scheme according to some embodiments of the present disclosure is shown;

[0016] Figure 6 A schematic diagram of another example of a memory optimization scheme according to some embodiments of the present disclosure is shown;

[0017] Figure 7 A schematic diagram of an architecture for an image generation model, which fundamentally discloses some embodiments, is shown.

[0018] Figure 8 A flowchart of a process for image generation according to some embodiments of the present disclosure is shown;

[0019] Figure 9 A flowchart of a process for image generation according to some embodiments of the present disclosure is shown;

[0020] Figure 10 A block diagram of an apparatus for image generation according to some embodiments of the present disclosure is shown;

[0021] Figure 11 A block diagram of an apparatus for image generation according to some embodiments of the present disclosure is shown; and

[0022] Figure 12 A block diagram of an apparatus capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation

[0023] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0024] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0025] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0026] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0027] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0028] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0029] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0030] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.

[0031] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0032] As used in this paper, the term "model" refers to a system that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that uses multiple layers of processing units to process inputs and provide corresponding outputs. In this paper, "model" may also be referred to as a "machine learning model," a "machine learning network," or simply a "network," and these terms are used interchangeably. A model can also include different types of processing units or networks.

[0033] As briefly mentioned earlier, image generation models are widely used in generative applications. Image generation models, such as text-to-image models, can generate images based on user-input text. Current image generation models perform well in generating low-resolution images (e.g., images with a pixel resolution of 256×256 or 512×512), producing the images users expect. However, they do not yet adequately meet user expectations in generating high-resolution images (e.g., images with a pixel resolution of 1024×1024 or 2048×2048), meaning the generated high-resolution images do not match human intent well. How to improve the model's performance in super-resolution tasks has become a pressing issue. In super-resolution tasks, the selection of the starting model and signal-to-noise ratio, as well as sampling strategies and memory optimization, are all problems that need to be addressed.

[0034] Embodiments of this disclosure propose a scheme for image generation. According to various embodiments of this disclosure, a trained first image generation model is obtained, the first image generation model being configured to generate images with a first resolution. A second image generation model is obtained by training the first image generation model using second training data, the second training data including images with a second resolution. The second image generation model is configured to generate images with a second resolution higher than the first resolution. The second image generation model is then trained using a second reward model.

[0035] In embodiments of this disclosure, a low-resolution image generation model is first obtained, and then trained to obtain a finely tuned high-resolution image generation model. Subsequently, a reward model is used to fine-tune the finely tuned high-resolution image generation model. This results in an image generation model finely tuned by the reward model, thereby improving the performance of the high-resolution image generation model and making it more aligned with user expectations. The resulting image generation model achieves better results in super-resolution tasks. Specifically, in some embodiments, the first trained image generation model is also trained using a reward model, ensuring that the final image generation model achieves the user's desired performance in super-resolution tasks.

[0036] Example Environment

[0037] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1As shown, model 130-1 with pre-training parameter values ​​and model 130-2 with post-training parameter values ​​can be collectively referred to as model 130 or referred to individually. Model 130 may be included in electronic device 140 and / or electronic device 150.

[0038] exist Figure 1 In environment 100, it is desirable to train and use a machine learning model (i.e., model 130) configured for various application environments. For example, in the case of an image generation model, an image corresponding to a text instruction can be generated based on a user-inputted text instruction.

[0039] like Figure 1 As shown, environment 100 includes electronic device 140 and electronic device 150. Electronic device 140 may contain a model training system, and electronic device 150 may contain a model application system. Figure 1 The upper part illustrates the model training phase, and the lower part illustrates the model application phase. Before training, the parameter values ​​of model 130 can have initial values ​​or pre-trained parameter values ​​obtained through a pre-training process. Model 130-1 can be trained via forward and backward propagation, during which the parameter values ​​of model 130-1 can be updated and adjusted. After training, model 130-2 is obtained. Model training can further include pre-training and fine-tuning. After pre-training, model 130-1 has generalization ability, such as the ability to process images based on input text instructions. Then, in the fine-tuning phase, the pre-trained model 130-1 is fine-tuned / refined for the downstream image generation task. At this point, the parameter values ​​of model 130-2 have been updated, and based on the updated parameter values, model 130-2 can be used to implement image processing tasks, such as image generation tasks, in the model application phase.

[0040] During the fine-tuning / adjustment phase of model training, model 130 can be trained using a model training system based on a training sample set 110 comprising multiple training samples 112. Each training sample 112 can involve a binary format. For example, for an image generation task, training sample 112 can include training input 120 and training output 122 from the image generation task. The training input for the image generation task can, for example, include training text and an image corresponding to the training text. Training sample 112, comprising training input 120 and training output 122, can be used to train model 130. Specifically, the training process can be performed iteratively using a large number of training samples. After training is complete, model 130 can include knowledge about the image generation task. During the model application phase, model 130 (at this point, model 130 has the trained parameter values) can be used to perform the corresponding task. For example, model input 142 from the image generation task can be received, and a corresponding model output 144 can be output.

[0041] exist Figure 1 In this context, electronic devices 140 and 150 can include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. Terminal devices can involve any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. Servers include, but are not limited to, mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0042] It should be understood that Figure 1 The components and arrangements shown in environment 100 are merely examples, and a computing system suitable for implementing the exemplary implementations described in this disclosure may include one or more different components, other components, and / or different arrangements. Implementations of this disclosure are not limited in this respect. Embodiments of this disclosure primarily relate to the training phase of an image generation model.

[0043] It should be understood that the structure and function of environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0044] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0045] Overall architecture of model training

[0046] Figure 2 An example architecture diagram of a training system 200 for a second image generation model according to some embodiments of the present disclosure is shown. Figure 2As shown, the training system 200 for the second image generation model can be implemented or included in the electronic device 140.

[0047] In some embodiments, the electronic device acquires a trained first image generation model 220-1 and trains the trained first image generation model 220-1 using second training data 210 to obtain a second image generation model 230.

[0048] In some embodiments, the trained first image generation model 220-1 is configured to generate images with a first resolution. The first resolution is, for example, a pixel resolution of 256×256 or a pixel resolution of 512×512. It should be understood that the specific values ​​of the resolutions listed herein are merely exemplary and are not intended to be limiting. In this disclosure, the first resolution is also referred to as a low resolution.

[0049] In some embodiments, the trained first image generation model 220-1 may have pre-trained parameter values ​​obtained through a pre-training process, or fine-tuned parameter values ​​obtained through a fine-tuning process, or parameter values ​​fine-tuned by human feedback obtained through a reward model training process. The process of obtaining the trained first image generation model 220-1 is described below. Figure 3 The description of that will not be repeated here.

[0050] In some embodiments, the trained first image generation model 220-1 is a model trained using a reward model, such as a model fine-tuned with human feedback. This results in the final second image generation model having better performance and effectiveness in super-resolution or high-resolution image generation tasks. In other words, using a model trained using a reward model, such as a model fine-tuned with human feedback, as the initial model for a super-resolution task can help make the model used for the super-resolution task more stable during the training process of high-resolution image generation or super-resolution tasks, and ensure the reliability of the model used for the super-resolution task.

[0051] In some embodiments, the second image generation model 230 is configured to generate an image with a second resolution. The second resolution is higher than the first resolution. The second resolution is, for example, a pixel resolution of 1024×1024 or a pixel resolution of 2048×2048. In this disclosure, the second resolution is also referred to as high resolution.

[0052] In some embodiments, the second image generation model 230 can be obtained by training the trained first image generation model 220-1 using the second training data 210. Figure 2In the example, the second training data 210 includes second text 212 and a second image 214 corresponding to the second text 212. The second image 214 is an image with a second resolution. The second text 212 can be descriptive text for the image generation target, such as "white kitten and black puppy". The second image 214 is, for example, a picture of a white kitten and a black puppy corresponding to "white kitten and black puppy".

[0053] The trained first image generation model 220-1 can generate a predicted image based on the second text 212. The parameters of the trained first image generation model 220-1 are then adjusted by comparing the predicted image with the second image 214. For example, a predicted image is generated based on the text "white kitten and black puppy," and then this predicted image is compared with the second image 214 (an image containing a white kitten and a black puppy), and the parameters of the trained first image generation model 220-1 are adjusted based on the comparison result.

[0054] In some embodiments, the trained first image generation model 220-1 can obtain a noisy image by diffusing noise onto the second image 214. Then, the first image generation model 220-1 denoises the obtained noisy image based on the second text 212 to obtain a predicted image. The parameters of the trained first image generation model 220-1 are then adjusted by comparing the noisy image and the predicted image.

[0055] It should be understood that training the trained first image generation model 220-1 using the second training data 210 is not limited to the methods described above. Various training methods currently available or to be developed in the art can be used to train the trained first image generation model 220-1 using the second training data 210 to obtain the second image generation model 230. In this disclosure, the second image generation model 230 is also referred to as a finely tuned high-resolution image generation model or a high-resolution finely tuned model.

[0056] After the electronic device obtains the second image generation model 230-0, it can train the second image generation model 230-0 using the second reward model 260 to obtain the second image generation model 230-1. In this disclosure, the second image generation model 230-1 is also referred to as a high-resolution image generation model fine-tuned by human feedback or a high-resolution human feedback fine-tuned model. In this disclosure, the second image generation model 230 and the second image generation model 230-1 can also be collectively referred to as the second image generation model. The third text 240 can be descriptive text for the image generation target. The third text 240 can be the same as or different from the second text 212. In some embodiments, the third text 240 can be a part of the second text 212. The third image 250 can be a predicted image generated by the second image generation model 230 corresponding to the third text 240.

[0057] The second reward model 260 can score the data pair consisting of the third text 240 and the third image 250 to obtain a second reward score 270. The electronic device can adjust the parameters of the second image generation model 230-0 based on the second reward score 270 to obtain the second image generation model 230-1. That is, the electronic device can use the trained second reward model 260 to fine-tune the second image generation model 230 to improve the performance of the second image generation model 230.

[0058] In some embodiments, for the same third text 240, a second image generation model 230 may generate multiple third images 250. In this case, a second reward model 260 may score each of the multiple third images 250 to obtain multiple second reward scores 270, and then the electronic device may fine-tune the second image generation model 230 based on the multiple second reward scores 270 to improve the performance of the second image generation model 230.

[0059] In some embodiments, the second reward model 260 may use a simple binary reward signal, such as using “+” or “-” symbols to represent the reward or punishment given, that is, the score of the reward model, for example, the second reward score 270 is 0 or 1.

[0060] In some embodiments, the second reward model 260 can be represented using integers between 0 and 5, i.e., the score of the reward model. For example, the second reward score 270 is an integer between 0 and 5, where 5 represents the highest reward and 0 represents the lowest reward. Such a reward signal allows the model to better understand the quality of the generated image and helps improve the model's performance in subsequent adjustment phases.

[0061] The second reward model can be implemented using any suitable network architecture. For example, the ALT CLIP (Adaptively Learned Text-Image Contrastive Learning) model can be used as the second reward model. The similarity score output by the ALT CLIP model typically refers to the degree of matching between the text description and the generated image.

[0062] The above describes using a trained first image generation model as a starting point to train a second image generation model. Below is an example implementation of training the first image generation model.

[0063] Figure 3 An example architecture diagram of a training system 300 for a first image generation model according to some embodiments of the present disclosure is shown. Figure 3 As shown, the training system 300 for the first image generation model can be implemented or included in the electronic device 140.

[0064] In some embodiments, the electronic device acquires an initial image generation model 320 and trains the initial image generation model 320 using first training data 310 to obtain a first image generation model 220-0.

[0065] In some embodiments, the initial image generation model 320 can be a pre-trained model. That is, the initial image generation model 320 can have pre-trained parameter values ​​obtained through a pre-training process.

[0066] In some embodiments, the first training data 310 includes first text 312 and a first image 314 corresponding to the first text 312. The first image 314 is an image with a first resolution. The first text 312 may be descriptive text for the image generation target, such as "white kitten and black puppy". The first image 314 may be, for example, a picture of a white kitten and a black puppy corresponding to "white kitten and black puppy".

[0067] In some embodiments, the initial image generation model 320 may generate a predicted image based on the first text 312, and then adjust the parameters of the initial image generation model 320 by comparing the predicted image with the first image 314. For example, a predicted image may be generated based on the text "white kitten and black puppy", and then the predicted image may be compared with the first image 314 (a picture with a white kitten and a black puppy), and the parameters of the initial image generation model 320 may be adjusted according to the comparison result.

[0068] In some embodiments, the initial image generation model 320 may perform diffusion noise addition based on the second image 214 to obtain a noisy image, then perform denoising on the noisy image based on the second text 212 to obtain a predicted image, and then adjust the parameters of the initial image generation model 320 by comparing the noisy image and the predicted image.

[0069] It should be understood that training the initial image generation model 320 using the first training data 310 is not limited to the methods described above. Various training methods currently available or to be developed in the art can be used to train the initial image generation model 320 using the first training data 310 to obtain the first image generation model 220-0. In this disclosure, the first image generation model 220-0 is also referred to as a fine-tuned low-resolution image generation model or a low-resolution fine-tuned model.

[0070] After the electronic device obtains the first image generation model 220-0, it trains the first image generation model 220-0 using the first reward model 360 to obtain the first image generation model 220-1. In this disclosure, the first image generation model 220-1 is also referred to as a low-resolution image generation model fine-tuned by human feedback or a low-resolution human feedback fine-tuning model. In this disclosure, the first image generation model 220-0 and the first image generation model 220-1 can also be collectively referred to as the first image generation model.

[0071] The fourth text 340 may be descriptive text for the image generation target. The fourth text 340 may be the same as or different from the first text 312. In some embodiments, the fourth text 340 may be a part of the first text 312. The fourth image 350 may be a predicted image generated by the first image generation model 220-0 corresponding to the fourth text 340.

[0072] The first reward model 360 can score the data pair consisting of the fourth text 340 and the fourth image 350 to obtain a first reward score 370. The electronic device can adjust the parameters of the first image generation model 220-0 based on the first reward score 370 to obtain the first image generation model 220-1. That is, the electronic device can use the trained first reward model 360 to fine-tune the first image generation model 220-0 to improve the performance of the second image generation model 220-0.

[0073] In some embodiments, for the same fourth text 340, a first image generation model 220-0 can generate multiple fourth images 350. In this case, a first reward model 360 can score each of the multiple fourth images 350 to obtain multiple first reward scores 370. Then, the electronic device fine-tunes the first image generation model 220-0 based on the multiple first reward scores 370 to improve the performance of the first image generation model 220-0.

[0074] In some embodiments, the first reward model 360 may use a simple binary reward signal, such as using “+” or “-” symbols to represent the reward or punishment given, that is, the score of the reward model, for example, the first reward score 370 is 0 or 1.

[0075] In some embodiments, the first reward model 360 can be represented using integers between 0 and 5, i.e., the score of the reward model. For example, the first reward score 370 is an integer between 0 and 5, where 5 represents the highest reward and 0 represents the lowest reward. Such a reward signal allows the model to better understand the quality of the generated image and helps improve the model's performance in subsequent adjustment phases.

[0076] The first reward model 360 can be implemented using any suitable network architecture. For example, the ALT CLIP model can be used as the first reward model 360, and the similarity score output by the ALT CLIP model usually refers to the degree of matching between the text description and the generated image.

[0077] In this embodiment, the super-resolution model is initialized by fine-tuning the low-resolution image generation model and then fine-tuning based on human feedback. Compared to directly fine-tuning the high-resolution model and then fine-tuning based on human feedback, this initialization method allows the super-resolution model to be trained more stably in subsequent super-resolution tasks and ensures the reliability of the generated image structure of the super-resolution model.

[0078] In some embodiments, the first image generation model (220-0, 220-1) and the second image generation model (230, 230-1) can employ a diffusion model, such as DDPM (Denoising Diffusion Probabilistic Models), Latent Diffusion Mode, or Stable Diffusion. The following section combines... Figure 4 An example architecture for the diffusion model is illustrated.

[0079] Figure 4 An architectural diagram of an example model for image generation according to some embodiments of this disclosure is shown. Figure 4 As shown, in some embodiments, the image generation model (e.g., a first image generation model or a second image generation model) includes: an image encoding network 430, a noise-adding network 440, a text encoding network 450, a noise-reducing network 460, and an image decoding network 470.

[0080] Image coding network 430 is used to encode the acquired input image 410 to obtain corresponding image features. In some embodiments, image coding network 430 may employ, but is not limited to, Variational Autoencoder (VAE), which maps the input image 410 to a latent feature space to obtain corresponding image features Z.

[0081] The noise-adding network 440 is used to diffuse and add noise to the image features Z, projecting the image features Z into the latent space to obtain the latent space vector, thereby obtaining the corresponding noisy image features, i.e., the noisy image features Z. T Where T represents the number of diffusion steps, or the number of time steps. That is, in the noisy network 440, image feature Z is generated by T diffusion steps. T Z T This represents the potential space value at time T.

[0082] In some embodiments, the noise-adding network 440 randomly adds Gaussian features to the image features Z. This process can be a fixed Markov chain process, which transforms the original data distribution into a normal distribution by continuously adding Gaussian noise.

[0083] A text encoding network 450 is used to encode the acquired descriptive text 420 to obtain corresponding text features. In some embodiments, the text encoding network 450 may employ, but is not limited to, a Contrastive Language-Image Pre-training (CLIP) model.

[0084] The denoising network 460 is used to refine the obtained noisy image features Z based on the obtained text features. T Denoising is performed to obtain the denoised image features Z'. In the denoising network 460, under the constraint of text features, the denoising process is used to denoise the noisy image features Z'. T T denoising predictions are performed to ultimately generate the latent space prediction vector Z', which is the predicted image feature Z'. During the denoising process, text features are used to constrain the noisy image feature Z. T The denoising process enables the denoising network 460 to output predicted image features Z' related to the input descriptive text 420 after T denoising operations.

[0085] The image decoding network 450 is used to decode the obtained denoised image features (i.e., the latent space prediction vector Z') to obtain the predicted image corresponding to the input text 420, i.e., the output image 480.

[0086] For diffusion models, the noise level associated with the noise addition and denoising processes directly affects the performance of the image generation model.

[0087] Continue to refer to Figure 2 and Figure 3 In some embodiments, a first signal-to-noise ratio (SNR) is used during the training of the first image generation model 220-0, while a second SNR is used during the training of the trained first image generation model 220-1 using the second training data 210. The second SNR is different from the first SNR.

[0088] In some embodiments, the second signal-to-noise ratio is less than the first signal-to-noise ratio. That is, during the training of the first image generation model 220-1 using the second training data 210, more noise can be added to the image. This allows the resulting second image generation model 230-0 to learn to add more details under high-noise conditions, thereby achieving better performance in super-resolution tasks.

[0089] In some embodiments, the ratio of the first signal-to-noise ratio to the second signal-to-noise ratio is positively correlated with the ratio of the second resolution to the first resolution. For example, if the first resolution is 512×512 and the second resolution is 1024×1024, the first signal-to-noise ratio is a, and the second signal-to-noise ratio is a / 4. It should be understood that the specific values ​​or ratios of resolution and noise listed herein are merely exemplary and are not intended to be limiting.

[0090] In some embodiments, a first signal-to-noise ratio (SNR) is used during the training of the first image generation model 220-0, while a third SNR is used during the training of the second image generation model 230-0 using the second reward model 260. The third SNR is different from the first SNR. The third SNR may be the same as or different from the second SNR.

[0091] In some embodiments, the third signal-to-noise ratio is less than the first signal-to-noise ratio. That is, during the training of the second image generation model 230-0 using the second reward model 260, more noise can be added to the image. This allows the resulting second image generation model 220-1 to learn to add more details under high-noise conditions, thereby achieving better performance in super-resolution tasks.

[0092] In some embodiments, the ratio of the first signal-to-noise ratio to the third signal-to-noise ratio is positively correlated with the ratio of the second resolution to the first resolution. For example, if the first resolution is 512×512 and the second resolution is 1024×1024, the first signal-to-noise ratio is a, and the third signal-to-noise ratio is a / 4. It should be understood that the specific values ​​or ratios of resolution and noise listed herein are merely exemplary and are not intended to be limiting.

[0093] For diffusion models, image noise gradually increases with time steps during the denoising stage, while it gradually decreases with time steps during the denoising stage. In super-resolution tasks, the stages related to image texture and detail are mainly those corresponding to the earlier time steps of the model. For example, if a diffusion model diffuses for 1000 steps, the first 500 steps of the denoising network may be more related to image type, while the last 500 steps are mainly related to image details and texture. Therefore, in super-resolution tasks, the focus is on the stages related to image details and texture, emphasizing sampling and optimization of these stages so that the second image generation model can focus on adding detail and texture information.

[0094] Continue to refer to Figure 2 In some embodiments of this disclosure, when training the second image generation model 230-0 using the second reward model 260, a set of time steps is sampled from multiple time steps according to a sampling strategy. This sampling strategy ensures that the sampling probability of time steps with high model noise levels (e.g., time steps near the predicted image feature Z' in the denoising stage) is greater than that of time steps with high noise levels (e.g., time steps near the noisy image feature Z' in the denoising stage). T The sampling probability at time steps. In some embodiments, a power-law sampling strategy is used when training the second image generation model 230-0 using the second reward model 260.

[0095] As briefly mentioned earlier, in some embodiments, the image generation model uses a diffusion model, and the training of this diffusion model is completed in the latent space, which makes the computational and storage requirements for training relatively small. However, when the second image generation model 230-0 is trained using the second reward model 260, the loss function needs to be calculated in the image space. High-resolution images make the computational and storage requirements for training very large, so memory optimization becomes a necessary operation.

[0096] Figure 5 A schematic diagram of an example of a memory optimization scheme according to some embodiments of this disclosure is shown. Figure 5 As shown, in some embodiments, the electronic device includes multiple processing units, such as processing units 1-n. When training the second image generation model 230-0 using the second reward model 260, the training data of the second image generation model 230-0 can be divided among the n processing units. During training, each of the n processing units processes a portion of the training data. Accordingly, during the parameter update phase, each of the n processing units updates its corresponding training data.

[0097] In some embodiments, the training data includes parameters of the second image generation model 230-0, such as weights w. In some embodiments, the training data may also include intermediate state values ​​of the training of the second image generation model 230-0, such as gradients and optimizer states.

[0098] As an example, the second image generation model 230-0 includes an n-layer network, where each layer contains weight parameters W1, W2, ..., Wn. Figure 5 As shown, when training the second image generation model 230-0 using the second reward model 260, during forward propagation, the weights W1, W2..., and Wn of each of the n layers of the network are assigned to processing units 1, 2,..., and n, respectively. During backpropagation, processing units 1, 2,..., and n process and / or store their respective gradients g1, g2..., gn. During the parameter update phase, processing units 1, 2,..., and n process and / or store their respective optimizer states S1, S2..., Sn, and weights W1-, W2-..., and Wn-. Thus, each processing unit only processes and stores a portion of the training data, significantly reducing the required GPU memory. Therefore, the problem of GPU memory explosion during training can be avoided. Simultaneously, it improves training stability and enables the training of larger-scale image generation models.

[0099] It should be understood that, Figure 5 The example shown is merely one instance of memory optimization and does not constitute a limitation of this disclosure. Other similar strategies can be employed in other embodiments of this disclosure. For example, the optimizer states of the second image generation model 230-0 can be distributed across multiple processing units. As another example, the optimizer states and gradients of the second image generation model 230-0 can be distributed across multiple processing units. As yet another example, such as... Figure 5 As shown, the weights, optimizer states, and gradients of the second image generation model 230-0 are divided into multiple processing units.

[0100] It should also be understood that the division of training data is not limited to Figure 5 The method shown can be varied, but can be any suitable method, such as processing and storing a set of training data on each processing unit, which is a subset of the total training data of the second image generation model.

[0101] Figure 6 A schematic diagram of another example of a memory optimization scheme according to some embodiments of the present disclosure is shown.

[0102] In some embodiments, to reduce the memory required to train the second image generation model 230-0 using the second reward model 260, or to train a larger-scale second image generation model 230-0, during the forward propagation of the second image generation model 230-0, the intermediate state values ​​of the first part of the intermediate state values ​​of the second image generation model are stored, but the intermediate state values ​​of the second part of the intermediate state values ​​of the second image generation model are not stored; while during the backward propagation, the intermediate state values ​​of the second part are determined based on the intermediate state values ​​of the first part.

[0103] like Figure 6 As shown, exemplarily, the second image generation model includes nodes 1-N, which generate activation values ​​a1 to an respectively during the training process. However, only a1, a3, ..., an are stored during forward propagation. During backward propagation, when a2, a4, ..., an-1 are needed, a2, a4, ..., an-1 are recalculated based on a1, a3, ..., an. This significantly reduces the required GPU memory because only a portion of the activation values ​​are stored, thus enabling the training of a larger-scale second image generation model.

[0104] It should be understood that, Figure 6 This explanation merely illustrates how to store the first part of the intermediate state values ​​of the second image generation model during forward propagation, without storing the second part of the intermediate state values. During backward propagation, the second part of the intermediate state values ​​is determined based on the first part of the intermediate state values, which does not constitute a limitation of this disclosure. This disclosure can use various suitable methods to store some intermediate state values ​​as needed. That is, which intermediate state values ​​of the second image generation model are stored and which are not are not limited to... Figure 6 The division shown is not a fixed method, but can be any suitable method.

[0105] Figure 7 A schematic diagram of an architecture for an image generation model, based on some embodiments fundamentally disclosed, is shown. For example... Figure 7 As shown, the model 700 for image generation can be implemented or included in the electronic device 150.

[0106] In some embodiments, the electronic device acquires descriptive text 710 for the image generation target, and then encodes the descriptive text 710 using the text encoding network 450 of the first image generation model 220-1 to obtain text features. The text features and random noise 720 are then input into the denoising network 460 of the first image generation model 220-1 to obtain predicted image features, which are then passed through the image decoding network 470 of the first image generation model 220-1 to obtain the first output image 730.

[0107] Next, the electronic device inputs the first output image 730 into the image encoding network 730 of the second image generation model 230-1 to obtain image features. Then, based on the image features, the second output image 740 is obtained through the noise-adding network 440, the noise-reducing network 460, and the image decoding network 470 of the second image generation model 230-1. The resolution of the second output image 740 is greater than the resolution of the first output image 730.

[0108] Example process

[0109] Figure 8 A flowchart of a process 800 for image generation according to some embodiments of the present disclosure is shown. Process 800 may be implemented or included at an electronic device 140. Reference is made below. Figure 8 Describe the process 800.

[0110] In box 810, a trained first image generation model is obtained, which is configured to generate an image with a first resolution.

[0111] In some embodiments, obtaining a trained first image generation model includes:

[0112] An initial image generation model is trained using first training data, the first training data including images having the first resolution; and

[0113] The initial image generation model is trained using the first reward model to obtain the first image generation model.

[0114] In box 820, a second image generation model is obtained by training a first image generation model using second training data, the second training data including images with a second resolution, and the second image generation model is configured to generate images with a second resolution higher than the first resolution.

[0115] In box 830, the second image generation model is trained using the second reward model.

[0116] In some embodiments, the first image generation model and the second image generation model each include a diffusion model, a first signal-to-noise ratio is used in the training of the first image generation model, and a second signal-to-noise ratio used in at least one of the following is less than the first signal-to-noise ratio:

[0117] The first image generation model is trained using the second training data, or

[0118] The second image generation model is trained using the second reward model.

[0119] In some embodiments, the ratio of the first signal-to-noise ratio to the second signal-to-noise ratio is positively correlated with the ratio of the second resolution to the first resolution.

[0120] In some embodiments, training a second image generation model using a second reward model includes:

[0121] The training data for the second image generation model is divided into multiple processing units, such that each processing unit processes a portion of the training data, which includes model parameters and intermediate training state values; and

[0122] The corresponding parts of the training data are updated in multiple processing units respectively.

[0123] In some embodiments, training a second image generation model using a second reward model includes:

[0124] During the forward propagation of the second image generation model, the intermediate state values ​​of the first part of the intermediate state values ​​of the second image generation model are stored, but the intermediate state values ​​of the second part of the intermediate state values ​​of the second image generation model are not stored; and

[0125] During the backpropagation process of the second image generation model, the intermediate state values ​​of the second part are determined based on the intermediate state values ​​of the first part.

[0126] In some embodiments, the second image generation model corresponds to a denoising process and a noise-adding process including multiple time steps, and training the second image generation model using a second reward model includes:

[0127] According to a preset sampling strategy, a set of time steps is sampled from multiple time steps, wherein the sampling strategy ensures that the sampling probability of time steps with low noise levels is greater than that of time steps with high noise levels; and

[0128] The second image generation model is trained using a second reward model based on noise addition and denoising operations in a set of time steps.

[0129] In some embodiments, when training a second image generation model using a second reward model, the model parameters are sampled using a power-law sampling strategy.

[0130] Figure 9 A flowchart of a process 900 for image generation according to some embodiments of the present disclosure is shown. Process 900 may be implemented or included at an electronic device 150. Reference is made below. Figure 9 Describe the process 900.

[0131] In box 910, obtain the descriptive text for the image generation target.

[0132] In box 920, a first image is generated based on the descriptive text using a first image generation model, the first image having the first resolution.

[0133] In box 930, based on the first image, a second image with a second resolution is generated using a second image generation model. The second resolution is greater than the first resolution. The second image generation model is trained according to the method of this disclosure.

[0134] Example devices and equipment

[0135] Figure 10 A schematic structural block diagram of an image generation apparatus 1000 according to certain embodiments of the present disclosure is shown. The apparatus 1000 may be implemented as or included in an electronic device 140. Various modules / components in the apparatus 1000 may be implemented by hardware, software, firmware, or any combination thereof.

[0136] like Figure 10 As shown, the apparatus 1000 includes an acquisition module 1010 configured to acquire a trained first image generation model, the first image generation model being configured to generate images with a first resolution. The apparatus 1000 also includes a first training module 1020 configured to train the first image generation model using second training data to obtain a second image generation model, the second training data including images with a second resolution, the second image generation model being configured to generate images with a second resolution higher than the first resolution. The apparatus 1000 also includes a second training module 1030 configured to train the second image generation model using a second reward model.

[0137] In some embodiments, the first image generation model and the second image generation model each include a diffusion model, the acquisition module 1010 is further configured to use a first signal-to-noise ratio (SNR) in the training of the first image generation model, the first training module 1020 is further configured to use a second SNR less than the first SNR in the training of the first image generation model using second training data, and / or the second training module 1030 is further configured to use a second SNR less than the first SNR in the training of the second image generation model using a second reward model.

[0138] In some embodiments, the ratio of the first signal-to-noise ratio to the second signal-to-noise ratio is positively correlated with the ratio of the second resolution to the first resolution.

[0139] In some embodiments, the second training module 1030 is further configured to:

[0140] The training data for the second image generation model is divided into multiple processing units, such that each processing unit processes a portion of the training data, which includes model parameters and intermediate training state values; and

[0141] The corresponding parts of the training data are updated in multiple processing units respectively.

[0142] In some embodiments, the second training module 1030 is further configured to:

[0143] During the forward propagation of the second image generation model, the intermediate state values ​​of the first part of the intermediate state values ​​of the second image generation model are stored, but the intermediate state values ​​of the second part of the intermediate state values ​​of the second image generation model are not stored; and

[0144] During the backpropagation process of the second image generation model, the intermediate state values ​​of the second part are determined based on the intermediate state values ​​of the first part.

[0145] In some embodiments, the second image generation model corresponds to a denoising process and a noise-adding process including multiple time steps, and the second training module 1030 is further configured to:

[0146] According to a preset sampling strategy, a set of time steps is sampled from multiple time steps, wherein the sampling strategy ensures that the sampling probability of time steps with low noise levels is greater than that of time steps with high noise levels; and

[0147] The second image generation model is trained using a second reward model based on noise addition and denoising operations in a set of time steps.

[0148] In some embodiments, the second training module 1030 is further configured to sample the model parameters using a power-law sampling strategy.

[0149] In some embodiments, the acquisition module 1010 is further configured to:

[0150] An initial image generation model is trained using first training data, the first training data including images having the first resolution; and

[0151] The initial image generation model is trained using the first reward model to obtain the first image generation model.

[0152] Figure 11 A schematic structural block diagram of an image generation apparatus 1100 according to certain embodiments of the present disclosure is shown. The apparatus 1100 may be implemented as or included in an electronic device 150. Various modules / components in the apparatus 1100 may be implemented by hardware, software, firmware, or any combination thereof.

[0153] like Figure 11 As shown, the device 1100 includes an acquisition module 1110 configured to acquire descriptive text for an image generation target. The device 1100 also includes a first image generation module 1120 configured to generate a first image based on the descriptive text using a first image generation model, the first image having a first resolution. The device 1100 further includes a second image generation module 1130 configured to generate a second image based on the first image using a second image generation model, the second resolution being greater than the first resolution. The second image generation model is based on... Figure 10 The device shown was trained.

[0154] Figure 12 A block diagram is shown illustrating an electronic device 1200 in which one or more embodiments of the present disclosure may be implemented. It should be understood that... Figure 12 The electronic device 1200 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 12 The electronic device 1200 shown can be used to achieve Figure 1 Electronic devices 110.

[0155] like Figure 12 As shown, electronic device 1200 is in the form of a general-purpose electronic device. Components of electronic device 1200 may include, but are not limited to, one or more processors or processing units 1210, memory 1220, storage device 1230, one or more communication units 1240, one or more input devices 1250, and one or more output devices 1260. Processing unit 1210 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 1220. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 1200.

[0156] Electronic device 1200 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 1200, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 1220 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 1230 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 1200.

[0157] Electronic device 1200 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 12 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 1220 may include computer program product 1225 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0158] The communication unit 1240 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 1200 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 1200 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0159] Input device 1250 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 1260 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 1200 can also communicate with one or more external devices (not shown) via communication unit 1240 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 1200, or with any device that enables electronic device 1200 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0160] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above.

[0161] According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transient computer-readable medium and includes computer-executable instructions that are executed by a processor to implement the methods described above.

[0162] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0163] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0164] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0165] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0166] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for image generation, comprising: Obtain a trained first image generation model, the first image generation model being configured to generate an image with a first resolution; A second image generation model is obtained by training the first image generation model using second training data, the second training data including images with a second resolution, the second image generation model being configured to generate images with the second resolution, and the second resolution being higher than the first resolution; as well as The second image generation model is trained using the second reward model.

2. The method according to claim 1, wherein the first image generation model and the second image generation model each comprise a diffusion model, a first signal-to-noise ratio is used in the training of the first image generation model, and a second signal-to-noise ratio is used that is less than the first signal-to-noise ratio in at least one of the following: The first image generation model is trained using the second training data, or The second image generation model is trained using the second reward model.

3. The method of claim 2, wherein the ratio of the first signal-to-noise ratio to the second signal-to-noise ratio is positively correlated with the ratio of the second resolution to the first resolution.

4. The method of claim 1, wherein training the second image generation model using the second reward model comprises: The training data for the second image generation model is divided into multiple processing units, such that each of the multiple processing units processes a portion of the training data, wherein the training data includes model parameters and intermediate state values ​​of the training. as well as The corresponding portions of the training data are updated by the multiple processing units respectively.

5. The method of claim 1, wherein training the second image generation model using the second reward model comprises: During the forward propagation of the second image generation model, the intermediate state values ​​of the first part of the intermediate state values ​​of the second image generation model are stored, but the intermediate state values ​​of the second part of the intermediate state values ​​of the second image generation model are not stored. as well as During the backpropagation process of the second image generation model, the intermediate state values ​​of the second part are determined based on the intermediate state values ​​of the first part.

6. The method of claim 1, wherein the second image generation model corresponds to a denoising process and a noise-adding process comprising multiple time steps, and training the second image generation model using the second reward model comprises: According to a preset sampling strategy, a set of time steps is sampled from the plurality of time steps, wherein the sampling strategy makes the sampling probability of time steps with low noise levels greater than the sampling probability of time steps with high noise levels; and Based on the noise addition and denoising operations in the set of time steps, the second image generation model is trained using the second reward model.

7. The method of claim 6, wherein when training the second image generation model using the second reward model, the model parameters are sampled using a power-law sampling strategy.

8. The method of claim 1, wherein obtaining the trained first image generation model comprises: An initial image generation model is trained using first training data, the first training data including images having the first resolution; as well as The initial image generation model is trained using the first reward model to obtain the first image generation model.

9. A method for generating an image, comprising: Obtain descriptive text for the target generated from the image; Based on the descriptive text, a first image is generated using a first image generation model, and the first image has the first resolution; as well as Based on the first image, a second image with a second resolution is generated using a second image generation model, wherein the second resolution is greater than the first resolution, and the second image generation model is trained by the method according to any one of claims 1-8.

10. An apparatus for generating an image, comprising: The acquisition module is configured to acquire a trained first image generation model, the first image generation model being configured to generate an image with a first resolution; A first training module is configured to train a first image generation model using second training data to obtain a second image generation model, wherein the second training data includes images with a second resolution, and the second image generation model is configured to generate images with the second resolution, wherein the second resolution is higher than the first resolution. as well as The second training module is configured to train the second image generation model using the second reward model.

11. An apparatus for generating an image, comprising: The acquisition module is configured to acquire descriptive text for the image generation target; A first image generation module is configured to generate a first image based on the description text using a first image generation model, wherein the first image has a first resolution. as well as The second image generation module is configured to generate a second image with a second resolution based on the first image using a second image generation model. The second resolution is greater than the first resolution. The second image generation model is trained by the apparatus according to claim 10.

12. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 9.

13. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 9.