Training method of image generation model, image generation method, device, equipment, storage medium and program product

By executing multiple iteration rounds and updating model parameters in the image generation model, the problem that the image generation model is prone to introduce errors in the image generation process is solved, and higher generation effect and consistency are achieved.

CN120219201APending Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510292852.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The image generation model based on the diffusion model is prone to introduce errors in the image generation process, resulting in deviations from the natural language description.

Method used

By performing multiple iteration rounds in the image generation model, the target image samples are gradually recovered from the noise, and the loss value is determined based on the image sample pairs of each iteration round, the parameters of the image generation model are updated.

Benefits of technology

It improves the generation effect of the image generation model, reduces the accumulation of errors, enhances the consistency between the generated image and the description text, and improves the image generation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219201A_ABST
    Figure CN120219201A_ABST
Patent Text Reader

Abstract

The invention provides an image generation model training method and device, an image generation method and device, equipment, a storage medium and a program product. The method comprises the following steps: executing at least one iteration round through an initial image generation model to generate a target image sample; based on the target image sample, determining a loss value for a preset annotation image sample for describing the text sample and the image sample pair of each iteration round; and based on the loss value, updating parameters of the initial image generation model to obtain a trained image generation model. The image generation effect can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method for training an image generation model, an image generation method, an apparatus, a device, a storage medium, and a program product. Background Art

[0002] In the related art, an image generation model based on a diffusion model (Diffusion Model) can generate images based on natural language descriptions. However, since each generated image requires multiple steps of processing to gradually recover a clear image from noise, errors may be introduced in each step, and these errors will be amplified as the steps accumulate, resulting in a deviation between the generated image and the natural language description. Summary of the Invention

[0003] Embodiments of this application provide a method for training an image generation model, an image generation method, an apparatus, a device, a storage medium, and a program product, which can improve the image generation effect.

[0004] The technical solution of the embodiments of this application is implemented as follows:

[0005] Embodiments of this application provide a method for training an image generation model, the method including:

[0006] Executing at least one iteration round through an initial image generation model to generate a target image sample, where in each of the at least one iteration round, the following steps are sequentially executed:

[0007] Denoising the image to be denoised based on a description text sample to obtain a denoised image,

[0008] Sampling the denoised image to obtain a plurality of sampled images,

[0009] Selecting an image sample pair from the plurality of sampled images, where the image sample pair includes a positive sampled image sample and a negative sampled image sample,

[0010] In the case that the iteration round does not meet the iteration end condition, returning to the step of denoising the image to be denoised based on the description text sample,

[0011] In the case that the iteration round meets the iteration end condition, using the denoised image as the target image sample, or using the positive sampled image sample as the target image sample,

[0012] Wherein, the image to be denoised is a noise image sample in the first iteration round, and in each iteration round after the first iteration round, includes any one of the image sample pair, the denoised image, and the positive sampled image sample obtained in the previous iteration round;

[0013] Determine a loss value based on the target image sample, the labeled image sample preset for the description text sample, and the image sample pairs of each iteration round;

[0014] Update the parameters of the initial image generation model based on the loss value to obtain a trained image generation model.

[0015] An embodiment of the present application provides an image generation method, and the method includes:

[0016] Obtain a description text;

[0017] Generate multiple generated images through an image generation model based on the description text, where the image generation model is trained through the image generation model training method provided by the embodiment of the present application.

[0018] An embodiment of the present application provides a training device for an image generation model, including:

[0019] A data processing module, configured to execute at least one iteration round through an initial image generation model to generate a target image sample, where in each of the at least one iteration rounds, the following steps are sequentially executed:

[0020] Denoise the image to be denoised based on the description text sample to obtain a denoised image,

[0021] Sample the denoised image to obtain multiple sampled images,

[0022] Select an image sample pair from the multiple sampled images, where the image sample pair includes a positive sampled image sample and a negative sampled image sample,

[0023] In the case that the iteration round does not meet the iteration end condition, return to the step of denoising the image to be denoised based on the description text sample,

[0024] In the case that the iteration round meets the iteration end condition, use the denoised image as the target image sample, or use the positive sampled image sample as the target image sample,

[0025] Wherein, the image to be denoised is a noise image sample in the first iteration round, and in each iteration round after the first iteration round, it includes any one of the image sample pair, the denoised image, and the positive sampled image sample obtained in the previous iteration round;

[0026] A training module, configured to determine a loss value based on the target image sample, the labeled image sample preset for the description text sample, and the image sample pairs of each iteration round;

[0027] The training module is further configured to update the parameters of the initial image generation model based on the loss value to obtain a trained image generation model.

[0028] An embodiment of the present application provides an image generation device, including:

[0029] A data acquisition module, configured to acquire a description text;

[0030] An image generation module, configured to generate images based on the description text through an image generation model to obtain a plurality of generated images, where the image generation model is trained by the image generation model training method provided by the embodiment of the present application.

[0031] An embodiment of the present application provides an electronic device, where the electronic device includes:

[0032] A memory, configured to store computer-executable instructions or a computer program;

[0033] A processor, configured to implement the image generation model training method provided by the embodiment of the present application or the image generation method provided by the embodiment of the present application when executing the computer-executable instructions or the computer program stored in the memory.

[0034] An embodiment of the present application provides a computer-readable storage medium, storing a computer program or computer-executable instructions, configured to implement the image generation model training method provided by the embodiment of the present application or the image generation method provided by the embodiment of the present application when being executed by a processor.

[0035] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions, where the computer program or computer-executable instructions implement the image generation model training method provided by the embodiment of the present application or the image generation method provided by the embodiment of the present application when being executed by a processor.

[0036] The embodiment of the present application has the following beneficial effects:

[0037] In the embodiments of the present application, by utilizing the characteristic that the image generation model generates the target image sample through at least one iteration round, that is, the denoised image of each iteration round is generated based on the denoised image of the previous iteration round, and finally the target image sample without noise is obtained. An image sample pair corresponding to each iteration round is constructed based on multiple sampled images of each iteration round (that is, from multiple sampled images of each iteration round, a positive sampled image sample and a negative sampled image sample are selected to construct an accurate sample pair), so that the image generation model can be trained based on the image sample pairs of each iteration round. Based on the target image sample, the labeled image sample preset for the description text sample, and the image sample pairs of each iteration round, a loss value is determined, and the parameters of the image generation model are updated based on the loss value, realizing stage-by-stage learning, which can improve the ability of the image generation model to distinguish positive and negative sampled image samples. When updating the image generation model based on the loss value, it can drive the parameters of the image generation model to generate in the direction consistent with the positive sampled image sample, thereby eliminating the error of the image generation model in each iteration round, improving the consistency between the target image sample and the description text sample, and ensuring the image generation accuracy. Description of the Drawings

[0038] Figure 1 is a schematic structural diagram of the image generation system architecture provided by the embodiments of the present application;

[0039] Figure 2A is a first schematic structural diagram of the server provided by the embodiments of the present application;

[0040] Figure 2B is a second schematic structural diagram of the server provided by the embodiments of the present application;

[0041] Figure 3 is a schematic diagram of the principle of the training method of the image generation model provided by the embodiments of the present application;

[0042] Figure 4A is a first flowchart of the training method of the image generation model provided by the embodiments of the present application;

[0043] Figure 4B is a second flowchart of the training method of the image generation model provided by the embodiments of the present application;

[0044] Figure 4C is a third flowchart of the training method of the image generation model provided by the embodiments of the present application;

[0045] Figure 4D is a fourth flowchart of the training method of the image generation model provided by the embodiments of the present application;

[0046] Figure 4EIt is the fifth process schematic diagram of the training method of the image generation model provided by the embodiments of the present application;

[0047] Figure 4F It is the sixth process schematic diagram of the training method of the image generation model provided by the embodiments of the present application;

[0048] Figure 4G It is the seventh process schematic diagram of the training method of the image generation model provided by the embodiments of the present application;

[0049] Figure 4H It is the eighth process schematic diagram of the training method of the image generation model provided by the embodiments of the present application;

[0050] Figure 5 It is the process schematic diagram of the image generation method provided by the embodiments of the present application;

[0051] Figure 6 It is the schematic diagram of the network structure of the image generation model provided by the embodiments of the present application;

[0052] Figure 7A It is the first schematic diagram of the feature fusion principle provided by the embodiments of the present application;

[0053] Figure 7B It is the second schematic diagram of the feature fusion principle provided by the embodiments of the present application;

[0054] Figure 7C It is the third schematic diagram of the feature fusion principle provided by the embodiments of the present application;

[0055] Figure 8A It is the first schematic diagram of the iteration process provided by the embodiments of the present application;

[0056] Figure 8B It is the second schematic diagram of the iteration process provided by the embodiments of the present application;

[0057] Figure 8C It is the third schematic diagram of the iteration process provided by the embodiments of the present application;

[0058] Figure 8D It is the fourth schematic diagram of the iteration process provided by the embodiments of the present application;

[0059] Figure 9 It is the schematic diagram of the acquisition principle of the second loss value provided by the embodiments of the present application;

[0060] Figure 10 It is the application process schematic diagram of the image generation method in the game scenario provided by the embodiments of the present application;

[0061] Figure 11 It is the schematic diagram of the image generation interface provided by the embodiments of the present application.

[0062] It should be noted that the above-mentioned "first" and "second" are only used to distinguish different solutions, and do not represent the distinction of the advantages or disadvantages of the solutions or the priority in the implementation process. Detailed implementation manners

[0063] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0064] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0065] In the following description, the terms "first / second / third" are only used to distinguish similar objects, and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0066] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0067] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as commonly understood by those of ordinary skill in the art to which the present application belongs. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application, and are not intended to limit the present application.

[0068] In the embodiments of the present application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations during actual application, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope authorized by laws and regulations and the personal information subject.

[0069] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described, and the nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0070] 1) Responsive to, used to represent the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more of the executed operations can be real-time or can have a set delay; without special instructions, there is no restriction on the execution order of multiple executed operations.

[0071] 2) Human-computer interaction interface, an interface for providing human-computer interaction functions / an interface for displaying image generation information.

[0072] For example, a graphical user interface (GUI) display, such as an augmented reality (AR) interface, a virtual reality (VR) interface, a voice user interface (VUI), an interactive projection interface (using projection technology to display information on a plane), an eye movement detection interface (an interface controlled by detecting the user's line of sight), a holographic interface (a three-dimensional hologram formed by projecting an image through holographic projection technology, and a stereoscopic image can be seen without wearing special glasses), a multimodal interface (an interactive interface that combines multiple interaction methods such as touch, vision, and hearing), a brain-machine interface (BMI) interface, etc.

[0073] 3) Description text (Prompt), which is a natural language instruction input by the user and is used to guide the generation of an image that meets expectations. It is the core bridge connecting human intentions and the image generation result and directly affects the quality, style, and content of the image. For example, the description text can be expressed as "On a peaceful summer evening, the sun is slowly setting, and the sky is dyed with a gradient of orange-red and purple. A vast wheat field is gently swaying in the breeze, and the wheat ears are shining with golden light under the setting sun. There is a winding dirt road in the middle of the wheat field, leading to a small wooden house in the distance, and there is curling smoke rising from the chimney of the wooden house. There are also several flying birds in the sky, flying into the distance."

[0074] 4) Annotated image sample, which refers to an image sample corresponding to the description content of the description text sample. For example, the annotated image sample corresponding to the above description text example can be a landscape painting full of the atmosphere of a summer evening, including the main elements (sky, wheat field, dirt road, small wooden house, flying birds, etc.) that appear in the description text, with warm and soft colors and rich details (such as the texture of the wheat field, the details of the wooden house, the dynamics of the flying birds, etc.).

[0075] 5) A noise image sample refers to a noise image obtained by adding noise to a labeled image sample. The noise can be Gaussian noise, and the noise image sample has the same size as the labeled image sample.

[0076] 6) A denoised image refers to the denoised image generated in each iteration round (or at least one time step) during the process of denoising a noise image sample for at least one iteration round.

[0077] 7) A positive sampling image sample refers to the sampling image with the highest quality parameter among the multiple sampling images corresponding to each iteration round. Among them, the quality parameter is an index used to quantify the aesthetic quality of an image (such as image composition, color, etc., which can be quantified by an aesthetic evaluation score) and the technical quality (such as image sharpness, noise level, etc.).

[0078] 8) A negative sampling image sample refers to the sampling image with the lowest quality parameter among the multiple sampling images corresponding to each iteration round.

[0079] 9) Diffusion Models is a generative model that simulates the process of data samples being gradually covered by noise (forward diffusion or diffusion process), and then learns how to reverse the removal of this noise to restore the original data (reverse diffusion or denoising process).

[0080] 10) Forward Process refers to the process of gradually adding Gaussian noise to the original data (such as a labeled image sample). After multiple steps, the data gradually becomes pure noise (a noise image sample). This process is similar to "destroying" the data, and finally obtains a random distribution independent of the original data.

[0081] 11) Reverse Process refers to the process of gradually removing noise and reconstructing high-quality samples from random noise. In each step of this process, the diffusion model attempts to recover some information of the original data (labeled image sample) from the current noisy data. After multiple iterations, a new sample (target image sample) similar to the original data is finally recovered.

[0082] 12) Transformer refers to a temporal model based on the self-attention mechanism. In the encoder part, it can effectively encode temporal information, and its processing ability for temporal information is much better than that of the Long Short-Term Memory (LSTM) network, and it is also fast. It is widely used in the fields of natural language processing, computer vision, machine translation, speech recognition, etc.

[0083] 13) The Monte Carlo method refers to a class of algorithms that rely on repeated random sampling to obtain numerical results.

[0084] 14) The Monte Carlo Tree Search (MCTS) is a heuristic search algorithm used in the decision-making process. The core idea is to explore possible action paths through random simulation (the Monte Carlo method) and a tree structure, gradually focusing on the optimal strategy.

[0085] 15) Generating an image, also known as a generative image, refers to using a machine learning model, especially a deep learning model, to learn the features and patterns of existing images (such as captured images), and then generating brand-new images based on these learned patterns. These images can be completely original or variants based on existing images.

[0086] In the related art, an image generation model based on the Diffusion Model can generate images based on natural language descriptions. However, since each generated image requires multiple steps of processing to gradually recover a clear image from noise, errors may be introduced at each step, and these errors will be amplified as the steps accumulate, resulting in a deviation between the generated image and the natural language description.

[0087] The embodiments of the present application provide a training method, an image generation method, a device, a device, a computer-readable storage medium, and a computer program product for an image generation model, which can improve the image generation effect. The following describes the exemplary applications of the electronic device provided by the embodiments of the present application. The electronic device provided by the embodiments of the present application can be implemented as various types of terminal devices such as a laptop computer, a tablet computer, a desktop computer, a set-top box, a smart phone, a smart speaker, a smart watch, a smart TV, a vehicle-mounted terminal, etc., or can also be implemented as a server.

[0088] See Figure 1 , Figure 1 which is a schematic structural diagram of the image generation system architecture provided by the embodiments of the present application. Figure 1 It involves a server 100, a terminal device 200, and a network 300. The terminal device 200 is connected to the server 100 through the network 300. Among them, the network 300 can be a wide area network, a local area network, or a combination of the two.

[0089] In some embodiments, the image generation method provided in the embodiment of the present application can be implemented collaboratively by a server and a terminal device. For example, the terminal device 200 sends the description text to the server 100, and the server 100 receives the description text and the image generation model trained by the training method of the image generation model provided in the embodiment of the present application, performs image generation based on the description text, obtains multiple generated images, and sends the multiple generated images to the terminal device 200. Among them, the training method of the image generation model provided in the embodiment of the present application can be executed by the terminal device or the server, or can be executed collaboratively by the terminal device and the server (such as the terminal device sends a description text sample, annotated image sample and noise image sample to the server for the server to train the image generation model).

[0090] The image generation model provided in the embodiments of the present application can be applied to various scenarios that require text image processing, such as creative design scenarios, game development scenarios, etc., as illustrated below.

[0091] 1) Creative design scenarios. For example, a terminal device receives a text description input by a user (such as “a cyberpunk-style future city”), initiates a generation request to the server, and the server generates a high-definition image through an image generation model and returns it to the terminal device for the user to edit or use directly.

[0092] 2) Game development. For example, the terminal device receives the text requirements of "Medieval Dragon Knight and Volcano Scene" submitted by the designer to the cloud collaboration platform. The server of the collaboration platform generates character / scene concept maps through the image generation model and synchronizes them to the terminal devices of team members to support real-time feedback and iteration.

[0093] 3) Business and marketing, for example, the server processes advertising copy (such as "Summer Beach Drinks") in batches through the image generation model, generates multiple versions of generated images and stores them in the cloud. Marketers browse cloud web pages or applications through terminal devices to preview generated images, select images and download them to the local computer for delivery.

[0094] 4) Personalized services, such as the server generating an image of the virtual image based on the user's description of "pink short hair, mechanical prosthetic eyes" and storing it in the user's account. The user can view the image of the virtual image through the terminal device and synchronize it to social media or game platforms.

[0095] In other embodiments, the image generation method provided in the embodiment of the present application can be implemented by a terminal device or a server alone. The terminal device 200 or the server 100 calls a local image generation model, performs image generation processing based on the description text through the image generation method provided in the embodiment of the present application, and obtains multiple generated images, wherein the image generation model is trained by the training method of the image generation model provided in the embodiment of the present application.

[0096] Here, the server 100 can be a single server. In this case, the training method and the image generation method of the image generation model provided in the embodiments of the present application can be implemented by the same server. The server 100 can also be a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. In this case, the training method and the image generation method of the image generation model provided in the embodiments of the present application can be implemented by different servers, and the embodiments of the present application do not make any limitations.

[0097] Taking the server for training the image generation model as an example, refer to Figure 2A , Figure 2A which is the first structural schematic diagram of the server provided in the embodiments of the present application. Figure 2A The server 100-1 shown in the figure includes at least one processor 110-1, a memory 130-1, and at least one network interface 120-1. Each component in the server 100-1 is coupled together through a bus system 140-1. It can be understood that the bus system 140-1 is used to implement the connection and communication between these components. In addition to the data bus, the bus system 140-1 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2A all kinds of buses are labeled as the bus system 140-1.

[0098] The processor 110-1 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a Digital Signal Processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0099] The memory 130-1 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disc drives, etc. The memory 130-1 optionally includes one or more storage devices that are physically located far from the processor 110-1.

[0100] The memory 130-1 includes a volatile memory or a non-volatile memory, and may also include both a volatile memory and a non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 130-1 described in the embodiments of the present application is intended to include any suitable type of memory.

[0101] In some embodiments, the memory 130-1 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which will be exemplarily described below.

[0102] The operating system 131-1 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0103] The network communication module 132-1 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 120-1. Exemplary network interfaces 120-1 include: Bluetooth, Wi-Fi (Wireless Fidelity), and Universal Serial Bus (USB), etc.;

[0104] In some embodiments, the device provided in the embodiments of the present application can be implemented in software. Figure 2A The training device 133 of the image generation model stored in the memory 130-1 is shown, which may be software in the form of programs and plugins, etc., and includes the following software modules: a data processing module 1331 and a training module 1332. These modules are logical, so they can be arbitrarily combined or further split according to the functions to be implemented. The functions of each module will be described below.

[0105] Taking the server for image generation as an example, see Figure 2B , Figure 2B is the second structural schematic diagram of the server provided by the embodiments of the present application. Figure 2B The server 100-2 shown includes: at least one processor 110-2, a memory 130-2, and at least one network interface 120-2. Each component in the server 100-2 is coupled together through a bus system 140-2. It can be understood that the bus system 140-2 is used to realize the connection and communication between these components. The bus system 140-2 includes, in addition to a data bus, a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 2BEach of the various buses is labeled as bus system 140-2. For the specific descriptions of the processor 110-2 and the memory 130-2, please refer to the above, and details will not be repeated here.

[0106] In some embodiments, the device provided by the embodiments of the present application can be implemented in software. Figure 2B An image generation device 134 stored in the memory 130-2 is shown. It can be software in the form of a program and a plug-in, etc., including the following software modules: a data acquisition module 1341 and an image generation module 1342. These modules are logical, so they can be combined arbitrarily or further split according to the functions to be implemented. The functions of each module will be described below.

[0107] In some embodiments, the terminal device or the server can implement the training method and the image generation method of the image generation model provided by the embodiments of the present application by running various computer-executable instructions or computer programs. For example, the computer-executable instructions can be commands at the microprogram level, machine instructions, or software instructions. The computer program can be a native program or a software module in the operating system; it can be a local (Native) application program (APPlication, APP), that is, a program that needs to be installed in the operating system to run; it can also be a small program that can be embedded in any APP, that is, a program that only needs to be downloaded to the browser environment to run. In short, the above computer-executable instructions can be instructions in any form, and the above computer programs can be application programs, modules, or plug-ins in any form.

[0108] In other embodiments, the device provided by the embodiments of the present application can be implemented in hardware. As an example, the device provided by the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the training method and the image generation method of the image generation model provided by the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (Application Specific Integrated Circuit, ASIC), digital signal processors (Digital Signal Processor, DSP), programmable logic devices (Programmable Logic Device, PLD), complex programmable logic devices (Complex Programmable Logic Device, CPLD), field-programmable gate arrays (Field-Programmable Gate Array, FPGA), or other electronic components.

[0109] Next, in combination with the exemplary applications and implementations of the server provided in the embodiments of the present application, taking the server as the execution subject, the training method of the image generation model provided in the embodiments of the present application will be described.

[0110] First, the basic training principle of the image generation model will be introduced. Refer to Figure 3 , Figure 3 which is a schematic diagram of the principle of the training method of the image generation model provided in the embodiments of the present application. First, after adding noise to the preset labeled image sample for the description text sample, a noisy image sample is obtained. Next, through the image generation model to be trained, at least one iteration round of processing is performed based on the noisy image sample and the description text sample. The processing of each iteration round includes: denoising the image to be denoised to obtain a denoised image; performing multiple samplings based on the conditional probability distribution of the denoised image to obtain multiple sampled images; screening the multiple sampled images to obtain an image sample pair, where the image sample pair includes a positive sampled image sample and a negative sampled image sample; using the denoised image of the last iteration round as the target image sample, or using the positive sampled image sample of the last iteration round as the target image sample. Among them, when the iteration round is the first iteration, the image to be denoised is the noisy image sample. When the iteration round is any iteration round after the first iteration, the image to be denoised can be the denoised image of the previous iteration round, or the image sample pair of the previous iteration round. Then, based on the target image sample and the labeled image sample, a first loss value is confirmed, and based on the image sample pair of each iteration round, a second loss value is determined. Finally, based on the first loss value and the second loss value, the parameters of the image generation model to be trained are updated to obtain the trained image generation model.

[0111] Next, the specific training method of the image generation model will be introduced. Refer to Figure 4A , Figure 4A which is the first process schematic diagram of the training method of the image generation model provided in the embodiments of the present application, and will be described in combination with Figure 4A the steps shown.

[0112] In step 101, through the initial image generation model, at least one iteration round is executed to generate a target image sample.

[0113] In some embodiments, refer to Figure 4B , in at least one iteration round, the following steps 1011 to 1015 are sequentially executed, which will be specifically described below.

[0114] In step 1011, the image to be denoised is denoised based on the description text sample to obtain a denoised image.

[0115] Here, the image to be denoised is a noisy image sample in the first iteration round, and in each iteration round after the first iteration round, it includes any one of the image sample pair, the denoised image, and the positive sampled image sample obtained in the previous iteration round.

[0116] Exemplarily, the core elements for describing a text sample may include the following parts:

[0117] The main object, such as a person, a scene, or an object (such as "a Shiba Inu wearing glasses"), etc., can refine the features of the main object, such as the color, shape, material, etc. of the main object (such as "the butterfly wings made of blue glass material").

[0118] The environment and background, such as time, place, lighting (such as "the desert at sunset, long shadows, warm tones"), weather, or atmosphere, etc. (such as "the neon city in the heavy rain, cyberpunk style").

[0119] The artistic style, such as the painting style (such as "ink painting", "cartoon", "Ukiyo-e print", etc.).

[0120] The composition and perspective, such as the lens type (such as "wide-angle lens", "bird's-eye view"), the aspect ratio of the picture (such as "a 16:9 movie screen", "square composition"), etc.

[0121] Additional modifiers, such as keywords for enhancing details (such as "ultra-high resolution", "rendered with Unreal Engine"), excluding unwanted content (such as "no text", "no blurred background"), etc.

[0122] In some embodiments, the noisy image sample is obtained by adding noise to a pre-set annotated image sample for the text sample.

[0123] Exemplarily, first, the annotated image sample is processed by a pre-trained image encoder (such as a Variational Auto-Encoder (VAE)) to perform image encoding on the annotated image sample, and the obtained image features are used as the annotated image features.

[0124] The variational autoencoder (VAE) can be trained as follows: The image encoder converts the input image into the mean and variance in the latent space to describe the distribution of latent variables, samples latent variables from this distribution (specific operation: mean + random noise × standard deviation) to ensure gradient propagation. The image decoder restores the latent variables to an image, calculates the pixel-level error (such as mean squared error) with the original image, which is called the reconstruction loss, and calculates the relative entropy (Kullback-Leibler Divergence, KL divergence) between the two. The total loss is the weighted sum of the reconstruction loss and the KL divergence. During training, the parameters of the image encoder and the image decoder are adjusted simultaneously through backpropagation to obtain the trained image encoder and image decoder.

[0125] Next, a seed is randomly generated, and a Gaussian noise matrix is generated based on the seed.

[0126] Here, the seed is the initial value of the random number generator, which is used to control the randomness of noise generation. The same noise will be generated based on the same seed, and different seeds will generate different noises.

[0127] Next, the Gaussian noise matrix is superimposed on the labeled image features to obtain the diffusion starting point.

[0128] See Figure 6 , Figure 6 which is a schematic diagram of the network structure of the image generation model provided in the embodiment of the present application. After image encoding the labeled image sample X w and superimposing the Gaussian noise matrix (i.e., the random noise map), the diffusion starting point Z0 is obtained. Based on the diffusion starting point Z0, forward diffusion is performed to obtain the noise image feature Z T .

[0129] Denote the labeled image feature output by the image encoder as E i , and denote the generated Gaussian noise matrix as ∈0. Then the diffusion starting point can be expressed by formula (1):

[0130] Z0 = E i + ∈0 (1)

[0131] Finally, based on the diffusion starting point, forward diffusion is performed to obtain the noise image feature, and the noise image feature is used as the noise image sample, or the image obtained by performing image decoding processing on the noise feature through the image decoder is used as the noise image sample.

[0132] Forward diffusion is performed for T time steps based on the diffusion starting point (i.e., superimposing T time steps to obtain the noise image sample Z of pure noise T),The transition probability distribution for each time step can be expressed as Equation (2):

[0133]

[0134] where, represents the normal distribution, I represents the identity matrix, and β t represents the noise level at the t-th step. The scheduling parameter α for each time step t = 1 - β t , where α t decays with the time step. For example, α1 = 0.9, α2 = 0.6, α3 = 0.5. The diffusion result at the t-th time step can be expressed as: where, The noise image feature Z T (i.e., the noise image sample) is obtained through forward diffusion over T time steps.

[0135] In some embodiments, referring to Figure 4C , denoising the image to be denoised based on the description text sample to obtain a denoised image can be achieved through the following steps 201 to 205, which will be specifically described below.

[0136] Exemplarily, referring to Figure 6 , denoising the image to be denoised can be achieved through a denoising network. The denoising network includes an encoder and a decoder. Among them, the encoder corresponds to Figure 6 the feature fusion - 1 and feature fusion - 2 with step - by - step downsampling shown in Figure 6 , and the decoder corresponds to the feature fusion - 3 and feature fusion - 4 with step - by - step upsampling shown in

[0137] In step 201, the description text sample is encoded for text features to obtain text features.

[0138] In some embodiments, first, the description text sample is tokenized to obtain a plurality of input units (tokens).

[0139] Exemplarily, punctuation marks such as spaces, full stops, commas, etc. can be used as tokenization markers to tokenize the description text sample (Tokenization), that is, the description text sample is segmented into a plurality of input units.

[0140] Exemplarily, the description text sample can be tokenized through a tokenizer based on Byte - Pair Encoding (BPE), a tokenizer based on Byte - Level Byte Pair Encoding, etc. The specific tokenizer (Tokenizer) or tokenization method used for tokenizing the description text sample is not limited in the embodiments of the present application.

[0141] Next, perform word embedding processing on multiple input units to obtain the word embedding features of each input unit.

[0142] Exemplarily, performing word embedding processing on multiple input units to obtain the word embedding features of each input unit can be achieved through the following steps: First, construct a vocabulary that contains all the words or tokens that appear in the text. For each word or token in the vocabulary, convert it into a vector of a fixed size. This process is called word embedding (Word Embeddings). The dimension of the word embedding is a hyperparameter, and values such as 50, 100, 300, etc. can be selected. The higher the embedding dimension, the richer the information that the model can capture, but the higher the computational cost. By training models such as Word2Vec and GloVe, the embedding matrix of words or tokens can be learned, so that each input unit can be converted into the corresponding word embedding feature through the word embedding model.

[0143] Next, perform position encoding processing on multiple input units to obtain the position encoding features of each input unit.

[0144] Exemplarily, performing position encoding (Position Embeddings) processing on multiple input units to obtain the position encoding features of each input unit, the position encoding processing can be generated through a fixed algorithm. For example, it can be implemented using a combination of sine and cosine functions. For the dimension index i of the position encoding feature of each position (the position of each input unit in the description text sample), (for example, the number of dimensions of the position encoding feature of each token is d, then the range value of the dimension index i is from 0 to (d - 1)), the position encoding of the dimension where i is even uses the sine function, and the position encoding of the dimension where i is odd uses the cosine function. The specific implementation method of the position encoding processing in the embodiments of the present application is not limited. Here, position encoding processing is performed because when converting the description text sample into a numerical representation (word embedding feature), the sequential information in the text will be lost. The purpose of position encoding is to encode the sequential information of the tokens.

[0145] Next, fuse the word embedding feature and the position encoding feature of each input unit to obtain the fused feature of each input unit.

[0146] Exemplarily, add the word embedding feature and the position encoding feature of each input unit to obtain the fused feature of each input unit.

[0147] Next, perform attention encoding processing (such as self-attention mechanism, multi-head attention mechanism, etc.) on the fused feature of each input unit to obtain the attention encoding feature of each input unit.

[0148] Exemplarily, before performing attention encoding, the fused features of each input unit can also be normalized (e.g., layer normalization (Layer Norm)) to obtain the normalized features of each input unit. For the normalized features of each input unit, the following attention encoding process is performed: First, a linear transformation is performed to generate three matrices: a query matrix (Query, Q), a key matrix (Key, K), and a value matrix (Value, V). The linear transformation is implemented by learnable weight matrices W Q , W K and W V . Next, the dot product of Q and K is calculated to obtain attention scores (here, the dot product operation is performed between the query matrix of the currently processed input unit and the key vectors of all input units (including the currently processed input unit) respectively to obtain the attention scores corresponding to all input units). Next, the attention scores are normalized, for example, by applying a normalization function (such as the softmax function) to make the attention scores become a probability distribution, and V is weighted using the attention probability distribution to obtain a rich representation containing the correlations of input units at different positions, that is, the attention encoding features of the input units.

[0149] Exemplarily, through a Transformer structure, the fused features of each input unit can be processed for attention encoding to obtain text features. Here, the Transformer structure can be stacked in multiple layers for attention encoding processing. The embodiments of the present application do not limit the number of layers of the specific Transformer structure.

[0150] Finally, the attention encoding features of each input unit are combined to obtain the text features describing the text sample.

[0151] Exemplarily, the attention encoding features of each input unit are concatenated into text features.

[0152] In step 202, the denoising image features of the image to be denoised are obtained.

[0153] In some embodiments, when the iteration round is the first iteration, the image to be denoised is a noisy image sample. When the iteration round is any iteration round after the first iteration, the image to be denoised is the positive sampling image sample and the negative sampling image sample of the previous iteration round.

[0154] Exemplarily, referring to Figure 8A , Figure 8A which is the first schematic diagram of the iterative process provided by the embodiments of the present application, Figure 8A shows the iterative process when T = 2 (performing 2 iteration rounds) and K = 3 (performing 3 samplings based on the conditional probability distribution of the denoised image at each stage). In the first iteration round, the denoising image features of the image to be denoised correspond toFigure 8A The noise image feature Z in T , in the last iteration round, the image to be denoised is the positive sampling image sample - 1 and the negative sampling image sample - 1, and the corresponding features of the image to be denoised correspond to Figure 8A the Z corresponding to the positive sampling image sample - 1 in T-1′ and the Z corresponding to the negative sampling image sample - 1 T-1′ .

[0155] In step 203, feature fusion is performed based on the text feature and the feature of the image to be denoised to obtain a fused feature.

[0156] Exemplarily, referring to Figure 7A , Figure 7A which is the first schematic diagram of the feature fusion principle provided by the embodiments of the present application. Taking the query matrix (Q, i.e., the feature of the image to be denoised) as the input, it is processed through Residual Block 1 (corresponding to steps 2031 to 2032 below), Spatial Transformer 1 (corresponding to step 2033 below), and Residual Block 2. The output of Residual Block 2 is used as the query matrix to be combined with the text feature (as the key vector and value vector) and input into Spatial Transformer 2 for cross - attention processing (corresponding to step 2034 below), and the processing result is downsampled (corresponding to step 2035 below) to obtain a fused feature.

[0157] Here, the network structure of the Residual Block can be referred to Figure 7B , and the network structure of the Spatial Transformer can be referred to Figure 7C , and the principle will be described in detail below.

[0158] In some embodiments, referring to Figure 4D , Figure 4C step 203 shown can be implemented through the following steps 2031 to 2035, which will be specifically described below.

[0159] In step 2031, the time step of the iteration round is embedded - encoded to obtain a time - step embedding.

[0160] It should be noted that the time step (t) here represents the stage where the current iteration round is in the Reverse Process.

[0161] In some embodiments, embedding encoding is used to map the original numerical data (i.e., time steps) into a low-dimensional vector space so that the image generation model can better process this data. Specifically, before embedding encoding, it is necessary to determine the number of embedding dimensions (a preset value), that is, the length of the feature vector to which the time steps will be mapped. Next, an embedding matrix is initialized. The size of the matrix is V×D, where V is the size of the input space (e.g., the total number of time steps), and D is the embedding dimension (i.e., the vector length of the time step embedding). The time steps are mapped to the corresponding row vectors in the embedding matrix to obtain the time step embeddings.

[0162] Exemplarily, assume the number of embedding dimensions is 2 and the embedding matrix is (Through a matrix learned by pre-training, the initialization of the embedding matrix can be random initialization). The embedding matrix represents mapping the time step "1" to the values "0.1, 0.2" corresponding to two dimensions, mapping the time step "2" to the values "0.3, 0.4" corresponding to two dimensions, mapping the time step "3" to the values "0.5, 0.6" corresponding to two dimensions. Using the value of the time step as an index, the corresponding row vector is looked up from the embedding matrix to obtain the time step embedding. For example, the time step embedding of the time step "1" can be represented as [0.1, 0.2].

[0163] In some other embodiments, the time steps can also be embedded and encoded through one-hot encoding to obtain the time step embeddings.

[0164] In some other embodiments, the sine representation of the time steps can also be obtained through sine encoding, and then feature mapping is performed through a Multilayer Perceptron (MLP) to obtain the time step embeddings.

[0165] In step 2032, feature residual processing is performed on the time step embeddings and the image features to be denoised to obtain residual features.

[0166] In some embodiments, performing feature residual processing on the time step embeddings and the image features to be denoised to obtain residual features can be achieved through Figure 7A the residual block 1 in. Refer to Figure 7B , Figure 7B is the second schematic diagram of the feature fusion principle provided by the embodiments of the present application, Figure 7B showing a schematic diagram of the network structure of the residual block.

[0167] First, mapping processing is performed on the time step embeddings to obtain the time step mapping features.

[0168] Exemplarily, refer to Figure 7B, the time-step embeddings are mapped through a Dense Layer to obtain time-step mapping features.

[0169] Next, the image features to be denoised are convolved to obtain first convolutional features.

[0170] Exemplarily, refer to Figure 7B , the image features to be denoised are convolved through a convolutional layer to obtain first convolutional features,

[0171] Next, the time-step mapping features and the first convolutional features are fused to obtain first fusion features.

[0172] Exemplarily, refer to Figure 7B , the time-step mapping features and the first convolutional features are added (element-wise addition) to obtain first fusion features. Here, the time-step mapping features and the first convolutional features have the same dimension, and the dimensions of the time-step mapping features and the first convolutional features can be unified into the same dimension through padding.

[0173] Next, the first fusion features and the image features to be denoised are fused to obtain second fusion features.

[0174] Exemplarily, refer to Figure 7B , a skip connection is performed on the first fusion features and the image features to be denoised, that is, the first fusion features and the image features to be denoised are element-wise added to obtain second fusion features

[0175] Next, the second fusion features are convolved to obtain second convolutional features.

[0176] Exemplarily, refer to Figure 7B , the second fusion features are convolved through a convolutional layer to obtain second convolutional features.

[0177] Finally, the second convolutional features are used as residual features.

[0178] In step 2033, self-attention encoding is performed on the residual features to obtain self-attention features.

[0179] In some embodiments, a linear transformation is performed on the residual features to obtain three matrices, namely the query matrix (Q), the key matrix (K), and the value matrix (V). The linear transformation is performed by a learnable weight matrix W QImplementation (i.e., Q = K = V). Next, calculate the dot product of Q and K to obtain the attention scores. Next, normalize the attention scores, for example, by applying a normalization function (such as the softmax function), to make the attention scores become a probability distribution (i.e., the attention weight matrix). Use the attention probability distribution to weight V to obtain the self-attention features.

[0180] Exemplarily, refer to Figure 7A , perform self-attention encoding on the residual features to obtain self-attention features, which can be achieved through the spatial transformer 1. After obtaining the self-attention features, the output of the spatial transformer 2 can be further processed through the residual block 2 (refer to the processing in step 2032). Then, use the output of the residual block 2 (the new self-attention features) as the input of the spatial transformer 2 (i.e., the input of step 2034).

[0181] In step 2034, perform cross-attention encoding on the self-attention features and the text features to obtain cross-attention features.

[0182] In some embodiments, performing cross-attention encoding on the self-attention features and the text features to obtain cross-attention features can be achieved in the following way:

[0183] First, perform convolution processing on the self-attention features to obtain the third convolution features.

[0184] Exemplarily, refer to Figure 7C , perform convolution processing on the output of the residual block 2 shown in Figure 7A through a convolutional layer to obtain the third convolution features.

[0185] Next, perform feature projection processing on the third convolution features to obtain the feature projection results.

[0186] Exemplarily, refer to Figure 7C , perform feature projection on the third convolution features (such as linear projection (LinearProjection), multi-layer perceptron (MLP) projection, etc.) to obtain the feature projection results.

[0187] For example, after multiplying the weight matrix of the linear layer with the third convolution features, add the product result to the bias matrix of the linear layer as the feature projection result.

[0188] Next, perform a linear transformation on the feature projection results to obtain the query matrix.

[0189] Exemplarily, refer to Figure 7C , perform a linear transformation on the feature projection results through the learnable weight matrix W of the dense layer - 1 Q , (such as multiplying W QPerform matrix multiplication with the feature projection result to obtain a query matrix (Q).

[0190] Next, perform two linear transformations on the text features to obtain a key matrix and a value matrix respectively.

[0191] For example, refer to Figure 7C , and perform a linear transformation on the text features through the learnable weight matrix W K of the dense layer - 2 to obtain a key matrix (K); perform a linear transformation on the text features through the learnable weight matrix W V of the dense layer - 3 to obtain a value matrix (V).

[0192] Next, perform a dot product on the query matrix and the key matrix to obtain a first dot product result.

[0193] For example, refer to Figure 7C , perform a dot product (Dot Product) on the query matrix and the key matrix to obtain a first dot product result.

[0194] Next, scale the first dot product result to obtain a scaled result.

[0195] For example, refer to Figure 7C , scale the first dot product result (for example, divide the first dot product result by a scaling factor where d k is the number of dimensions of the query matrix and the key matrix) to obtain a scaled result.

[0196] Next, normalize the scaled result to obtain an attention weight matrix.

[0197] For example, refer to Figure 7C , normalize the first dot product result through the softmax function to obtain an attention weight matrix.

[0198] Next, perform a dot product on the attention weight matrix and the value matrix to obtain a second dot product result.

[0199] For example, refer to Figure 7C , perform a dot product (matrix multiplication) on the attention weight matrix and the value matrix to obtain a second dot product result.

[0200] Finally, perform a convolution operation on the second dot product result to obtain cross - attention features.

[0201] For example, refer to Figure 7C , perform a convolution operation on the second dot product result through a convolutional layer to obtain cross - attention features.

[0202] In step 2035, downsample the cross - attention features to obtain fused features.

[0203] In some embodiments, the cross-attention features are downsampled through a pooling operation to obtain fused features.

[0204] Exemplarily, the pooling operation can be max pooling, average pooling, etc., and the specific pooling operation is not limited in the embodiments of the present application.

[0205] In other embodiments, the cross-attention features are downsampled through strided convolution to obtain fused features, that is, a convolution operation with a stride greater than 1 is used, and downsampling can be achieved while convolving.

[0206] It should be noted that steps 2031 to 2035 correspond to Figure 6 the processing of feature fusion - 1 shown in, and the processing of feature fusion - 2, feature fusion - 3, and feature fusion - 4 is similar to that of feature fusion - 1 (see Figure 7A the processing in, the difference being that in the processing of feature fusion - 3 and feature fusion - 4, downsampling is changed to upsampling (e.g., achieved through deconvolution processing), and no repeated description is provided here.

[0207] Continuing to refer to Figure 4C , in step 204, a conditional probability distribution is predicted based on the fused features to obtain a conditional probability distribution.

[0208] In some embodiments, the fused features are feature-mapped to obtain a conditional probability distribution.

[0209] Exemplarily, the fused features are convolved through a convolutional layer (such as a convolutional layer with a kernel size of 1×1) to map the fused features into a conditional probability distribution.

[0210] Exemplarily, the obtaining of the conditional probability distribution is the inverse process of formula (2) and can be represented by formula (3):

[0211]

[0212] where the mean μ θ (Z t ,t) and the variance ∑ θ (Z t ,t) are learned through a denoising network, that is, when predicting a given Z t , the mean and variance of Z t-1 are predicted.

[0213] In step 205, a Gaussian distribution is sampled based on the conditional probability distribution to obtain a denoised image.

[0214] Continuing with the above example, according to the conditional probability distribution p θ (Z t-1 |Z t )'s mean and variance, perform random sampling following a Gaussian distribution to obtain the denoised image Z t-1 .

[0215] For example, according to the mean and variance, generate random numbers that conform to a Gaussian distribution at each pixel position of the denoised image, and use the random numbers at each pixel position as the pixel values of the sampled image at the corresponding pixel position to obtain the denoised image.

[0216] It should be noted that during the training process of the image generation model, training can be based on multiple noise image samples. For the convenience of description, only the processing of one noise image sample is exemplarily described, and it will not be repeated below.

[0217] Continuing to refer to Figure 4B , in step 1012, sample the denoised image to obtain multiple sampled images.

[0218] In some embodiments, when the iteration round is the first iteration round, the image to be denoised is a noise image sample. When the iteration round is any iteration round after the first iteration round, the image to be denoised is the positive sampled image sample and the negative sampled image sample of the previous iteration round; sample based on the conditional probability distribution of at least one denoised image of the iteration round to obtain multiple sampled images of the iteration round, where the denoised image of the last iteration round is the target image sample.

[0219] Exemplarily, let t be an integer that decreases from T, the minimum value of t is 1, and T is the total number of multiple iteration rounds. Perform the following processing on the t-th iteration round: denoise the image to be denoised in the t-th iteration round to obtain the denoised image of the t-th iteration round; when t = T, the image to be denoised in the t-th iteration round (i.e., the first iteration round) is a noise image sample. When t ≤ T - 1, the image to be denoised in the t-th iteration round is the positive sampled image sample and the negative sampled image sample of the previous iteration round; sample the denoised image of the t-th iteration round to obtain multiple sampled images of the t-th iteration round, where the denoised image of the last iteration round is the target image sample after removing noise.

[0220] Exemplarily, referring to Figure 8A , Figure 8A shows the iteration process when T = 2 (performing 2 iteration rounds) and K = 3 (sampling 3 times based on the conditional probability distribution of the denoised image in each stage). In the first iteration round, first, perform denoising on Z T (the image to be denoised) once to obtain Z T-1 (the denoised image), and next, based on ZT-1 Perform 3 samplings on the conditional probability distribution (or distribution parameters) to obtain 3 Zs T-1′ , and finally, determine the image sample pairs from the 3 Zs T-1′ (corresponding to S1 in Figure 8A to obtain the positive sampled image sample - 1 and the negative sampled image sample - 1). In the second iteration round, that is Figure 8A the last iteration round in , first, perform denoising on the positive sampled image sample - 1 and the negative sampled image sample - 2 (the image to be denoised) once respectively to obtain the Zs corresponding to the positive sampled image sample - 1 and the negative sampled image sample - 2 respectively T-2 (denoised images), next, based on the conditional probability distribution of each Z T-2 perform 3 samplings to obtain 3 Zs corresponding to each Z T-2 , and finally, determine the image sample pairs from a total of 6 Zs T-2′ (corresponding to S2 in T-2′ to obtain the positive sampled image sample - 2 and the negative sampled image sample - 2). Take the Z Figure 8A obtained by denoising the positive sampled image sample - 1 (Z T-1′ ) in the last iteration round as the target image sample (the Z T-2 here is the feature vector, and the target image sample Y T-2 can be obtained by decoding Z Figure 6 through the image decoder shown in T-2 ). w )

[0221] By performing denoising on the image to be denoised in each iteration round, the noise in the image is gradually reduced, thereby improving the quality of the generated image. By sampling based on the conditional probability distribution of the denoised image, multiple sampled images with different features can be generated, increasing the diversity of image samples and helping the image generation model learn a more comprehensive data distribution

[0222] In some embodiments, when the iteration round is the first iteration round, perform multiple Gaussian distribution samplings based on the conditional probability distribution of the denoised image to obtain multiple sampled images, where the denoised image is obtained by denoising the noise image sample as the image to be denoised; when the iteration round is any iteration round after the first iteration round, perform multiple Gaussian distribution samplings based on the conditional probability distribution of the first denoised image to obtain multiple sampled images of the first denoised image, and perform multiple Gaussian distribution samplings based on the conditional probability distribution of the second denoised image to obtain multiple sampled images of the second denoised image, where the first denoised image is the image obtained by denoising the positive sampled image sample as the image to be denoised, and the second denoised image is the image obtained by denoising the negative sampled image sample as the image to be denoised

[0223] For example, refer to Figure 8A , in the first iteration round, after performing denoising on Z T (image to be denoised) once to obtain Z T-1 (denoised image), based on the conditional probability distribution of Z T-1 , perform 3 samplings (corresponding to the multiple Gaussian distribution samplings above), to obtain 3 Z T-1′ (that is, one Gaussian sampling obtains one sampled image). In the second iteration round, that is Figure 8A , in the last iteration round in T-2 , perform 3 samplings based on the conditional probability distribution of Z T-2′ (the first denoised image) obtained by denoising the positive sampled image sample - 1, to obtain 3 Z T-2 of the first denoised image, and perform 3 samplings based on the conditional probability distribution of Z T-2′ (the second denoised image) obtained by denoising the negative sampled image sample - 1, to obtain 3 Z

[0224] In some embodiments, refer to Figure 4E , perform multiple Gaussian distribution samplings based on the conditional probability distribution of the denoised image to obtain multiple sampled images, which can be implemented through the following steps 301 to 303, and the following is a specific description.

[0225] In step 301, according to the mean and variance of the conditional probability distribution of the denoised image, generate random numbers that conform to the Gaussian distribution at each pixel position of the denoised image.

[0226] Continuing the above example, according to the mean (μ θ (Z t ,t)) and variance (∑ θ (Z t ,t)) in formula (3), generate random numbers that conform to the Gaussian distribution at each pixel position of the denoised image.

[0227] For example, random numbers can be obtained by using the Inverse Transform Method. The probability density function of the random number x can be expressed as formula (4):

[0228]

[0229] The cumulative distribution function (CDF) can be calculated through f(x), and then the inverse function F -1 (u) of the CDF can be obtained. Generate a random variable U ∼ Uniform(0,1), that is, generate a random variable U that is uniformly distributed in the interval [0, 1]. The random number x = F -1(U).

[0230] In step 302, the random number at each pixel position of the denoised image is used as the pixel value of the sampled image at the corresponding pixel position.

[0231] In some embodiments, if the generation range of the random number does not match the range of image pixel values (for example, a negative value or a non-integer value is generated), adjustments such as normalization and scaling are required to ensure that the pixel value is between 0 and 255 (applicable to images with 8-bit gray scale).

[0232] In step 303, a sampled image is generated based on the pixel values of the sampled image at the corresponding pixel positions.

[0233] In some embodiments, the random number is used as the pixel value of the sampled image at the corresponding pixel position to obtain the sampled image.

[0234] In other embodiments, when the iteration round is the first iteration round, the image to be denoised is a noise image sample. When the iteration round is any iteration round after the first iteration round, the image to be denoised is each sampled image of the previous iteration round; sampling is performed based on each denoised image of the iteration round to obtain multiple sampled images of the iteration round, where the denoised image of the last iteration round is the target image sample.

[0235] Exemplarily, refer to Figure 8B , Figure 8B which is the second schematic diagram of the iterative process provided by the embodiments of the present application, Figure 8B showing the iterative process when T = 2 (performing 2 iteration rounds) and K = 3 (performing 3 samplings based on the conditional probability distribution of the denoised image at each stage). In the first iteration round, first, Z T (the image to be denoised) is denoised once to obtain Z T-1 (the denoised image). Next, 3 samplings are performed based on the conditional probability distribution of Z T-1 to obtain 3 Z T-1′ . Finally, an image sample pair is determined from the 3 Z T-1′ (corresponding to S1 in Figure 8B ). The Z T-1′ with the highest quality parameter is used as the positive sampled image sample, and the Z T-1′ with the lowest quality parameter is used as the negative sampled image sample). In the second iteration round, that is, Figure 8B the last iteration round in T-1′ (the image to be denoised) is denoised once respectively to obtain the Z T-1′ respectively corresponding Z T-2 (the denoised image). Next, based on each Z T-2Perform 3 samplings on the conditional probability distribution to obtain each Z T-2 The corresponding 3 Zs T-2′ , finally, from the 9 Zs T-2′ Determine the image sample pair (corresponding to Figure 8B S2 in, and take the Z with the highest quality parameter T-2′ As the positive sampling image sample, and take the Z with the lowest quality parameter T-2′ As the negative sampling image sample). Take the positive sampling image sample of the last iteration round (i.e., Figure 8B The 9 Zs shown in T-2′ The Z with the highest quality parameter T-2′ ) corresponding Z T-2 As the target image sample (the Z T-2 Feature vector, which can be obtained by Figure 6 The image decoder shown in decodes Z T-2 To obtain the target image sample Y w ), or take the positive sampling image sample of the last iteration round as the target image sample.

[0236] In some other embodiments, when the iteration round is the first iteration round, the image to be denoised is a noise image sample. When the iteration round is any iteration round after the first iteration round, the image to be denoised is the positive sampling image sample of the previous iteration round; sample based on the denoised image of the iteration round to obtain multiple sampled images of the iteration round, where the denoised image of the last iteration round is the target image sample.

[0237] Exemplarily, refer to Figure 8C , Figure 8C Is the third schematic diagram of the iterative process provided by the embodiments of the present application, Figure 8C Shows the iterative process when T = 2 (performing 2 iteration rounds) and K = 3 (performing 3 samplings based on the conditional probability distribution of the denoised image in each stage). In the first iteration round, first, perform 1 denoising on Z T (The image to be denoised) to obtain Z T-1 (The denoised image). Next, perform 3 samplings based on the conditional probability distribution of Z T-1 To obtain 3 Zs T-1′ , finally, determine the image sample pair from the 3 Zs T-1′ (Corresponding to Figure 8C In S1 to obtain the positive sampling image sample - 1 and the negative sampling image sample - 1). In the second iteration round, that is Figure 8C The last iteration round in, first, perform 1 denoising on the positive sampling image sample - 1 (the image to be denoised) to obtain the Z corresponding to the positive sampling image sample - 1 T-2 (The denoised image). Next, based on Z T-2Perform 3 samplings on the conditional probability distribution to obtain 3 Zs T-2′ , finally, determine the image sample pairs from the 3 Zs T-2′ (corresponding to S2 in Figure 8C ) to obtain the positive sampled image sample - 2 and the negative sampled image sample - 2). Take the Z of the last iteration round T-2 as the target image sample, where the Z Figure 6 can be decoded by the image decoder shown in T-2 to obtain the target image sample Y w .

[0238] In some other embodiments, when the iteration round is the first iteration round, the image to be denoised is a noise image sample, and when the iteration round is any iteration round after the first iteration round, the image to be denoised is the denoised image of the previous iteration round; sample based on the denoised image of the iteration round to obtain multiple sampled images of the iteration round, where the positive sampled image sample of the last iteration round is the target image sample.

[0239] For example, refer to Figure 8D , Figure 8D is the fourth schematic diagram of the iteration process provided by the embodiments of the present application Figure 8D which shows the iteration process when T = 2 (perform 2 iteration rounds) and K = 3 (perform 3 samplings on the conditional probability distribution of the denoised image based on each stage). In the first iteration round, first, perform 1 denoising on Z T (the image to be denoised) to obtain Z T-1 (the denoised image), next, perform 3 samplings on the conditional probability distribution of Z T-1 to obtain 3 Zs T-1′ , finally, determine the image sample pairs from the 3 Zs T-1′ (corresponding to S1 in Figure 8D ) to obtain the positive sampled image sample - 1 and the negative sampled image sample - 1). In the 2nd iteration round, that is Figure 8D the last iteration round in T-1 (the image to be denoised) to obtain Z T-2 (the denoised image), next, perform 3 samplings on the conditional probability distribution of Z T-2 to obtain 3 Zs T-2′ , finally, determine the image sample pairs from the 3 Zs T-2′ (corresponding to S2 in Figure 8D ) to obtain the positive sampled image sample - 2 and the negative sampled image sample - 2). Take the positive sampled image sample - 2 of the last iteration round as the target image sample.

[0240] Through steps 1011 to 1012, the feature of generating images based on at least one iteration round is utilized, that is, the image generation result of each iteration round is generated based on the image generation result of the previous iteration round. Finally, the target image sample with noise removed is obtained. Multiple possible sampled images are generated for each iteration round, and the search tree construction of the image generation model for multi-step image generation is realized, so as to generate more accurate image sample pairs corresponding to each iteration round respectively, which are then used for the training of the image generation model.

[0241] Continue to refer to Figure 4B , in step 1013, an image sample pair is selected from multiple sampled images, where the image sample pair includes a positive sampled image sample and a negative sampled image sample.

[0242] In some embodiments, when the iteration round is the first iteration round, the following processing is performed: obtain the first quality parameter of each sampled image, use the sampled image with the highest first quality parameter among the multiple sampled images as the positive sampled image sample of the iteration round, and use the sampled image with the lowest first quality parameter among the multiple sampled images as the negative sampled image sample of the iteration round; when the iteration round is any iteration round after the first iteration round, the following processing is performed: obtain the second quality parameter of the sampled image of each first denoised image, and obtain the third quality parameter of the sampled image of each second denoised image. Use the sampled image with the highest second quality parameter among the multiple sampled images of the first denoised image as the positive sampled image sample of the iteration round, and use the sampled image with the lowest third quality parameter among the multiple sampled images of the second denoised image as the negative sampled image sample of the iteration round.

[0243] Exemplarily, refer to Figure 8A , when it is the first iteration round, among the 3 Zs obtained by sampling 3 times based on the conditional probability distribution of Z T-1 , use the Z with the highest first quality parameter T-1′ as the positive sampled image sample, and use the Z with the lowest first quality parameter T-1′ as the negative sampled image sample (corresponding to the positive sampled image sample -1 and negative sampled image sample -1 in T-1′ ). When it is the second iteration round, that is Figure 8A the last iteration round in Figure 8A ), sample 3 times based on the conditional probability distribution of each Z T-2 to obtain 3 Zs corresponding to each Z T-2 . Determine the image sample pair from the 6 Zs T-2′ T-2′ . Specifically, among the 3 Zs T-2′ sampled from the Z corresponding to the positive sampled image sample -1 T-2 , use the Z with the highest second quality parameter T-2′T-2′ As a positive sampling image sample (corresponding to the positive sampling image sample - 2 in Figure 8A ), and taking the Z obtained by sampling the negative sampling sample - 1 T-2 Among the 3 Zs obtained by sampling, T-2′ the Z with the lowest third quality parameter in T-2′ is used as a negative sampling image sample (corresponding to the negative sampling image sample - 2 in Figure 8A ).

[0244] In some embodiments, referring to Figure 4F , when the iteration round is the first iteration round, obtaining the first quality parameter of each sampling image can be implemented through the following steps 401 to 402, which are specifically described below.

[0245] In step 401, perform sampling image feature encoding on the sampling image to obtain sampling image features.

[0246] In some embodiments, through the image feature extraction network of the pre - trained evaluation model, perform sampling image feature encoding on the sampling image to obtain sampling image features.

[0247] Exemplarily, the image feature extraction network can be a convolutional neural network (such as a Residual Network (ResNet), etc.), and can be expressed as: f(x)=ResNet(x), where x represents the sampling image and f(x) represents the sampling image features.

[0248] In step 402, perform feature mapping on the sampling image features to obtain the first quality parameter of the sampling image.

[0249] In some embodiments, through the fully connected (FC) layer of the pre - trained evaluation model, perform feature mapping on the sampling image features to obtain the first quality parameter of the sampling image.

[0250] Exemplarily, performing feature mapping on the sampling image features to obtain the first quality parameter of the sampling image can be expressed as: α(x)=FC(f(x)), where α(x) represents the first quality parameter, f(x) represents the sampling image features, and FC represents the fully connected layer.

[0251] Exemplarily, the evaluation model is trained as follows: Obtain image samples and the annotation parameters corresponding to the image samples (or called aesthetic scores, which can be manually annotated). Here, publicly available datasets can be used, such as the Aesthetic Visual Analysis (AVA) dataset. Use a pre-trained image feature extraction network (or called backbone network, such as ResNet-152) to extract the features of the images, predict the aesthetic scores (i.e., quality parameters) of the image samples through a fully connected network, calculate the mean square error between the predicted aesthetic scores and the annotated aesthetic scores to obtain a loss value, obtain the gradient information of the loss value for each parameter of the evaluation model through the backpropagation algorithm, and update the parameters of the evaluation model using the obtained gradient information according to the gradient descent optimization algorithm (such as batch gradient descent, stochastic gradient descent, etc.). Repeat the above process until a certain number of iterations are reached or the evaluation model converges, thereby obtaining the trained evaluation model.

[0252] In some other embodiments, the evaluation model can use pre-trained aesthetic evaluation models, such as the Transformer for image quality (TRIQ) model, the Image Quality Transformer (IQT) model, the Human Preference Score (HPS) model, etc. The embodiments of the present application do not limit the specific evaluation model.

[0253] In some other embodiments, the pixel mean (average value of image pixels), standard deviation (i.e., the degree of dispersion of image pixel gray values relative to the mean), entropy (average information amount of the image), average gradient, etc. of the sampled image can also be used as the first quality parameter of the sampled image.

[0254] It should be noted that the methods for obtaining the second quality parameter, the third quality parameter, and the methods for obtaining the fourth quality parameter, the fifth quality parameter, and the sixth quality parameter below are the same as the implementation methods for obtaining the first quality parameter, and will not be repeated here.

[0255] Since the quality parameter is based on image features, it provides a relatively objective and accurate image quality evaluation method, rather than based on subjective visual judgment. Compared with the manual annotation method, it can reduce the influence of subjective factors and obtain accurate quality parameters.

[0256] In some other embodiments, from the multiple sampled images included in the image generation results of each iteration round, an image sample pair corresponding to each iteration round is selected, and it can also be implemented by performing the following processing for each iteration round: obtaining a fourth quality parameter of each sampled image of the iteration round; using the sampled image with the highest fourth quality parameter among the multiple sampled images as the positive sampled image sample of the iteration round, and using the sampled image with the lowest fourth quality parameter among the multiple sampled images as the negative sampled image sample of the iteration round.

[0257] For example, see Figure 8A , in the first iteration round, use Z with the highest fourth quality parameter T-1′ as the positive sampled image sample, and use Z with the lowest fourth quality parameter T-1′ as the negative sampled image sample (corresponding to Figure 8A S1 in to obtain the positive sampled image sample -1 and the negative sampled image sample -1), in the second iteration round, that is Figure 8A the last iteration round in, use Z with the highest fourth quality parameter T-2′ as the positive sampled image sample (corresponding to Figure 8A the positive sampled image sample -2 in), and use Z with the lowest fourth quality parameter T-2′ as the negative sampled image sample (corresponding to Figure 8A the negative sampled image sample -2 in).

[0258] In some other embodiments, from the multiple sampled images included in the image generation results of each iteration round, an image sample pair corresponding to each iteration round is selected, and it can also be implemented in the following way: obtaining a fifth quality parameter of each sampled image of the iteration round; using the sampled image with the highest fifth quality parameter among the multiple sampled images as the positive sampled image sample of the iteration round, and using the sampled image with the lowest fifth quality parameter among the multiple sampled images as the negative sampled image sample of the iteration round.

[0259] For example, see Figure 8C , in the first iteration round, use Z with the highest fifth quality parameter T-1′ as the positive sampled image sample, and use Z with the lowest fifth quality parameter T-1′ as the negative sampled image sample (corresponding to Figure 8C the positive sampled image sample -1 and the negative sampled image sample -1 in), in the second iteration round, that is Figure 8C the last iteration round in, use Z with the highest fifth quality parameter T-2′ as the positive sampled image sample (corresponding to Figure 8C the positive sampled image sample -2 in), and use Z with the lowest fifth quality parameter T-2′ as the negative sampled image sample (corresponding to Figure 8C the negative sampled image sample -2 in).

[0260] Through steps 1011 to 1013, by constructing a search tree and screening the sampled images based on quality parameters, better samples in terms of statistical probability are generated (i.e., the positive sampled image samples and negative sampled image samples in each iteration round obtained through evaluation by quality parameters), and image sample pairs are constructed, so that there are fewer deformed images in the training samples, the learning upper limit of the image generation model is improved, and thus the image generation effect is enhanced.

[0261] Continue to refer to Figure 4B , in step 1014, when the iteration round does not meet the iteration end condition, return to the step of denoising the image to be denoised based on the description text sample.

[0262] In some embodiments, the iteration end condition includes any one of the following: the current iteration round reaches a preset number of iterations; the quality parameter of the denoised image obtained in the current iteration round is greater than or equal to a preset quality parameter threshold; the quality parameter of the positive sampled image sample obtained in the current iteration round is greater than or equal to a preset quality parameter threshold, where the acquisition of the quality parameter can refer to the descriptions in steps 401 to 402 above.

[0263] Exemplarily, when the iteration round does not reach the preset number of iterations (such as 20 times), transfer to the process of denoising the image to be denoised based on the description text sample (i.e., step 1011 above).

[0264] In step 1015, when the iteration round meets the iteration end condition, use the denoised image as the target image sample, or use the positive sampled image sample as the target image sample.

[0265] Continuing the above example, when the iteration round reaches the preset number of iterations (such as 20 times), use any one of the denoised image or the positive sampled image sample obtained in the last iteration round as the target image sample.

[0266] Continue to refer to Figure 4A In step 102, based on the target image sample, the labeled image sample preset for the description text sample, and the image sample pairs in each iteration round, determine the loss value.

[0267] In some embodiments, the loss value includes a first loss value and a second loss value. Based on the target image sample and the labeled image sample preset for the description text sample, determine the first loss value, and based on the image sample pairs in each iteration round, determine the second loss value.

[0268] In some embodiments, based on a target image sample and an annotated image sample preset for a description text sample, a first loss value can be determined in the following manner: obtaining the pixel difference between the target image sample and the annotated image sample at each pixel position; obtaining the sum of the squares of the pixel differences at each pixel position, and dividing the sum of the squares by the total number of pixels to obtain the first loss value.

[0269] Exemplarily, the obtaining of the first loss value (mean squared error) can be represented by formula (5):

[0270]

[0271] where X W_i represents the pixel value of the annotated image sample X W at the i-th position, and Y W_i represents the pixel value of the target image sample Y W at the i-th position.

[0272] By supervising the target image sample and the real image (annotated image sample) with the first loss value, the target image sample is precisely aligned with the real image in terms of pixel values (i.e., the optimization objective of the first loss value is to minimize pixel error). Additionally, by directly comparing pixel differences, resource consumption is reduced.

[0273] In some embodiments, referring to Figure 9 , Figure 9 is a schematic diagram of the principle for obtaining the second loss value provided in an embodiment of the present application. For multiple iteration rounds (corresponding to Figure 9 t = T to t = 1 in Figure 6 ), the denoising network (referring to the denoising network shown in Figure 9 ) is used to denoise the image to be denoised in each iteration round to obtain a denoised image, multiple sampled images are obtained by sampling based on the conditional probability distribution of the denoised image, and an image sample pair for each iteration round (corresponding to Figure 9 image sample pair_T to image sample pair_1 in

[0274] is determined. The score of each image sample pair (or called reward score, corresponding to Figure 9 score_T to score_1 in

[0275] is obtained, and the average value of the scores of multiple iteration rounds is used as the second loss value, thereby performing backpropagation update on the denoising network based on the second loss value. The second loss value obtained based on the image sample pair in each iteration round enables stage-by-stage learning, which can enhance the ability of the image generation model to distinguish positive and negative sampled image samples, such that when updating the image generation model based on the second loss value, it can drive the parameters of the image generation model to generate in a direction consistent with the positive sampled image sample.

[0275] In some embodiments, referring toFigure 4G , based on the image sample pairs of each iteration round, determine the second loss value, which can be achieved by performing the following steps 1021 to 1025 on the image generation results of each iteration round. The following is a specific description.

[0276] In step 1021, obtain the sixth quality parameter of each sampled image, and obtain the first average value of the sixth quality parameters of multiple sampled images.

[0277] In some embodiments, for the implementation of obtaining the sixth quality parameter of each sampled image, reference can be made to the description of steps 401 to 402 above, which will not be elaborated here.

[0278] In step 1022, obtain the difference between the sixth quality parameter of the positive sampled image sample and the sixth quality parameter of the negative sampled image sample.

[0279] In some embodiments, obtain the difference between the highest sixth quality parameter (i.e., the sixth quality parameter of the positive sampled image sample) and the lowest sixth quality parameter (i.e., the sixth quality parameter of the negative sampled image sample).

[0280] In step 1023, obtain the ratio of the logarithm of the total time steps of at least one iteration round to the time steps of the iteration round, and obtain the square root of the ratio.

[0281] In some embodiments, obtain the square root of the ratio of the logarithm of the total time steps to the time steps of the current iteration round.

[0282] In step 1024, add the first average value, the product of the preset constant parameter and the square root, and the difference to obtain the score of the image sample pair.

[0283] In some embodiments, the score of the image sample pair can be represented by formulas (6) and (7):

[0284]

[0285] Where, represents the score of the positive sampled image sample in the image sample pair at the i-th step, represents the score of the negative sampled image sample in the image sample pair at the i-th step, represents the first average value of the fifth quality parameters of multiple sampled images at the i-th step, c is a constant (for example, c = 2 is taken), N represents the total time steps of multiple iteration rounds, n i represents the time steps of the current iteration round, d i represents the difference between the highest fifth quality parameter and the lowest fifth quality parameter at the i-th step.

[0286] In step 1025, obtain a second average value of the scores of the image sample pairs corresponding to at least one iteration round as the second loss value.

[0287] In some embodiments, obtain an average value of the scores of the positive sampled image samples corresponding to at least one iteration round (i.e., ) as the second loss value.

[0288] In other embodiments, obtain an average value of the scores of the sampled image sample pairs corresponding to at least one iteration round (i.e., sum) as the second loss value.

[0289] Through steps 1021 to 1025, an image generation model for generating images for at least one iteration round is implemented. The quality of the image sample pairs is evaluated in each iteration round (i.e., the quality parameters are obtained), and then the second loss value is obtained. Thus, for the multi-step image generation mode, the model training of multi-step reinforcement learning is adaptively performed using the second loss value to improve the performance of the image generation model.

[0290] Continue to refer to Figure 4A , in step 103, update the parameters of the initial image generation model based on the loss value to obtain a trained image generation model.

[0291] In some embodiments, refer to Figure 4H , Figure 4A shown in step 103, which can be implemented through the following steps 1031 to 1032. The following is a specific description.

[0292] In step 1031, obtain the product of the preset weight parameter and the first loss value, and obtain the sum of the product and the second loss value, and use the sum as the combined loss value.

[0293] In some embodiments, the combined loss value can be expressed by formula (8):

[0294] LOSS = L2 + aL1 (8)

[0295] where LOSS represents the combined loss value, L1 and L2 respectively represent the first loss value and the second loss value, and a represents the preset weight parameter, for example, a = 0.1.

[0296] In step 1032, update the parameters of the image generation model to be trained based on the combined loss value to obtain a trained image generation model.

[0297] In some embodiments, the gradient information of the combined loss value with respect to each parameter of the image generation model is obtained through the backpropagation algorithm. According to the gradient descent optimization algorithm (such as batch gradient descent, stochastic gradient descent, etc.), the obtained gradient information is used to update the parameters of the image generation model. The above process is repeated until a certain number of iterations is reached or the image generation model converges, thereby obtaining the trained image generation model.

[0298] In other embodiments, based on the second loss value, the parameters of the image generation model to be trained are updated to obtain the trained image generation model.

[0299] Through steps 101 to 103, an image generation model for the image generation mode of multiple iteration rounds (multiple time steps) is implemented. Image sample pairs for each iteration round are constructed in a targeted manner, and the quality of the image sample pairs is evaluated (i.e., the quality parameters are obtained) in each iteration round. Furthermore, the second loss value is obtained. Thus, for the multi-step image generation mode, the model training of multi-step reinforcement learning is adaptively performed using the second loss value. Through this step-by-step learning method, the generation effect of each step of the image generation model is improved, thereby improving the matching degree between the image generation model and the training data and enhancing the final image generation effect.

[0300] In other embodiments, after each training is completed, a batch of new images can be generated using the current generation model as the pseudo-labeled image samples corresponding to the description text samples. Thus, the image generation model is jointly trained in combination with the original training data. By dynamically expanding the scale of the training data during the training process, the diversity of the training data is improved, the model's over-reliance on the original training data is prevented, and the overfitting of the generation model to the original training data is reduced, achieving the beneficial effect of improving the generalization ability and robustness of the image generation model.

[0301] Next, the image generation method provided in the embodiments of the present application will be described in combination with the exemplary applications and implementations of the server provided in the embodiments of the present application. Taking the server as the execution subject, the image generation method provided in the embodiments of the present application will be described. Refer to Figure 5 , Figure 5 is a schematic flowchart of the image generation method provided in the embodiments of the present application, and will be described in combination with Figure 5 the steps shown.

[0302] In step 501, a description text is obtained.

[0303] In some embodiments, the text instruction directly input by the user on the front-end interface (such as a web page, application, etc.) is obtained as the description text (corresponding to the description text sample above). It is also possible to convert the user's voice instruction into text through Automatic Speech Recognition (ASR) as the description text. The embodiments of the present application do not limit the specific implementation manner of obtaining the description text.

[0304] For example, refer to Figure 11 , in response to the triggering operation on the text input control 001, the description text is obtained.

[0305] In step 502, multiple generated images are obtained by generating images based on the description text through the image generation model, where the image generation model is trained by the image generation model training method provided by the embodiments of the present application.

[0306] In some embodiments, multiple noise maps conforming to the Gaussian distribution (corresponding to the noise image sample above) are obtained, and image generation processing is performed on each noise map and the description text for at least one iteration round to obtain multiple generated images (that is, one noise map corresponds to one generated image).

[0307] For example, in at least one iteration round, the following steps are sequentially executed: denoising the image to be denoised based on the description text to obtain a denoised image; in the case where the iteration round does not meet the iteration end condition, return to the step of denoising the image to be denoised based on the description text; in the case where the iteration round meets the iteration end condition, use the denoised image as the generated image, where when the iteration round is the first iteration, the image to be denoised is the noise map, and when the iteration round is any iteration after the first iteration, the image to be denoised is the denoised image of the previous iteration round.

[0308] Here, the implementation manner of denoising the image to be denoised based on the description text to obtain a denoised image can refer to the description of steps 201 to 205 above, and will not be elaborated here.

[0309] For example, refer to Figure 11 , in response to the triggering operation on the image generation control 002, multiple generated images are obtained by generating images based on the description text through the image generation model.

[0310] For example, refer to Figure 11 , in response to the triggering operation on the image download control 003, the generated images are downloaded.

[0311] The image generation method provided by the embodiments of the present application can be applied to various scenarios that require image generation (text-to-image), such as creative design scenarios, game development scenarios, etc.

[0312] The following describes the application of the image generation method provided in the embodiments of the present application in the game development scenario.

[0313] In the related art, based on natural language descriptions, realistic images are generated, and then images are selected manually or using metrics. For example, after generating an image, the generation result is returned to the user, and the evaluation is obtained through user selection. For example, for the sentence in a martial arts novel "The spring water drips down, and the water droplets strike the mountain rock. The clear sound breaks the silence among the pine trees and also touches the heart of the sleepless person in the small pine hut.", a series of generated pictures are produced through 10 random generations. The best and worst images are obtained through manual selection or according to a certain strategy (such as text-image similarity), and positive and negative sample pairs are constructed for reinforcement learning. However, for the generated images produced through multiple steps, only the positive and negative sample pairs of the last step are trained, resulting in a mismatch between training and application, making the application effect of the image generation model not good. In addition, by generating training images offline, after the model learns these data, the effect is still not good. That is, because the data is generated offline and there is no new and better data to learn, the generalization ability of the image generation model is poor.

[0314] The embodiments of the present application provide a training method and an image generation method for an image generation model, which can solve the above problems:

[0315] 1) Online generation of samples: Through an online generation method, new training data is generated according to the current image generation model after each training of the image generation model, enriching the trainable data and making it possible to increase the data upper limit.

[0316] 2) Generating sample pairs based on the Monte Carlo search tree (corresponding to steps 1011 to 1015 above): Compared with randomly generating two samples to construct a sample pair or using the best and worst images of the last step as a sample pair, the embodiments of the present application use a tree structure to generate multiple possible sampled images, and screen from the multiple sampled images to construct a more accurate sample pair (corresponding to the image sample pair above), so as to train the image generation model with high-quality data.

[0317] 3) Improving the generation effect of each step of the image generation model through a step-level learning method (corresponding to steps 102 and 103 above): Considering that the generated image is obtained through multiple-step iteration, sampling is performed based on the conditional probability distribution of the intermediate data of each step to obtain step-level samples (corresponding to multiple sampled images in each iteration round above), and the step-level samples are evaluated (corresponding to obtaining the quality parameters of each sampled image above), so as to learn through the step-level samples to improve the matching degree between the image generation model and the training data, thereby improving the application effect.

[0318] Next, a specific description will be made in conjunction with Figure 10 the flowchart shown in

[0319] Step 601, call the game database.

[0320] In some embodiments, the game database includes a variety of game image samples, such as game scenes, game characters, props, etc., and corresponding game description text samples for each game image sample. The game description text samples include game settings and descriptions of the game image samples.

[0321] Exemplarily, the game settings may include: worldview (such as the time when the game takes place (e.g., future, medieval) and space (e.g., Earth, alien, magic continent)); character settings (such as the identity of the character (e.g., warrior, detective), motivation (revenge, exploration), affiliated forces, etc.); art style, such as pixel art, realistic, cartoon rendering, steampunk, etc.; color and material, such as the main color (e.g., dark and depressing), material details (metallic luster, fabric texture).

[0322] Step 602, call the training of the image generation model.

[0323] In some embodiments, the image generation model is trained by the training method of the image generation model provided in the embodiments of the present application, and the trained image generation model is deployed to the image generation service.

[0324] Exemplarily, an existing game scene, game character, and corresponding description can be used to train a base generation model (such as the Stable Diffusion model, etc.), and then the training method of the generation model provided in the embodiments of the present application is used for secondary training to obtain the final image generation model.

[0325] Exemplarily, see Figures 8A to 8C , in order to obtain a step-level training sample pair that is more suitable for optimizing the image generation model in the multi-step single-result image generation model (that is, the image generation model has undergone T-step forward calculations with T>1 and finally produces 1 result, that is, only one generated image), and to ensure that the training sample pairs obtained in each step have high quality (corresponding to the above-mentioned image sample pairs, including positive sampling image samples and negative sampling image samples), the embodiments of the present application adopt the method of Monte Carlo tree search (MCTS) to select sample pairs to obtain as optimal training sample pairs as possible in a limited number of samplings.

[0326] Exemplarily, the embodiments of the present application expand the process of multi-step denoising and sampling of the Stable Diffusion model and finally generating an image, mainly including: saving the intermediate results of multi-step denoising (such as Z T 、Z T-1 、Z T-2Evaluate the final sampled image.

[0327] For example, see Figure 8A , the process of generating training sample pairs is as follows:

[0328] First, for Z T Perform a denoising inference to get Z T-1 (Here, Z is generated T-1 is a random process, because Z T-1 is randomly sampled from the conditional probability distribution), and then based on Z T-1 The conditional probability distribution of is sampled K times (e.g., 5 times) to obtain multiple Z T-1′ , then for each Z T-1′ The samples are evaluated (they can be decoded to generate images, and the images are evaluated by using evaluation indicators to obtain evaluation scores, i.e., the quality parameters mentioned above). Next, the samples are screened, and the two samples with the highest and lowest evaluation scores are retained. Then, the Z of each retained sample is T-1′ The sample undergoes a denoising inference to obtain Z T-2 , and then perform K sampling to obtain multiple Z T-2′ , and repeat in sequence until the specified number of steps S (such as 20 steps).

[0329] The reason why the samples of each step are screened is that multi-step reinforcement learning requires rewards for multiple steps separately. However, due to the large number of samples and limited computing resources, it is impossible to retain all samples. Therefore, a search tree is established to screen samples and generate training sample pairs to obtain the scores of each step for multi-step reinforcement learning.

[0330] For reinforcement learning, enough data is needed to fine-tune the model. However, the method of generating an image through multiple steps of denoising in related technologies can only generate a pair of training samples. The data volume cannot meet the training requirements. In addition, since the image generation process has not been explored many times, it is impossible to cover as many possible training sample pairs as possible, resulting in reinforcement learning quickly failing to learn valid data. The above process can generate training sample pairs in each step, so that the amount of sample data is increased to T times (that is, each step in the T steps has a pair of training samples, instead of only the last step having a training sample pair).

[0331] For example, the training of the image generation model can include two stages, namely fine-tuning training and reinforcement learning training.

[0332] For example, when performing fine-tuning training on the scene character images of a certain game, game images and game-related description samples are obtained as image-text pairs to construct training data. A total of M rounds (e.g., 10) of iteration are performed on the full training data (i.e., N image-text pairs). In each round of iteration, the full image data is divided into multiple batches according to a preset batch size. The model is updated once for each batch of data. When all batches have been trained once, that is, when all samples have been trained once in the model, it is called one round of iteration.

[0333] The fine-tuning training process is as follows:

[0334] 1) Parameter initialization before the first batch training in the first round: For the image encoder (VAE), text encoder, and denoising network (U-Net), the parameters of a pre-trained model (such as stable-diffusion v1-5) can be adopted, and in this training, only the parameters of the U-Net need to be updated, and the others are not updated. The learning rate of 0.0004 is adopted for initialization. After every 5 rounds of learning, the learning rate becomes 0.1 times the original, and a total of 10 rounds of training are performed.

[0335] 2) Extract image-text sample pairs. The number of extracted pairs is denoted as the batch size (Batch Size, BS). The image-text sample pairs are input into the model, and the following processing is performed: For each image, a latent space representation Ei is generated through the encoder. BS seeds are randomly generated, and BS noise maps (with the same dimension as Z0) are generated. This map is superimposed on the representation Ei as Z0, and Z0 then undergoes a diffusion process to generate Z T (as the original input for subsequent U-Net denoising).

[0336] 3) The text description is passed through the text encoder to obtain text features, and the text features are input into the image generation model (the text features are used as K and V inputs), and for Z T T times of denoising U-Net forward calculations are performed under the KV constraint. After the first forward calculation, Z T-1 is obtained, and finally, after T times, the U-Net outputs the generated image (or called the predicted image).

[0337] 4) Calculate the loss: Calculate the generation loss (i.e., the MSE loss between the predicted image and the original image), and calculate the total loss of the samples in this batch.

[0338] 6) Adopt the stochastic gradient descent method to backpropagate the total loss into the model to obtain the gradient of the model parameters (U-Net) and update the parameters.

[0339] 7) Complete the training of all N / BS batches, and end one training epoch (Epoch).

[0340] During reinforcement learning training:

[0341] For each text-image pair in the training data of the fine-tuning training stage, perform the following processing: Input the text into the image generation model obtained from the previous fine-tuning training. Generate step-level reinforcement learning training sample pairs according to the above MCTS method, that is, 1 training sample pair (positive and negative samples) per step. A total of T training sample pairs and the scores of each training sample pair are generated for each text-image pair. Each training sample pair is input into the image generation model for reinforcement learning training. A total of N iterations (epochs, such as 10 epochs) are performed. Generate a reinforcement learning loss (i.e., the second loss value) based on the scores of the above training sample pairs, and feedback it to the image generation model through backpropagation to optimize the U-Net model. Here, the method for obtaining the scores of the training sample pairs can refer to the descriptions of formulas (6) and (7) in step 1024 above.

[0342] The training process is as follows:

[0343] 1) First, extract BS text-image pairs from the full set of text-image pairs (e.g., BS = 1)

[0344] 2) Input each text-image pair into the model (forward calculation), predict the noise, and generate the final generated image. During this process, collect reinforcement learning information (generate training sample pairs during the T-step process according to the above sample generation method and generate scores for each training sample pair according to the above method).

[0345] 3) Calculate the reinforcement learning loss (i.e., the second loss value) based on the T training sample pairs of each text-image pair.

[0346] 4) Calculate the model generation loss (i.e., the first loss value) for each text-image pair.

[0347] 5) Calculate the total loss of each text-image pair: Loss = L2 + a * L1.

[0348] 6) Perform model backpropagation, using the stochastic gradient descent method, to backpropagate the loss into the model to obtain the gradients of the model parameters (U-Net) and update the parameters.

[0349] 7) Completing the training of all text-image pairs is regarded as completing one iteration (epoch).

[0350] Exemplarily, the reinforcement learning loss is to calculate the reward score for each sample in the training sample pair (reward model, where (positive sample reward score) is used as the score of this training sample pair, collect the sum of all scores generated by the T-step samples and take the average (total score / T) to obtain the reinforcement learning loss generated by the T-step samples corresponding to the image in this text-image pair.

[0351] Exemplarily, for the implementation manner of training the image generation model, reference may be made to the descriptions of steps 101 to 104 above, which will not be elaborated here.

[0352] Step 603: Invoke the user input method selection interface.

[0353] In some embodiments, the user selects an input method, and an input interface is invoked according to the input method for the user to input game description text.

[0354] Exemplarily, the input method may be text box input, voice input, preset template selection, etc. The embodiments of the present application do not limit the specific input method of the user.

[0355] Step 604: Invoke the input interface.

[0356] In some embodiments, in response to receiving a trigger operation input by the user, an input interface is displayed for the user to input game description text.

[0357] Exemplarily, the game description text may be expressed as: {World view: {Time: 2145, Space: Future city, Technological background: Popularity of cybernetic body modification}, Character setting: {Name: ***, Identity: Scientist}, Visual description: {Style: Cyberpunk, Key elements: [Left half-brain transparent cybernetic body shell, Right arm mantis blade, Neon purple glowing belt], Material keywords: [Brushed metal, Matte rubber, Glass reflection]}.

[0358] Exemplarily, the input interface may give examples to help the user understand how to write effective description text. When the user's input is empty or contains illegal characters, corresponding error prompts are given.

[0359] Exemplarily, the input interface may invoke a game setting keyword library (such as preset "Cyberpunk, Steampunk, Fantasy", etc.) to display drop-down recommendations for the user to construct game description text.

[0360] Exemplarily, the number of generated images to be output may be set through the image generation quantity control on the input interface.

[0361] Step 605: Invoke the image generation service.

[0362] In some embodiments, in response to receiving the game description text input by the user on the input interface, a trained image generation model is invoked to generate a preset number of generated images.

[0363] Here, for the implementation manner of generating a preset number of generated images through the image generation model, reference may be made to the description of step 502 above, which will not be elaborated here.

[0364] Step 606: Output and display.

[0365] In some embodiments, in response to receiving a plurality of generated images generated by an image generation service, each generated image is displayed on a display interface.

[0366] Step 607: Select an image according to a user instruction.

[0367] In some embodiments, in response to a user's image selection instruction (such as clicking on a generated image in an image generation interface), processing such as exporting the selected image by the user or generating an image sharing link is performed.

[0368] Through steps 601 to 607, automatic generation of game character / scene images is achieved, reducing the art production cost, supporting the rapid output or verification of the creativity of game development teams, achieving the beneficial effect of shortening the cycle of creative implementation. In addition, game players can also create high-quality game wallpapers or community content, and through the dissemination of community content, the popularity of the game can also be increased, achieving the beneficial effect of game promotion.

[0369] Next, the exemplary structure of the software module implementation of the training device 133 of the image generation model provided in the embodiments of the present application will be continued. In some embodiments, as Figure 2A shown, the software module in the training device 133 of the image generation model stored in the memory 130-1 may include:

[0370] A data processing module 1331, configured to perform at least one iteration round through an initial image generation model to generate a target image sample, wherein the following steps are sequentially performed in each of the at least one iteration round:

[0371] Denoise the image to be denoised based on a description text sample to obtain a denoised image.

[0372] Sample the denoised image to obtain a plurality of sampled images.

[0373] Select an image sample pair from the plurality of sampled images, where the image sample pair includes a positive sampled image sample and a negative sampled image sample.

[0374] In the case where the iteration round does not meet the iteration end condition, return to the step of denoising the image to be denoised based on the description text sample.

[0375] In the case where the iteration round meets the iteration end condition, use the denoised image as the target image sample, or use the positive sampled image sample as the target image sample.

[0376] Among them, the image to be denoised is a noise image sample in the first iteration round, and in each iteration round after the first iteration round, it includes any one of the image sample pair, the denoised image, and the positive sampled image sample obtained in the previous iteration round.

[0377] The training module 1332 is configured to determine a loss value based on the target image sample, the labeled image sample preset for the description text sample, and the image sample pair in each iteration round.

[0378] In some embodiments, the training module 1332 is further configured to update the parameters of the initial image generation model based on the loss value to obtain a trained image generation model.

[0379] In some embodiments, when the iteration round is the first iteration round, the image to be denoised is the noise image sample; when the iteration round is any iteration round after the first iteration round, the image to be denoised is the positive sampled image sample and the negative sampled image sample in the previous iteration round. The data processing module 1331 is further configured to sample based on the conditional probability distribution of at least one denoised image in the iteration round to obtain multiple sampled images in the iteration round, where the denoised image in the last iteration round is the target image sample.

[0380] In some embodiments, the data processing module 1331 is further configured to, when the iteration round is the first iteration round, perform multiple Gaussian distribution samplings based on the conditional probability distribution of the denoised image to obtain the multiple sampled images, where the denoised image is obtained by denoising the noise image sample as the image to be denoised; when the iteration round is any iteration round after the first iteration round, perform multiple Gaussian distribution samplings based on the conditional probability distribution of the first denoised image to obtain multiple sampled images of the first denoised image, and perform multiple Gaussian distribution samplings based on the conditional probability distribution of the second denoised image to obtain multiple sampled images of the second denoised image, where the first denoised image is an image obtained by denoising the positive sampled image sample as the image to be denoised, and the second denoised image is an image obtained by denoising the negative sampled image sample as the image to be denoised.

[0381] In some embodiments, the data processing module 1331 is further configured to generate a random number conforming to the Gaussian distribution at each pixel position of the denoised image according to the mean and variance of the conditional probability distribution of the denoised image; use the random number at each pixel position of the denoised image as the pixel value of the sampled image at the corresponding pixel position; and generate a sampled image based on the pixel value of the sampled image at the corresponding pixel position.

[0382] In some embodiments, the data processing module 1331 is further configured to, when the iteration round is the first iteration round, perform the following processing: obtain a first quality parameter of each of the sampled images, use the sampled image with the highest first quality parameter among the multiple sampled images as the positive sampled image sample for the iteration round, and use the sampled image with the lowest first quality parameter among the multiple sampled images as the negative sampled image sample for the iteration round; when the iteration round is any iteration round after the first iteration round, perform the following processing: obtain a second quality parameter of the sampled image of each of the first denoised images, and obtain a third quality parameter of the sampled image of each of the second denoised images, use the sampled image with the highest second quality parameter among the multiple sampled images of the first denoised image as the positive sampled image sample for the iteration round, and use the sampled image with the lowest third quality parameter among the multiple sampled images of the second denoised image as the negative sampled image sample for the iteration round.

[0383] In some embodiments, the data processing module 1331 is further configured to perform sampled image feature encoding on the sampled images to obtain sampled image features; perform feature mapping on the sampled image features to obtain the first quality parameter of the sampled images.

[0384] In some embodiments, the data processing module 1331 is further configured to perform the following processing for each iteration round: obtain a fourth quality parameter of each of the sampled images for the iteration round; use the sampled image with the highest fourth quality parameter among the multiple sampled images as the positive sampled image sample for the iteration round, and use the sampled image with the lowest fourth quality parameter among the multiple sampled images as the negative sampled image sample for the iteration round.

[0385] In some embodiments, the data processing module 1331 is further configured to perform text feature encoding on the description text sample to obtain text features; obtain the to-be-denoised image features of the to-be-denoised image; perform feature fusion based on the text features and the to-be-denoised image features to obtain fused features; perform conditional probability distribution prediction based on the fused features to obtain a conditional probability distribution; perform Gaussian distribution sampling based on the conditional probability distribution to obtain the denoised image.

[0386] In some embodiments, the data processing module 1331 is further configured to perform embedding encoding on the time steps of the iteration round to obtain time step embeddings; perform feature residual processing on the time step embeddings and the image features to be denoised to obtain residual features; perform self-attention encoding on the residual features to obtain self-attention features; perform cross-attention encoding on the self-attention features and the text features to obtain cross-attention features; and perform downsampling on the cross-attention features to obtain the fusion features.

[0387] In some embodiments, when the iteration round is the first iteration round, the image to be denoised is the noise image sample. When the iteration round is any iteration round after the first iteration round, the image to be denoised is the positive sampling image sample of the previous iteration round. The data processing module 1331 is further configured to sample based on the conditional probability distribution of the denoised image of the iteration round to obtain multiple sampled images of the iteration round, where the denoised image of the last iteration round is the target image sample.

[0388] In some embodiments, when the iteration round is the first iteration round, the image to be denoised is the noise image sample. When the iteration round is any iteration round after the first iteration round, the image to be denoised is the denoised image of the previous iteration round. The data processing module 1331 is further configured to sample based on the conditional probability distribution of the denoised image of the iteration round to obtain multiple sampled images of the iteration round, where the positive sampling image sample of the last iteration round is the target image sample.

[0389] In some embodiments, the data processing module 1331 is further configured to obtain the fifth quality parameter of each of the sampled images of the iteration round; use the sampled image with the highest fifth quality parameter among the multiple sampled images as the positive sampling image sample of the iteration round, and use the sampled image with the lowest fifth quality parameter among the multiple sampled images as the negative sampling image sample of the iteration round.

[0390] In some embodiments, the loss value includes a first loss value. The training module 1332 is further configured to obtain the pixel difference between the target image sample and the annotated image sample at each pixel position; obtain the sum of the squares of the pixel differences at each pixel position, and divide the sum of the squares by the total number of pixels to obtain the first loss value.

[0391] In some embodiments, the loss value includes a second loss value. The training module 1332 is further configured to perform the following processing on the image generation result of each iteration round: obtain the sixth quality parameter of each of the sampled images, and obtain the first average value of the sixth quality parameters of the multiple sampled images; obtain the difference between the sixth quality parameter of the positive sampled image sample and the sixth quality parameter of the negative sampled image sample; obtain the ratio of the logarithm of the total time steps of the at least one iteration round to the time steps of the iteration round, and obtain the square root of the ratio; add the first average value, the product of a preset constant parameter and the square root, and the difference to obtain the score of the image sample pair; obtain the second average value of the scores of the image sample pairs corresponding to the at least one iteration round as the second loss value.

[0392] In some embodiments, the training module 1332 is further configured to obtain the product of a preset weight parameter and the first loss value, and obtain the sum of the product and the second loss value, and use the sum as the combined loss value; based on the combined loss value, update the parameters of the image generation model to be trained to obtain the trained image generation model.

[0393] Next, the exemplary structure of the image generation device 134 provided in the embodiments of the present application as a software module will be continued. In some embodiments, as Figure 2B shown, the software module in the image generation device 134 stored in the memory 130-2 may include:

[0394] A data acquisition module 1341, configured to acquire a description text.

[0395] An image generation module 1342, configured to generate images based on the description text through an image generation model to obtain a plurality of generated images, where the image generation model is trained by the image generation model training method provided in the embodiments of the present application.

[0396] The embodiments of the present application provide a computer program product, which includes a computer program or computer executable instructions, and the computer program or computer executable instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer executable instructions from the computer-readable storage medium, and the processor executes the computer executable instructions, so that the electronic device executes the image generation model training method or the image generation method described above in the embodiments of the present application.

[0397] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will be caused to execute the training method or image generation method of the image generation model provided by the embodiment of the present application. For example, as Figure 4A shown in the training method of the image generation model, or Figure 5 shown in the image generation method.

[0398] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.

[0399] In some embodiments, the computer-executable instructions may be in the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0400] As an example, the computer-executable instructions may or may not correspond to files in the file system, and may be stored as part of a file storing other programs or data. For example, they may be stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (for example, files storing one or more modules, subroutines, or code portions).

[0401] As an example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed at multiple locations and interconnected by a communication network.

[0402] In summary, through the embodiments of the present application, by utilizing the characteristic that the image generation model generates the target image sample through at least one iteration round, that is, the denoised image of each iteration round is generated based on the denoised image of the previous iteration round, and finally the target image sample with noise removed is obtained. Image sample pairs corresponding to each iteration round are constructed from multiple sampled images of each iteration round (that is, from the multiple sampled images of each iteration round, a positive sampled image sample and a negative sampled image sample are selected to construct an accurate sample pair), so that the image generation model can be trained based on the image sample pairs of each iteration round. Based on the target image sample, the labeled image sample preset for the description text sample, and the image sample pairs of each iteration round, a loss value is determined to update the parameters of the image generation model based on the loss value, realizing stage-by-stage learning, capable of enhancing the ability of the image generation model to distinguish positive and negative sampled image samples, such that when updating the image generation model based on the loss value, it can drive the parameters of the image generation model to generate in a direction consistent with the positive sampled image sample, thereby eliminating the error of the image generation model in each iteration round, enhancing the consistency between the target image sample and the description text sample, and ensuring the image generation accuracy.

[0403] As described above, the above are only the embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A training method for an image generation model, characterized in that: The method comprises: At least one iteration is performed through the initial image generation model to generate a target image sample, wherein the following steps are performed in sequence in the at least one iteration: The denoised image is denoised based on the description text sample to obtain the denoised image. Sampling the denoised image to obtain a plurality of sampled images, Selecting an image sample pair from the plurality of sampled images, wherein the image sample pair includes a positive sampled image sample and a negative sampled image sample, When the iteration round does not meet the iteration end condition, return to the step of denoising the denoised image based on the description text sample, When the iteration round satisfies the iteration end condition, the denoised image is used as the target image sample, or the positive sampling image sample is used as the target image sample, The image to be denoised is a noise image sample in the first iteration round, and each iteration round after the first iteration round includes any one of the image sample pair obtained in the previous iteration round, the denoised image, and the positive sampling image sample; Determine a loss value based on the target image sample, the annotated image sample preset for the description text sample, and the image sample pair of each iteration round; Based on the loss value, the parameters of the initial image generation model are updated to obtain a trained image generation model.

2. The method according to claim 1, characterized in that When the iteration round is the first iteration round, the image to be denoised is the noise image sample; when the iteration round is any iteration round after the first iteration round, the image to be denoised is the positive sampling image sample and the negative sampling image sample of the previous iteration round; The step of sampling the denoised image to obtain a plurality of sampled images comprises: Sampling is performed based on the conditional probability distribution of at least one denoised image of the iteration round to obtain a plurality of sampled images of the iteration round, wherein the denoised image of the last iteration round is the target image sample.

3. The method according to claim 2, characterized in that The sampling based on the conditional probability distribution of at least one denoised image of the iteration round to obtain a plurality of sampled images of the iteration round includes: When the iteration round is the first iteration round, multiple Gaussian distribution sampling is performed based on the conditional probability distribution of the denoised image to obtain the multiple sampled images, wherein the denoised image is obtained by performing the denoising on the noise image sample as the image to be denoised; When the iteration round is any iteration round after the first iteration round, the Gaussian distribution sampling is performed multiple times based on the conditional probability distribution of the first denoised image to obtain multiple sampled images of the first denoised image, and the Gaussian distribution sampling is performed multiple times based on the conditional probability distribution of the second denoised image to obtain multiple sampled images of the second denoised image, The first denoised image is an image obtained by denoising the positively sampled image samples as the image to be denoised, and the second denoised image is an image obtained by denoising the negatively sampled image samples as the image to be denoised.

4. The method according to claim 3, characterized in that The performing multiple Gaussian distribution sampling based on the conditional probability distribution of the denoised image to obtain the multiple sampled images includes: Generating a random number conforming to a Gaussian distribution at each pixel position of the denoised image according to the mean and variance of the conditional probability distribution of the denoised image; Using the random number at each pixel position of the denoised image as the pixel value of the sampled image at the corresponding pixel position; A sampled image is generated based on the pixel values ​​of the sampled image at corresponding pixel positions.

5. The method according to claim 3 or 4, characterized in that: The selecting of image sample pairs from the plurality of sampled images comprises: When the iteration round is the first iteration round, the following processing is performed: obtaining a first quality parameter of each of the sampled images, Using the sampling image with the highest first quality parameter among the multiple sampling images as the positive sampling image sample of the iteration round, and using the sampling image with the lowest first quality parameter among the multiple sampling images as the negative sampling image sample of the iteration round; When the iteration round is any iteration round after the first iteration round, the following processing is performed: obtaining a second quality parameter of each sampled image of the first denoised image, and obtaining a third quality parameter of each sampled image of the second denoised image, The sampling image with the highest second quality parameter among multiple sampling images of the first denoised image is used as the positive sampling image sample of the iterative round, and the sampling image with the lowest third quality parameter among multiple sampling images of the second denoised image is used as the negative sampling image sample of the iterative round.

6. The method according to claim 5, characterized in that When the iteration round is the first iteration round, obtaining the first quality parameter of each sampled image includes: Performing sampling image feature encoding on the sampling image to obtain sampling image features; Feature mapping is performed on the features of the sampled image to obtain a first quality parameter of the sampled image.

7. The method according to claim 1 or 2, characterized in that: The selecting of image sample pairs from the plurality of sampled images comprises: The following processing is performed for each of the iteration rounds: Acquire a fourth quality parameter of each of the sampled images in the iteration round; The sampling image with the highest fourth quality parameter among the multiple sampling images is used as the positive sampling image sample of the iteration round, and the sampling image with the lowest fourth quality parameter among the multiple sampling images is used as the negative sampling image sample of the iteration round.

8. The method according to any one of claims 2 to 7, characterized in that: The method of performing denoising on the image to be denoised based on the description text sample to obtain the denoised image comprises: Performing text feature encoding on the description text sample to obtain text features; Acquire the image features to be denoised of the image to be denoised; Perform feature fusion based on the text feature and the image feature to be denoised to obtain a fused feature; Perform conditional probability distribution prediction based on the fusion features to obtain conditional probability distribution; Gaussian distribution sampling is performed based on the conditional probability distribution to obtain the denoised image.

9. The method according to claim 8, characterized in that The step of fusing features based on the text features and the features of the image to be denoised to obtain fused features includes: Embedding the time steps of the iteration round to obtain time step embedding; Performing feature residual processing on the time step embedding and the image features to be denoised to obtain residual features; Performing self-attention encoding on the residual feature to obtain a self-attention feature; Performing cross-attention encoding on the self-attention feature and the text feature to obtain a cross-attention feature; The cross-attention feature is downsampled to obtain the fused feature.

10. The method according to claim 1, characterized in that When the iteration round is the first iteration round, the image to be denoised is the noise image sample; when the iteration round is any iteration round after the first iteration round, the image to be denoised is the positive sampling image sample of the previous iteration round; The step of sampling the denoised image to obtain a plurality of sampled images comprises: Sampling is performed based on the conditional probability distribution of the denoised image of the iteration round to obtain a plurality of sampled images of the iteration round, wherein the denoised image of the last iteration round is the target image sample.

11. The method according to claim 1, characterized in that: When the iteration round is the first iteration round, the image to be denoised is the noisy image sample; when the iteration round is any iteration round after the first iteration round, the image to be denoised is the denoised image of the previous iteration round; The step of sampling the denoised image to obtain a plurality of sampled images comprises: Sampling is performed based on the conditional probability distribution of the denoised image of the iteration round to obtain a plurality of sampled images of the iteration round, wherein the positive sampled image sample of the last iteration round is the target image sample.

12. The method according to claim 10 or 11, characterized in that: The selecting of image sample pairs from the plurality of sampled images comprises: Acquire a fifth quality parameter of each of the sampled images in the iteration round; The sampling image with the highest fifth quality parameter among the multiple sampling images is used as the positive sampling image sample of the iteration round, and the sampling image with the lowest fifth quality parameter among the multiple sampling images is used as the negative sampling image sample of the iteration round.

13. The method according to any one of claims 1 to 12, characterized in that: The loss value includes a first loss value, and the loss value is determined based on the target image sample, the labeled image sample preset for the description text sample, and the image sample pair of each iteration round, including: Obtaining a pixel difference between the target image sample and the labeled image sample at each pixel position; The sum of squares of pixel differences at each of the pixel positions is obtained, and the sum of squares is divided by the total number of pixels to obtain the first loss value.

14. The method according to any one of claims 1 to 13, characterized in that: The loss value includes a second loss value, and the loss value is determined based on the target image sample, the labeled image sample preset for the description text sample, and the image sample pair of each iteration round, including: The following processing is performed for each of the iteration rounds: Obtaining a sixth quality parameter of each of the sampled images, and obtaining a first average value of the sixth quality parameters of the plurality of sampled images; Obtaining a difference between a sixth quality parameter of the positively sampled image sample and a sixth quality parameter of the negatively sampled image sample; Obtaining a ratio between the logarithm of the total time steps of the at least one iteration round and the time steps of the iteration round, and obtaining a square root of the ratio; Adding the product of the first average value, a preset constant parameter and the square root and the difference value to obtain a score of the image sample pair; Obtain a second average value of the scores of the image sample pairs respectively corresponding to the at least one iteration round as the second loss value.

15. The method according to any one of claims 1 to 12, characterized in that: The loss value includes a first loss value and a second loss value, and the updating of the parameters of the initial image generation model based on the loss value to obtain the trained image generation model includes: Obtaining a product of a preset weight parameter and the first loss value, and obtaining a sum of the product and the second loss value, and using the sum as a combined loss value; Based on the combined loss value, the parameters of the image generation model to be trained are updated to obtain the trained image generation model.

16. An image generation method, characterized in that: The method comprises: Get the description text; An image generation model is used to generate images based on the description text to obtain a plurality of generated images, wherein the image generation model is trained by the method described in any one of claims 1 to 15.

17. A training device for an image generation model, characterized in that: The device comprises: The data processing module is used to perform at least one iteration round through the initial image generation model to generate a target image sample, wherein the following steps are performed in sequence in the at least one iteration round: The denoised image is denoised based on the description text sample to obtain the denoised image. Sampling the denoised image to obtain a plurality of sampled images, Selecting an image sample pair from the plurality of sampled images, wherein the image sample pair includes a positive sampled image sample and a negative sampled image sample, When the iteration round does not meet the iteration end condition, return to the step of denoising the denoised image based on the description text sample, When the iteration round satisfies the iteration end condition, the denoised image is used as the target image sample, or the positive sampling image sample is used as the target image sample, The image to be denoised is a noise image sample in the first iteration round, and each iteration round after the first iteration round includes any one of the image sample pair obtained in the previous iteration round, the denoised image, and the positive sampling image sample; A training module, used for determining a loss value based on the target image sample, the annotated image sample preset for the description text sample, and the image sample pair of each iteration round; The training module is also used to update the parameters of the initial image generation model based on the loss value to obtain a trained image generation model.

18. An electronic device, characterized in that: The electronic device comprises: A memory for storing computer executable instructions or computer programs; A processor, for implementing the training method of the image generation model described in any one of claims 1 to 15, or implementing the image generation method described in claim 16, when executing the computer executable instructions or computer program stored in the memory.

19. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, they implement the training method of the image generation model described in any one of claims 1 to 15, or implement the image generation method described in claim 16.

20. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, they implement the training method of the image generation model described in any one of claims 1 to 15, or implement the image generation method described in claim 16.

Citation Information

Cited By

  • Training method of multimedia resource generation model and multimedia resource generation method

    CN120494015A