Training method of image generation model and image generation method and device

By acquiring training samples for the target image generation task, utilizing the noise prediction effects of the teacher and student models, and combining prediction techniques, the problem of insufficient accuracy of existing large-scale image generation models under different constraints is solved, thereby improving the generation quality and applicability of the model.

CN121147342APending Publication Date: 2025-12-16BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511220427.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively train large-scale image generation models to generate high-quality images under various constraints, resulting in insufficient accuracy and applicability of the models.

Method used

By acquiring the first and second training samples corresponding to the target image generation task, noise prediction is performed using the teacher model to obtain the noise latent representation. Combined with the noise prediction results of the initial student model, the student model is trained, and the noise latent representations under different constraints are fused to train the model.

Benefits of technology

The noise prediction performance of the image generation model under different constraints was achieved, improving the accuracy and applicability of the image generation model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121147342A_ABST
    Figure CN121147342A_ABST
Patent Text Reader

Abstract

The invention discloses a training method of an image generation model and an image generation method and device, relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models and the like, and can be applied to scenes such as content generation based on artificial intelligence. According to the specific implementation scheme, a first training sample and a second training sample corresponding to a target image generation task are acquired; wherein the second training sample is obtained by replacing a target constraint condition in the first training sample with an empty condition; performing noise prediction by using a teacher model according to the first training sample and the second training sample to obtain a first noise potential representation and a second noise potential representation; according to the first training sample, performing noise prediction by using an initial student model to obtain a third noise potential representation; and training the initial student model according to the first noise potential representation, the second noise potential representation and the third noise potential representation to obtain a target image generation model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to the technical field of computer vision, deep learning, large model, etc., and can be applied to scenarios such as content generation based on artificial intelligence. Specifically, the present application relates to a training method of an image generation model, an image generation method and device. BACKGROUND

[0002] A large model refers to a machine learning model with large-scale parameters and complex computing structure, which can process massive data and complete various complex tasks such as natural language processing, computer vision, speech recognition, image generation, etc. With the development of artificial intelligence, the application of large models is becoming more and more widespread, such as text-to-image and image-to-image using large models. SUMMARY

[0003] The present application provides a training method of an image generation model, an image generation method and device. The specific solutions are as follows:

[0004] According to an aspect of the present application, a training method of an image generation model is provided, comprising:

[0005] obtaining a first training sample and a second training sample corresponding to a target image generation task; wherein the second training sample is obtained by replacing a target constraint condition in the first training sample with an empty condition;

[0006] respectively according to the first training sample and the second training sample, performing noise prediction using a teacher model to obtain a first noise latent representation and a second noise latent representation;

[0007] according to the first training sample, performing noise prediction using an initial student model to obtain a third noise latent representation;

[0008] according to the first noise latent representation, the second noise latent representation and the third noise latent representation, training the initial student model to obtain a target image generation model.

[0009] According to another aspect of the present application, a training device of an image generation model is provided, comprising:

[0010] a first obtaining module configured to obtain a first training sample and a second training sample corresponding to a target image generation task; wherein the second training sample is obtained by replacing a target constraint condition in the first training sample with an empty condition;

[0011] a first prediction module configured to respectively according to the first training sample and the second training sample, perform noise prediction using a teacher model to obtain a first noise latent representation and a second noise latent representation;

[0012] a second prediction module, configured to perform noise prediction on the first training sample by using an initial student model to obtain a third noise latent representation;

[0013] a training module, configured to train the initial student model according to the first noise latent representation, the second noise latent representation and the third noise latent representation to obtain a target image generation model.

[0014] According to another aspect of the present application, an image generation method is provided, comprising:

[0015] obtaining an image generation constraint input by a user;

[0016] generating a target image by using an image generation model according to the image generation constraint; wherein the image generation model is trained by the training method according to the above aspect.

[0017] According to another aspect of the present application, an image generation apparatus is provided, comprising:

[0018] an obtaining module, configured to obtain an image generation constraint input by a user;

[0019] a generating module, configured to generate a target image by using an image generation model according to the image generation constraint; wherein the image generation model is trained by the training method according to the above aspect.

[0020] According to another aspect of the present application, an electronic device is provided, comprising:

[0021] at least one processor; and

[0022] a memory connected with the at least one processor; wherein

[0023] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to the above aspect.

[0024] According to another aspect of the present application, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method according to the above aspect.

[0025] According to another aspect of the present application, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the steps of the method according to the above aspect.

[0026] It is to be understood that the details set forth herein do not limit the scope of the embodiments of the application. Other embodiments of the application will be readily apparent to those skilled in the art from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0027] The accompanying drawings are included to provide a further understanding of the application, and are incorporated in and constitute a part of this application. In the drawings:

[0028] Figure 1 A flowchart of a training method of an image generation model according to an embodiment of the application is shown in FIG. 1.

[0029] Figure 2 A flowchart of a training method of an image generation model according to another embodiment of the application is shown in FIG. 2.

[0030] Figure 3 A flowchart of a training method of an image generation model according to another embodiment of the application is shown in FIG. 3.

[0031] Figure 4 A flowchart of a training process of an image generation model according to an embodiment of the application is shown in FIG. 4.

[0032] Figure 5 A flowchart of an image generation method according to an embodiment of the application is shown in FIG. 5.

[0033] Figure 6 A flowchart of an image generation method according to another embodiment of the application is shown in FIG. 6.

[0034] Figure 7 A structural diagram of a training device of an image generation model according to an embodiment of the application is shown in FIG. 7.

[0035] Figure 8 A structural diagram of an image generation device according to an embodiment of the application is shown in FIG. 8.

[0036] Figure 9 A block diagram of an electronic device for implementing the training method of the image generation model according to an embodiment of the application is shown in FIG. 9. DETAILED DESCRIPTION

[0037] Exemplary embodiments of the application are described herein with reference to the accompanying drawings, which are included to provide a thorough understanding of the embodiments of the application. The embodiments of the application, however, can take many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the application to those skilled in the art. As such, the embodiments of the application are not intended to be limiting, but rather, the intent is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the application.

[0038] It should be noted that the acquisition, storage, use, processing and the like of data in the technical solutions of the present application comply with relevant provisions of national laws and regulations and do not violate public order and good customs.

[0039] The training method of the image generation model, the image generation method, the device, the electronic equipment and the storage medium of the embodiments of the present application are described below with reference to the accompanying drawings.

[0040] Figure 1 The flowchart of the training method of the image generation model provided by an embodiment of the present application is shown.

[0041] The training method of the image generation model of the embodiments of the present application can be executed by the training device of the image generation model of the embodiments of the present application, which can be configured in an electronic equipment.

[0042] The electronic equipment can be any device with computing capability, such as a personal computer, a mobile terminal, a server, etc. The mobile terminal can be a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, etc. hardware devices with various operating systems, touch screens and / or display screens.

[0043] As shown in the figure, the training method of the image generation model includes: Figure 1

[0044] Step 101, obtaining the first training sample and the second training sample corresponding to the target image generation task.

[0045] The target image generation task can be one of a plurality of image generation tasks.

[0046] For example, the first training sample and the second training sample corresponding to the target image generation task can be training samples of any training round.

[0047] For example, the image generation task corresponding to the current training round can be one of the text-to-image task and the image-to-image task.

[0048] For example, if the target image generation task is the text-to-image task, the first training sample and the second training sample are text-to-image data, and if the target image generation task is the image-to-image task, the first training sample and the second training sample are image-to-image data.

[0049] In the present application, the first training sample can include one or more constraint conditions. The constraint condition is used to constrain the image generation result, so that the generated image meets the constraint condition.

[0050] ​For example, if the target image generation task is a text-to-image task, the constraint condition included in the first training sample is a text condition, and if the target image generation task is an image-to-image task, the constraint condition in the first training sample can include one or more of a text condition and an image condition.

[0051] The text condition is used to provide semantic constraints and can describe the content, style, or attributes of the target image to be generated, guiding the model to generate an image that meets the semantics.

[0052] The image condition is used to provide a source image and control the structure, layout, style, etc. of the generated image.

[0053] For a text-to-image task, the text condition is, for example, "a cat wearing a hat, watercolor style".

[0054] For an image-to-image task, the text condition is, for example, "adjust the color of the clothes of the person in the given image to red", and the image condition is the given image.

[0055] As can be seen, the constraint condition included in the first training sample in the present application can be a text condition or an image condition, or a text condition and an image condition.

[0056] In the present application, the second training sample can be obtained by replacing the target constraint condition in the first training sample with a null condition.

[0057] The target constraint condition can be one or more of the constraint conditions included in the first training sample. For example, the target constraint condition can be a text condition or an image condition, or a text condition and an image condition.

[0058] For a text-to-image task, the constraint condition in the first training sample is a text condition, and the second training sample can be obtained by replacing the text condition with a null condition.

[0059] For an image-to-image task, if the constraint condition in the first training sample includes a text condition and an image condition, the second training sample can be obtained by replacing the image condition in the first training sample with a null condition.

[0060] For example, in addition to including a constraint condition, the first training sample can also include a time step, a noise image, etc. The second training sample in the present application includes the same content as the first training sample except that the null condition included therein is different from the target constraint condition in the first training sample.

[0061] Step 102: respectively according to the first training sample and the second training sample, performing noise prediction by using a teacher model to obtain a first noise latent representation and a second noise latent representation.

[0062] The teacher model can be a large model that can perform multiple image generation tasks, such as completing a text-to-image task, i.e., generating an image according to a text condition, or completing an image-to-image task, i.e., generating a new image according to a text condition and an image condition, etc.

[0063] For example, the teacher model can be a model for predicting noise.

[0064] In this application, the teacher model can be used to predict noise based on the first training sample to obtain the first noise latent representation, and the teacher model can be used to predict noise based on the second training sample to obtain the second noise latent representation.

[0065] For example, the first training sample includes a first time step, a noisy image, a constraint condition required for a target image generation task, etc. The noisy image at the first time step can be mapped to a latent space to obtain a noisy latent representation of the noisy image. The encoding vector at the first time step is concatenated with the noisy latent representation in the channel dimension. Then, the concatenated latent representation is fused with the encoding vector of the constraint condition through an interactive attention mechanism. Then, the fused latent representation is input to the teacher model for noise prediction to obtain the first noise latent representation.

[0066] The first noise latent representation is an estimation of the noise distribution in the noisy image at the first time step based on the first training sample by the teacher model.

[0067] For example, if the second training sample includes a noisy image, a constraint condition required for a target image generation task, etc., the noisy latent representation of the noisy image can be fused with the encoding vector of the constraint condition through an interactive attention mechanism. Then, the fused latent representation is input to the teacher model for noise prediction to obtain the second noise latent representation.

[0068] For example, if the second training sample further includes a constraint condition other than the target constraint condition, the second noise latent representation can be obtained by using a method similar to obtaining the first noise latent representation, and therefore will not be described here.

[0069] The second noise latent representation is an estimation of the noise distribution in the noisy image at the first time step based on the second training sample by the teacher model.

[0070] Step 103, according to the first training sample, using the initial student model to predict noise to obtain a third noise latent representation.

[0071] The initial student model is an image generation model, the parameter size of the initial student model can be smaller than the parameter size of the teacher model, and the initial student model can be a model for predicting noise.

[0072] The initial student model can be a student model before the first parameter adjustment or a student model trained in a previous training round.

[0073] In the present application, the noise image of the first time step can be mapped to the latent space to obtain a noisy latent representation of the noise image, the encoding vector of the first time step is spliced with the noisy latent representation in the channel dimension, and then the spliced latent representation is fused with the encoding vector of the constraint condition through the interaction attention mechanism, and then the fused latent representation is input to the student model for noise prediction to obtain a third noise latent representation.

[0074] The third noise latent representation is an estimated noise distribution in the noise image of the first time step based on the first training sample by the initial student model.

[0075] In step 104, the initial student model is trained according to the first noise latent representation, the second noise latent representation and the third noise latent representation to obtain a target image generation model.

[0076] In the present application, the first noise latent representation and the second noise latent representation can be fused to obtain a fused noise latent representation, and the loss can be determined according to the difference between the fused noise latent representation and the third noise latent representation, and the parameters of the initial student model can be adjusted according to the loss, and the parameter model after the parameter adjustment is continuously trained until the training end condition is met to obtain the target image generation model.

[0077] The training end condition can be that the current training round reaches a set number of times, or the consistency loss is less than a loss threshold.

[0078] The target image generation model in the present application can perform various image generation tasks, such as text-to-image tasks and image-to-image tasks.

[0079] In the present application, the first training sample and the second training sample corresponding to the target image generation task are obtained, the teacher model is used to perform noise prediction according to the first training sample and the second training sample respectively to obtain noise prediction results under the target constraint condition and under the condition without target constraint, and the initial student model is used to perform noise prediction according to the first training sample, and the initial student model is trained according to the noise prediction results of the teacher model under the target constraint condition and under the condition without target constraint, and the noise prediction result of the initial student model, which can make the student model learn information under different constraint conditions and improve the accuracy and applicability of the image generation model.

[0080] Figure 2A flowchart of a training method of an image generation model according to another embodiment of the present application is shown.

[0081] As shown in Figure 2 The training method of the image generation model includes the following steps.

[0082] In step 201, first training samples and second training samples corresponding to a target image generation task are obtained.

[0083] In some embodiments, the training process of the initial student model can include multiple training rounds, and the target image generation task is associated with the training rounds, that is, each training round has an associated target image generation task.

[0084] For example, each training round can correspond to an image generation task, that is, each training round trains an image generation task.

[0085] In addition, the image generation tasks of different training rounds can be the same or different, so that the initial student model can be trained in combination with multiple image generation tasks to obtain a target image generation model, thereby improving the accuracy and robustness of the model.

[0086] For example, the target image generation task of each training round can be determined, and the first training samples and the second training samples matched with the target image generation task of each training round can be obtained.

[0087] For example, the image generation tasks of each training round can be pre-set, such as text-to-image tasks for odd rounds and image-to-image tasks for even rounds, so that the image generation task of any training round can be determined according to whether the training round is odd or even.

[0088] For example, a task can also be randomly selected from multiple image generation tasks as the target image generation task of any training round.

[0089] Therefore, by training the student model using training samples corresponding to the image generation task associated with each training round, the pertinence of model training can be improved.

[0090] In step 202, first noise latent representations and second noise latent representations are obtained by using the teacher model to perform noise prediction according to the first training samples and the second training samples, respectively.

[0091] In step 203, third noise latent representations are obtained by using the initial student model to perform noise prediction according to the first training samples.

[0092] In the present application, steps 201-203 can adopt any of the implementation manners of the embodiments of the present application, and therefore will not be described here.

[0093] Step 204, fusing the first noise latent representation and the second noise latent representation to obtain a fused noise latent representation.

[0094] In some embodiments, the first noise latent representation and the second noise latent representation can be added to obtain the fused noise latent representation.

[0095] In some embodiments, a first weight of the first noise latent representation and a second weight of the second noise latent representation can be determined, and the first noise latent representation and the second noise latent representation can be weighted and summed according to the first weight and the second weight to obtain the fused noise latent representation. Thus, the weights of the noise latent representation under the constraint condition and the noise latent representation under the empty condition can be controlled according to actual needs, and different training needs can be met.

[0096] In some embodiments, the target intensity range can be determined from the candidate intensity range corresponding to the plurality of condition types according to the target constraint condition, and the target guidance intensity can be determined from the target intensity range, and the first noise latent representation and the second noise latent representation can be fused according to the target guidance intensity to obtain the fused noise latent representation. Thus, the target intensity range is determined based on the target constraint condition, and the target guidance intensity is selected from the target intensity range, which can improve the diversity of the target guidance intensity. The first noise latent representation and the second noise latent representation are fused based on the target guidance intensity, which can meet the fusion needs of different weights of the constraint condition.

[0097] The plurality of condition types can include text conditions, image conditions, etc., and different condition types have corresponding intensity ranges. The intensity range refers to the value range of the guidance intensity, and the guidance intensity is used to control the weight of the corresponding condition type. For example, the intensity range corresponding to the text condition is [4, 6], and the intensity range corresponding to the image condition is [1, 3],

[0098] The target guidance intensity can be used to control the weight of the target constraint condition. The greater the target guidance intensity, the stronger the controllability of the generated image and the lower the diversity.

[0099] It should be noted that the intensity ranges corresponding to different condition types can be the same or different, which can be determined according to actual needs, and the present application does not limit this.

[0100] For example, the target condition type to which the target constraint condition belongs can be determined from the plurality of condition types, and the target intensity range can be determined according to the candidate intensity range corresponding to the target condition type. For example, the candidate intensity range corresponding to the target condition type can be determined as the target intensity range, or a sub-intensity range in the candidate intensity range corresponding to the target condition type can be determined as the target intensity range.

[0101] As an example, for the text-to-image task, if the constraint condition in the first training sample is a text condition, and the target constraint condition is also a text condition, the intensity range corresponding to the text condition can be determined as the target intensity range.

[0102] As an example, for the image-to-image task, if the constraint condition in the first training sample includes a text condition and an image condition, and the target constraint condition is an image condition, the intensity range corresponding to the image condition can be determined as the target intensity range.

[0103] As an example, for the image-to-image task, if the constraint condition in the first training sample is a text condition, and the image condition is empty, the target constraint condition is a text condition, and the intensity range corresponding to the text condition can be determined as the target intensity range.

[0104] As an example, for the image-to-image task, if the constraint condition in the first training sample includes a text condition and an image condition, and the target constraint condition is a text condition, the intensity range corresponding to the text condition can be determined as the target intensity range.

[0105] As an example, for the image-to-image task, if the constraint condition in the first training sample is an image condition, and the text condition is empty, the target constraint condition is an image condition, and the intensity range corresponding to the image condition can be determined as the target intensity range.

[0106] Therefore, according to the candidate intensity range corresponding to the target condition type to which the target constraint condition type belongs, the target intensity range is determined, which can improve the accuracy of the target intensity range.

[0107] For example, the target intensity range is divided into a plurality of sub-intensity ranges, and the target guidance intensity is obtained by sampling in the target intensity range according to the probability of the plurality of sub-intensity ranges.

[0108] As an example, the target intensity range can be divided into a plurality of sub-intensity ranges, and the target guidance intensity can be obtained by sampling in the target intensity range according to the probability of the plurality of sub-intensity ranges.

[0109] For example, according to the target guidance intensity, the first noise latent representation and the second noise latent representation can be fused by the following method: the sum of a preset value and the target guidance intensity can be determined, the fifth noise latent representation can be determined according to the product between the sum and the first noise latent representation, the sixth noise latent representation can be determined according to the product between the target guidance intensity and the second noise latent representation, and the fused noise latent representation can be determined according to the difference between the fifth noise latent representation and the sixth noise latent representation.

[0110] The preset value can be 1 or other values, which are not limited.

[0111] As an example, if the image generation task of the current round is the text-to-image task, the constraint condition included in the first training sample is the text condition, and the target constraint condition is the text condition, the calculation formula of the fused noise latent representation is shown in the following formula (1):

[0112]

[0113] wherein, is the first noise latent representation output by the teacher model when the input information is the text condition c, the image condition is empty, the time step is t, and the noise image z t . is the second noise latent representation output by the teacher model when the input information is that both the text condition and the image condition are empty, the time step is t, and the noise image z t ; and w2 is the guidance strength corresponding to the text condition.

[0114] As an example, if the image generation task of the current round is the image-to-image task, the constraint condition included in the first training sample is the text condition and the image condition, and the target constraint condition is the image condition, the calculation formula of the fused noise latent representation is shown in the following formula (2):

[0115]

[0116] wherein, ε θ (z t ,c,t,i1,i2,...,i n ) is the first noise latent representation output by the teacher model when the input information is the text condition c, the image condition is i1,i2,...,i n , the time step is t, and the noise image z t . is the second noise latent representation output by the teacher model when the input information is the text condition c, the image condition is empty, the time step is t, and the noise image z t ; and w1 is the guidance strength corresponding to the image condition.

[0117] As an example, if the image generation task of the current round is the image-to-image task, the constraint condition included in the first training sample is the text condition, and the target constraint condition is the text condition, the calculation formula of the fused noise latent representation is shown in the following formula (3):

[0118]

[0119] wherein, is the first noise latent representation output by the teacher model when the input information is the text condition c, the image condition is empty, the time step is t, and the noise image z tIn this case, the first noise latent representation of the teacher model output; The input information is empty, both text and image conditions are empty, the time step is t, and the image is noisy z. t In the case of w2, the second noise latent representation is output by the teacher model; w2 is the guidance intensity corresponding to the text condition.

[0120] As an example, if the image generation task in the current round is a graph-to-graph task, and the constraints in the first training sample are text conditions and image conditions, and the target constraint is a text condition, then the calculation formula for the fused noise latent representation is as follows (4):

[0121]

[0122] Where, ε θ (z t ,c,t,i1,i2,...,i n ) is given by input information as text condition c, and image conditions as i1, i2, ..., i n The time step is t, and the noisy image z t In this case, the first noise latent representation of the teacher model output; The input information is text, the condition is empty, and the image condition is i1, i2, ..., i n The time step is t, and the noisy image z t In the case of w2, the second noise latent representation is output by the teacher model; w2 is the guidance intensity corresponding to the text condition.

[0123] As an example, if the image generation task in the current round is a graph-to-graph task, and the constraints included in the first training sample are image conditions and the target constraints are image conditions, then the calculation formula for the fused noise latent representation is as follows (5):

[0124]

[0125] in, The text condition is empty, and the image condition is i1, i2, ..., i n The time step is t, and the noisy image z t In this case, the first noise latent representation of the teacher model output; The input information is empty, both text and image conditions are empty, the time step is t, and the image is noisy z. t In the case of the second noise latent representation output by the teacher model, w1 is the guidance intensity corresponding to the image condition.

[0126] Thus, the first noise latent representation and the second noise latent representation are fused based on the target guidance intensity, so that the controllability and diversity of image generation can be controlled by controlling the weight of the target constraint condition by using the target guidance intensity.

[0127] In step 205, the initial student model is trained according to the fused noise latent representation and the third noise latent representation to obtain the target image generation model.

[0128] In some embodiments, the fused noise latent representation can be denoised to obtain a first image latent representation, and the third noise latent representation can be denoised to obtain a second image latent representation. Then, the initial student model is trained according to the first image latent representation and the second image latent representation to obtain the target image generation model.

[0129] Wherein, the first image latent representation is decoded to obtain an image generated by the teacher model based on the first training sample and the second training sample; and the second image latent representation is decoded to obtain an image generated by the initial student model based on the second training sample.

[0130] Thus, by denoising the fused noise latent representation and the third noise latent representation respectively, and training the initial student model based on the denoised image latent representations, the accuracy of the student model in estimating noise can be improved.

[0131] For example, the first denoising strategy can be used to denoise the fused noise latent representation, and the second denoising strategy can be used to denoise the third noise latent representation. The first denoising strategy and the second denoising strategy can be related to the noise adding strategy of the noise image. In addition, the first denoising strategy and the second denoising strategy can be the same or different, and no limitation is made in this regard.

[0132] For example, the first training sample can include a first time step, and the third noise latent representation is a noisy latent representation of a noise image at the first time step, which is obtained by using the initial student model to predict noise. In order to train the student model consistently, noise can be added to the first image latent representation by using the noise adding strategy to restore a noisy latent representation of a noise image at a second time step, and the initial student model is used to predict noise based on the noisy latent representation at the second time step to obtain a fourth noise latent representation. The fourth noise latent representation is denoised to obtain a third image latent representation, and the initial student model is trained according to the difference between the second image latent representation and the third image latent representation to obtain the target image generation model.

[0133] The second time step is a previous time step of the first time step, that is, the noisy latent representation of the noisy image of the previous time step of the first time step is restored by adding noise to the latent representation of the image generated based on the teacher model.

[0134] As an example, the initial student model can be used to perform noise prediction according to the noisy latent representation of the second time step, in combination with the encoding vector of the second time step and the encoding vector of the constraint condition in the first training sample, to obtain a fourth noisy latent representation.

[0135] As an example, the consistency loss can be determined according to the difference between the second image latent representation corresponding to the current time step and the third image latent representation corresponding to the previous time step, and the parameters of the initial student model are adjusted according to the consistency loss to obtain the target image generation model.

[0136] For example, the distance between the second image latent representation and the third image latent representation can be calculated, and the consistency loss can be determined according to the distance.

[0137] Therefore, by adding noise to the image latent representation of the image generated by the teacher model to restore the noisy latent representation of the previous time step, performing noise prediction on the initial student model according to the noisy latent representation of the previous time step to obtain the image latent representation corresponding to the previous time step, and performing consistency training on the initial student model using the image latent representation corresponding to the current time step and the image latent representation corresponding to the previous time step, the target image generation model is obtained. Thus, the processing steps of the target image generation model can be reduced and the image generation speed can be improved by performing consistency training on the initial student model according to the images generated by the initial student model at adjacent time steps. In the embodiment of the present application, the fusion noisy latent representation is obtained by fusing the noisy latent representation of the teacher model under the target constraint condition and the noisy latent representation of the teacher model without the target constraint condition, and the initial student model is trained based on the fusion noisy latent representation and the third noisy latent representation. The conditional information corresponding to the target constraint condition of the teacher model, such as text information and image information, can be distilled into the student model, improving the accuracy of the target generation model and the quality of the generated image.

[0138] Figure 3 The flowchart of the training method of the image generation model provided by another embodiment of the present application is shown.

[0139] As Figure 3 shown, the training method of the image generation model includes:

[0140] Step 301: According to the second probability of each image generation task, a target image generation task is determined from each image generation task.

[0141] The second probability of one image generation task can refer to a probability that the image generation task is selected as the target image generation task.

[0142] In the present application, each image generation task can include a text-to-image task, an image-to-image task, and the like. In addition, the second probabilities of different image generation tasks can be the same or different, and the present application does not limit this.

[0143] In the present application, the second probabilities of each image generation task can be obtained, and one image generation task can be selected from each image generation task as a target image generation task according to the second probabilities of each image generation task.

[0144] Taking any training round as an example, one image generation task can be selected from each image generation task as a target image generation task of any training round according to the second probabilities of each image generation task.

[0145] Therefore, by determining the target image generation task according to the second probabilities of each image generation task, not only can the proportion requirement of each image generation task in distillation training be met, different training requirements can be met, but also the accuracy of the model can be improved by training the initial student model in combination with multiple image generation tasks.

[0146] In step 302, a first training sample is obtained according to the target image generation task.

[0147] In the present application, different image generation tasks can have corresponding training sets, and the first training sample can be obtained from the training set of the target image generation task.

[0148] For example, the first training sample of any training round can be obtained from the training set of the target image generation task.

[0149] In step 303, a target constraint condition is determined from constraint conditions contained in the first training sample.

[0150] In some embodiments, part or all of the constraint conditions contained in the first training sample can be used as the target constraint condition.

[0151] For example, if the target image generation task is a text-to-image task and the constraint condition in the first training sample is a text condition, the text condition can be used as the target constraint condition.

[0152] For example, the target image generation task is the image-to-image task. If the constraint condition contained in the first training sample is one, the constraint condition can be used as the target constraint condition. If the constraint condition contained in the first training sample is multiple, the target constraint condition can be determined from the multiple constraint conditions according to the first probability of the multiple constraint conditions. In this way, the target constraint condition can be enriched according to the probability of the multiple constraint conditions, covering different model input conditions, and then enriching the second training sample.

[0153] The first probability can refer to the probability that the constraint condition is determined as the target constraint condition.

[0154] For the image-to-image task, if the text condition is included in the first training sample and the image condition is empty, the text condition can be directly determined as the target constraint condition. If the text condition and the image condition are included in the first training sample, it can be determined whether the text condition or the image condition is used as the target constraint condition according to the first probability of the text condition and the image condition. If the image condition is included in the first training sample and the text condition is empty, the image condition can be directly determined as the target constraint condition.

[0155] In step 304, the target constraint condition in the first training sample is replaced with the empty condition to obtain the second training sample.

[0156] In this application, the target constraint condition in the first training sample can be replaced with the empty condition, and the other contents in the first training sample remain unchanged to obtain the second training sample.

[0157] For example, the target constraint condition can be replaced with a zero vector to obtain the second training sample.

[0158] In some embodiments, different image generation tasks can have corresponding training sets. In the training set of the text-to-image task, the training sample can contain a text condition. The training set of the image-to-image task has multiple types of training samples. Different types of training samples contain different constraint conditions. The first training sample can be selected from the training set according to the third probability of different types of training samples, and the target constraint condition in the first training sample can be determined. The target constraint condition in the first training sample is replaced with the empty condition to obtain the second training sample.

[0159] The third probability can refer to the probability that the corresponding type of training sample is selected as the first training sample.

[0160] Exemplarily, the plurality of types of training samples can include the following three cases: containing both the text condition and the image condition, containing the text condition and the empty image condition, and containing the image condition and the empty text condition. If the first training sample contains both the text condition and the image condition, the text condition or the image condition can be determined as the target constraint condition; if the first training sample contains the text condition as the constraint condition, the text condition can be determined as the target constraint condition; and if the first training sample contains only the image condition, the text condition can be determined as the target constraint condition.

[0161] As an example, if the current task is the image-to-image task, considering that the text condition and the image condition can both be empty conditions, the following four cases can be divided:

[0162] Case a: the condition part is the text condition + the image condition, and the no-condition part is the text condition + the empty image condition;

[0163] Case b: the condition part is the text condition + the empty image condition, and the no-condition part is the empty text condition + the empty image condition;

[0164] Case c: the condition part is the text condition + the image condition, and the no-condition part is the empty text condition + the image condition;

[0165] Case d: the condition part is the empty text condition + the image condition, and the no-condition part is the empty text condition + the empty image condition.

[0166] Among them, the condition part corresponds to the constraint condition case in the first training sample, and the no-condition part corresponds to the constraint condition case in the second training sample.

[0167] In the above image-to-image task, the above four cases can be trained with corresponding probabilities, and each training round can be trained in one case.

[0168] For example, in the training set corresponding to the image-to-image task, the training samples contain both the text condition and the image condition, if the image generation task of the current training round is the image-to-image task, the original training sample can be selected from the training set, and one case can be determined from the four cases according to the probability of the above four cases. The original training sample selected is processed to obtain the first training sample and the second training sample corresponding to the case.

[0169] Step 305, respectively according to the first training sample and the second training sample, the first image latent representation and the second image latent representation are obtained by using the teacher model.

[0170] Step 306, according to the first training sample, the third image latent representation is obtained by using the initial student model.

[0171] Step 307, training the initial student model according to the first image latent representation, the second image latent representation and the third image latent representation to obtain a target image generation model.

[0172] In the present application, steps 306-307 can adopt any of the implementation manners of the embodiments of the present application, and therefore will not be described here.

[0173] In the embodiments of the present application, the target image generation task is determined from the image generation tasks according to the second probability of each image generation task, and the first training sample is obtained according to the target image generation task, and the second training sample is obtained according to the first training sample. Therefore, the student model can be trained through multiple image generation tasks according to the probability of the image generation task, which not only meets the training needs of different image generation tasks, but also improves the robustness and applicability of the image generation model. Moreover, by replacing the target constraint condition in the first training sample with the empty condition to obtain the second training sample, the second training sample can be enriched to realize sample diversification.

[0174] For ease of understanding, the following will be described in conjunction with Figure 4 . Figure 4 A schematic diagram of a training process of an image generation model provided by an embodiment of the present application.

[0175] As Figure 4 shown, the encoding vector x0 of the original image is added with noise through forward diffusion to obtain the noisy latent representation x n+k of the noisy image. n+k Based on the noisy latent representation x n+k , the image latent representation is obtained by using the student model. Moreover, based on the noisy latent representation x n+k , the image latent representation is obtained by using the teacher model. By adding noise to the image latent representation , the noisy latent representation x of the noisy image of the previous time step is obtained. Based on the noisy latent representation x , the image latent representation is obtained by using the student model. According to the image latent representation and the image latent representation , the loss is calculated, the parameters of the student model are adjusted according to the loss, until the training end condition is met, and finally the image generation model is obtained.

[0176] Wherein, the noise can be predicted by using the student model to obtain the noise latent representation, and the image latent representation Figure 4 (not shown) is obtained by denoising the noise latent representation.

[0177] In addition, the teacher model can be based on the noisy latent representation xn+k Conditional and unconditional noise prediction is performed, and the noise latent representations obtained are fused based on the guidance intensity, and the noise latent representations obtained after fusion are denoised to obtain image latent representations Figure 4 The conditional corresponds to the constraint condition contained in the first training sample, and the unconditional corresponds to the constraint condition contained in the second training sample.

[0178] The image generation model training method of the present application is suitable for consistency distillation training of image generation models. By introducing guidance intensity corresponding to text conditions and guidance intensity corresponding to image conditions, a multi-task data mixing method is used to select a distillation mode based on guidance intensity fusion for training at each training step. The training process is stable, and the quality of the generated images can be improved.

[0179] In addition, since consistency distillation training is usually performed after model fine-tuning, the fine-tuned model already has the ability to generate fine images, and the resolution of the images generated by the model is relatively high, such as reaching 1024 levels. Therefore, the data set for distillation training can use high-resolution, high-quality fine images.

[0180] Since consistency distillation training needs to solve ordinary differential equations based on the teacher model, the accuracy of the solution affects the effect of distillation training, so the training set for distillation training can be similar to the data domain of the training set of the teacher model, to ensure the accuracy of the ordinary differential equation solution of the teacher model.

[0181] To achieve the above embodiments, an image generation method is further provided in the embodiments of the present application. Figure 5 The flowchart of the image generation method provided by an embodiment of the present application is shown.

[0182] As shown in Figure 5 The image generation method comprises:

[0183] Step 501, obtaining the image generation constraint condition input by the user.

[0184] The image generation constraint condition can be a constraint condition used to constrain image generation. For example, the image generation constraint condition can be one or more of a text condition and an image condition.

[0185] In the present application, the user can input the image generation constraint condition in the dialogue interface of the image generation model client, and the client can send the obtained image generation constraint condition to the server, so that the server generates an image according to the image generation constraint condition.

[0186] For example, the user inputs the image generation constraint condition as "generate a landscape picture with flowers, trees and grass".​

[0187] Step 502, generating the target image according to the image generation constraint condition and using the image generation model.

[0188] The target image generation model can be trained by using the training method described in any of the above embodiments.

[0189] In this application, the image generation constraint condition can be encoded to obtain an encoding vector, and the encoding vector of the image generation constraint condition is input into the image generation model to generate the target image.

[0190] In the embodiments of this application, the image generation model is trained by using the above training method according to the image constraint condition input by the user to generate the image, which can improve the accuracy and quality of image generation and meet different image generation requirements.

[0191] Figure 6 The flowchart of the image generation method provided by another embodiment of this application is shown.

[0192] As shown in the figure, the image generation method comprises: Figure 6

[0193] Step 601, obtaining the image generation constraint condition input by the user.

[0194] In this application, step 601 can be implemented by using any of the embodiments of this application, and therefore will not be described here.

[0195] Step 602, performing noise prediction by using the image generation model according to the image generation constraint condition to obtain a first target latent representation.

[0196] In this application, the image generation constraint condition can be encoded to obtain an encoding vector of the image generation constraint condition, and the noise prediction is performed by using the image generation model according to the encoding vector of the image generation constraint condition to obtain the first target latent representation.

[0197] Step 603, replacing the image generation constraint condition by using the null condition to obtain a replaced constraint condition.

[0198] The replaced constraint condition can be one or more, which is not limited.

[0199] For example, the image generation constraint condition comprises a text condition and an image condition, and then three replaced constraint conditions can be obtained: text condition+null image condition, null text condition+image condition, and null text condition+null image condition.

[0200] Step 604, performing noise prediction by using the image generation model according to the replaced constraint condition to obtain a second target latent representation.​

[0201] In the present application, for each replaced constraint condition, the noise can be predicted by using the image generation model according to the encoding vector of the replaced constraint condition, to obtain a second target latent representation.

[0202] In step 605, the first target latent representation and the second target latent representation are fused to obtain a target fusion latent representation.

[0203] In some embodiments, the first target latent representation and the second target latent representation can be added to obtain the target fusion latent representation.

[0204] In some embodiments, the guidance intensity related to the replaced condition in the image generation constraint condition can be determined, and the first target latent representation and the second target latent representation are fused according to the guidance intensity to obtain the target fusion latent representation.

[0205] For example, if the image generation constraint condition includes a text condition and an image condition, the first guidance intensity corresponding to the text condition and the second guidance intensity corresponding to the image condition can be determined, and the first guidance intensity and the second guidance intensity are used to fuse the first target latent representation and the second target latent representation to obtain the target fusion latent representation.

[0206] As an example, the target fusion latent representation can be determined by using the following formula (6):

[0207]

[0208] In the formula (6), w1 is the second guidance intensity, and w2 is the first guidance intensity.

[0209] It can be understood that the formula is also applicable to the case where the image generation constraint condition is a text condition or an image condition. For example, if the image generation constraint condition is a text condition, only the first guidance intensity needs to be determined, and the target fusion latent representation is determined according to the first guidance intensity.

[0210] Therefore, in the case where the image generation task includes a text condition and an image condition, the guidance intensities corresponding to the two constraint conditions can be determined, and the target latent representations under different constraint conditions are fused according to these guidance intensities, which can improve the accuracy of the target fusion latent representation, and thus the quality of the generated image based on the target fusion latent representation can be improved.

[0211] In step 606, a target image is generated according to the target fusion latent representation.

[0212] In the present application, the target fusion latent representation can be denoised to obtain a denoised latent representation, and the denoised latent representation is decoded to obtain the target image.

[0213] In the embodiments of the present application, the constraint condition for image generation is replaced by using the empty condition to obtain a replaced constraint condition, noise prediction is performed using an image generation model based on the constraint condition for image generation and the replaced constraint condition, corresponding target noise latent representation is obtained, and the target latent representations under different constraint conditions are fused, which can improve the accuracy of the noise latent representation, thereby improving the accuracy of the generated image based on the fused latent representation, and further improving the image generation quality.

[0214] To achieve the above-mentioned embodiments, the embodiments of the present application also provide a training device of an image generation model. Figure 7 The structure diagram of the training device of the image generation model provided by an embodiment of the present application is shown.

[0215] As shown in Figure 7 The training device 700 of the image generation model includes:

[0216] The first acquisition module 710 is configured to acquire a first training sample and a second training sample corresponding to a target image generation task, wherein the second training sample is obtained by replacing a target constraint condition in the first training sample with an empty condition.

[0217] The first prediction module 720 is configured to perform noise prediction using a teacher model according to the first training sample and the second training sample respectively, and acquire a first noise latent representation and a second noise latent representation.

[0218] The second prediction module 730 is configured to perform noise prediction using an initial student model according to the first training sample, and acquire a third noise latent representation.

[0219] The training module 740 is configured to train the initial student model according to the first noise latent representation, the second noise latent representation and the third noise latent representation, and obtain a target image generation model.

[0220] Optionally, the training module 740 is configured to:

[0221] fuse the first noise latent representation and the second noise latent representation to obtain a fused noise latent representation;

[0222] train the initial student model according to the fused noise latent representation and the third noise latent representation to obtain the target image generation model.

[0223] Optionally, the training module 740 is configured to:

[0224] de-noise the fused noise latent representation to obtain a first image latent representation.

[0225] de-noising the third noise latent representation to obtain a second image latent representation;

[0226] training the initial student model according to the first image latent representation and the second image latent representation to obtain the target image generation model.

[0227] Optionally, the first training sample includes a first time step, and the training module 740 is configured to:

[0228] adding noise to the first image latent representation to obtain a noisy latent representation corresponding to a second time step; the second time step is a time step preceding the first time step;

[0229] performing noise prediction on the noisy latent representation by using the initial student model to obtain a fourth noise latent representation;

[0230] de-noising the fourth noise latent representation to obtain a third image latent representation;

[0231] training the initial student model according to a difference between the second image latent representation and the third image latent representation to obtain the target image generation model.

[0232] Optionally, the training module 740 is configured to:

[0233] determining a target intensity range from candidate intensity ranges corresponding to a plurality of condition types according to the target constraint condition;

[0234] determining a target guidance intensity from the target intensity range; the target guidance intensity is used to control a weight of the target constraint condition;

[0235] fusing the first noise latent representation and the second noise latent representation according to the target guidance intensity to obtain the fused noise latent representation.

[0236] Optionally, the training module 740 is configured to:

[0237] determining a target condition type to which the target constraint condition belongs from the plurality of condition types;

[0238] determining the target intensity range according to a candidate intensity range corresponding to the target condition type.

[0239] Optionally, the training module 740 is configured to:

[0240] determining a fifth noise latent representation according to a product of a sum of a preset value and the target guidance intensity and the first noise latent representation.

[0241] determine a sixth noise latent representation according to a product between the target guidance strength and the second noise latent representation;

[0242] determine the fusion noise latent representation according to a difference between the fifth noise latent representation and the sixth noise latent representation.

[0243] Optionally, the first obtaining module 710 is configured to:

[0244] obtain the first training sample matched with the target image generation task;

[0245] determine the target constraint condition from constraint conditions contained in the first training sample;

[0246] replace the target constraint condition in the first training sample with the empty condition to obtain the second training sample.

[0247] Optionally, the first obtaining module 710 is configured to:

[0248] in response to the constraint conditions contained in the first training sample being multiple, determine the target constraint condition from the multiple constraint conditions according to a first probability of the multiple constraint conditions.

[0249] Optionally, the training process of the initial student model includes multiple training rounds, and the target image generation task is associated with the training rounds.

[0250] Optionally, the apparatus can further include:

[0251] a second obtaining module configured to obtain a second probability of each image generation task;

[0252] a determining module configured to determine the target image generation task from the each image generation task according to the second probability of the each image generation task.

[0253] It should be noted that the above explanation of the training method of the image generation model is also applicable to the training apparatus of the image generation model of this embodiment, and thus will not be described here again.

[0254] In the embodiments of the present application, the first training sample and the second training sample corresponding to the target image generation task are obtained, the teacher model is used to perform noise prediction according to the first training sample and the second training sample respectively, the noise prediction result under the target constraint condition and the noise prediction result under the non-target constraint condition are obtained, and the initial student model is used to perform noise prediction according to the first training sample. According to the noise prediction result of the teacher model under the target constraint condition and the noise prediction result under the non-target constraint condition, and the noise prediction result of the initial student model, the initial student model is trained, which can make the student model learn the information of different constraint conditions, and improve the accuracy and applicability of the image generation model.

[0255] To implement the above-mentioned embodiments, the embodiments of the present application also provide an image generation device. Figure 8 The structure diagram of the image generation device provided by an embodiment of the present application is shown.

[0256] As shown in Figure 8 The image generation device 800 includes:

[0257] The acquisition module 810 is configured to acquire an image generation constraint condition input by a user.

[0258] The generation module 820 is configured to generate a target image by using an image generation model according to the image generation constraint condition, wherein the image generation model is trained by using the training method according to any one of the above-mentioned embodiments.

[0259] Optionally, the generation module 820 is configured to:

[0260] perform noise prediction by using the image generation model according to the image generation constraint condition, to obtain a first target latent representation;

[0261] replace the image generation constraint condition with an empty condition, to obtain a replaced constraint condition;

[0262] perform noise prediction by using the image generation model according to the replaced constraint condition, to obtain a second target latent representation;

[0263] fuse the first target latent representation and the second target latent representation, to obtain a target fused latent representation;

[0264] generate the target image according to the target fused latent representation.

[0265] Optionally, the generation module 820 is configured to:

[0266] determine a first guidance intensity corresponding to the text condition and a second guidance intensity corresponding to the image condition;

[0267] According to the first guidance intensity and the second guidance intensity, the first target potential representation and the second target potential representation are fused to obtain a target fused potential representation.

[0268] It should be noted that the foregoing explanation of the image generation method embodiment is also applicable to the training device of the image generation model of the embodiment, and thus will not be described here again.

[0269] In the embodiments of the present application, the image generation model trained by the training method according to the image constraint condition input by the user can generate images, which can improve the accuracy and quality of image generation and meet different image generation requirements.

[0270] According to the embodiments of the present application, the present application also provides an electronic device, a readable storage medium and a computer program product.

[0271] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.

[0272] As shown in Figure 9 The device 900 includes a computing unit 901 that can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 902 or a computer program loaded into a RAM (Random Access Memory) 903 from a storage unit 908. Various programs and data required for the operation of the device 900 can also be stored in the RAM 903. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An I / O (Input / Output) interface 905 is also connected to the bus 904.

[0273] A plurality of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0274] The computing unit 901 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 901 performs various methods and processes described above, such as the training method of the image generation model. For example, in some embodiments, the training method of the image generation model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded to the RAM 903 and executed by the computing unit 901, one or more steps of the training method of the image generation model described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the training method of the image generation model by other any appropriate means, such as by means of firmware.

[0275] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a Field Programmable Gate Array (FPGA), an Application-Specific Integrated Circuit (ASIC), an Application Specific Standard Product (ASSP), a System on a Chip (SOC), a Complex Programmable Logic Device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0276] Program code for carrying out methods of the present application can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0277] In the context of this application, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include a linearly-programmed electronic storage, a portable computer diskette, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory), or a flash memory, an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0278] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0279] The systems and techniques described here can be implemented in a computing system that includes a back-end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front-end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.

[0280] The computer system can include clients and servers. This relationship can be

[0281] It should be noted that the electronic device for implementing the image generation method of the embodiments of the present application is similar in structure to the electronic device shown in the above figure, and therefore will not be described here.

[0282] According to the embodiments of the present application, the present application also provides a computer program product, when the instruction processor in the computer program product executes, executes the training method of the image generation model or the image generation method proposed in the above embodiments of the present application.

[0283] It should be understood that the steps can be reordered, added or deleted using the various forms of flow shown above. For example, the steps described in the present application can be executed in parallel, sequentially or in different order, as long as the desired results of the technical solutions disclosed in the present application can be achieved, which is not limited herein.

[0284] The above detailed description does not constitute a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A training method for an image generation model, comprising: Obtain the first training sample and the second training sample corresponding to the target image generation task; wherein, the second training sample is obtained by replacing the target constraint in the first training sample with an empty condition; Based on the first training sample and the second training sample respectively, noise prediction is performed using the teacher model to obtain the first noise latent representation and the second noise latent representation; Based on the first training sample, noise prediction is performed using the initial student model to obtain a third latent noise representation. The initial student model is trained based on the first noise latent representation, the second noise latent representation, and the third noise latent representation to obtain the target image generation model.

2. The method as described in claim 1, wherein, The step of training the initial student model based on the first noise latent representation, the second noise latent representation, and the third noise latent representation to obtain the target image generation model includes: The first noise latent representation and the second noise latent representation are fused to obtain a fused noise latent representation; The initial student model is trained based on the fused noise latent representation and the third noise latent representation to obtain the target image generation model.

3. The method as described in claim 2, wherein, The step of training the initial student model based on the fused noise latent representation and the third noise latent representation to obtain the target image generation model includes: The fused noise latent representation is denoised to obtain the first image latent representation; The third noise latent representation is denoised to obtain the second image latent representation; The initial student model is trained based on the first image latent representation and the second image latent representation to obtain the target image generation model.

4. The method of claim 3, wherein, The first training sample includes a first time step. The step of training the initial student model based on the first image latent representation and the second image latent representation to obtain the target image generation model includes: Noise is added to the latent representation of the first image to obtain the noisy latent representation corresponding to the second time step; wherein, the second time step is the time step preceding the first time step; Based on the noisy latent representation, noise prediction is performed using the initial student model to obtain the fourth noise latent representation; The fourth noise latent representation is denoised to obtain the third image latent representation; The initial student model is trained based on the difference between the second image latent representation and the third image latent representation to obtain the target image generation model.

5. The method of claim 2, wherein, The step of fusing the first noise latent representation and the second noise latent representation to obtain a fused noise latent representation includes: Based on the target constraints, the target intensity range is determined from the candidate intensity ranges corresponding to multiple condition types; A target guidance intensity is determined from the target intensity range; wherein, the target guidance intensity is used to control the weight of the target constraint conditions; Based on the target guidance intensity, the first noise latent representation and the second noise latent representation are fused to obtain the fused noise latent representation.

6. The method of claim 5, wherein, The step of determining the target intensity range from candidate intensity ranges corresponding to multiple condition types based on the target constraint conditions includes: Determine the target condition type to which the target constraint belongs from the plurality of condition types; The target intensity range is determined based on the candidate intensity range corresponding to the target condition type.

7. The method of claim 5, wherein, The step of fusing the first noise latent representation and the second noise latent representation according to the target guidance intensity to obtain the fused noise latent representation includes: The fifth noise latent representation is determined by multiplying the sum of the preset value and the target guidance intensity with the first noise latent representation; A sixth noise latent representation is determined based on the product between the target guidance strength and the second noise latent representation; The fused noise latent representation is determined based on the difference between the fifth noise latent representation and the sixth noise latent representation.

8. The method of claim 1, wherein, The acquisition of the first and second training samples corresponding to the target image generation task includes: Obtain the first training sample that matches the target image generation task; The target constraint is determined from the constraints contained in the first training sample; The target constraint in the first training sample is replaced using the empty condition to obtain the second training sample.

9. The method of claim 8, wherein, The target image generation task is a graph-to-graph task, and determining the target constraints from the constraints contained in the first training samples includes: In response to the fact that the first training sample contains multiple constraints, the target constraint is determined from the multiple constraints based on the first probability of the multiple constraints.

10. The method according to any one of claims 1-9, wherein, The initial student model training process includes multiple training rounds, and the target image generation task is associated with the training rounds.

11. The method of any one of claims 1-9, further comprising: Obtain the second probability for each image generation task; The target image generation task is determined from the image generation tasks based on the second probability of each image generation task.

12. An image generation method, comprising: Obtain the image input from the user and generate the constraints; Based on the image generation constraints, a target image is generated using an image generation model; wherein the image generation model is trained using the method according to any one of claims 1-11.

13. The method of claim 12, wherein, The step of generating the target image using an image generation model based on the image generation constraints includes: Based on the image generation constraints, noise prediction is performed using the image generation model to obtain a first target potential representation. The image generation constraints are replaced using empty conditions to obtain the replaced constraints; Based on the replaced constraints, noise prediction is performed using the image generation model to obtain the potential representation of the second target. The first target potential representation and the second target potential representation are fused to obtain the target fused potential representation; The target image is generated based on the target fusion latent representation.

14. The method of claim 13, wherein, The image generation constraints include text conditions and image conditions. The fusion of the first target latent representation and the second target latent representation to obtain a target fused latent representation includes: Determine the first guidance intensity corresponding to the text condition and the second guidance intensity corresponding to the image condition; Based on the first guidance strength and the second guidance strength, the first target potential representation and the second target potential representation are fused to obtain the target fused potential representation.

15. A training device for an image generation model, comprising: The first acquisition module is used to acquire the first training sample and the second training sample corresponding to the target image generation task; wherein the second training sample is obtained by replacing the target constraint in the first training sample with an empty condition; The first prediction module is used to perform noise prediction using the teacher model based on the first training sample and the second training sample respectively, and to obtain the first noise latent representation and the second noise latent representation. The second prediction module is used to perform noise prediction based on the first training sample and the initial student model to obtain a third noise latent representation. The training module is used to train the initial student model based on the first noise latent representation, the second noise latent representation, and the third noise latent representation to obtain the target image generation model.

16. An image generation apparatus, comprising: The acquisition module is used to acquire the image generation constraints input by the user; A generation module is used to generate a target image based on the image generation constraints and using an image generation model; wherein the image generation model is trained by the method according to any one of claims 1-11.

17. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-14.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-14.

19. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-14.