Image generation method and device

By adjusting the denoising time series in the diffusion model to reduce the difference between the added noise and the denoising samples, the data distribution inconsistency caused by denoising error in the prior art is solved, the model performance is improved and more accurate target images are generated.

CN116468623BActive Publication Date: 2025-05-06ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310268935.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-14
Publication Date
2025-05-06
Estimated Expiration
2043-03-14

AI Technical Summary

Technical Problem

During the iteration process, the existing diffusion model has denoising errors due to algorithm errors and acceleration methods, resulting in inconsistent data distributions during the noise addition and denoising process of parameters at the same time, affecting the model performance.

Method used

By obtaining noise information and inputting the image generation model, the target time series is used to reduce the sample difference between the noise-added sample in the noise-added sample sequence and the denoised sample in the denoised sample sequence as the training target, and the denoised time series is adjusted so that it is approximately the same as the noise-added time series.

Benefits of technology

The denoising error is reduced, so that the image generation model is consistent in the data distribution under the same time parameters during the training phase, thereby improving model performance and making the generated target image more accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116468623B_ABST
    Figure CN116468623B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide an image generation method and device, wherein the image generation method includes: obtaining noise information; inputting the noise information into an image generation model to obtain a target image, wherein the image generation model is used to process the noise information based on a target time series, the target time series is obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as a training target, and the target time series includes time parameters for sequentially processing the noise information. The time series is adjusted so that the data distribution of the image generation model in the training stage under the same time parameters in the process of stepwise denoising is approximately the same as the data distribution in the process of stepwise denoising, thereby reducing denoising errors and making the intermediate processing results obtained when processing the noise information more accurate, thereby ensuring the performance of the image generation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of model training technology, and in particular to image generation methods. Background Art

[0002] The diffusion model is a generative model that can gradually destroy the image distribution into Gaussian noise by constructing a Markov chain, and then gradually denoise the generated image through network learning of the reverse distribution. The diffusion model has achieved very impressive results in a variety of tasks, including but not limited to multimodal generation tasks such as text generation and image generation.

[0003] When the image distribution is gradually destroyed into Gaussian noise, Gaussian noise is further added to the current noisy data at each step of the iteration process, so that the iterative process of gradual noisy addition determines the data distribution at each time parameter. Similarly, stepwise denoising, as the inverse process of stepwise noisy addition, is also an iterative process, and the data distribution at each time parameter is also determined during the iterative process. However, due to the errors in the algorithm of the iterative process, and the current iterative process is often implemented through acceleration methods, further denoising errors will be caused, resulting in different data distributions in the stepwise noisy process and the stepwise denoising process under the same time parameter, affecting the performance of the diffusion model. Summary of the invention

[0004] In view of this, the embodiments of this specification provide three image generation methods. One or more embodiments of this specification also involve three image generation devices, a data processing method for image generation, a data processing device for image generation, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects existing in the prior art.

[0005] According to a first aspect of an embodiment of this specification, there is provided an image generation method, comprising:

[0006] Obtain noise information;

[0007] The noise information is input into an image generation model to obtain a target image, wherein the image generation model is used to process the noise information based on a target time series, the target time series is obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as a training goal, and the target time series includes time parameters for processing the noise information in sequence.

[0008] According to a second aspect of the embodiments of this specification, there is provided an image generating device, including:

[0009] An acquisition module, configured to acquire noise information;

[0010] An input module is configured to input the noise information into an image generation model to obtain a target image, wherein the image generation model is used to process the noise information based on a target time series, the target time series is obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as a training target, and the target time series includes time parameters for processing the noise information in sequence.

[0011] According to a third aspect of the embodiments of this specification, there is provided an image generation method, including:

[0012] Obtaining an image generation request, wherein the image generation request carries text information;

[0013] Inputting the text information into an image generation model to obtain a target image, wherein the image generation model is used to process the noise information based on a target time series, the target time series is obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as a training target, and the target time series includes time parameters for sequentially processing the noise information;

[0014] The target image is rendered.

[0015] According to a fourth aspect of the embodiments of this specification, there is provided an image generating device, including:

[0016] An acquisition module is configured to acquire an image generation request, wherein the image generation request carries text information;

[0017] An input module is configured to input the text information into an image generation model to obtain a target image, wherein the image generation model is used to process the noise information based on a target time series, the target time series is obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as a training target, and the target time series includes time parameters for sequentially processing the noise information;

[0018] The rendering module is configured to render the target image.

[0019] According to a fifth aspect of the embodiments of this specification, there is provided an image generation method, which is applied to a cloud-side device, including:

[0020] An image generation request sent by a receiving end-side device, wherein the image generation request carries noise information;

[0021] Inputting the noise information into an image generation model to obtain a target image, wherein the image generation model is used to process the noise information based on a target time series, the target time series is obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as a training target, and the target time series includes time parameters for sequentially processing the noise information;

[0022] The target image is sent to the client device for rendering.

[0023] According to a sixth aspect of the embodiments of this specification, there is provided an image generation device, which is applied to a cloud-side device, including:

[0024] A receiving module, configured to receive an image generation request sent by a terminal device, wherein the image generation request carries noise information;

[0025] An input module is configured to input the noise information into an image generation model to obtain a target image, wherein the image generation model is used to process the noise information based on a target time series, the target time series is obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as a training target, and the target time series includes time parameters for sequentially processing the noise information;

[0026] The sending module is configured to send the target image to the terminal device for rendering.

[0027] According to a seventh aspect of the embodiments of this specification, there is provided a data processing method for image generation, which is applied to a cloud-side device, and includes:

[0028] Determine a sample image set, wherein the sample image set includes a plurality of sample images;

[0029] Extract any one sample image from the sample image set, perform noise processing on the sample image, and obtain a noisy sample sequence corresponding to the noisy time series, wherein the noisy sample sequence includes at least two noisy samples;

[0030] The last noisy sample in the noisy sample sequence is used as a target noisy sample, and based on the denoising time sequence, the target noisy sample is denoised to obtain a denoised sample sequence corresponding to the denoising time sequence, wherein the initial denoising time sequence is in reverse order to the denoising time sequence;

[0031] Determine a sample difference between a noisy sample in the noisy sample sequence and a denoised sample in the denoised sample sequence corresponding to the noisy sample, and adjust the denoised time series according to the sample difference to obtain an adjusted denoised time series;

[0032] Continue to perform the step of extracting any one sample image from the sample image set until an image generation model that meets the training stop condition is obtained, and determine the denoised time series corresponding to the image generation model that meets the training stop condition as the target time series.

[0033] According to an eighth aspect of the embodiments of this specification, there is provided a data processing device for image generation, which is applied to a cloud-side device, including:

[0034] A determination module is configured to determine a sample image set, wherein the sample image set includes a plurality of sample images;

[0035] A noise adding module is configured to extract any sample image from the sample image set, perform noise adding processing on the sample image, and obtain a noise adding sample sequence corresponding to the noise adding time series, wherein the noise adding sample sequence includes at least two noise adding samples;

[0036] a denoising module configured to take the last denoised sample in the denoised sample sequence as a target denoised sample, and denoise the target denoised sample based on a denoised time sequence to obtain a denoised sample sequence corresponding to the denoised time sequence, wherein an initial denoised time sequence is in reverse order to the denoised time sequence;

[0037] an adjustment module configured to determine a sample difference between a noisy sample in the noisy sample sequence and a denoised sample in the denoised sample sequence corresponding to the noisy sample, and adjust the denoised time series according to the sample difference to obtain an adjusted denoised time series;

[0038] The execution module is configured to continue to execute the step of extracting any one sample image from the sample image set until an image generation model that meets the training stop condition is obtained, and the denoised time series corresponding to the image generation model that meets the training stop condition is determined as the target time series.

[0039] According to a ninth aspect of the embodiments of this specification, a computing device is provided, including:

[0040] Memory and processor;

[0041] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above method are implemented.

[0042] According to the tenth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and the instructions implement the steps of the above method when executed by a processor.

[0043] According to an eleventh aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above method.

[0044] An embodiment of the present specification provides an image generation method, which obtains noise information; inputs the noise information into an image generation model to obtain a target image, wherein the image generation model is used to process the noise information based on a target time series, the target time series is obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as a training goal, and the target time series includes time parameters for processing the noise information in sequence.

[0045] The above method obtains a target time series by taking reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as the training target, and processes the noise information based on the target time series so that the sample difference between the noisy sample sequence and the denoised sample sequence is minimized, thereby adjusting the target time series, so that the data distribution of the image generation model in the process of gradual noisy and the data distribution of the process of gradual denoising under the same time parameters in the training stage are approximately the same, thereby reducing the denoising error, making the intermediate processing result obtained when processing the noise information more accurate, thereby ensuring the performance of the image generation model, and further making the target image obtained by the image generation model more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 This is a schematic diagram of an application scenario of an image generation method provided by an embodiment of this specification;

[0047] Figure 2 is a flow chart of an image generation method provided by an embodiment of this specification;

[0048] Figure 3 is a schematic diagram of a noise-added sample sequence and a noise-removed sample sequence in an image generation method provided by an embodiment of the present specification;

[0049] Figure 4 is a schematic diagram of an image generation method provided by an embodiment of this specification;

[0050] Figure 5 is a processing flow chart of an image generation method provided by an embodiment of this specification;

[0051] Figure 6 is a structural schematic diagram of an image generating device provided by an embodiment of this specification;

[0052] Figure 7is a flow chart of an image generation method provided according to an embodiment of the present specification;

[0053] Figure 8 is a structural schematic diagram of an image generating device provided by an embodiment of this specification;

[0054] Fig. 9 is a flow chart of an image generation method provided according to an embodiment of the present specification;

[0055] Fig.10 is a structural schematic diagram of an image generating device provided by an embodiment of this specification;

[0056] Fig.11 is a flow chart of a data processing method for image generation provided according to one embodiment of the present specification;

[0057] Fig.12 is a structural schematic diagram of a data processing device for image generation provided by an embodiment of this specification;

[0058] Fig.13 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION

[0059] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.

[0060] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0061] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0062] First, the terms involved in one or more embodiments of this specification are explained.

[0063] Generative model: refers to a model that can randomly generate observation data.

[0064] Diffusion process: A random process that gradually adds noise to the data, with an inverse process to restore the data distribution. Can be used to build generative models.

[0065] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0066] In this specification, three image generation methods are provided. This specification also involves three image generation devices, a data processing method for image generation, a data processing device for image generation, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0067] See also Figure 1 , Figure 1 A schematic diagram of an application scenario of an image generation method provided according to an embodiment of the present specification is shown.

[0068] Figure 1 The cloud-side device 102 and the terminal-side device 104 are included, wherein the cloud-side device 102 can be understood as a cloud server. Of course, in another feasible solution, the cloud-side device 102 can also be replaced by a physical server; the terminal-side device 104 includes but is not limited to a desktop computer, a laptop computer, etc.; for ease of understanding, in the embodiments of this specification, the cloud-side device 102 is a cloud server and the terminal-side device 104 is a laptop computer for detailed introduction.

[0069] In specific implementation, the cloud-side device 102 may be deployed with an image generation model. The user may send noise information to the cloud-side device 102 through the terminal-side device 104, and the cloud-side device 102 sends the noise information to the image generation model to obtain a target image and send it to the terminal-side device 104. The terminal-side device 104 may render the target image and display it to the user, thereby realizing image generation based on noise information.

[0070] The image generation model can be obtained by training based on the noisy sample sequence and the denoised sample sequence. Specifically, the target time series can be obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as the training target. The image processing model can process the noise information based on the target time series to obtain the target image, thereby ensuring the accuracy of the obtained target image.

[0071] See also Figure 2 , Figure 2 A flowchart of an image generation method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0072] Step 202: Obtain noise information.

[0073] Specifically, the noise information may be information obtained from a noise information database.

[0074] Step 204: Input the noise information into an image generation model to obtain a target image, wherein the image generation model is used to process the noise information based on a target time series, the target time series is obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as a training goal, and the target time series includes time parameters for processing the noise information in sequence.

[0075] Specifically, after obtaining the noise information, the noise information may be input into an image generation model to obtain a target image.

[0076] The target image can be understood as an image generated after the noise information is processed by the image generation model, and the target image is an image that conforms to the data distribution. The image generation model can be a diffusion model. The noisy sample sequence can be understood as the sample data included in the noisy process in the image generation model training process. The sample data can be, for example, a sample image and a noisy sample. The sample image is used as the initial sample of the noisy process. The noisy sample is used as the noisy sample obtained by noisy processing the sample image. The noisy sample sequence can be a noisy sample sequence obtained by noisy processing the sample image based on the noisy time series. The denoised sample sequence can be understood as the sample data included in the denoising process in the image generation model training process. It can be a denoised sample sequence obtained by denoising the last noisy sample in the denoised sample sequence based on the denoising time series. Then, accordingly, the denoised sample sequence can include the last noisy sample and the denoised sample. The last noisy sample is used as the initial sample in the denoising process, and the denoised sample is used as the denoised sample obtained by denoising the initial sample. The denoised time series can be opposite to the denoised time series. For example, if the noisy time series is "t0, t1, t2, t3, t4", then the corresponding denoised time series is "t4, t3, t2, t1, t0". Usually, when training an image generation model, the noisy time series and the denoised time series can be determined based on the final noisy samples. The target time series can be understood as the adjusted denoised time series. The sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence can be understood as the difference in sample distribution between the noisy samples and the denoised samples.

[0077] The target time series may include at least two time parameters, and the time parameter can be understood as the time step for processing the noise information. For example, the noise information is processed based on the time parameter t2 to obtain the intermediate information after denoising corresponding to the time parameter t1, and t1 and t2 are both time parameters. Specifically, when the image generation model is used to process the noise information, the noise information can usually be processed by the neural network in the image generation model. The time parameter t2 and the noise information can be input into the neural network to obtain the intermediate information corresponding to the time parameter t1 output by the neural network. The information corresponding to the next time parameter can be predicted based on the time parameter t1 and the intermediate information.

[0078] In the specific implementation, before the noise information is input into the image generation model, in order to ensure the performance of the image generation model in processing noisy images in the application stage, the image generation model needs to be pre-trained first. Specifically, the denoising time series of the image generation model in the denoising process in the training stage needs to be adjusted so that the denoised sample sequence and the noisy sample sequence are approximately the same. The pre-training steps of the image generation model include the following steps one to five.

[0079] Step 1: Determine a sample image set, where the sample image set includes a plurality of sample images.

[0080] The sample image set may be understood as a sample set for pre-training an image generation model, wherein a plurality of sample images are stored in the sample image set. The sample image set may be stored in a cloud-side device, such as a cloud server. The sample image may be understood as an original image without adding noise.

[0081] Based on this, a sample image set can be obtained from the cloud-side device.

[0082] Step 2: extract any one sample image from the sample image set, perform noise processing on the sample image, and obtain a noisy sample sequence corresponding to the noisy time series, wherein the noisy sample sequence includes at least two noisy samples;

[0083] Specifically, after the sample image set is determined, any sample image can be extracted from the sample image set, and noise processing is performed on the sample image to obtain a noisy sample sequence corresponding to the noisy time series.

[0084] Among them, performing noise processing on the sample image can be understood as adding Gaussian noise to the sample image. The noise-added time series can be a time parameter of the image generation model set manually. Alternatively, the noise-added time series can also be constructed according to a preset time interval. Specifically, at least two noise-added time parameters can be determined according to the preset time interval, and the noise-added time series can be constructed according to the at least two noise-added time parameters. When the image generation model is trained using a cloud-side device, the construction of the noise-added time series can be implemented by the cloud-side device, or it can be constructed by the image generation model, and the embodiments of this specification are not limited here.

[0085] For example, a sample image X0 can be extracted from the sample image set. The sample image X0 is subjected to noise processing to obtain a noise sample sequence "X0, X1, X2, X3" corresponding to the noise time series "t0, t1, t2, t3". Among them, t0 can be understood as the starting point in the process of noise addition to the sample image. The noise time parameter t0 corresponds to the noise sample X0, the noise time parameter, the noise time parameter t1 corresponds to the noise sample X1, the noise time parameter t2 corresponds to the noise sample X2, and the noise time parameter t3 corresponds to the noise sample X3.

[0086] In a specific implementation, the step of performing noise processing on the sample image to obtain a noise sample sequence corresponding to the noise time series includes:

[0087] Performing noise processing on the sample image, and determining a noise sample corresponding to any noise time parameter in the noise time series according to the noise result;

[0088] A noisy sample sequence is constructed according to the noisy samples corresponding to any one of the noisy time parameters.

[0089] Among them, any noise adding time parameter in the noise adding time series can be understood as each noise adding time parameter in the noise adding time series.

[0090] Based on this, the sample image can be denoised, and the denoised sample corresponding to each denoising time parameter in the denoising time series can be determined according to the denoising result, and the denoised sample sequence can be constructed according to the denoising sample corresponding to each denoising time parameter.

[0091] Continuing with the above example, the noise adding time sequence is “t0, t1, t2, t3”, the sample image X0 (whose corresponding noise adding time parameter is t0) can be continuously noised, and the noise adding sample X1 corresponding to the noise adding time parameter t1, the noise adding sample X2 corresponding to the noise adding time parameter t2, and the noise adding sample X3 corresponding to the noise adding time parameter t3 can be extracted from the noise adding result of the continuous noise adding process on the sample image X0. Then the obtained noise adding sample sequence is “X0, X1, X2, X3”.

[0092] In addition, the sample image X0 (whose corresponding noise adding time parameter is t0) can be noised to obtain the noise adding sample X1 corresponding to the noise adding time parameter t1, and then the noise adding sample X1 can be noised to obtain the noise adding sample X2 corresponding to the noise adding time parameter t2, and then the noise adding sample X2 can be noised to obtain the noise adding sample X3 corresponding to the noise adding time parameter t3, then the noise adding sample sequence is "X0, X1, X2, X3".

[0093] In summary, by constructing a noisy sample sequence, since the noisy process as a forward process of the diffusion process is accurate and there is no numerical error, the noisy sample sequence can provide reference data for the subsequent adjustment of the denoising time series.

[0094] Step 3: taking the last noisy sample in the noisy sample sequence as the target noisy sample, and denoising the target noisy sample based on the denoising time sequence to obtain a denoised sample sequence corresponding to the denoising time sequence, wherein the initial denoising time sequence is in reverse order to the denoising time sequence;

[0095] Among them, the last noisy sample in the noisy sample sequence can be understood as the noisy sample corresponding to the last noisy time parameter in the noisy sample sequence. For example, for the aforementioned noisy sample sequence "the first noisy sample, the second noisy sample and the third noisy sample", the last noisy sample is the third noisy sample. The last noisy sample can be the final noisy sample obtained after gradually noisy the sample image.

[0096] The initial denoising time series can be understood as the first denoising time series after denoising the sample image, that is, the denoising time series before adjustment. The initial denoising time series is in reverse order to the denoising time series, which can be understood as the order of the denoising time parameters in the initial denoising time series is opposite to the order of the denoising time parameters in the denoising time series.

[0097] Since the initial denoising time series and the denoising time series are opposite, the initial denoising time series can be determined according to the denoising time series, and the last denoising sample in the denoising sample sequence is taken as the target denoising sample. The target denoising sample can be understood as the initial denoising sample in the denoising process. Based on the denoising time series, the initial denoising sample (i.e., the target denoising sample) is denoised to obtain the denoising sample sequence corresponding to the denoising time series.

[0098] Continuing with the above example, according to the denoising time series "t0, t1, t2, t3", the opposite denoising time series is determined to be "t3, t2, t1, t0", and the last denoising sample X3 in the denoising sample sequence "X0, X1, X2, X3" is used as the target denoising sample, that is, as the initial denoising sample X3 in the denoising process, based on the denoising time series "t3, t2, t1, t0", the initial denoising sample X3 (which corresponds to the denoising time parameter t3 and also corresponds to the denoising time parameter t3) is denoised to obtain the denoising sample sequence "X3, X'2, X'1, X'0" corresponding to the denoising time series "t3, t2, t1, t0".

[0099] In a specific implementation, the target noisy sample is denoised based on the denoised time series to obtain a denoised sample sequence corresponding to the denoised time series, including:

[0100] According to the denoising execution order, in the denoising time sequence, determining a target denoising time parameter;

[0101] Based on the target denoising time parameter, the sample to be processed is denoised to obtain a denoised sample corresponding to the next denoising time parameter of the target denoising time parameter, wherein for the first denoising time parameter in the denoising time sequence, the sample to be processed is the target noisy sample; for the mth denoising time parameter in the denoising time sequence, the sample to be processed is the denoised sample corresponding to the m-1th denoising time parameter, where m is greater than 1;

[0102] A denoised sample sequence is constructed according to the denoised samples corresponding to each denoised time parameter in the denoised time sequence.

[0103] Wherein, m is an integer. The denoising execution order can be understood as the order in which the initial denoising samples are denoised during the denoising process. For example, for the denoising time sequence "t3, t2, t1, t0", the target denoising time parameter is determined in the denoising time sequence according to the denoising execution order, which can be understood as determining the denoising time parameters t3, t2, t1, t0 in sequence according to the denoising order. That is, the target denoising node can be understood as any denoising time parameter in the denoising time sequence. The next denoising time parameter of the target denoising time parameter can be understood as the next denoising time parameter after the target denoising time parameter in the denoising time sequence. For example, for the denoising time sequence "t3, t2, t1, t0", when the target denoising time parameter is t3, the next denoising time parameter of the target denoising time parameter is t2; when the target denoising time parameter is t2, the next denoising time parameter of the target denoising time parameter is t1.

[0104] The samples to be processed can be understood as samples that need to be denoised at the target denoising time parameter. However, for the denoising time parameters at different times, the samples to be processed are also different. For example, for the first denoising time parameter t3 in the denoising time series, the sample to be processed is the last denoised sample in the denoised sample series (that is, the target denoised sample or the initial denoised sample). For another example, for the second denoising time parameter t2 in the denoising time series, the sample to be processed is the denoised sample corresponding to the denoising time parameter t2 obtained based on the denoising time parameter t3 and the initial denoising sample.

[0105] Using the above example, the denoising time sequence is "t3, t2, t1, t0". According to the denoising execution order, the target denoising time parameter t3 in the denoising time sequence is determined. According to the target denoising time parameter t3, the initial denoising sample X3 is denoised to obtain the denoised sample X'2 corresponding to the next denoising time parameter t2 of the target denoising time parameter t3. According to the denoising execution order, the target denoising time parameter t2 in the denoising time sequence is determined. According to the target denoising time parameter t2, the denoising sample X'2 is denoised to obtain the denoised sample X'1 corresponding to the next denoising time parameter t1 of the target denoising time parameter t2. According to the denoising execution order, the target denoising time parameter determined is t1. According to the target denoising time parameter t1, the denoising sample X'1 is denoised to obtain the denoised sample X'0 corresponding to the next denoising time parameter t0 of the target denoising time parameter t1. And according to the denoising samples corresponding to each denoising time parameter, a denoising sample sequence is constructed as "X3, X'2, X'1, X'0". It can be understood that the denoising sample sequence "X3, X'2, X'1, X'0" corresponds to the denoising time sequence "t3, t2, t1, t0".

[0106] In addition, since the sample image is used as the starting point in the noise sample sequence, the sample image should correspond to the last denoised sample in the denoised sample sequence. And the target noise sample is used as the starting point in the denoised sample sequence, and the target noise sample should correspond to the last noise sample in the noise sample sequence. Therefore, in order to determine the correspondence between the noise sample and the denoised sample, the noise sample sequence includes the sample image and the noise sample, and the denoised sample sequence may include the target noise sample and the denoised sample.

[0107] In summary, by constructing a denoised sample sequence, a data basis is provided for the subsequent adjustment of the denoised time series.

[0108] In specific implementation, when denoising the samples to be processed, it can be implemented according to the sampling of data distribution. The specific implementation method is as follows:

[0109] The denoising process is performed on the samples to be processed based on the target denoising time parameter to obtain denoised samples corresponding to the next denoising time parameter of the target denoising time parameter, including:

[0110] Based on the target denoising time parameter, denoising is performed on the sample to be processed to obtain a data distribution corresponding to the denoised sample, wherein the denoised sample is a denoised sample corresponding to the next denoising time parameter of the target denoising time parameter;

[0111] The data distribution corresponding to the denoising sample is sampled to obtain a denoising sample corresponding to a next denoising time parameter of the target denoising time parameter.

[0112] Specifically, when the target denoising time parameter is the first denoising time parameter in the denoising time sequence, the initial denoising sample (target denoising sample) is denoised based on the first denoising time parameter to obtain the data distribution corresponding to the second denoising sample, and then the data distribution corresponding to the second denoising sample is sampled to obtain the second denoising sample. When the target denoising time parameter is the nth denoising time parameter, the nth denoising sample corresponding to the nth denoising time parameter can be denoised based on the nth denoising time parameter to obtain the data distribution corresponding to the n+1th denoising sample, and then the data distribution corresponding to the n+1th denoising sample is sampled to obtain the n+1th denoising sample, where n is greater than 1 and n is an integer.

[0113] Continuing with the above example, when the target denoising time parameter is the first denoising time parameter t3 in the denoising time series, the initial denoising sample X3 can be denoised according to the first denoising time parameter t3 to obtain the data distribution corresponding to the second denoising sample X'2 (the second denoising sample X'2 corresponds to the second denoising time parameter t2), and then the data distribution corresponding to the second denoising sample X'2 is sampled to obtain the second denoising sample X'2 corresponding to the second denoising time parameter t2.

[0114] When n is 2, for the second denoising time parameter t2, the second denoising sample X'2 can be denoised according to the second denoising time parameter t2 to obtain the data distribution corresponding to the third denoising sample X'1, and then the data distribution corresponding to the second denoising sample X'1 is sampled to obtain the third denoising sample X'1 corresponding to the third denoising time parameter t1.

[0115] In summary, by sampling the data distribution, denoising of the samples to be processed can be achieved.

[0116] Step 4: determining a sample difference between a noisy sample in the noisy sample sequence and a denoised sample in the denoised sample sequence corresponding to the noisy sample, and adjusting the denoised time series according to the sample difference to obtain an adjusted denoised time series;

[0117] Among them, determining the denoised samples in the denoised sample sequence and the denoised samples corresponding to the denoised samples in the denoised sample sequence can be understood as the denoised samples and denoised samples obtained under the same denoising time parameters and denoising time parameters. For example, for the denoised time sequence "t0, t1, t2, t3", the corresponding denoised sample sequence is "X0, X1, X2, X3", and the denoised time sequence "t3, t2, t1, t0", the corresponding denoised sample sequence is "X3, X'2, X'1, X'0", where X0 as the initial image sample corresponds to the denoised sample X'0, the denoised sample X1 corresponds to the denoised sample X'1, and the denoised sample X2 corresponds to the denoised sample X'2. The denoising time parameter t0 corresponds to the sample image X0, the denoising time parameter t1 corresponds to the denoised sample X1, the denoising time parameter t2 corresponds to the denoised sample X2, and the denoising time parameter t3 corresponds to the denoised sample X3. The denoising time parameter t3 corresponds to the target noisy sample X3, the denoising time parameter t2 corresponds to the denoised sample X'2, the denoising time parameter t1 corresponds to the denoised sample X'1, and the denoising time parameter t0 corresponds to the denoised sample X'0.

[0118] In specific implementation, after adjusting the last denoising time parameter in the denoising time sequence, the denoising sample sequence may be updated, and then the next denoising time parameter may be adjusted according to the updated denoising sample sequence. The specific implementation method is as follows:

[0119] The step of determining a sample difference between a noisy sample in the noisy sample sequence and a denoised sample in the denoised sample sequence corresponding to the noisy sample, and adjusting the denoised time series according to the sample difference to obtain an adjusted denoised time series includes:

[0120] Determine the i-th denoising time parameter in the denoising time series, where i starts from 2 and is an integer;

[0121] According to the i-th denoising time parameter, determining in the denoising sample sequence a denoised sample corresponding to the i-th denoising time parameter, and determining in the denoising sample sequence a denoised sample corresponding to the i-th denoising time parameter;

[0122] Determine the similarity between the sample features of the noise-added sample and the sample features of the denoised sample, and adjust the i-1th denoising time parameter according to the similarity to obtain an adjusted denoising time series;

[0123] Determining whether the i-th denoising time parameter is the last denoising time parameter in the denoising time sequence;

[0124] If not, then determine the denoised sample sequence corresponding to the adjusted denoised time sequence, i is incremented by 1, and continue to perform the step of determining the i-th denoising time parameter in the denoised time sequence;

[0125] If so, the adjusted denoising time series is determined according to the adjusted i-1th denoising time parameter.

[0126] Among them, since the denoising time series and the denoising time series are opposite, then according to the i-th denoising time parameter, the corresponding denoising time parameter can be determined in the denoising sample sequence, and according to the corresponding denoising time parameter, the denoising sample corresponding to the i-th denoising time parameter is determined. And according to the i-th denoising time parameter, the corresponding denoising sample is determined in the denoising sample sequence. Determine the similarity between the sample features of the denoising sample and the sample features of the denoising sample, and adjust the i-1th denoising time parameter according to the similarity.

[0127] Since the denoised sample corresponding to the second denoising time parameter is obtained based on the first denoising time parameter and the initial denoising sample, the first denoising time parameter can be adjusted by the sample difference between the denoised sample and the denoised sample corresponding to the second denoising time parameter. Specifically, the second denoising time parameter in the denoised time sequence can be determined. Since the denoised time sequence is opposite to the denoised time sequence, the denoised sample corresponding to the second denoising time parameter can be determined in the denoised sample sequence according to the second denoising time parameter, and the denoised sample corresponding to the second denoising time parameter can be determined in the denoised sample sequence. The similarity between the sample feature of the denoised sample and the sample feature of the denoised sample is determined, and the first denoising time parameter is adjusted according to the similarity to obtain the adjusted denoised time sequence. It is determined whether the second denoising time parameter is the last denoising time parameter in the denoised time sequence. If so, the adjustment is terminated, and the adjusted denoised time sequence is constructed according to the adjusted denoising time parameter.

[0128] If not, the target noisy sample is denoised, the denoised sample sequence corresponding to the adjusted denoised time sequence is determined, and then the second denoising time parameter in the adjusted denoised time sequence is determined, and the above steps are continued until the determined denoising time parameter is the last denoising time parameter in the denoised time sequence.

[0129] Using the above example, for the denoised time series "t0, t1, t2, t3", the denoised time series "t3, t2, t1, t0", the denoised sample sequence "X0, X1, X2, X3" and the denoised sample sequence "X3, X'2, X'1, X'0", the second denoising time parameter t2 can be determined in the denoised time series, and the denoised sample X2 corresponding to t2 and the denoised sample X'2 corresponding to t2 can be determined. The similarity between the sample features of the denoised sample X2 and the sample features of the denoised sample X'2 is calculated, and the first denoising time parameter t3 is adjusted according to the similarity to obtain the adjusted first denoising time parameter t'3. Since the second denoising time parameter t2 is not the last denoising time parameter in the denoising time series, the adjusted denoising time series "t'3, t2, t1, t0" is constructed according to the adjusted first denoising time parameter t'3, and the target noisy sample is denoised to obtain the denoised sample sequence "X3, X"2, X"1, X"0" corresponding to the adjusted denoising time series "t'3, t2, t1, t0", and the third denoising time parameter t1 in the adjusted denoising time series is determined, and the noisy sample X1 and the denoised sample X"1 corresponding to t1 are determined, and the similarity between the sample features of the noisy sample X1 and the sample features of the denoised sample X"1 is calculated, and the second denoising time parameter t2 is adjusted according to the similarity to obtain the adjusted second denoising time parameter t'2. The third The denoising time parameter t1 is not the last denoising time parameter in the denoising time series. Therefore, according to the adjusted first denoising time parameter t'3 and the adjusted second denoising time parameter t'2, the adjusted denoising time series "t'3, t'2, t1, t0" and the denoising sample sequence "X3, X"'2, X"'1, X"'0" corresponding to the adjusted denoising time series "t'3, t'2, t1, t0" are constructed, and the fourth denoising time parameter t0 is continued to be determined, and the denoised sample X0 (i.e., the sample image) and the denoised sample X"'0 corresponding to t0 are determined, and the similarity between the sample features of the denoised sample X0 and the sample features of the denoised sample X"'0 is calculated. The third denoising time parameter t1 is adjusted according to the similarity to obtain the adjusted third denoising time parameter t'1. The fourth denoising time parameter t0 is the last denoising time parameter in the denoising time sequence. Therefore, the adjusted denoising time sequence can be determined as “t′3, t′2, t′1, t0” according to the adjusted denoising time parameter.

[0130] Specifically, Figure 3 FIG. 1 is a schematic diagram showing a noise-added sample sequence and a noise-removed sample sequence in an image generation method provided by an embodiment of the present specification. Figure 3As shown, the denoising sample sequence includes "sample image X0, denoising sample X1 and denoising sample X2", the sample image X0 (corresponding to the denoising time parameter t0) is denoised to obtain the denoised sample X1 corresponding to the denoising time parameter t1, and the denoising sample X1 is denoised to obtain the denoised sample X2 corresponding to the denoising time parameter t2. Correspondingly, the denoising sample sequence includes "target denoising sample X2, denoising sample X11, denoising sample X00", the target denoising sample X2 is denoised at the denoising time parameter t2 to obtain the denoising sample X11 corresponding to the denoising time parameter t1, and the denoising sample X11 is processed at the denoising time parameter t1 to obtain the denoising sample X00 corresponding to the denoising time parameter t0. For the same denoising time node and denoising time node t1, the denoising sample X11 corresponding to the denoising time parameter t1 can be determined, and the denoising sample X1 corresponding to the denoising time parameter t1 can be determined. The denoising time parameter t2 is adjusted according to the sample difference between the denoised sample X1 and the denoised sample X11.

[0131] In a specific implementation, adjusting the i-1th denoising time parameter according to the sample characteristics of the denoised sample and the sample characteristics of the denoised sample to obtain an adjusted denoised sample sequence includes:

[0132] Calculating the similarity between the sample features of the noise-added sample and the sample features of the denoised sample;

[0133] Adjusting the i-1th denoising time parameter according to the similarity to obtain an adjusted i-1th denoising time parameter;

[0134] According to the (i-1)th denoising time parameter, an adjusted denoising time series is obtained.

[0135] Specifically, the similarity between the sample features of the noisy sample and the sample features of the denoised sample can be calculated, and the denoising time parameter can be adjusted according to the similarity to obtain the adjusted denoising time parameter, thereby obtaining the adjusted denoising time series.

[0136] For example, for the denoising time series "t3, t2, t1, t0", the first denoising time parameter t3 is adjusted to obtain the adjusted first denoising time parameter t'3, then the adjusted denoising time series is "t'3, t2, t1, t0".

[0137] In practical applications, the denoising time parameter can be optimized based on the distance between the sample features of the noisy samples and the sample features of the denoised samples. Figure 4 , Figure 4 A schematic diagram of an image generating method provided by an embodiment of the present specification is shown.

[0138] Figure 4 In the process 402, the actual denoising process before adjusting the denoising time parameters is performed. θ , process 404 is the ideal denoising process, and process 406 is the actual denoising process after adjusting the denoising time parameter. θ,τ , 400 is the denoised sample Xt i , represents t i The corresponding denoised samples, 4042 is the denoised sample Xt i-1 , represents t i-1 The corresponding denoised sample, 4044 is the denoised sample Xt i-2, Represents t i-2 The corresponding denoised sample, 4046 is the denoised sample Xt i-1 , represents the time parameter t in the actual denoising process after adjusting the denoising time parameter i-1 The corresponding denoised sample, t i , t i-1 , t i-2 represents the denoising time parameter, ranging from 400 to 4042, indicating the data distribution q(x ti-1 |x ti ), from 4042 to 4044 represents the data distribution q(x ti-2 |x ti -1), 4022 is the denoised sample x'ti-1 in the actual denoising process before adjusting the denoising time parameter, and 4062 is the denoised sample x'ti-1 in the actual denoising process after adjusting the denoising time parameter. It can be understood that the distance between the denoised sample 4046 in process 404 and the iterative result 4062 in process 406 is the smallest, that is, they are approximately the same. For each iteration (i.e., each denoising) in the actual denoising process, the training f θ,τ The distance between the iterative result 4062 of the optimized time parameters (the iterative result of process 406) and the denoised sample 4042 in process 404 can complete the alignment operation of the denoising time parameters.

[0139] Step 5: Continue to execute the step of extracting any one sample image from the sample image set until an image generation model that meets the training stop condition is obtained, and determine the denoised time series corresponding to the image generation model that meets the training stop condition as the target time series.

[0140] Specifically, after obtaining the adjusted denoised time series, the aforementioned steps 1 to 4 can be continued to be performed, and the image generation model can be continuously trained until an image generation model that satisfies the training stop condition is obtained. Here, the denoised time series corresponding to the image generation model that satisfies the training stop condition is the denoised time series adjusted by the aforementioned steps 1 to 5, and the denoised time series corresponding to the image generation model that satisfies the training stop condition is finally obtained as the target time series. It can be understood that after obtaining the denoised time series after the first adjustment, any sample image can be extracted, and the sample image can be denoised to obtain a denoised sample sequence corresponding to the denoised time series, and the last denoised sample in the denoised sample sequence can be used as the target denoised sample, and the target denoised sample can be denoised based on the denoised time series after the first adjustment to obtain a denoised sample sequence corresponding to the denoised time series after the first adjustment. Based on the sample difference between the denoised sample in the denoised sample sequence and the denoised sample corresponding to the denoised sample in the denoised sample sequence, the denoised time series after the first adjustment is adjusted to obtain a denoised time series after the second adjustment. Then continue to perform the aforementioned steps 1 to 4.

[0141] The training stop condition can be understood as reaching a preset number of iterations or the loss value of the model reaching a preset loss value threshold.

[0142] In summary, the above method obtains the target time series by taking reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as the training target, and processes the noise information based on the target time series to minimize the sample difference between the noisy sample sequence and the denoised sample sequence, thereby adjusting the target time series, so that the data distribution of the image generation model in the process of gradual noisy and the data distribution of the process of gradual denoising under the same time parameters in the training stage are approximately the same, thereby reducing the denoising error, making the intermediate processing result obtained when processing the noise information more accurate, thereby ensuring the performance of the image generation model, and further making the target image obtained by the image generation model more accurate.

[0143] The following combination Figure 5 , taking the application of the image generation method provided in this specification in generating images from text as an example, the image generation method is further described. Figure 5 A processing flow chart of an image generation method provided by an embodiment of the present specification is shown, which specifically includes the following steps.

[0144] Step 502: the terminal device receives the text information input by the user in the image generation request upload box, receives the user's upload instruction for the image generation request, and sends the image generation request to the cloud device.

[0145] The terminal device displays an image generation request upload box to the user. The image generation request carries text information.

[0146] Specifically, the user clicks the control "OK" on the display interface of the terminal device, and the terminal device determines the upload instruction of the user for the image generation request based on the user's click instruction.

[0147] Step 504: The cloud-side device receives the image generation request and determines the text information.

[0148] Step 506: The cloud-side device inputs the determined text information into the image generation model to obtain a target image corresponding to the text information.

[0149] The image generation model here may be the image generation model after adjusting the denoising time parameters. The specific adjustment process is the same as that described above and will not be repeated here.

[0150] Specifically, after the text information is input into the image generation model, the text information can be input into a trained text encoder, the text encoder is used to extract the text features of the text information, and the text features are mapped to image features through a neural network, and finally a target image that conforms to the semantics of the text information is obtained based on the image features.

[0151] Step 508: The cloud-side device sends the target image to the terminal-side device.

[0152] Step 510: The terminal device renders the target image and displays it to the user in an output result display box.

[0153] In addition, the terminal-side device can receive the pre-trained image generation model sent by the cloud-side device, and use the image generation model to process text information on the terminal-side device.

[0154] In summary, the above method obtains the target time series by taking reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as the training target, and processes the noise information based on the target time series to minimize the sample difference between the noisy sample sequence and the denoised sample sequence, thereby adjusting the target time series, so that the data distribution of the image generation model in the process of gradual noisy and the data distribution of the process of gradual denoising under the same time parameters in the training stage are approximately the same, thereby reducing the denoising error, making the intermediate processing result obtained when processing the noise information more accurate, thereby ensuring the performance of the image generation model, and further making the target image obtained by the image generation model more accurate.

[0155] Corresponding to the above method embodiment, this specification also provides an image generating device embodiment, Figure 6FIG. 1 shows a schematic diagram of the structure of an image generating device provided by an embodiment of the present specification. Figure 6 As shown, the device comprises:

[0156] An acquisition module 602 is configured to acquire noise information;

[0157] The input module 604 is configured to input the noise information into an image generation model to obtain a target image, wherein the image generation model is used to process the noise information based on a target time series, the target time series is obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as a training target, and the target time series includes time parameters for processing the noise information in sequence.

[0158] In an optional embodiment, the device further includes a training module configured to:

[0159] Determine a sample image set, wherein the sample image set includes a plurality of sample images;

[0160] Extract any one sample image from the sample image set, perform noise processing on the sample image, and obtain a noisy sample sequence corresponding to the noisy time series, wherein the noisy sample sequence includes at least two noisy samples;

[0161] The last noisy sample in the noisy sample sequence is used as a target noisy sample, and based on the denoising time sequence, the target noisy sample is denoised to obtain a denoised sample sequence corresponding to the denoising time sequence, wherein the initial denoising time sequence is in reverse order to the denoising time sequence;

[0162] Determine a sample difference between a noisy sample in the noisy sample sequence and a denoised sample in the denoised sample sequence corresponding to the noisy sample, and adjust the denoised time series according to the sample difference to obtain an adjusted denoised time series;

[0163] Continue to perform the step of extracting any one sample image from the sample image set until an image generation model that meets the training stop condition is obtained, and determine the denoised time series corresponding to the image generation model that meets the training stop condition as the target time series.

[0164] In an optional embodiment, the training module is further configured to:

[0165] Performing noise processing on the sample image, and determining a noise sample corresponding to any noise time parameter in the noise time series according to the noise result;

[0166] A noisy sample sequence is constructed according to the noisy samples corresponding to any one of the noisy time parameters.

[0167] In an optional embodiment, the training module is further configured to:

[0168] According to the denoising execution order, in the denoising time sequence, determining a target denoising time parameter;

[0169] Based on the target denoising time parameter, the sample to be processed is denoised to obtain a denoised sample corresponding to the next denoising time parameter of the target denoising time parameter, wherein for the first denoising time parameter in the denoising time sequence, the sample to be processed is the target noisy sample; for the mth denoising time parameter in the denoising time sequence, the sample to be processed is the denoised sample corresponding to the m-1th denoising time parameter, where m is greater than 1;

[0170] A denoised sample sequence is constructed according to the denoised samples corresponding to each denoised time parameter in the denoised time sequence.

[0171] In an optional embodiment, the training module is further configured to:

[0172] Based on the target denoising time parameter, denoising is performed on the sample to be processed to obtain a data distribution corresponding to the denoised sample, wherein the denoised sample is a denoised sample corresponding to the next denoising time parameter of the target denoising time parameter;

[0173] The data distribution corresponding to the denoising sample is sampled to obtain a denoising sample corresponding to a next denoising time parameter of the target denoising time parameter.

[0174] In an optional embodiment, the training module is further configured to:

[0175] Determine the i-th denoising time parameter in the denoising time series, where i starts from 2 and is an integer;

[0176] According to the i-th denoising time parameter, determining in the denoising sample sequence a denoised sample corresponding to the i-th denoising time parameter, and determining in the denoising sample sequence a denoised sample corresponding to the i-th denoising time parameter;

[0177] Determine the similarity between the sample features of the noise-added sample and the sample features of the denoised sample, and adjust the i-1th denoising time parameter according to the similarity to obtain an adjusted denoising time series;

[0178] Determining whether the i-th denoising time parameter is the last denoising time parameter in the denoising time sequence;

[0179] If not, then determine the denoised sample sequence corresponding to the adjusted denoised time sequence, i is incremented by 1, and continue to perform the step of determining the i-th denoising time parameter in the denoised time sequence;

[0180] If so, the adjusted denoising time series is determined according to the adjusted i-1th denoising time parameter.

[0181] In an optional embodiment, the training module is further configured to:

[0182] Calculating the similarity between the sample features of the noise-added sample and the sample features of the denoised sample;

[0183] Adjusting the i-1th denoising time parameter according to the similarity to obtain an adjusted i-1th denoising time parameter;

[0184] According to the (i-1)th denoising time parameter, an adjusted denoising time series is obtained.

[0185] In summary, the above-mentioned device obtains a target time series by taking reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as the training target, and processes the noise information based on the target time series so that the sample difference between the noisy sample sequence and the denoised sample sequence is minimized, thereby adjusting the target time series, so that the data distribution of the image generation model in the process of gradual noisy and the data distribution of the process of gradual denoising under the same time parameters in the training stage are approximately the same, thereby reducing the denoising error, making the intermediate processing result obtained when processing the noise information more accurate, thereby ensuring the performance of the image generation model, and further making the target image obtained by the image generation model more accurate.

[0186] The above is a schematic scheme of an image generating device of this embodiment. It should be noted that the technical scheme of the image generating device and the technical scheme of the above image generating method belong to the same concept, and the details not described in detail in the technical scheme of the image generating device can be referred to the description of the technical scheme of the above image generating method.

[0187] See also Figure 7 , Figure 7 A flowchart of an image generation method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0188] Step 702: Obtain an image generation request, wherein the image generation request carries text information;

[0189] Step 704: inputting the text information into an image generation model to obtain a target image, wherein the image generation model is used to process the noise information based on a target time series, the target time series is obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as a training target, and the target time series includes time parameters for sequentially processing the noise information;

[0190] Step 706: Render the target image.

[0191] The above method obtains a target time series by taking reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as the training target, and processes the noise information based on the target time series so that the sample difference between the noisy sample sequence and the denoised sample sequence is minimized, thereby adjusting the target time series, so that the data distribution of the image generation model in the process of gradual noisy and the data distribution of the process of gradual denoising under the same time parameters in the training stage are approximately the same, thereby reducing the denoising error, making the intermediate processing result obtained when processing the noise information more accurate, thereby ensuring the performance of the image generation model, and further making the target image obtained by the image generation model more accurate.

[0192] Corresponding to the above method embodiment, this specification also provides an image generating device embodiment, Figure 8 FIG. 1 shows a schematic diagram of the structure of an image generating device provided by an embodiment of the present specification. Figure 8 As shown, the device comprises:

[0193] The acquisition module 802 is configured to acquire an image generation request, wherein the image generation request carries text information;

[0194] An input module 804 is configured to input the text information into an image generation model to obtain a target image, wherein the image generation model is used to process the noise information based on a target time series, the target time series is obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as a training target, and the target time series includes time parameters for sequentially processing the noise information;

[0195] The rendering module 806 is configured to render the target image.

[0196] The above-mentioned device obtains a target time series by taking reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as the training target, and processes the noise information based on the target time series so that the sample difference between the noisy sample sequence and the denoised sample sequence is minimized, thereby adjusting the target time series, so that the data distribution of the image generation model in the process of gradual noisy and the data distribution of the process of gradual denoising under the same time parameters in the training stage are approximately the same, thereby reducing the denoising error, making the intermediate processing result obtained when processing the noise information more accurate, thereby ensuring the performance of the image generation model, and further making the target image obtained by the image generation model more accurate.

[0197] See also Fig. 9 , Fig. 9 A flowchart of an image generation method provided according to an embodiment of the present specification is shown, which is applied to a cloud-side device and specifically includes the following steps.

[0198] Step 902: receiving an image generation request sent by a terminal device, wherein the image generation request carries noise information;

[0199] Step 904: inputting the noise information into an image generation model to obtain a target image, wherein the image generation model is used to process the noise information based on a target time series, the target time series is obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as a training target, and the target time series includes time parameters for sequentially processing the noise information;

[0200] Step 906: Send the target image to the terminal device for rendering.

[0201] The above method obtains a target time series by taking reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as the training target, and processes the noise information based on the target time series so that the sample difference between the noisy sample sequence and the denoised sample sequence is minimized, thereby adjusting the target time series, so that the data distribution of the image generation model in the process of gradual noisy and the data distribution of the process of gradual denoising under the same time parameters in the training stage are approximately the same, thereby reducing the denoising error, making the intermediate processing result obtained when processing the noise information more accurate, thereby ensuring the performance of the image generation model, and further making the target image obtained by the image generation model more accurate.

[0202] Corresponding to the above method embodiments, this specification also provides an image generating device embodiment, which is applied to a cloud-side device. Fig.10FIG. 1 shows a schematic diagram of the structure of an image generating device provided by an embodiment of the present specification. Fig.10 As shown, the device comprises:

[0203] The receiving module 1002 is configured to receive an image generation request sent by a terminal device, wherein the image generation request carries noise information;

[0204] An input module 1004 is configured to input the noise information into an image generation model to obtain a target image, wherein the image generation model is used to process the noise information based on a target time series, the target time series is obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as a training target, and the target time series includes time parameters for sequentially processing the noise information;

[0205] The sending module 1006 is configured to send the target image to the terminal device for rendering.

[0206] The above-mentioned device obtains a target time series by taking reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as the training target, and processes the noise information based on the target time series so that the sample difference between the noisy sample sequence and the denoised sample sequence is minimized, thereby adjusting the target time series, so that the data distribution of the image generation model in the process of gradual noisy and the data distribution of the process of gradual denoising under the same time parameters in the training stage are approximately the same, thereby reducing the denoising error, making the intermediate processing result obtained when processing the noise information more accurate, thereby ensuring the performance of the image generation model, and further making the target image obtained by the image generation model more accurate.

[0207] See also Fig.11 , Fig.11 A flowchart of a data processing method for image generation provided according to an embodiment of the present specification is shown, which is applied to a cloud-side device and specifically includes the following steps.

[0208] Step 1102: determining a sample image set, wherein the sample image set includes a plurality of sample images;

[0209] Step 1104: extract any one sample image from the sample image set, perform noise processing on the sample image, and obtain a noisy sample sequence corresponding to the noisy time series, wherein the noisy sample sequence includes at least two noisy samples;

[0210] Step 1106: taking the last noisy sample in the noisy sample sequence as a target noisy sample, and denoising the target noisy sample based on the denoising time sequence to obtain a denoised sample sequence corresponding to the denoising time sequence, wherein the initial denoising time sequence is in reverse order to the denoising time sequence;

[0211] Step 1108: determining a sample difference between a noisy sample in the noisy sample sequence and a denoised sample in the denoised sample sequence corresponding to the noisy sample, and adjusting the denoised time series according to the sample difference to obtain an adjusted denoised time series;

[0212] Step 1110: Continue to execute the step of extracting any one sample image from the sample image set until an image generation model that meets the training stop condition is obtained, and determine the denoising time series corresponding to the image generation model that meets the training stop condition as the target time series.

[0213] The above method obtains a target time series by taking reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as the training target, and processes the noise information based on the target time series so that the sample difference between the noisy sample sequence and the denoised sample sequence is minimized, thereby adjusting the target time series, so that the data distribution of the image generation model in the process of gradual noisy and the data distribution of the process of gradual denoising under the same time parameters in the training stage are approximately the same, thereby reducing the denoising error, making the intermediate processing result obtained when processing the noise information more accurate, thereby ensuring the performance of the image generation model, and further making the target image obtained by the image generation model more accurate.

[0214] Corresponding to the above method embodiments, this specification also provides an embodiment of a data processing device for image generation, which is applied to a cloud-side device. Fig.12 FIG. 1 is a schematic diagram showing a data processing device for image generation provided by an embodiment of the present specification. Fig.12 As shown, the device comprises:

[0215] A determination module 1202 is configured to determine a sample image set, where the sample image set includes a plurality of sample images;

[0216] The noise adding module 1204 is configured to extract any sample image from the sample image set, perform noise adding processing on the sample image, and obtain a noise adding sample sequence corresponding to the noise adding time series, wherein the noise adding sample sequence includes at least two noise adding samples;

[0217] The denoising module 1206 is configured to use the last denoised sample in the denoised sample sequence as a target denoised sample, and denoise the target denoised sample based on a denoised time sequence to obtain a denoised sample sequence corresponding to the denoised time sequence, wherein the initial denoised time sequence is in reverse order to the denoised time sequence;

[0218] The adjustment module 1208 is configured to determine a sample difference between a noisy sample in the noisy sample sequence and a denoised sample in the denoised sample sequence corresponding to the noisy sample, and adjust the denoised time series according to the sample difference to obtain an adjusted denoised time series;

[0219] The execution module 1210 is configured to continue to execute the step of extracting any one sample image from the sample image set until an image generation model that meets the training stop condition is obtained, and determine the denoising time series corresponding to the image generation model that meets the training stop condition as the target time series.

[0220] The above-mentioned device obtains a target time series by taking reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as the training target, and processes the noise information based on the target time series so that the sample difference between the noisy sample sequence and the denoised sample sequence is minimized, thereby adjusting the target time series, so that the data distribution of the image generation model in the process of gradual noisy and the data distribution of the process of gradual denoising under the same time parameters in the training stage are approximately the same, thereby reducing the denoising error, making the intermediate processing result obtained when processing the noise information more accurate, thereby ensuring the performance of the image generation model, and further making the target image obtained by the image generation model more accurate.

[0221] Fig.13 The block diagram of a computing device 1300 according to one embodiment of the present specification is shown. The components of the computing device 1300 include but are not limited to a memory 1310 and a processor 1320. The processor 1320 is connected to the memory 1310 via a bus 1330, and the database 1350 is used to store data.

[0222] The computing device 1300 also includes an access device 1340 that enables the computing device 1300 to communicate via one or more networks 1360. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1340 may include one or more of any type of network interface (e.g., a network interface card (NIC)) of wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.

[0223] In one embodiment of the present application, the above components of the computing device 1300 and Fig.13 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Fig.13 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of the present application. Those skilled in the art may add or replace other components as needed.

[0224] The computing device 1300 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 1300 may also be a mobile or stationary server.

[0225] The processor 1320 is used to execute the following computer executable instructions, which implement the steps of the above method when executed by the processor.

[0226] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above method belong to the same concept, and the details not described in detail in the technical scheme of the computing device can be referred to the description of the technical scheme of the above method.

[0227] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which can implement the steps of the above method when executed by a processor.

[0228] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above method belong to the same concept, and the details not described in detail in the technical scheme of the storage medium can be referred to the description of the technical scheme of the above method.

[0229] An embodiment of the present specification further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above method.

[0230] The above is an illustrative solution of a computer program of this embodiment. It should be noted that the technical solution of the computer program and the technical solution of the above method belong to the same concept, and the details not described in detail in the technical solution of the computer program can be referred to the description of the technical solution of the above method.

[0231] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0232] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0233] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0234] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0235] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A method for generating an image, comprising: Obtain noise information; The noise information is input into an image generation model to obtain a target image, wherein the image generation model is used to process the noise information based on a target time series, the target time series is obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as a training target, the target time series includes time parameters for sequentially processing the noise information, the denoised sample sequence is obtained by denoising the last noisy sample in the noisy sample sequence based on the denoised time series, the initial denoised time series is in reverse order with the denoised time series, and the denoised time series is obtained by adjusting the initial denoised time series based on the sample difference.

2. The method according to claim 1, before inputting the noise information into the image generation model, further comprising: Determine a sample image set, wherein the sample image set includes a plurality of sample images; Extract any one sample image from the sample image set, perform noise processing on the sample image, and obtain a noisy sample sequence corresponding to the noisy time series, wherein the noisy sample sequence includes at least two noisy samples; The last noisy sample in the noisy sample sequence is used as a target noisy sample, and based on the denoising time sequence, the target noisy sample is denoised to obtain a denoised sample sequence corresponding to the denoising time sequence, wherein the initial denoising time sequence is in reverse order to the denoising time sequence; Determine a sample difference between a noisy sample in the noisy sample sequence and a denoised sample in the denoised sample sequence corresponding to the noisy sample, and adjust the denoised time series according to the sample difference to obtain an adjusted denoised time series; Continue to perform the step of extracting any one sample image from the sample image set until an image generation model that meets the training stop condition is obtained, and determine the denoised time series corresponding to the image generation model that meets the training stop condition as the target time series.

3. The method according to claim 2, wherein the step of performing noise processing on the sample image to obtain a noise sample sequence corresponding to the noise time series comprises: Performing noise processing on the sample image, and determining a noise sample corresponding to any noise time parameter in the noise time series according to the noise result; A noisy sample sequence is constructed according to the noisy samples corresponding to any one of the noisy time parameters.

4. The method according to claim 2, wherein the step of denoising the target noisy sample based on the denoised time series to obtain a denoised sample sequence corresponding to the denoised time series comprises: According to the denoising execution order, in the denoising time sequence, determining a target denoising time parameter; Based on the target denoising time parameter, the sample to be processed is denoised to obtain a denoised sample corresponding to the next denoising time parameter of the target denoising time parameter, wherein for the first denoising time parameter in the denoising time sequence, the sample to be processed is the target noisy sample; for the mth denoising time parameter in the denoising time sequence, the sample to be processed is the denoised sample corresponding to the m-1th denoising time parameter, where m is greater than 1; A denoised sample sequence is constructed according to the denoised samples corresponding to each denoised time parameter in the denoised time sequence.

5. The method according to claim 4, wherein the denoising process is performed on the sample to be processed based on the target denoising time parameter to obtain the denoised sample corresponding to the next denoising time parameter of the target denoising time parameter, comprising: Based on the target denoising time parameter, denoising is performed on the sample to be processed to obtain a data distribution corresponding to the denoised sample, wherein the denoised sample is a denoised sample corresponding to the next denoising time parameter of the target denoising time parameter; The data distribution corresponding to the denoising sample is sampled to obtain a denoising sample corresponding to a next denoising time parameter of the target denoising time parameter.

6. The method according to claim 2, wherein determining a sample difference between a noisy sample in the noisy sample sequence and a denoised sample in the denoised sample sequence corresponding to the noisy sample, and adjusting the denoised time series according to the sample difference to obtain an adjusted denoised time series comprises: Determine the i-th denoising time parameter in the denoising time series, where i starts from 2 and is an integer; According to the i-th denoising time parameter, determining in the denoising sample sequence a denoised sample corresponding to the i-th denoising time parameter, and determining in the denoising sample sequence a denoised sample corresponding to the i-th denoising time parameter; Determine the similarity between the sample features of the noise-added sample and the sample features of the denoised sample, and adjust the i-1th denoising time parameter according to the similarity to obtain an adjusted denoising time series; Determining whether the i-th denoising time parameter is the last denoising time parameter in the denoising time sequence; If not, then determine the denoised sample sequence corresponding to the adjusted denoised time sequence, i is incremented by 1, and continue to perform the step of determining the i-th denoising time parameter in the denoised time sequence; If so, the adjusted denoising time series is determined according to the adjusted i-1th denoising time parameter.

7. The method according to claim 6, wherein determining the similarity between the sample features of the noisy sample and the sample features of the denoised sample, and adjusting the i-1th denoising time parameter according to the similarity to obtain the adjusted denoising time series comprises: Calculating the similarity between the sample features of the noise-added sample and the sample features of the denoised sample; Adjusting the i-1th denoising time parameter according to the similarity to obtain an adjusted i-1th denoising time parameter; According to the (i-1)th denoising time parameter, an adjusted denoising time series is obtained.

8. A method for generating an image, comprising: Obtaining an image generation request, wherein the image generation request carries text information; Input the text information into an image generation model to obtain a target image, wherein the image generation model is used to process noise information based on a target time series, the target time series is obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as a training target, the target time series includes time parameters for sequentially processing the noise information, the denoised sample sequence is obtained by denoising the last noisy sample in the noisy sample sequence based on the denoised time series, the initial denoised time series is in reverse order with the denoised time series, and the denoised time series is obtained by adjusting the initial denoised time series based on the sample difference; The target image is rendered.

9. An image generation method, applied to a cloud-side device, comprising: An image generation request sent by a receiving end-side device, wherein the image generation request carries noise information; The noise information is input into an image generation model to obtain a target image, wherein the image generation model is used to process the noise information based on a target time series, the target time series is obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as a training target, the target time series includes time parameters for sequentially processing the noise information, the denoised sample sequence is obtained by denoising the last noisy sample in the noisy sample sequence based on the denoised time series, the initial denoised time series is in reverse order with the denoised time series, and the denoised time series is obtained by adjusting the initial denoised time series based on the sample difference; The target image is sent to the client device for rendering.

10. A data processing method for image generation, applied to a cloud-side device, comprising: Determine a sample image set, wherein the sample image set includes a plurality of sample images; Extract any one sample image from the sample image set, perform noise processing on the sample image, and obtain a noisy sample sequence corresponding to the noisy time series, wherein the noisy sample sequence includes at least two noisy samples; The last noisy sample in the noisy sample sequence is used as a target noisy sample, and based on the denoising time sequence, the target noisy sample is denoised to obtain a denoised sample sequence corresponding to the denoising time sequence, wherein the initial denoising time sequence is in reverse order to the denoising time sequence; Determine a sample difference between a noisy sample in the noisy sample sequence and a denoised sample in the denoised sample sequence corresponding to the noisy sample, and adjust the denoised time series according to the sample difference to obtain an adjusted denoised time series; Continue to perform the step of extracting any one sample image from the sample image set until an image generation model that meets the training stop condition is obtained, and determine the denoised time series corresponding to the image generation model that meets the training stop condition as the target time series.

11. An image generating device, comprising: An acquisition module, configured to acquire noise information; An input module is configured to input the noise information into an image generation model to obtain a target image, wherein the image generation model is used to process the noise information based on a target time series, the target time series is obtained by reducing the sample difference between the noisy samples in the noisy sample sequence and the corresponding denoised samples in the denoised sample sequence as a training target, the target time series includes time parameters for sequentially processing the noise information, the denoised sample sequence is obtained by denoising the last noisy sample in the noisy sample sequence based on the denoised time series, the initial denoised time series is in reverse order with the denoised time series, and the denoised time series is obtained by adjusting the initial denoised time series based on the sample difference.

12. A computing device comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 10 are implemented.

13. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.

14. A computer program product, characterized in that The method comprises computer instructions which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Image processing model training method and device

    CN115424088A