Image generation method and device, equipment and storage medium

By acquiring the image unit sequence and combining the alternate execution of cache pruning steps and complete inference steps, the problem of the bottleneck of computing complexity and inference speed in the image generation task is solved, achieving more efficient image generation.

CN119941551APending Publication Date: 2025-05-06ALIPAY (HANGZHOU) INFORMATION TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411972412.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Diffusion Transformer limits its wide deployment in practical applications due to the computational complexity and inference speed bottlenecks in image generation tasks.

Method used

By acquiring the sequence of image units corresponding to the potential noise image and determining the target image characteristics of the image unit sequence in the feature processing block set, combining the alternating execution of the cache pruning step and the complete inference step, redundant calculations are reduced and inference speed is improved.

Benefits of technology

It effectively improves the inference speed of the image generation model, balances the acceleration effect and generation quality, and reduces the error introduced by the cache.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941551A_ABST
    Figure CN119941551A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image generation method and device, equipment and a storage medium, and the method comprises the steps: dividing a time step into a complete reasoning step and a cache trimming step in an image denoising process, trimming a part of image units in the cache trimming step, and carrying out the replacement through employing cache image features, thereby reducing the reasoning times of the image units, and improving the image denoising efficiency. The problem of redundant calculation caused by the fact that the number of image units is large and multiple times of reasoning are needed in the reasoning process of the image generation model is solved, so that the reasoning speed of the image generation model is increased, meanwhile, errors caused by cache introduction are reduced in combination with a complete reasoning step, and the acceleration effect and the generation quality are balanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to an image generation method, device, equipment and storage medium. Background Art

[0002] In recent years, diffusion models have shown great potential in generation tasks and have become one of the mainstream technologies in the fields of image, video, text and other data generation. In particular, the introduction of Diffusion Transformer (DiT) marks the combination of diffusion models and Transformer architecture, which has brought generation models to new heights in generation quality and scalability. Compared with the traditional U-Net architecture, DiT can handle complex generation tasks more finely with its powerful modeling capabilities, especially in high-resolution image generation.

[0003] Although the Diffusion Transformer performs well in generation tasks, its high computational complexity and inference speed bottleneck severely limit its widespread deployment in practical applications. Summary of the invention

[0004] The main purpose of this specification is to provide an image generation method, device, equipment and storage medium, aiming to improve the inference speed of the image generation model. The technical solution is as follows:

[0005] In a first aspect, an embodiment of the present specification provides an image generation method, comprising:

[0006] Acquire an image unit sequence corresponding to a potential noise image; the image unit sequence includes a plurality of image units;

[0007] If the time step type of the time step is a cache pruning step, then the first image feature corresponding to the image unit sequence in each processing block in the feature processing block set is obtained, and the target image feature of the image unit sequence is determined based on each of the first image features; each of the processing blocks includes at least one first processing block that determines the first image feature corresponding to the image unit sequence in the first processing block based on a second image feature and a third image feature, wherein the second image feature is a historical cache feature of the image unit in the target feature processing block at the previous time step, and the third image feature is an image feature of the image unit generated by a previous processing block of the target feature processing block;

[0008] If the time step type of the time step is a complete inference step, obtaining a fourth image feature corresponding to the image unit sequence in each of the processing blocks in the feature processing block set, and determining a target image feature of the image unit sequence based on each of the fourth image features; the fourth image feature is determined based on each of the third image features;

[0009] Determining image noise data corresponding to the time step based on the target image feature;

[0010] The potential noise image is denoised based on the image noise data to generate a target image corresponding to the time step.

[0011] In a second aspect, the embodiments of this specification provide an image generation model training method, including:

[0012] The sample potential noise image is converted into a sample image unit sequence by using a unitization module in an image generation model; the sample potential noise image is generated based on the sample image;

[0013] determining a total number of sample time steps and a time step type for each of the sample time steps;

[0014] If the time step type of the sample time step is a cache pruning step, then the first sample image feature corresponding to the sample image unit sequence in each processing block in the feature processing block set is obtained, and the sample target image feature of the image unit sequence is determined based on each of the first sample image features; each of the processing blocks includes at least one first processing block that determines the first sample image feature corresponding to the sample image unit sequence in the first processing block based on the second sample image feature and the third sample image feature, the second image feature is a historical cache feature of the image unit in the target feature processing block at the previous time step, and the third image feature is an image feature of the image unit generated by a previous processing block of the target feature processing block;

[0015] If the time step type of the sample time step is a complete inference step, obtaining a fourth sample image feature corresponding to the sample image unit sequence in each of the processing blocks in the feature processing block set, and determining a target sample image feature of the sample image unit sequence based on each of the fourth sample image features; the fourth sample image feature is determined based on each of the third sample image features;

[0016] Determining sample image noise data corresponding to the sample time step based on the sample target image feature;

[0017] Performing denoising on the sample potential noise image based on the sample image noise data to generate a sample generated image corresponding to the sample time step;

[0018] If the number of the sample time step does not reach the total number of sample time steps, the sample time step is updated, and the sample generated image is used as a sample potential noise image, and the step of converting the sample potential noise image into a sample image unit sequence by using a unitization module in the image generation model is executed until the number of the sample time step reaches the total number of sample time steps, thereby obtaining a sample generated image;

[0019] Based on a first preset loss function, a first predicted loss value of the sample image and the sample generated image is determined, and based on the first predicted loss value, supervised training is performed on the image generation model and model parameters of the image generation model are iteratively updated until the first predicted loss value converges, thereby obtaining a trained image generation model.

[0020] In a third aspect, an embodiment of the present specification provides an image generating device, including:

[0021] An acquisition unit, used to acquire an image unit sequence corresponding to a potential noise image; the image unit sequence includes a plurality of image units;

[0022] A first inference unit is used for obtaining, if the time step type of the time step is a cache pruning step, a first image feature corresponding to the image unit sequence in each processing block in a feature processing block set, and determining a target image feature of the image unit sequence based on each of the first image features; each of the processing blocks includes at least one first processing block that determines the first image feature corresponding to the image unit sequence in the first processing block based on a second image feature and a third image feature, wherein the second image feature is a historical cache feature of the image unit in the target feature processing block at the previous time step, and the third image feature is an image feature of the image unit generated by a previous processing block of the target feature processing block;

[0023] a second inference unit, configured to obtain, if the time step type of the time step is a complete inference step, a fourth image feature corresponding to the image unit sequence in each of the processing blocks in the feature processing block set, and determine a target image feature of the image unit sequence based on each of the fourth image features; the fourth image feature is determined based on each of the third image features;

[0024] A noise determination unit, configured to determine image noise data corresponding to the time step based on the target image feature;

[0025] A denoising processing unit is used to perform denoising on the potential noise image based on the image noise data to generate a target image corresponding to the time step.

[0026] In a fourth aspect, an embodiment of the present specification provides an image generating device, including:

[0027] A unitization unit, used to convert a sample potential noise image into a sample image unit sequence by using a unitization module in an image generation model; the sample potential noise image is generated based on the sample image;

[0028] a parameter determination unit, used to determine the total number of sample time steps and the time step type of each of the sample time steps;

[0029] A first training unit is used for obtaining, if the time step type of the sample time step is a cache pruning step, a first sample image feature corresponding to the sample image unit sequence in each processing block in a feature processing block set, and determining a sample target image feature of the image unit sequence based on each of the first sample image features; each of the processing blocks includes at least one first processing block that determines the first sample image feature corresponding to the sample image unit sequence in the first processing block based on a second sample image feature and a third sample image feature, wherein the second image feature is a historical cache feature of the image unit in the target feature processing block at the previous time step, and the third image feature is an image feature of the image unit generated by a previous processing block of the target feature processing block;

[0030] a second training unit, for obtaining, if the time step type of the sample time step is a complete inference step, fourth sample image features corresponding to the sample image unit sequence in each of the processing blocks in the feature processing block set, and determining a target sample image feature of the sample image unit sequence based on each of the fourth sample image features; the fourth sample image feature is determined based on each of the third sample image features;

[0031] A sample noise determination unit, configured to determine sample image noise data corresponding to the sample time step based on the sample target image feature;

[0032] A sample denoising unit, configured to perform denoising processing on the sample potential noise image based on the sample image noise data, and generate a sample generated image corresponding to the sample time step;

[0033] A time step updating unit, configured to update the sample time step if the number of the sample time step does not reach the total number of the sample time steps, and use the sample generated image as a sample potential noise image, and proceed to execute the step of converting the sample potential noise image into a sample image unit sequence by using a unitization module in the image generation model, until the number of the sample time step reaches the total number of the sample time steps, thereby obtaining a sample generated image;

[0034] A parameter updating unit is used to determine a first predicted loss value of the sample image and the sample generated image based on a first preset loss function, perform supervised training on the image generation model based on the first predicted loss value, and iteratively update the model parameters of the image generation model until the first predicted loss value converges, thereby obtaining a trained image generation model.

[0035] In a fifth aspect, an embodiment of the present specification provides an electronic device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps of the above method when executed by the processor.

[0036] In a sixth aspect, an embodiment of the present specification provides a storage medium having a computer program stored thereon, and the computer program implements the steps of the above method when executed by a processor.

[0037] In a seventh aspect, an embodiment of the present specification provides a computer program product, comprising: a computer program, when the computer program is executed by a processor of an electronic device, the processor can at least implement the methods described in the first aspect to the second aspect.

[0038] In an embodiment of the present specification, by obtaining an image unit sequence corresponding to a potential noise image, a target image feature of the image unit sequence is obtained through a feature processing block set, wherein the image unit sequence includes multiple image units, and if the time step type of the time step is a cache pruning step, each processing block in the feature processing block set of the step can process and obtain the first image feature of the image unit sequence, and there is at least one first processing block, and the first processing block confirms its first image feature based on the second image feature and the third image feature, wherein the second image feature is a historical cache feature of the image unit in the first processing block in the target feature processing block of the previous time step, and the third image feature is an image feature of the image unit generated by the previous processing block of the first processing block, and the first processing block re-infers part of the image units in the image unit sequence based on the third image feature, and obtains the first image feature of the first processing block in combination with the second image feature; if the time step is a complete inference step, each processing block in the feature processing block set is used to generate a fourth image feature of the image unit sequence, and the target image feature of the image unit sequence is determined based on the fourth image feature generated by each processing block, wherein the fourth image feature is inferred by the processing block based on the third image feature of the image unit generated by the previous processing block. By dividing the time step into a complete inference step and a cache pruning step, in the cache pruning step, a part of the image units are pruned and replaced with cached image features, which reduces the number of inference times for the image units and solves the problem of redundant calculation caused by the large number of image units and the need for multiple inferences in the image generation model inference process, thereby improving the inference speed of the image generation model. At the same time, the complete inference step is combined to reduce the error introduced by the cache, balancing the acceleration effect and generation quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0040] Figure 1 is a scene schematic diagram of an image generation method provided in an embodiment of this specification;

[0041] Figure 2 A flowchart of an image generation method is provided for the embodiments of this specification;

[0042] Figure 3 is an example schematic diagram of an image generation method provided in an embodiment of this specification;

[0043] Figure 4It is a flowchart of an image generation method provided in an embodiment of this specification;

[0044] Figure 5 It is a flowchart of an image generation method provided in an embodiment of this specification;

[0045] Figure 6 It is a flowchart of an image generation method provided in an embodiment of this specification;

[0046] Figure 7 is an example schematic diagram of an image generation method provided in an embodiment of this specification;

[0047] Figure 8 It is a flowchart of an image generation model training method provided in an embodiment of this specification;

[0048] Fig. 9 It is a schematic diagram of a model structure of an image generation model training method provided in an embodiment of this specification;

[0049] Fig.10 It is a flowchart of an image generation model training method provided in an embodiment of this specification;

[0050] Fig.11 It is a flowchart of an image generation method provided in an embodiment of this specification;

[0051] Fig.12 is a structural schematic diagram of an image generating device provided in an embodiment of this specification;

[0052] Fig.13 is a structural schematic diagram of an image generating device provided in an embodiment of this specification;

[0053] Fig.14 It is a structural schematic diagram of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0054] The following will be combined with the drawings in the embodiments of this specification to clearly and completely describe the technical solutions in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.

[0055] In related technologies, diffusion models are a generative model inspired by non-equilibrium thermodynamics. The core idea is to gradually add noise to the data by simulating the diffusion process, and then learn to reverse the process to construct the required data samples from the noise. Diffusion Transformer is a diffusion model combined with the Transformer architecture for image and video generation tasks. It can efficiently capture dependencies in the data and generate high-quality results.

[0056] Image generation models such as Diffusion Transformer divide noisy images into multiple image units. Each image unit extracts a feature when passing through a feature processing block to predict noise. When passing through N feature processing blocks in a time step, the number of extracted features will be multiplied by N, resulting in an increase in the amount of calculation. In addition, the attention mechanism of Diffusion Transformer has quadratic complexity, which means that as the number of image units increases, the computational cost increases exponentially. In addition, the multi-step reasoning process further exacerbates this problem, resulting in the model's reasoning speed being much lower than the actual application requirements when generating high-resolution images and long videos.

[0057] Based on the above problems, an image generation method of an embodiment of the present specification is proposed, which reduces redundant calculations by caching and pruning image units in intermediate steps, considers the importance of different image units in the feature processing block, and caches them at the image unit level, that is, not all image units that require reasoning of the feature processing block are pruned, but some of them are pruned, and by alternating between executing cache pruning steps and complete reasoning steps, the acceleration effect and generation quality are balanced.

[0058] Please also see Figure 1, a scene schematic diagram of an image generation method is provided in an embodiment of the present specification. The image generation device provided in the embodiment of the present specification can be a terminal device such as a mobile phone, a computer, a tablet computer, a smart watch or a vehicle-mounted device, or can be a module in the terminal device for implementing the image generation method. The image generation device can obtain an image unit sequence corresponding to a potential noise image when receiving an image generation instruction from a user. If the time step type of the time step is a cache pruning step, the first image feature corresponding to the image unit sequence in each processing block in the feature processing block set is obtained, and a target image feature of the image unit sequence is determined based on each first image feature. If the time step type of the time step is a complete inference step, the fourth image feature corresponding to the image unit sequence in each processing block in the feature processing block set is obtained, and a target image feature of the image unit sequence is determined based on each fourth image feature. Image noise data corresponding to the time step is determined based on the target image feature, and denoising is performed on the potential noise image based on the image noise data to generate a target image corresponding to the time step.

[0059] Optionally, the image generation device can perform image generation model training. The image generation device can use the unitization module in the image generation model to convert the sample potential noise image into a sample image unit sequence; the sample potential noise image is generated based on the sample image, and the total number of sample time steps and the time step type of each sample time step are determined. If the time step type of the sample time step is a cache pruning step, the first sample image feature corresponding to the sample image unit sequence in each processing block in the feature processing block set is obtained, and the sample target image feature of the image unit sequence is determined based on each first sample image feature; each processing block includes at least one first processing block that determines the first sample image feature corresponding to the sample image unit sequence in the first processing block based on the second sample image feature and the third sample image feature, the second image feature is the historical cache feature of the image unit in the target feature processing block in the previous time step, and the third image feature is the image feature of the image unit generated by the previous processing block of the target feature processing block. If the time step type of the sample time step is a complete inference step, the fourth sample image feature corresponding to the sample image unit sequence in each processing block in the feature processing block set is obtained. The image feature determines the target sample image feature of the sample image unit sequence based on each fourth sample image feature; the fourth sample image feature is determined based on each third sample image feature, and the sample image noise data corresponding to the sample time step is determined based on the sample target image feature, and the sample potential noise image is denoised based on the sample image noise data to generate a sample generated image corresponding to the sample time step. If the number of the sample time step does not reach the total number of sample time steps, the sample time step is updated, and the sample generated image is used as the sample potential noise image, and the step of converting the sample potential noise image into a sample image unit sequence using the unitization module in the image generation model is executed until the number of the sample time step reaches the total number of sample time steps to obtain the sample generated image, and the first predicted loss value of the sample image and the sample generated image is determined based on the first preset loss function, and the image generation model is supervised and trained based on the first predicted loss value and the model parameters of the image generation model are iteratively updated until the first predicted loss value converges to obtain the trained image generation model.

[0060] It should be noted that the image generation device used for image generation and the image generation device used for model training can be the same device or different devices. Preferably, the image generation method and the model training method are implemented on different image generation devices.

[0061] The image generation method provided in this specification is described in detail below in conjunction with specific embodiments.

[0062] See also Figure 2 , is a flow chart of an image generation method provided in the embodiment of this specification. Figure 2As shown, the method in the embodiment of this specification may include the following steps S102-S110.

[0063] S102, obtaining an image unit sequence corresponding to a potential noise image; the image unit sequence includes a plurality of image units;

[0064] The image generation method in the embodiment of this specification is mainly a method for generating images using conditional information, and can also be used to generate videos, etc. It can be understood that when generating a video, it is necessary to generate images frame by frame first, and then splice these images into a video in chronological order. Therefore, the same method can be used to generate videos. Specifically, after the user inputs the conditional information of the required target image to the image generation device, the image generation device can generate a random noise image (usually Gaussian noise). The pixel values ​​of this image are completely random. Then, the random noise image at the pixel level is compressed to a low-dimensional potential feature through an encoder to obtain a potential noise image (noise-added latent variable), which can reduce the computational complexity of the model. Then, the image unit sequence of the potential noise image is obtained. Exemplarily, based on the unitization module (Patchify), the image is split to obtain an image unit sequence, and the image is regarded as a series of patches / units (patch) to capture the global dependencies between them.

[0065] See also Figure 3 , which is an example schematic diagram of an image generation method provided in the embodiments of this specification. The potential noise image is divided into blocks (image units) of fixed size, each block is linearly embedded, position coding is added, and then the obtained vector sequence is input into the feature processing block. By dividing the potential noise image, the potential noise image can be converted from a two-dimensional feature to a one-dimensional sequence (image unit sequence), thereby obtaining a series of image units (tokens).

[0066] S104, if the time step type of the time step is a cache pruning step, obtaining first image features corresponding to the image unit sequence in each processing block in the feature processing block set, and determining target image features of the image unit sequence based on each of the first image features;

[0067] It can be understood that in order to reduce the noise of the potential noise image, it is necessary to reduce the noise step by step through multiple time steps. At each time step, it is necessary to process at least one feature processing block to extract image features, so as to predict the noise data in the image. In one embodiment of the present specification, the time step is divided into two types: cache pruning step and complete reasoning step. When executing each time step, different reasoning actions are performed according to the different time step types. When the time step type is a cache pruning step, each processing block includes at least one first processing block that determines the first image feature corresponding to the image unit sequence in the first processing block based on the second image feature and the third image feature. The second image feature is the historical cache feature of the image unit in the target feature processing block of the previous time step, and the third image feature is the image feature of the image unit generated by the previous processing block of the target feature processing block. That is, in at least one processing block, the first image feature corresponding to the image unit sequence is composed of two parts, one part is the historical cache image feature, and the other part is the image feature inferred by the processing block.

[0068] It is understandable that if all the image unit sequences in a processing block use the historical cache features inferred by the processing block at the previous time step, it may lead to insufficient update of some important image units, thus affecting the generation quality. Therefore, in order to ensure the reasoning effect, considering the different importance of each image unit in different processing blocks at different time steps, it can be determined that the importance of a part of the image units input to the feature processing block is lower than that of the remaining image units, and the image units with relatively lower importance use the historical image features of the processing block cached at the previous time step as the second image features of the processing block at the current time step.

[0069] Exemplarily, a processing block, specifically a DiT block, includes at least an MSA (multi-head self-attention module) and an MLP (multi-layer perceptron module). For details, please refer to the model architecture of DiT in the related art. In each feature processing block, the processed tokens (each token corresponds to a patch) representation, that is, the image unit features of the updated image unit sequence, will be passed to the next block. Each block operates and updates the representation of tokens independently without sharing internal states. Finally, when all blocks have completed processing, the output of the last block will contain the final representation of all tokens, that is, the target image features.

[0070] Optionally, additional conditional information (conditional information of different modalities, etc.) can be embedded when the image unit sequence is input to the processing block. The conditional information here can include time step information and category labels (text information). The conditional information (time step embedding) and other conditions (such as category or text embedding) are represented as two independent embedding vectors, which are appended to the beginning of the input image unit sequence to form an expanded input sequence.

[0071] S106, if the time step type of the time step is a complete inference step, obtaining a fourth image feature corresponding to the image unit sequence in each of the processing blocks in the feature processing block set, and determining a target image feature of the image unit sequence based on each of the fourth image features;

[0072] In one embodiment of the present specification, when the time step type of the time step is a complete inference step, a feature processing block is used to infer each image unit in the image unit sequence, and a fourth image feature corresponding to the image unit sequence in each processing block is obtained, and a target image feature of the image unit sequence is determined based on the fourth image feature. The fourth image feature generated by the processing block is determined based on each third image feature, that is, inferred based on the image feature of the image unit in the previous processing block of the processing block.

[0073] S108, determining image noise data corresponding to the time step based on the target image feature;

[0074] In one embodiment of the present specification, image noise data is determined based on the target image features obtained by processing. Specifically, after the last feature processing block, the target image features corresponding to the generated image unit sequence need to be decoded into the following two outputs: noise "Noise prediction" and the corresponding covariance matrix, as image noise data, the noise prediction and covariance matrix help the DiT model to more accurately capture the relationship between noise and image structure during the denoising process.

[0075] S110, performing denoising processing on the potential noise image based on the image noise data to generate a target image corresponding to the time step.

[0076] In one embodiment of the present specification, based on the predicted image noise data, the potential noise image is denoised to obtain a target image after noise reduction. It is understandable that the model needs to gradually reduce the noise through multiple time step iterations and adjust the pixel values ​​of the potential noise image until a clear image is finally restored. In each round of denoising, the model will make adjustments based on the image noise data.

[0077] In the embodiment of the present specification, by obtaining the image unit sequence corresponding to the potential noise image, if the time step type of the time step is a cache pruning step, the first image feature corresponding to the image unit sequence in each processing block in the feature processing block set is obtained, and the target image feature of the image unit sequence is determined based on each of the first image features. If the time step type of the time step is a complete reasoning step, the fourth image feature corresponding to the image unit sequence in each of the processing blocks in the feature processing block set is obtained, and the target image feature of the image unit sequence is determined based on each of the fourth image features. The image noise data corresponding to the time step is determined based on the target image feature, and the potential noise image is denoised based on the image noise data to generate the target image corresponding to the time step. In the cache pruning step, the inference is accelerated by skipping unimportant calculations by using the cache feature, and the importance of different image units in the feature processing block is considered, and the cache is performed at the image unit level, that is, all image units in the image unit sequence are not pruned, and the acceleration effect and the generation quality are balanced by alternately executing the cache pruning step and the complete reasoning step.

[0078] See also Figure 4 , is a flow chart of an image generation method provided in the embodiment of this specification. Figure 4 As shown, the method in the embodiment of this specification may include the following steps S202-S204.

[0079] S202, obtaining a target time step of the number of buffer intervals per interval after the start time step, and determining the time step types of the start time step and the target time step as a complete reasoning step;

[0080] S204: Determine time steps other than the complete time steps as cache pruning steps.

[0081] In one embodiment of the present specification, in order to reasonably allocate the alternating execution of cache pruning steps and complete reasoning steps, the image generation quality is avoided to be affected by pruning. At least the time step type of the starting time step is determined as a complete reasoning step, so that a complete reasoning can be performed at the beginning of reasoning to ensure the generation quality. Among them, the starting time step includes at least the first time step. In a feasible implementation, the starting time step can also be a preset number of time steps starting from the first time step. For example, steps 1-5 can be determined as the starting time step. Afterwards, according to the number of cache intervals, the target time step after the starting time step is determined, and the time step type of the target time step is determined as a complete reasoning step. Exemplarily, assuming that the number of cache intervals is 4, the total number of time steps is 20, and the starting time step is 1-5, then the remaining complete reasoning steps are 10, 15 and 20. The time step type of the time step between the complete reasoning steps is a cache pruning step.

[0082] In another feasible implementation, the time step type of the end time step may also be determined as a complete time step, and the end time step includes at least the last time step. For example, assuming that the total number of time steps is 1000, steps 1-5 and steps 995-1000 may be set as complete reasoning steps.

[0083] As another example, assuming that the number of cache intervals is 3, the total number of time steps is 20, the starting time step is 1-5, and the ending time step is 15-20, the remaining complete time steps are 9 and 13. Although the number of cache intervals between step 13 and the ending time step 15 does not reach 3, the remaining complete time steps from the starting time step to the ending time step can be determined according to the number of cache intervals.

[0084] The number of cache intervals represents the number of cache pruning steps performed before a full inference step. A larger number of cache intervals means more cache feature reuse, which can speed up inference; while a smaller number of cache intervals means more frequent full inference to reduce the error caused by caching. Therefore, the number of cache intervals can be selected according to actual needs.

[0085] Optionally, the cache interval number includes a first cache interval number and a second cache interval number, and the method further includes:

[0086] S2022, if the number of steps of the complete reasoning step is less than or equal to the first preset number of steps, determining the number of cache intervals to be the first number of cache intervals;

[0087] S2024: If the number of complete reasoning steps is greater than the first preset number of steps, determine the cache interval number to be the second cache interval number.

[0088] In one embodiment of the present specification, since there is a higher correlation between early time steps in the reasoning process, and the correlation between later time steps is relatively small, different cache interval numbers can be used to control the frequency of executing complete reasoning steps. Among them, the first preset number of steps is selected as the milestone time step for dividing the early reasoning and the late reasoning. It can be understood that starting from a certain milestone time step, the model gradually enters the later denoising step, and the correlation between the image units is low. The number of cache intervals before the first preset number of steps is determined as the first cache interval number, and the number of cache intervals after the first preset number of steps is determined as the second cache interval number. Starting from Gaussian noise, the high correlation between early time steps (i.e., the image unit changes less) is used to set a larger cache interval, the first cache interval number K1, to reduce redundant calculations of early steps; starting from the first preset number of steps, the correlation between image units is low, and the cache interval is adjusted to a smaller value, the second cache interval number K2, so as to perform complete reasoning steps more frequently and reduce the risk of error accumulation.

[0089] Specifically, the hyperparameter search on the validation set can be used to find the appropriate number of cache intervals K1, K2 and the first preset number of steps. That is, the images in the validation set are taken as input, and different numbers of cache intervals and first preset numbers of steps are designed. The above reasoning method is used to generate the final denoised image, and the similarity between the final denoised image and the image in the validation set is evaluated, so as to select the first number of cache intervals, the second number of cache intervals and the first preset number of steps that meet the generation requirements.

[0090] In the embodiment of the specification, by obtaining the target time step of each cache interval number after the start time step, the time step type of the start time step and the target time step is determined as a complete reasoning step, and the time step other than the complete time step is determined as a cache pruning step. By setting the number of cache intervals between complete reasoning steps, it is ensured that a complete reasoning step is executed after each cache pruning step of the cache interval number is executed, avoiding the cache pruning operation at each time step, maximizing the use of the cache, and accelerating multi-step reasoning.

[0091] See also Figure 5 , is a flow chart of an image generation method provided in the embodiment of this specification. Figure 5 As shown, the method in the embodiment of this specification may include the following steps S302-S316.

[0092] S302, segmenting the potential noise image according to a preset segmentation size to obtain a plurality of image units;

[0093] In one embodiment of the present specification, the acquired potential noise image can be first segmented into multiple image units, and the size of each image unit is a preset segmentation size. By performing image segmentation processing, the computational complexity can be reduced while improving image details and quality. Among them, the preset segmentation size determines the size and number of the image units, thereby affecting the overall computational amount of the image generation model. Exemplarily, the preset segmentation size can be 2*2.

[0094] S304, determining position identification information of the image unit in the potential noise image;

[0095] In one embodiment of the present specification, after the image unit is obtained, since the image unit will be converted into a one-dimensional sequence and spliced ​​into the processing block, in order to be able to learn and restore the position of each image unit in the image, Positional Embeddings must be added for position marking. For example, the classic non-learning sin&cosine position encoding technology can be used to obtain the position identification information of each image unit.

[0096] S306, generating an image unit sequence corresponding to the potential noise image based on the multiple image units and the position identification information;

[0097] In one embodiment of the present specification, an image unit sequence corresponding to the potential noise image is generated according to each image unit and its corresponding position identification information.

[0098] S308, if the processing block is the first processing block, generating a first image feature corresponding to the image unit sequence in the first processing block based on the second image feature and the third image feature;

[0099] In one embodiment of the present specification, each image unit in the image unit sequence in each time step needs to be processed by multiple processing blocks, the structures of different processing blocks are consistent, but the parameters of each processing block are different, and the influence of different processing blocks on the model generation output changes dynamically. Therefore, applying a pruning cache strategy in each processing block is not always the most appropriate choice. After determining that the time step type is a cache pruning step, a first processing block that needs cache pruning and a second processing block that does not need cache pruning can also be determined from the feature processing block. Among them, the second processing block is a feature processing block other than the first processing block in the feature processing block.

[0100] If the processing block is determined to be the first processing block, the first image feature of the image unit sequence in the first processing block is determined according to the second image feature and the third image feature. The first processing block obtains the second image unit feature of a part of the image units and reuses this part, and the first processing block obtains the third image feature of another part of the image units, that is, the image feature inferred in the previous processing block of the first processing block, and updates the image feature of this part of the image units based on the third image feature.

[0101] Optionally, in one embodiment, if the processing block is a first processing block, generating a first image feature corresponding to the image unit sequence in the first processing block based on the second image feature and the third image feature includes:

[0102] S3082, if the processing block is the first processing block, determining the inference image unit and the cache image unit in the image unit sequence;

[0103] In one embodiment of the present specification, after determining the first processing block, the cache image units that need to be pruned and the reasoning image units that need to be reasoned in the first processing block are determined. Exemplarily, the importance of each image unit can be evaluated, so that the cache image units and reasoning image units therein are determined according to the importance. It is understandable that there can be multiple first processing blocks, and the cache image units and reasoning image units therein need to be determined separately for each first processing block.

[0104] S3084, obtaining a second image feature of the inference image unit;

[0105] In one embodiment of the present specification, the second image feature of the inference image unit is obtained, that is, the image feature inferred by the inference image unit in the previous processing block is obtained as the feature input of the inference image unit of the current first processing block.

[0106] S3086, using the first processing block to generate an inference image feature of the inference image unit based on the second image feature;

[0107] In one embodiment of the present specification, a first processing block is used to process a second image feature of an inference image unit in an image feature sequence to obtain an inference image feature.

[0108] S3088, obtaining a third image feature of the cache image unit;

[0109] In one embodiment of the present specification, a historical cache feature of a cache image unit in the first processing block inferred by the first processing block in the last inference time step is obtained, and the historical cache feature is used as the third image feature of the cache image unit in the first processing block.

[0110] S30810: Determine the inferred image feature and the third image feature of the cached image unit as the first image feature corresponding to the image unit sequence in the first processing block.

[0111] In one embodiment of the present specification, the first image feature corresponding to the image unit sequence in the first processing block is constituted by the inferred image feature and the third image feature.

[0112] S310, if the processing block is a second processing block, generating a first image feature corresponding to the image unit sequence in the second processing block based on the third image feature;

[0113] In one embodiment of the present specification, if the processing block is the second processing block, a first image feature corresponding to the image unit sequence in the second processing block is generated based on the third image feature. That is, in the cache pruning step, for the second processing block, the first image feature is re-inferred based on the image feature generated by each image unit in the image unit sequence in the previous processing block.

[0114] Optionally, in one embodiment, determining a target image feature of the image unit sequence based on each of the first image features includes:

[0115] If the processing block is the last processing block in the feature processing block set, the first image feature of the processing block is used to determine the target image feature of the image unit sequence.

[0116] In the processing block set, each processing block updates the image features of the image unit sequence and gives the updated image features to the next processing block until the last processing block, which determines the updated result as the target image feature. If the processing block is not the last processing block, the image features of the image unit sequence are continuously updated according to the processing block type.

[0117] S312, determining image noise data corresponding to the time step based on the target image feature;

[0118] S314, performing denoising processing on the potential noise image based on the image noise data to generate a target image corresponding to the time step;

[0119] For details, please refer to steps S108-S1010 in the above-mentioned embodiment of the specification, which will not be elaborated here.

[0120] S316: If the number of the time step does not reach the second preset number of steps, the time step is updated, and the target image is used as a potential noise image, and the step of obtaining an image unit sequence corresponding to the potential noise image is executed until the number of the time step reaches the second preset number of steps.

[0121] In one embodiment of the present specification, multiple steps of denoising are required to obtain the final target image. Therefore, after obtaining the target image corresponding to a time step, a determination is made based on the number of time steps. If the number of time steps has not reached the second preset number of steps, the time step is updated, that is, the time step is updated to the next time step, and then the target image obtained by processing this time step is used as a potential noise image to perform the steps of obtaining the corresponding image unit sequence and processing through the feature processing block set until the number of time steps reaches the second preset number of steps. The second preset number of steps is the total number of steps, which can be set according to actual needs.

[0122] In an embodiment of the present specification, the acquired potential noise image is segmented according to a preset segmentation size to obtain multiple image units, position identification information of the image units in the potential noise image is determined, and an image unit sequence corresponding to the potential noise image is generated based on the multiple image units and the position identification information. If the processing block is the first processing block, a first image feature corresponding to the image unit sequence in the first processing block is generated based on the second image feature and the third image feature. If the processing block is the second processing block, a first image feature corresponding to the image unit sequence in the second processing block is generated based on the third image feature. Image noise data corresponding to a time step is determined based on the target image feature. The potential noise image is denoised based on the image noise data to generate a target image corresponding to the time step. If the number of the time step does not reach the second preset number of steps, the time step is updated, and the target image is used as the potential noise image, and the step of obtaining the image unit sequence corresponding to the potential noise image is executed until the number of the time step reaches the second preset number of steps. When the time step type is a cache pruning step, the feature processing block is divided into two types: the first processing block and the second processing block. The cache pruning strategy is implemented for the first processing block, that is, the first processing block includes a cache image unit and an inference image unit, the cache image unit uses the second image feature, and the inference image unit is re-inferred based on the corresponding third image feature. By further determining the types of different processing blocks and selectively implementing the cache strategy on the processing blocks, unnecessary calculations can be adaptively reduced in different time steps and processing blocks, significantly accelerating the inference speed.

[0123] See also Figure 6 , is a flow chart of an image generation method provided in the embodiment of this specification. Figure 6 As shown, the method in the embodiment of this specification may include the following steps S402-S410:

[0124] S402, determining a first importance score of each image unit in the image unit sequence in each of the processing blocks;

[0125] In one embodiment of the present specification, when the time step type is a cache pruning step, the first importance score of each image unit in the image unit sequence in each processing block of the step is determined. It is understandable that each image unit may have different importance in different processing blocks at different time steps, and the importance refers to the contribution of updating the image unit in the processing block of the time step to the final denoising effect. The evaluation of the importance score of each image unit can be achieved by pre-training an importance prediction model.

[0126] S404, determining a second importance score of each of the processing blocks based on the first importance score;

[0127] In one embodiment of the present specification, after determining the first importance score of each image unit in the processing block, the second importance score of the processing block may be further determined according to each first importance score.

[0128] Optionally, in one embodiment, the method further includes the following steps S4042-S4044:

[0129] S4042, determining the sum of first importance scores of the image unit sequences in each of the processing blocks;

[0130] In one embodiment of the present specification, when determining the second importance score according to the first importance score, the sum of the first importance scores of each image unit in the processing block may be determined first.

[0131] S4044: Determine a second importance score of each of the processing blocks based on the sum of the first importance scores.

[0132] In one embodiment of the present specification, the second importance score of each processing block is determined according to the sum of the first importance scores.

[0133] Exemplarily, the importance score of the processing block is calculated by averaging the first importance scores of all image units in each processing block by means of a global average preset pruning rate (Global Averaging) preset pruning rate:

[0134]

[0135] in, is the processing block f l The importance score of is the importance score of the i-th image unit in the l-th processing block, and n is the number of image units in the processing block.

[0136] S406, determining a first processing block and a second processing block from the feature processing block set based on the second importance score;

[0137] In one embodiment of the present specification, the contribution of the entire processing block to the current time step update can be evaluated through the second importance score, thereby determining whether pruning is required for the processing block. It is understandable that if the importance of all image units in a processing block is high, then there is no need to perform a cache operation for the processing block, and all reasoning needs to be re-performed. Therefore, based on the level of the second importance score, the first processing block and the second processing block can be determined from the feature processing block set. Exemplarily, a score threshold can be predetermined, and a processing block below the score threshold is determined as the first processing block, and a processing block greater than or equal to the score threshold is determined as the second processing block.

[0138] S408, if the processing block is a first processing block, based on a first importance score of each image unit in the image unit sequence in the first processing block and a preset pruning rate, determining an inference image unit and a cache image unit corresponding to the first processing block in the image unit sequence;

[0139] In one embodiment of the present specification, the inference image units and cache image units in the first processing block can be determined based on the first importance score of each image unit in the image unit sequence in the first processing block and the preset pruning rate. By setting the preset pruning rate, the image units with lower first importance scores can be screened out as cache image units. It can be understood that when the first importance score is low, it indicates that the update of the image unit does not bring a relatively large contribution to the update of the first processing block, so the historical cache features of the image unit in the first processing block in the previous time step can be used without repeated reasoning. Specifically, the image units in each block are sorted according to their importance scores. Generally, the image units with lower importance scores are less important and can be considered for pruning.

[0140] Assume that the preset cropping rate is the preset pruning rate r, that is, the image units in each block are to be cropped by the preset pruning rate r. You can calculate the number of image units that need to be cropped in each block:

[0141] Cropping amount = total number of image units in the block × r Cropping amount = total number of image units in the block × r.

[0142] For example, if there are 1000 image units in a block, and the crop rate is set to 0.3 (i.e., 30% cropping), then the number of image units that need to be cropped is 300. According to the sorted importance scores, the least important image units with the preset crop rate r% are selected for cropping. If there are 1000 image units in a block, and 300 image units are cropped, you can select the 300 image units with the lowest importance scores and mark them as cropping targets, i.e., cached image units, and the remaining image units are inference image units.

[0143] S410, if the processing block is a first processing block, determining the inference image unit and the cache image unit corresponding to each of the first processing blocks based on the first importance score of each image unit in the image unit sequence in the first processing block and the total number of cache units;

[0144] In one embodiment of the present specification, a total number of cache units can be set for each first processing block, that is, how many cache image units are needed in this processing block. The total number of cache units can be the same number set for all first processing blocks, or different total numbers of cache units can be set for different time steps or different first processing blocks. The image units are sorted from low to high according to the first importance score, and the image units that meet the total number of cache units are selected. Assuming that the total number of cache units is 300 and the image unit sequence is 1000, the 300 image units with the lowest importance scores are selected as cache image units.

[0145] Optionally, in one embodiment, the method may further include the following steps S4102-S4108:

[0146] S4102: Determine the total number of cache units of each of the first processing blocks based on the second importance score of each of the first processing blocks.

[0147] In one embodiment of the present specification, the total number of cache units can be determined based on the second importance score of the first processing block. For each block i, the blocks need to be sorted based on the importance score. A processing block with a high importance score indicates that the processing block is critical to model reasoning. Usually, we choose to delete tokens with a low importance score. Assume that a total of N tokens need to be pruned. total Crop image units. The sequence of image units in each processing block is the same. Assume that each block contains the same number of image units, denoted as N block The "total importance score" of each block indicates the importance of the block to the overall model. i Denotes the total importance score of block i. Compute the cropping ratio for each processed block. The cropping ratio of each block should be inversely proportional to its importance score - blocks with high importance should be cropped less, while blocks with low importance should be cropped more.

[0148] Calculate the total importance score of all blocks:

[0149] Where n is the number of blocks, S i is the total importance score of the ith block.

[0150] Calculate the crop ratio r for each block i , indicating the unit ratio that block i should be cropped to:

[0151]

[0152] Here, S total -S i is the “importance loss” of block i, that is, the proportion of block i that needs to be cropped.

[0153] Next, based on the cropping ratio r of each block i , we can calculate the number of cells that need to be cropped for each block. Since the total number of image cells in each block is N block is the same, the number of cells cut that is:

[0154]

[0155] Optionally, in one embodiment, based on the first importance score of each of the image units in the first processing block and the number of cache units, determining the inference image unit and the cache image unit corresponding to each of the first processing blocks includes the following steps S4104-S4108:

[0156] S4104, dividing the image units in the image unit sequence according to the grid size to obtain a plurality of grids;

[0157] In one embodiment of the present specification, when determining the cached image units, in order to avoid over-concentrating on trimming a certain area of ​​the image, the image units can be first divided into grids, and then trimmed evenly in each grid. The grid size of each grid is fixed, so each grid obtained by the division includes the same number of image units. For example, a 4*4 image can be divided into four 2*2 grids, each of which has 4 image units.

[0158] S4106, determining the number of unit caches of each of the grids based on the total number of cache units;

[0159] Specifically, the number of unit caches in each grid may be evenly divided according to the number of grids to determine how many image units need to be pruned in each grid, that is, how many image units are determined as cached image units.

[0160] S4108, based on the first importance score of each of the image units in the first processing block, respectively determine the target image units in each of the grids, and determine the target image units corresponding to each of the grids as the cache image units;

[0161] In one embodiment of the present specification, a target image unit is determined in each grid according to the first importance score of each image unit in the first processing block and the unit cache quantity, and the cache image unit of the first processing block is determined according to the target image unit of each grid. The target image unit corresponding to any one of the multiple grids is selected from the image unit sequence of the target grid based on the unit cache quantity in the target grid, and the importance score of the target image unit is not higher than the importance scores of other image units in the image unit sequence except the target image unit. Please refer to Figure 7, is a schematic diagram of an example of an image generation method provided in the embodiments of this specification. Figure 7 As shown, each square represents an image unit, and the number in the square represents the importance score of the image unit. The 4*4 image unit sequence is divided into 4 2*2 grids. Assuming that the number of unit caches in each grid is 2, two target image units with lower scores can be determined from each grid in the image unit according to the size of the first importance score, which are represented by dark squares in the figure. It can be understood that in some grids, there may be image units with the same importance scores. For example, if the first importance scores of the image units in a grid are 1, 3, 3, and 3 respectively, then one of the three image units with a score of 3 can be randomly determined as the target image unit. It should be noted that the image unit sequence is actually a one-dimensional sequence. The position of each image unit can be determined based on the position identification information. Therefore, the grid can be divided according to the position identification information. Figure 7 In order to facilitate the description of the role of grid processing, the cache image unit selection effect is displayed in the form of a two-dimensional sequence.

[0162] In the embodiment of the present specification, by determining the first importance score of each image unit in the image unit sequence in each processing block, determining the second importance score of each processing block based on the first importance score, and determining the first processing block and the second processing block from the feature processing block set based on the second importance score, if the processing block is the first processing block, based on the first importance score of each image unit in the image unit sequence in the first processing block and the preset pruning rate, the inference image unit and the cache image unit corresponding to the first processing block are determined in the image unit sequence, or based on the first importance score of each image unit in the image unit sequence in the first processing block and the total number of cache units, the inference image unit and the cache image unit corresponding to each first processing block are determined. By determining the first importance score of each image unit in the image unit sequence in different processing blocks at the time step, evaluating the second importance score of the processing block according to the first importance score, and then determining the type of each processing block in the processing block set, fine control of important image units and processing blocks is achieved, global consistency and accuracy in the generation process are guaranteed, and generation quality is maintained while accelerating inference. Furthermore, by dividing the image units in the image unit sequence according to the grid size, each grid screens the target image units of the unit cache quantity according to the first importance score to obtain the cached image units. The grid-based image unit pruning scheme can prevent pruning from being overly concentrated in certain areas, and further maintain the global consistency and quality of the generation.

[0163] See also Figure 8 , is a flow chart of a method for training an image generation model according to an embodiment of this specification. Figure 8As shown, the method of the embodiment of this specification may include the following steps S502-S516.

[0164] S502, using a unitization module in an image generation model to convert a sample potential noise image into a sample image unit sequence; the sample potential noise image is generated based on the sample image;

[0165] In one embodiment of the present specification, the image generation model is a DiT model, including a unitization module, a position encoding module, an embedding module, a feature processing module (including a feature processing block set), a normalization module, and an arrangement and recombination module. The unitization module is used to divide the potential noise image into image units; the position encoding module is used to obtain the position identification information of the divided image units; the embedding module is used to obtain the representation of the input conditional information, such as time step information and other image description information, wherein the time step information (for example, time step t) is used to control the degree of denoising; the feature processing module is used to determine the image features according to the embedding of the image units and the conditional information. Please refer to Fig. 9 , a model structure diagram of an image generation model training method is provided for the embodiment of this specification. The image generation model divides the (sample) potential noise image into sample image units through a unitization module, and further obtains the position identification information through a position embedding module to obtain a sample image unit sequence. The time step information and condition information are input into the embedding module together and are represented as two independent embedding vectors. These vectors are attached to the beginning of the input sample image unit sequence to form an extended input sequence. Then, the target image features are obtained by processing through a feature processing block set (including N processing blocks), and converted into predicted noise data and corresponding covariance matrix through a normalization module and a permutation and recombination module as sample image noise data.

[0166] Optionally, before step S502, the image generation model may be initialized, a sample image is obtained, and a sample potential noise image of the sample image is generated by using an encoder in the image generation model. The image generation model may also include an encoder for generating a sample potential noise image. It is understandable that when training the image generation model, noise may be first added to the sample image by an encoder and mapped to a latent space to obtain a sample potential noise image, which is then converted into a sample image unit by a unitization module. The position identification information of the sample image unit is obtained by a position encoding module to obtain a sample image unit sequence.

[0167] S504, determining the total number of sample time steps and the time step type of each sample time step;

[0168] In one embodiment of the present specification, since model denoising requires denoising one by one through multiple time steps, the total number of sample time steps and the time step type of each sample time step can be preset. The time step type includes a complete reasoning step and a cache pruning step. Two-Phase Round-Robin (TPRR) is a time step scheduling strategy specifically proposed for accelerating image generation models provided in an embodiment of the present specification, which aims to balance reasoning acceleration and generation quality loss by alternating the execution of complete reasoning steps (I-steps) and cache pruning steps (P-steps). The core idea of ​​this strategy is to maximize the use of cache to accelerate multi-step reasoning by scheduling reasoning steps in stages to avoid executing cache pruning operations at each time step. In the complete reasoning step, no cache or pruning operations are used, but the model reasoning is performed completely. These steps are used to reduce the errors introduced by the cache to ensure that the model can still obtain high-quality generation results at critical time steps; after each execution of K cache pruning steps (P-steps), the model will execute a complete reasoning step. I-steps calculate the features of all tokens without caching or pruning any tokens; the cached inference step speeds up inference by skipping unimportant calculations using cached features. In each P-step, the model uses the features of the image units cached previously for inference, reducing the number of image units that need to be calculated. K is a parameter that controls the frequency of cache usage, representing the number of P-steps executed before executing the complete inference step (I-steps). The type of each sample time step can be determined based on K.

[0169] S506, if the time step type of the sample time step is a cache pruning step, obtaining first sample image features corresponding to the sample image unit sequence in each processing block in the feature processing block set, and determining sample target image features of the image unit sequence based on each of the first sample image features;

[0170] Specifically, similar to the use process of the above-mentioned image generation method, in the training process of the image generation model, if the time step type of the sample time step is determined to be a cache pruning step, then in at least one sample processing block, the first image feature corresponding to the sample image unit sequence is composed of two parts, one part of the image unit directly uses the sample image feature of the historical cache, and the other part is the sample image feature inferred by the processing block based on the sample image feature of the image unit generated by the previous processing block.

[0171] Among them, each processing block includes at least one first processing block that determines the first sample image feature corresponding to the sample image unit sequence in the first processing block based on the second sample image feature and the third sample image feature, the second image feature is the historical cache feature of the image unit in the target feature processing block in the previous time step, and the third image feature is the image feature of the image unit generated by the previous processing block of the target feature processing block.

[0172] S508, if the time step type of the sample time step is a complete inference step, obtaining a fourth sample image feature corresponding to the sample image unit sequence in each of the processing blocks in the feature processing block set, and determining a target sample image feature of the sample image unit sequence based on each of the fourth sample image features;

[0173] In one embodiment of the present specification, when the time step type of the sample time step is a complete inference step, the sample processing block is used to infer each sample image unit in the sample image unit sequence, and the fourth sample image feature corresponding to the sample image unit sequence in each sample processing block is obtained, and the target sample image feature of the sample image unit sequence is determined based on the fourth sample image feature. The fourth sample image feature is determined based on each third sample image feature.

[0174] S510, determining sample image noise data corresponding to the sample time step based on the sample target image feature;

[0175] In one embodiment of the present specification, the sample target image features are processed by a normalization module and an arrangement and recombination module to obtain predicted sample image noise data.

[0176] S512, performing denoising processing on the sample potential noise image based on the sample image noise data to generate a sample generated image corresponding to the sample time step;

[0177] Specifically, at each sample time step, a denoised sample image is generated based on the predicted sample image noise data prediction.

[0178] S514, if the number of the sample time step does not reach the total number of sample time steps, then update the sample time step, and use the sample generated image as a sample potential noise image, and proceed to execute the step of converting the sample potential noise image into a sample image unit sequence using a unitization module in the image generation model, until the number of the sample time step reaches the total number of sample time steps, thereby obtaining a sample generated image;

[0179] Specifically, each time step (t) corresponds to a denoising stage, and finally a clear image is obtained through the back diffusion process. If the number of the current sample time step does not reach the total number of sample time steps, the sample time step is updated, and the current sample generated image is used as the sample potential noise image to execute the step of converting the sample potential noise image into a sample image unit sequence using the unitization module in the image generation model, until the number of the sample time step reaches the total number of sample time steps, the denoising process is completed, and the sample generated image is obtained.

[0180] S516, determining a first predicted loss value of the sample image and the sample generated image based on a first preset loss function, performing supervised training on the image generation model based on the first predicted loss value and iteratively updating model parameters of the image generation model until the first predicted loss value converges, thereby obtaining a trained image generation model.

[0181] Specifically, after a round of training is completed (that is, after the total number of sample time steps is reached), a sample image and sample condition information can be input, and a denoised image is obtained through the image generation model, and the denoised image is compared with the sample image to determine the model generation effect. For example, the mean square error (MSE) loss is usually used to measure the difference between the model output and the real image (sample image) to optimize the model parameters.

[0182] Optionally, in one embodiment, a LoRa model can be added for fine-tuning, that is, when executing the cache pruning step in the processing block, the input of this step is completely inferred synchronously through the LoRa model, and the image features of the complete inference generated by the LoRa model are compared with the image features obtained by performing cache pruning, so that the image features after cache pruning can still be close to the inference results of the complete inference step, improving the repair prediction ability of the image generation model, that is, when not reasoning about a part of the image units in the image, a better output result can be obtained, further improving the prediction accuracy of the image generation model. And after improving the prediction accuracy of the model, it can be considered to increase the number of pruning units, thereby improving the model reasoning speed.

[0183] In an embodiment of the present specification, a sample potential noise image is converted into a sample image unit sequence by using a unitization module in an image generation model, the total number of sample time steps and the time step type of each sample time step are determined, if the time step type of the sample time step is a cache pruning step, the first sample image feature corresponding to the sample image unit sequence in each processing block in the feature processing block set is obtained, and the sample target image feature of the image unit sequence is determined based on each first sample image feature, if the time step type of the sample time step is a complete inference step, the fourth sample image feature corresponding to the sample image unit sequence in each processing block in the feature processing block set is obtained, the target sample image feature of the sample image unit sequence is determined based on each fourth sample image feature, and the sample image feature corresponding to the sample time step is determined based on the sample target image feature. The method uses a sample image noise data, performs denoising on the sample potential noise image based on the sample image noise data, generates a sample generated image corresponding to the sample time step, updates the sample time step if the number of the sample time step does not reach the total number of the sample time step, and uses the sample generated image as the sample potential noise image, and then executes the step of converting the sample potential noise image into a sample image unit sequence using the unitization module in the image generation model until the number of the sample time step reaches the total number of the sample time step, obtains the sample generated image, determines the first predicted loss value of the sample image and the sample generated image based on the first preset loss function, performs supervised training on the image generation model based on the first predicted loss value, and iteratively updates the model parameters of the image generation model until the first predicted loss value converges, and obtains the trained image generation model. By adopting the image unit pruning strategy when training the image generation model, the number of inferences for some image units is reduced, the speed of the image generation model is improved, and the quality of the generated image can be guaranteed by continuously optimizing the loss function.

[0184] See also Fig.10 , is a flow chart of a method for training an image generation model according to an embodiment of this specification. Fig.10 As shown, the method in the embodiment of this specification may include the following steps S602-S608.

[0185] S602, obtaining a sample potential noise image and a sample image unit sequence corresponding to the sample potential noise image;

[0186] In one embodiment of the present specification, the importance scores of image units in different sample processing blocks at different time steps can be predicted by a cache predictor, and the cache predictor can be trained before training the image generation model. In each reasoning time step, the cache predictor accepts the input image unit sequence in each processing block and calculates the importance score (weight) of each image unit, which is used to interpolate the trimmed and untrimmed states of the image unit to ensure that differentiability is maintained during the training process.

[0187] The structure of the cache predictor may include multiple processing blocks (ie, a set of feature processing blocks), and the output of the last processing block is mapped through a linear layer and a sigmoid function to obtain an importance score.

[0188] S604, obtaining a first sample image feature of a target sample image unit generated by a target processing block in a first time step and a second sample image feature of the target sample image unit generated by the target processing block in a second time step;

[0189] In one embodiment of the present specification, the first time step is any time step in the sample time step; the second time step is the next time step of the first time step; the target processing block is any processing block in the feature processing block set; the target sample image unit is any sample image unit in the sample image unit sequence. For each target sample image unit, the sample image features obtained by complete inference in adjacent time steps are obtained, and the importance score of the target image unit can be determined by comparing the differences in sample images in adjacent time steps. The target image unit with a high importance score has a large feature gap between adjacent time steps.

[0190] To learn g by gradient descent θ , control the image generation model to perform complete reasoning at time step t+1 and cache the current Blockf l Output At time step t, the model performs forward propagation normally and also calculates the normal output when the image unit is not pruned.

[0191] S606, generating a sample prediction importance score of the target sample image unit in the first time step based on a cache predictor;

[0192] Specifically, the sample prediction importance score is obtained by inputting the target sample image unit in the first time step into the cache predictor, wherein the sample prediction importance score may be in the interval of (0, 1).

[0193] S608, determining a superimposed sample image feature of the target sample image unit based on the sample prediction importance score, the first sample image feature, and the second sample image feature;

[0194] Specifically, the superimposed sample image feature of the target sample image unit is determined according to the sample prediction importance score ω, the output first sample image feature and the output second sample image feature.

[0195]

[0196] The output of the superposition is expressed as:

[0197]

[0198] Final Output It can be considered as an intermediate state between pruning and unpruning.

[0199] S610, determining a second predicted loss value of the superimposed sample image feature and the first sample image feature based on a second preset loss function, training the cache predictor based on the predicted loss value and iteratively updating the model parameters of the cache predictor until the second predicted loss value converges, thereby obtaining a trained cache predictor.

[0200] In one embodiment of the present specification, when training the cache predictor, two steps t+1 and t are executed, the t+1 step and the t step are added together, and a weight is assigned to calculate the superposition output so that the result of the superposition output is consistent with the result of only performing the t step reasoning. At this time, the assigned weight is accurate. Through this interpolation mechanism, the output of the cache predictor can maintain differentiability. It can be understood that if the reasoning result of the t step is similar to the reasoning result of the t+1 step, it means that the importance weight of the current image unit is low. Based on the above method and the second preset loss function, model training is performed until the cache predictor converges, that is, the second prediction loss value converges. Exemplarily, the second preset loss function can be the mean square error (MSE), and the cache predictor is trained by minimizing the mean square error (MSE) of the generated output so that it can accurately predict which image units should be pruned while keeping the generated quality unaffected.

[0201] Optionally, after the cache predictor is obtained through training, if the time step type of the sample time step is a cache pruning step, obtaining the first sample image feature corresponding to the sample image unit sequence in each processing block in the feature processing block set includes the following steps S602-S610:

[0202] S602, if the time step type of the sample time step is a cache pruning step, using the cache predictor to determine a first sample importance score of each sample image unit in the sample image unit sequence in each of the processing blocks;

[0203] Specifically, each image unit may have different importance in different processing blocks at different time steps. If the time step type of the sample time step is a cache pruning step, a cache predictor is used to determine the first sample importance score of each sample image unit in the sample image unit sequence in each processing block.

[0204] S604, determining a second sample importance score of each of the sample processing blocks based on the first sample importance score;

[0205] Specifically, after determining the first sample importance score of each sample image unit in the sample processing block, the second sample importance score of each processing block is determined according to the first sample importance score.

[0206] S606, determining a first sample processing block and a second sample processing block from the processing blocks based on the second sample importance score;

[0207] Specifically, according to the second sample importance score of the processing block, the first sample processing block and the second sample processing block are selected, the first sample processing block is the processing block that needs to execute the cache pruning strategy, and the second sample processing block is the processing block that does not need to execute the cache pruning strategy. Among them, the processing block with a higher second sample importance score can be determined as the second sample processing block to ensure the accuracy of reasoning.

[0208] S608, if the sample processing block is a first sample processing block, generating a first sample image feature corresponding to the sample image unit sequence in the first sample processing block based on the second sample image feature and the third sample image feature;

[0209] S610: If the processing block is a second processing block, generating a first sample image feature corresponding to the image unit sequence in the second sample processing block based on the third sample image feature.

[0210] For details, please refer to steps S506-S508 of the above-mentioned embodiment of the specification, which will not be elaborated here.

[0211] It is understandable that after the cache predictor is trained, the cache predictor can be a part of the image generation model and added to the image generation model for selecting a processing block and an image unit (token).

[0212] See also Fig.11, a flowchart of an image generation method is provided for an embodiment of the present specification. Taking a time step as an example, when a time step starts, the time step type of the current time step is determined according to the time step selection strategy, including two types: Full (complete reasoning step) and Reuse (cache pruning step). The time step selection strategy can be specifically implemented by pre-configuring the determination logic of the time step type, and the model makes its own judgment according to the configuration logic. Specifically, by setting the preset start and end time steps as complete reasoning steps, the middle time steps are determined according to the preset cache interval, and a complete reasoning step is executed every time the preset cache interval number of cache pruning steps are executed. The model can record each time step type and infer the time step type of the next time step. When the time step type is a cache pruning step, the first importance score of each token in each block is determined according to the cache predictor, and the second importance score of the block is calculated. The block selection strategy is to screen the blocks (first processing blocks) that need to perform cache pruning according to the second importance score. The token pruning & reuse module determines the cache image units and reasoning image units in each first processing block, and obtains the cache features of the cache image units processed by each first processing block in the previous time step, and inputs them into the feature processing block set. A pruning strategy is adopted for the cache image units in the first processing block, that is, the cache features are not used for reasoning. The reasoning image units in the first processing block are re-reasoned according to the data of the previous processing block. For the second processing block, the image unit sequence is fully reasoned until the last block L outputs the target image features of the image unit sequence as the final output result of the time step, and the result of the time step is cached in the cache feature.

[0213] In an embodiment of the present specification, a sample prediction importance score of the target sample image unit in the first time step is generated according to a cache predictor, and a superimposed sample image feature of the target sample image unit is determined based on the sample prediction importance score, the first sample image feature and the second sample image feature output by the same processing block in adjacent time steps. The superimposed sample image feature is compared with the first sample image feature to determine the prediction accuracy of the importance score. By learning the importance of each image unit in different time steps, it is possible to help the model optimize the calculation process, reduce unnecessary calculations, and achieve better image generation effects.

[0214] The following will be combined with the attached Figure 12-13 , the image generation device provided in the embodiment of this specification is introduced in detail. It should be noted that the attached Figure 12-13 The image generating device in the embodiment of the present invention is used to execute the image generating device in the embodiment of the present invention. Figure 1-Figure 11 For the convenience of explanation, only the part related to the embodiment of this specification is shown. For the specific technical details not disclosed, please refer to this specification. Figure 1-Figure 11 The embodiment shown.

[0215] See also Fig.12 , which shows a schematic diagram of the structure of an image generation device provided by an exemplary embodiment of the present specification. The image generation device can be implemented as all or part of the device through software, hardware or a combination of both. The device 1 includes an acquisition unit 11, a first reasoning unit 12, a second reasoning unit 13, a noise determination unit 14 and a denoising processing unit 15.

[0216] An acquisition unit 11 is used to acquire an image unit sequence corresponding to a potential noise image; the image unit sequence includes a plurality of image units;

[0217] A first reasoning unit 12 is used for obtaining, if the time step type of the time step is a cache pruning step, the first image feature corresponding to the image unit sequence in each processing block in the feature processing block set, and determining the target image feature of the image unit sequence based on each of the first image features; each of the processing blocks includes at least one first processing block that determines the first image feature corresponding to the image unit sequence in the first processing block based on the second image feature and the third image feature, wherein the second image feature is a historical cache feature of the image unit in the target feature processing block at the previous time step, and the third image feature is an image feature of the image unit generated by a previous processing block of the target feature processing block;

[0218] The second inference unit 13 is used for obtaining the fourth image feature corresponding to the image unit sequence in each of the processing blocks in the feature processing block set if the time step type of the time step is a complete inference step, and determining the target image feature of the image unit sequence based on each of the fourth image features; the fourth image feature is determined based on each of the third image features;

[0219] A noise determination unit 14, configured to determine image noise data corresponding to the time step based on the target image feature;

[0220] The denoising processing unit 15 is used to perform denoising processing on the potential noise image based on the image noise data to generate a target image corresponding to the time step.

[0221] Optionally, the acquisition unit 11 is further used to acquire a target time step of the number of buffer intervals per interval after the start time step, and determine the time step types of the start time step and the target time step as a complete reasoning step;

[0222] Time steps other than the complete time steps are determined as cache pruning steps.

[0223] Optionally, the cache interval number includes a first cache interval number and a second cache interval number; the acquisition unit 11 is further configured to determine that the cache interval number is the first cache interval number if the number of steps of the complete reasoning step is less than or equal to a first preset number of steps;

[0224] If the number of the complete reasoning steps is greater than the first preset number of steps, the number of buffer intervals is determined to be a second number of buffer intervals.

[0225] Optionally, the first inference unit 12 is specifically configured to generate, if the processing block is a first processing block, a first image feature corresponding to the image unit sequence in the first processing block based on the second image feature and the third image feature;

[0226] If the processing block is a second processing block, a first image feature corresponding to the image unit sequence in the second processing block is generated based on the third image feature.

[0227] Optionally, the first inference unit 12 is specifically configured to determine the target image feature of the image unit sequence using the first image feature of the processing block if the processing block is the last processing block in the feature processing block set.

[0228] Optionally, the first inference unit 12 is specifically configured to determine the inference image unit and the cache image unit in the image unit sequence if the processing block is the first processing block;

[0229] Acquire a second image feature of the inference image unit;

[0230] generating, using the first processing block, an inference image feature of the inference image unit based on the second image feature;

[0231] Acquire a third image feature of the cache image unit;

[0232] The inferred image feature and the third image feature of the cached image unit are determined as the first image feature corresponding to the image unit sequence in the first processing block.

[0233] Optionally, the first reasoning unit 12 is specifically used to determine a first importance score of each image unit in the image unit sequence in each of the processing blocks;

[0234] Determining a second importance score for each of the processing blocks based on the first importance score;

[0235] The first processing block and the second processing block are determined from the set of feature processing blocks based on the second importance score.

[0236] Optionally, the first reasoning unit 12 is specifically used to determine the sum of first importance scores of the image unit sequences in each of the processing blocks;

[0237] A second importance score of each of the processing blocks is determined based on the sum of the first importance scores.

[0238] Optionally, if the processing block is a first processing block, the first inference unit 12 is specifically used to determine the inference image unit and the cache image unit corresponding to the first processing block in the image unit sequence based on a first importance score and a preset pruning rate of each image unit in the image unit sequence in the first processing block.

[0239] Optionally, if the processing block is a first processing block, the first inference unit 12 is specifically used to determine the inference image units and cache image units corresponding to each of the first processing blocks based on the first importance score of each image unit in the image unit sequence in the first processing block and the total number of cache units.

[0240] Optionally, the first reasoning unit 12 is specifically configured to determine the total number of cache units of each of the first processing blocks based on the second importance score of each of the first processing blocks.

[0241] Optionally, the first inference unit 12 is specifically used to divide the image units in the image unit sequence according to the grid size to obtain a plurality of grids; wherein each grid includes the same number of image units;

[0242] Determine the number of unit caches of each of the grids based on the total number of cache units;

[0243] Based on the first importance score of each of the image units in the first processing block, respectively determine the target image units in each of the grids, and determine the target image units corresponding to each of the grids as the cache image units;

[0244] The target image unit corresponding to any target grid among the multiple grids is selected from the image unit sequence of the target grid based on the unit cache quantity in the target grid, and the importance score of the target image unit is not higher than the importance scores of other image units in the image unit sequence except the target image unit.

[0245] Optionally, the denoising processing unit 15 is specifically used to segment the potential noise image according to a preset segmentation size to obtain a plurality of image units;

[0246] Determining position identification information of the image unit in the potential noise image;

[0247] An image unit sequence corresponding to the potential noise image is generated based on the multiple image units and the position identification information.

[0248] If the number of the time step does not reach the second preset number of steps, the time step is updated, and the target image is used as a potential noise image to execute the step of obtaining an image unit sequence corresponding to the potential noise image until the number of the time step reaches the second preset number of steps.

[0249] Further, see Attachment Fig.13 The image generating device shown includes a unitization unit 21, a parameter determination unit 22, a first training unit 23, a second training unit 24, a sample noise determination unit 25, a sample denoising unit 26, a time step updating unit 27 and a parameter updating unit 28.

[0250] A unitization unit 21, configured to convert a sample potential noise image into a sample image unit sequence by using a unitization module in an image generation model; the sample potential noise image is generated based on the sample image;

[0251] A parameter determination unit 22, used to determine the total number of sample time steps and the time step type of each sample time step;

[0252] The first training unit 23 is used for obtaining the first sample image feature corresponding to the sample image unit sequence in each processing block in the feature processing block set if the time step type of the sample time step is a cache pruning step, and determining the sample target image feature of the image unit sequence based on each of the first sample image features; each of the processing blocks includes at least one first processing block that determines the first sample image feature corresponding to the sample image unit sequence in the first processing block based on the second sample image feature and the third sample image feature, wherein the second image feature is a historical cache feature of the image unit in the target feature processing block at the previous time step, and the third image feature is an image feature of the image unit generated by a previous processing block of the target feature processing block;

[0253] The second training unit 24 is used for obtaining the fourth sample image feature corresponding to the sample image unit sequence in each of the processing blocks in the feature processing block set if the time step type of the sample time step is a complete inference step, and determining the target sample image feature of the sample image unit sequence based on each of the fourth sample image features; the fourth sample image feature is determined based on each of the third sample image features;

[0254] A sample noise determination unit 25, configured to determine sample image noise data corresponding to the sample time step based on the sample target image feature;

[0255] A sample denoising unit 26, configured to perform denoising processing on the sample potential noise image based on the sample image noise data, and generate a sample generated image corresponding to the sample time step;

[0256] A time step updating unit 27 is used to update the sample time step if the number of the sample time step does not reach the total number of the sample time steps, and use the sample generated image as a sample potential noise image, and proceed to execute the step of converting the sample potential noise image into a sample image unit sequence using a unitization module in the image generation model, until the number of the sample time step reaches the total number of the sample time steps, thereby obtaining a sample generated image;

[0257] A parameter updating unit 28 is used to determine a first predicted loss value of the sample image and the sample generated image based on a first preset loss function, perform supervised training on the image generation model based on the first predicted loss value, and iteratively update the model parameters of the image generation model until the first predicted loss value converges, thereby obtaining a trained image generation model.

[0258] Optionally, the device 2 further includes a cache prediction unit 29, which is specifically used to obtain a first sample image feature of a target sample image unit generated by a target processing block in a first time step and a second sample image feature of the target sample image unit generated by the target processing block in a second time step; the first time step is any time step in the sample time steps; the second time step is a time step next to the first time step; the target processing block is any processing block in a feature processing block set; the target sample image unit is any sample image unit in the sample image unit sequence;

[0259] generating a sample prediction importance score for the target sample image unit in the first time step based on a cache predictor;

[0260] Determining a superimposed sample image feature of the target sample image unit based on the sample prediction importance score, the first sample image feature, and the second sample image feature;

[0261] Based on a second preset loss function, a second predicted loss value of the superimposed sample image feature and the first sample image feature is determined, and based on the predicted loss value, the cache predictor is trained and the model parameters of the cache predictor are iteratively updated until the second predicted loss value converges, thereby obtaining a trained cache predictor.

[0262] Optionally, the first training unit 23 is specifically configured to determine, if the time step type of the sample time step is a cache pruning step, a first sample importance score of each sample image unit in the sample image unit sequence in each of the processing blocks using the cache predictor;

[0263] Determining a second sample importance score for each of the processing blocks based on the first sample importance score;

[0264] determining a first sample processing block and a second sample processing block from the processing blocks based on the second sample importance score;

[0265] If the sample processing block is a first sample processing block, generating a first sample image feature corresponding to the sample image unit sequence in the first sample processing block based on the second sample image feature and the third sample image feature;

[0266] If the processing block is a second processing block, a first sample image feature corresponding to the image unit sequence in the second sample processing block is generated based on the third sample image feature.

[0267] It should be noted that the image generation device provided in the above embodiment only uses the division of the above functional modules as an example when executing the image generation method and the image generation model training method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the image generation device and the image generation method and the model training method embodiments provided in the above embodiment belong to the same concept. The implementation process is detailed in the method embodiment and will not be repeated here.

[0268] The serial numbers of the embodiments of the present specification are for description only and do not represent the advantages and disadvantages of the embodiments. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0269] The present specification also provides a storage medium having a computer program stored thereon. When the computer program is executed by a processor, the above-mentioned Figure 1-Figure 7 The image generation method of the embodiment shown or Figure 8-Figure 11 The image generation model training method of the embodiment shown in the figure can be found in the specific execution process. Figure 1-Figure 11 The specific description of the illustrated embodiment will not be repeated here.

[0270] Please refer to Fig.14, which shows a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of this specification. The electronic device in this specification may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, the memory 120, the input device 130, and the output device 140 may be connected via the bus 150.

[0271] The processor 110 may include one or more processing cores. The processor 110 uses various interfaces and lines to connect various parts of the entire electronic device, and executes various functions and processes data of the terminal 100 by running or executing instructions, programs, code sets or instruction sets stored in the memory 120, and calling data stored in the memory 120. Optionally, the processor 110 can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 110 can integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user pages, and applications; the GPU is responsible for rendering and drawing display content; and the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 110, but may be implemented separately through a communication chip.

[0272] The memory 120 may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory 120 includes a non-transitory computer-readable medium (Non-Transitory Computer-Readable Storage Medium). The memory 120 may be used to store instructions, programs, codes, code sets or instruction sets. The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc. The operating system may be an Android system, including a system deeply developed based on the Android system, an IOS system developed by Apple, including a system deeply developed based on the IOS system or other systems.

[0273] The memory 120 can be divided into an operating system space and a user space. The operating system runs in the operating system space, and native and third-party applications run in the user space. In order to ensure that different third-party applications can achieve good operating results, the operating system allocates corresponding system resources to different third-party applications. However, the requirements for system resources in different application scenarios in the same third-party application are also different. For example, in the local resource loading scenario, the third-party application has higher requirements for disk reading speed; in the animation rendering scenario, the third-party application has higher requirements for GPU performance. The operating system and third-party applications are independent of each other, and the operating system often cannot perceive the current application scenario of the third-party application in a timely manner, resulting in the operating system being unable to perform targeted system resource adaptation according to the specific application scenario of the third-party application.

[0274] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to open up data communication between third-party applications and the operating system so that the operating system can obtain the current scenario information of third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.

[0275] The input device 130 is used to receive input commands or data, and includes but is not limited to a keyboard, a mouse, a camera, a microphone, or a touch device. The output device 140 is used to output commands or data, and includes but is not limited to a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 are touch screen displays.

[0276] The touch display screen can be designed as a full screen, a curved screen or a special-shaped screen. The touch display screen can also be designed as a combination of a full screen and a curved screen, or a combination of a special-shaped screen and a curved screen, which is not limited in the embodiments of this specification.

[0277] In addition, those skilled in the art can understand that the structure of the electronic device shown in the above drawings does not constitute a limitation on the electronic device, and the electronic device may include more or fewer components than shown, or combine certain components, or arrange the components differently. For example, the electronic device also includes components such as a radio frequency circuit, an input unit, a sensor, an audio circuit, a WiFi module, a power supply, and a Bluetooth module, which will not be described in detail here.

[0278] exist Fig.14 In the electronic device shown, the processor 110 may be used to call a computer application stored in the memory 120 and specifically perform the following operations:

[0279] Acquire an image unit sequence corresponding to a potential noise image; the image unit sequence includes a plurality of image units;

[0280] If the time step type of the time step is a cache pruning step, then the first image feature corresponding to the image unit sequence in each processing block in the feature processing block set is obtained, and the target image feature of the image unit sequence is determined based on each of the first image features; each of the processing blocks includes at least one first processing block that determines the first image feature corresponding to the image unit sequence in the first processing block based on a second image feature and a third image feature, wherein the second image feature is a historical cache feature of the image unit in the target feature processing block at the previous time step, and the third image feature is an image feature of the image unit generated by a previous processing block of the target feature processing block;

[0281] If the time step type of the time step is a complete inference step, obtaining a fourth image feature corresponding to the image unit sequence in each of the processing blocks in the feature processing block set, and determining a target image feature of the image unit sequence based on each of the fourth image features; the fourth image feature is determined based on each of the third image features;

[0282] Determining image noise data corresponding to the time step based on the target image feature;

[0283] The potential noise image is denoised based on the image noise data to generate a target image corresponding to the time step.

[0284] In one embodiment, after acquiring the image unit sequence corresponding to the potential noise image, the processor 110 further performs the following operations:

[0285] Obtain a target time step of the number of buffer intervals per interval after the start time step, and determine the time step types of the start time step and the target time step as a complete reasoning step;

[0286] Time steps other than the complete time steps are determined as cache pruning steps.

[0287] In one embodiment, the cache interval number includes a first cache interval number and a second cache interval number;

[0288] The processor 110 further performs the following operations before obtaining the target time step of the number of buffer intervals per interval after the start time step and determining the time step types of the start time step and the target time step as complete reasoning steps:

[0289] If the number of steps of the complete reasoning step is less than or equal to the first preset number of steps, determining the number of cache intervals to be the first number of cache intervals;

[0290] If the number of the complete reasoning steps is greater than the first preset number of steps, the number of buffer intervals is determined to be a second number of buffer intervals.

[0291] In one embodiment, when the processor 110 executes the acquisition of the first image feature corresponding to the image unit sequence in each processing block in the feature processing block set, the processor 110 specifically performs the following operations:

[0292] If the processing block is a first processing block, generating a first image feature corresponding to the image unit sequence in the first processing block based on the second image feature and the third image feature;

[0293] If the processing block is a second processing block, a first image feature corresponding to the image unit sequence in the second processing block is generated based on the third image feature.

[0294] In an embodiment, when determining the target image feature of the image unit sequence based on each of the first image features, the processor 110 specifically performs the following operations:

[0295] If the processing block is the last processing block in the feature processing block set, the first image feature of the processing block is used to determine the target image feature of the image unit sequence.

[0296] In one embodiment, when the processor 110 generates the first image feature corresponding to the image unit sequence in the first processing block based on the second image feature and the third image feature if the processing block is the first processing block, the processor 110 specifically performs the following operations:

[0297] If the processing block is the first processing block, determining the inference image unit and the cache image unit in the image unit sequence;

[0298] Acquire a second image feature of the inference image unit;

[0299] generating, using the first processing block, an inference image feature of the inference image unit based on the second image feature;

[0300] Acquire a third image feature of the cache image unit;

[0301] The inferred image feature and the third image feature of the cached image unit are determined as the first image feature corresponding to the image unit sequence in the first processing block.

[0302] In one embodiment, the processor 110 may also be used to call a computer application stored in the memory 120 and specifically perform the following operations:

[0303] Determining a first importance score of each image unit in the image unit sequence in each of the processing blocks;

[0304] Determining a second importance score for each of the processing blocks based on the first importance score;

[0305] The first processing block and the second processing block are determined from the set of feature processing blocks based on the second importance score.

[0306] In one embodiment, when the processor 110 determines the second importance score of each processing block based on the first importance score, the processor 110 specifically performs:

[0307] Determining a sum of first importance scores of the image unit sequences in each of the processing blocks;

[0308] A second importance score of each of the processing blocks is determined based on the sum of the first importance scores.

[0309] In one embodiment, when the processor 110 determines the inference image unit and the cache image unit in the image unit sequence if the processing block is the first processing block, the processor 110 specifically performs the following operations:

[0310] If the processing block is the first processing block, based on a first importance score and a preset pruning rate of each image unit in the image unit sequence in the first processing block, an inference image unit and a cache image unit corresponding to the first processing block are determined in the image unit sequence.

[0311] In one embodiment, when the processor 110 determines the inference image unit and the cache image unit in the image unit sequence if the processing block is the first processing block, the processor 110 specifically performs the following operations:

[0312] If the processing block is a first processing block, based on a first importance score of each image unit in the image unit sequence in the first processing block and the total number of cache units, the inference image units and cache image units corresponding to each of the first processing blocks are determined.

[0313] In one embodiment, the processor 110 may also be used to call a computer application stored in the memory 120 and specifically perform the following operations:

[0314] The total number of cache units of each of the first processing blocks is determined based on the second importance score of each of the first processing blocks.

[0315] In one embodiment, when the processor 110 determines the inference image unit and the cache image unit corresponding to each of the first processing blocks based on the first importance score of each of the image units in the first processing block and the number of cache units, the processor 110 specifically performs the following operations:

[0316] Dividing the image units in the image unit sequence according to the grid size to obtain a plurality of grids; wherein each grid includes the same number of image units;

[0317] Determine the number of unit caches of each of the grids based on the total number of cache units;

[0318] Based on the first importance score of each of the image units in the first processing block, respectively determine the target image units in each of the grids, and determine the target image units corresponding to each of the grids as the cache image units;

[0319] The target image unit corresponding to any target grid among the multiple grids is selected from the image unit sequence of the target grid based on the unit cache quantity in the target grid, and the importance score of the target image unit is not higher than the importance scores of other image units in the image unit sequence except the target image unit.

[0320] In one embodiment, when executing to obtain the image unit sequence corresponding to the potential noise image, the processor 110 specifically performs the following operations:

[0321] Segmenting the potential noise image according to a preset segmentation size to obtain a plurality of image units;

[0322] Determining position identification information of the image unit in the potential noise image;

[0323] An image unit sequence corresponding to the potential noise image is generated based on the multiple image units and the position identification information.

[0324] In one embodiment, after performing denoising processing on the potential noise image based on the image noise data to generate the target image corresponding to the time step, the processor 110 further performs the following operations:

[0325] If the number of the time step does not reach the second preset number of steps, the time step is updated, and the target image is used as a potential noise image to execute the step of obtaining an image unit sequence corresponding to the potential noise image until the number of the time step reaches the second preset number of steps.

[0326] In one embodiment, the processor 110 may also be used to call a computer application stored in the memory 120 and specifically perform the following operations:

[0327] The sample potential noise image is converted into a sample image unit sequence by using a unitization module in an image generation model; the sample potential noise image is generated based on the sample image;

[0328] determining a total number of sample time steps and a time step type for each of the sample time steps;

[0329] If the time step type of the sample time step is a cache pruning step, then the first sample image feature corresponding to the sample image unit sequence in each processing block in the feature processing block set is obtained, and the sample target image feature of the image unit sequence is determined based on each of the first sample image features; each of the processing blocks includes at least one first processing block that determines the first sample image feature corresponding to the sample image unit sequence in the first processing block based on the second sample image feature and the third sample image feature, the second image feature is a historical cache feature of the image unit in the target feature processing block at the previous time step, and the third image feature is an image feature of the image unit generated by a previous processing block of the target feature processing block;

[0330] If the time step type of the sample time step is a complete inference step, obtaining a fourth sample image feature corresponding to the sample image unit sequence in each of the processing blocks in the feature processing block set, and determining a target sample image feature of the sample image unit sequence based on each of the fourth sample image features; the fourth sample image feature is determined based on each of the third sample image features;

[0331] Determining sample image noise data corresponding to the sample time step based on the sample target image feature;

[0332] Performing denoising on the sample potential noise image based on the sample image noise data to generate a sample generated image corresponding to the sample time step;

[0333] If the number of the sample time step does not reach the total number of sample time steps, the sample time step is updated, and the sample generated image is used as a sample potential noise image, and the step of converting the sample potential noise image into a sample image unit sequence by using a unitization module in the image generation model is executed until the number of the sample time step reaches the total number of sample time steps, thereby obtaining a sample generated image;

[0334] Based on a first preset loss function, a first predicted loss value of the sample image and the sample generated image is determined, and based on the first predicted loss value, supervised training is performed on the image generation model and model parameters of the image generation model are iteratively updated until the first predicted loss value converges, thereby obtaining a trained image generation model.

[0335] In one embodiment, the processor 110 may also be used to call a computer application stored in the memory 120 and specifically perform the following operations:

[0336] Acquire a first sample image feature of a target sample image unit generated by a target processing block in a first time step and a second sample image feature of the target sample image unit generated by the target processing block in a second time step; the first time step is any time step in the sample time steps; the second time step is the next time step of the first time step; the target processing block is any processing block in a feature processing block set; the target sample image unit is any sample image unit in the sample image unit sequence;

[0337] generating a sample prediction importance score for the target sample image unit in the first time step based on a cache predictor;

[0338] Determining a superimposed sample image feature of the target sample image unit based on the sample prediction importance score, the first sample image feature, and the second sample image feature;

[0339] Based on a second preset loss function, a second predicted loss value of the superimposed sample image feature and the first sample image feature is determined, and based on the predicted loss value, the cache predictor is trained and the model parameters of the cache predictor are iteratively updated until the second predicted loss value converges, thereby obtaining a trained cache predictor.

[0340] In one embodiment, when the processor 110 executes the process of obtaining the first sample image feature corresponding to the sample image unit sequence in each processing block in the feature processing block set if the time step type of the sample time step is a cache pruning step, the processor 110 specifically performs the following operations:

[0341] If the time step type of the sample time step is a cache pruning step, determining a first sample importance score of each sample image unit in the sample image unit sequence in each of the processing blocks using the cache predictor;

[0342] Determining a second sample importance score for each of the processing blocks based on the first sample importance score;

[0343] determining a first sample processing block and a second sample processing block from the processing blocks based on the second sample importance score;

[0344] If the sample processing block is a first sample processing block, generating a first sample image feature corresponding to the sample image unit sequence in the first sample processing block based on the second sample image feature and the third sample image feature;

[0345] If the processing block is a second processing block, a first sample image feature corresponding to the image unit sequence in the second sample processing block is generated based on the third sample image feature.

[0346] In an embodiment of the present specification, by obtaining an image unit sequence corresponding to a potential noise image, if the time step type of a time step is a cache pruning step, then obtaining the first image feature corresponding to the image unit sequence in each processing block in a feature processing block set, determining the target image feature of the image unit sequence based on each of the first image features, and if the time step type of a time step is a complete reasoning step, then obtaining the fourth image feature corresponding to the image unit sequence in each of the processing blocks in the feature processing block set, determining the target image feature of the image unit sequence based on each of the fourth image features, determining the image noise data corresponding to the time step based on the target image feature, denoising the potential noise image based on the image noise data, and generating the target image corresponding to the time step. In the cache pruning step, the reasoning is accelerated by skipping unimportant calculations by using cache features, and considering the importance of different image units in the feature processing block, caching is performed at the image unit level, that is, not pruning all image units in the image unit sequence, and balancing the acceleration effect and generation quality by alternately executing cache pruning steps and complete reasoning steps.

[0347] Furthermore, by obtaining the target time step of each cache interval after the start time step, the time step type of the start time step and the target time step is determined as a complete reasoning step, and the time step other than the complete time step is determined as a cache pruning step. By setting the number of cache intervals between complete reasoning steps, it is ensured that a complete reasoning step is executed after each cache pruning step of the cache interval number is executed, avoiding the cache pruning operation at each time step, maximizing the use of the cache, and accelerating multi-step reasoning.

[0348] Furthermore, the acquired potential noise image is segmented according to a preset segmentation size to obtain a plurality of image units, position identification information of the image units in the potential noise image is determined, and an image unit sequence corresponding to the potential noise image is generated based on the plurality of image units and the position identification information; if the processing block is a first processing block, a first image feature corresponding to the image unit sequence in the first processing block is generated based on the second image feature and the third image feature; if the processing block is a second processing block, a first image feature corresponding to the image unit sequence in the second processing block is generated based on the third image feature; image noise data corresponding to the time step is determined based on the target image feature; denoising is performed on the potential noise image based on the image noise data to generate a target image corresponding to the time step; if the number of the time step does not reach the second preset number of steps, the time step is updated, and the target image is used as the potential noise image, and the step of obtaining the image unit sequence corresponding to the potential noise image is executed until the number of the time step reaches the second preset number of steps. When the time step type is a cache pruning step, the feature processing block is divided into two types: the first processing block and the second processing block. The cache pruning strategy is implemented for the first processing block, that is, the first processing block includes a cache image unit and an inference image unit, the cache image unit uses the second image feature, and the inference image unit is re-inferred based on the corresponding third image feature. By further determining the types of different processing blocks and selectively implementing the cache strategy on the processing blocks, unnecessary calculations can be adaptively reduced in different time steps and processing blocks, significantly accelerating the inference speed.

[0349] Furthermore, by determining the first importance score of each image unit in the image unit sequence in each processing block, determining the second importance score of each processing block based on the first importance score, and determining the first processing block and the second processing block from the feature processing block set based on the second importance score, if the processing block is the first processing block, based on the first importance score of each image unit in the image unit sequence in the first processing block and the preset pruning rate, the inference image unit and the cache image unit corresponding to the first processing block are determined in the image unit sequence, or, based on the first importance score of each image unit in the image unit sequence in the first processing block and the total number of cache units, the inference image unit and the cache image unit corresponding to each first processing block are determined. By determining the first importance score of each image unit in the image unit sequence in different processing blocks at the time step, evaluating the second importance score of the processing block according to the first importance score, and then determining the type of each processing block in the processing block set, fine control of important image units and processing blocks is achieved, global consistency and accuracy in the generation process are guaranteed, and generation quality is maintained while accelerating inference. Furthermore, by dividing the image units in the image unit sequence according to the grid size, each grid screens the target image units of the unit cache quantity according to the first importance score to obtain the cached image units. The grid-based image unit pruning scheme can prevent pruning from being overly concentrated in certain areas, and further maintain the global consistency and quality of the generation.

[0350] Furthermore, a unitization module in an image generation model is used to convert a sample potential noise image into a sample image unit sequence, and the total number of sample time steps and the time step type of each sample time step are determined. If the time step type of the sample time step is a cache pruning step, a first sample image feature corresponding to the sample image unit sequence in each processing block in a feature processing block set is obtained, and a sample target image feature of the image unit sequence is determined based on each first sample image feature. If the time step type of the sample time step is a complete inference step, a fourth sample image feature corresponding to the sample image unit sequence in each processing block in the feature processing block set is obtained, and a target sample image feature of the sample image unit sequence is determined based on each fourth sample image feature. The number of sample image noises corresponding to the sample time step is determined based on the sample target image feature. According to the sample image noise data, the sample potential noise image is denoised, and the sample generated image corresponding to the sample time step is generated. If the number of the sample time step does not reach the total number of the sample time step, the sample time step is updated, and the sample generated image is used as the sample potential noise image, and the step of converting the sample potential noise image into a sample image unit sequence by using the unitization module in the image generation model is executed until the number of the sample time step reaches the total number of the sample time step, and the sample generated image is obtained. The first predicted loss value of the sample image and the sample generated image is determined based on the first preset loss function, and the image generation model is supervised and trained based on the first predicted loss value, and the model parameters of the image generation model are iteratively updated until the first predicted loss value converges, and the trained image generation model is obtained. By adopting the image unit pruning strategy when training the image generation model, the number of inferences for some image units is reduced, the speed of the image generation model is improved, and the quality of the generated image can be guaranteed by continuously optimizing the loss function.

[0351] Furthermore, by generating a sample prediction importance score of the target sample image unit in the first time step according to a cache predictor, determining the superimposed sample image features of the target sample image unit based on the sample prediction importance score, the first sample image features and the second sample image features output by the same processing block in adjacent time steps, and comparing the superimposed sample image features with the first sample image features to determine the prediction accuracy of the importance score, and by learning the importance of each image unit in different time steps, it is possible to help the model optimize the calculation process, reduce unnecessary calculations, and achieve better image generation effects.

[0352] In addition, an embodiment of the present specification provides a computer program product, wherein the computer program product includes a computer program, and when the computer program is executed by a processor of an electronic device, the processor can at least implement the above-mentioned Figures 1 to 11 The methods provided in the illustrated embodiments.

[0353] A person skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0354] The above disclosure is only the preferred embodiment of this specification, which certainly cannot be used to limit the scope of rights of this specification. Therefore, equivalent changes made according to the claims of this specification are still within the scope covered by this specification.

Claims

1. A method for generating an image, the method comprising: Acquire an image unit sequence corresponding to a potential noise image; the image unit sequence includes a plurality of image units; If the time step type of the time step is a cache pruning step, then the first image feature corresponding to the image unit sequence in each processing block in the feature processing block set is obtained, and the target image feature of the image unit sequence is determined based on each of the first image features; each of the processing blocks includes at least one first processing block that determines the first image feature corresponding to the image unit sequence in the first processing block based on a second image feature and a third image feature, wherein the second image feature is a historical cache feature of the image unit in the target feature processing block at the previous time step, and the third image feature is an image feature of the image unit generated by a previous processing block of the target feature processing block; If the time step type of the time step is a complete inference step, obtaining a fourth image feature corresponding to the image unit sequence in each of the processing blocks in the feature processing block set, and determining a target image feature of the image unit sequence based on each of the fourth image features; the fourth image feature is determined based on each of the third image features; Determining image noise data corresponding to the time step based on the target image feature; The potential noise image is denoised based on the image noise data to generate a target image corresponding to the time step.

2. The method according to claim 1, after obtaining the image unit sequence corresponding to the potential noise image, further comprising: Obtain a target time step of the number of buffer intervals per interval after the start time step, and determine the time step types of the start time step and the target time step as a complete reasoning step; Time steps other than the complete time steps are determined as cache pruning steps.

3. The method of claim 2, wherein the cache interval number comprises a first cache interval number and a second cache interval number; The step of obtaining the target time step of the number of buffer intervals per interval after the start time step and determining the time step types of the start time step and the target time step as complete reasoning steps further includes: If the number of steps of the complete reasoning step is less than or equal to the first preset number of steps, determining the number of cache intervals to be the first number of cache intervals; If the number of the complete reasoning steps is greater than the first preset number of steps, the number of buffer intervals is determined to be a second number of buffer intervals.

4. The method according to claim 1, wherein obtaining the first image feature corresponding to the image unit sequence in each processing block in the feature processing block set comprises: If the processing block is a first processing block, generating a first image feature corresponding to the image unit sequence in the first processing block based on the second image feature and the third image feature; If the processing block is a second processing block, a first image feature corresponding to the image unit sequence in the second processing block is generated based on the third image feature.

5. The method according to claim 4, wherein determining the target image feature of the image unit sequence based on each of the first image features comprises: If the processing block is the last processing block in the feature processing block set, the first image feature of the processing block is used to determine the target image feature of the image unit sequence.

6. The method according to claim 4, wherein if the processing block is a first processing block, generating a first image feature corresponding to the image unit sequence in the first processing block based on the second image feature and the third image feature comprises: If the processing block is the first processing block, determining the inference image unit and the cache image unit in the image unit sequence; Acquire a second image feature of the inference image unit; generating, using the first processing block, an inference image feature of the inference image unit based on the second image feature; Acquire a third image feature of the cache image unit; The inferred image feature and the third image feature of the cached image unit are determined as the first image feature corresponding to the image unit sequence in the first processing block.

7. The method of claim 6, further comprising: Determining a first importance score of each image unit in the image unit sequence in each of the processing blocks; Determining a second importance score for each of the processing blocks based on the first importance score; The first processing block and the second processing block are determined from the set of feature processing blocks based on the second importance score.

8. The method of claim 7, wherein determining the second importance score of each processing block based on the first importance score comprises: Determining a sum of first importance scores of the image unit sequences in each of the processing blocks; A second importance score of each of the processing blocks is determined based on the sum of the first importance scores.

9. The method of claim 7, wherein if the processing block is the first processing block, determining the inference image unit and the cache image unit in the image unit sequence comprises: If the processing block is the first processing block, based on a first importance score and a preset pruning rate of each image unit in the image unit sequence in the first processing block, an inference image unit and a cache image unit corresponding to the first processing block are determined in the image unit sequence.

10. The method according to claim 7, wherein if the processing block is the first processing block, determining the inference image unit and the cache image unit in the image unit sequence comprises: If the processing block is a first processing block, based on a first importance score of each image unit in the image unit sequence in the first processing block and the total number of cache units, the inference image units and cache image units corresponding to each of the first processing blocks are determined.

11. The method of claim 10, further comprising: The total number of cache units of each of the first processing blocks is determined based on the second importance score of each of the first processing blocks.

12. The method according to claim 10, wherein determining the inference image unit and the cache image unit corresponding to each of the first processing blocks based on the first importance score of each of the image units in the first processing blocks and the number of cache units comprises: Dividing the image units in the image unit sequence according to the grid size to obtain a plurality of grids; wherein each grid includes the same number of image units; Determine the number of unit caches of each of the grids based on the total number of cache units; Based on the first importance score of each of the image units in the first processing block, respectively determine the target image units in each of the grids, and determine the target image units corresponding to each of the grids as the cache image units; The target image unit corresponding to any target grid among the multiple grids is selected from the image unit sequence of the target grid based on the unit cache quantity in the target grid, and the importance score of the target image unit is not higher than the importance scores of other image units in the image unit sequence except the target image unit.

13. The method according to claim 1, wherein obtaining a sequence of image units corresponding to a potential noise image comprises: Segmenting the potential noise image according to a preset segmentation size to obtain a plurality of image units; Determining position identification information of the image unit in the potential noise image; An image unit sequence corresponding to the potential noise image is generated based on the multiple image units and the position identification information.

14. The method according to claim 1, after performing denoising on the potential noise image based on the image noise data to generate the target image corresponding to the time step, further comprising: If the number of the time step does not reach the second preset number of steps, the time step is updated, and the target image is used as a potential noise image to execute the step of obtaining an image unit sequence corresponding to the potential noise image until the number of the time step reaches the second preset number of steps.

15. A method for training an image generation model, the method comprising: The sample potential noise image is converted into a sample image unit sequence by using a unitization module in an image generation model; the sample potential noise image is generated based on the sample image; determining a total number of sample time steps and a time step type for each of the sample time steps; If the time step type of the sample time step is a cache pruning step, then the first sample image feature corresponding to the sample image unit sequence in each processing block in the feature processing block set is obtained, and the sample target image feature of the image unit sequence is determined based on each of the first sample image features; each of the processing blocks includes at least one first processing block that determines the first sample image feature corresponding to the sample image unit sequence in the first processing block based on the second sample image feature and the third sample image feature, the second image feature is a historical cache feature of the image unit in the target feature processing block at the previous time step, and the third image feature is an image feature of the image unit generated by a previous processing block of the target feature processing block; If the time step type of the sample time step is a complete inference step, obtaining a fourth sample image feature corresponding to the sample image unit sequence in each of the processing blocks in the feature processing block set, and determining a target sample image feature of the sample image unit sequence based on each of the fourth sample image features; the fourth sample image feature is determined based on each of the third sample image features; Determining sample image noise data corresponding to the sample time step based on the sample target image feature; Performing denoising on the sample potential noise image based on the sample image noise data to generate a sample generated image corresponding to the sample time step; If the number of the sample time step does not reach the total number of sample time steps, the sample time step is updated, and the sample generated image is used as a sample potential noise image, and the step of converting the sample potential noise image into a sample image unit sequence by using a unitization module in the image generation model is executed until the number of the sample time step reaches the total number of sample time steps, thereby obtaining a sample generated image; Based on a first preset loss function, a first predicted loss value of the sample image and the sample generated image is determined, and based on the first predicted loss value, supervised training is performed on the image generation model and model parameters of the image generation model are iteratively updated until the first predicted loss value converges, thereby obtaining a trained image generation model.

16. The method of claim 15, further comprising: Acquire a first sample image feature of a target sample image unit generated by a target processing block in a first time step and a second sample image feature of the target sample image unit generated by the target processing block in a second time step; the first time step is any time step in the sample time steps; the second time step is the next time step of the first time step; the target processing block is any processing block in a feature processing block set; the target sample image unit is any sample image unit in the sample image unit sequence; generating a sample prediction importance score for the target sample image unit in the first time step based on a cache predictor; Determining a superimposed sample image feature of the target sample image unit based on the sample prediction importance score, the first sample image feature, and the second sample image feature; Based on a second preset loss function, a second predicted loss value of the superimposed sample image feature and the first sample image feature is determined, and based on the predicted loss value, the cache predictor is trained and the model parameters of the cache predictor are iteratively updated until the second predicted loss value converges, thereby obtaining a trained cache predictor.

17. The method of claim 16, wherein if the time step type of the sample time step is a cache pruning step, obtaining the first sample image feature corresponding to the sample image unit sequence in each processing block in the feature processing block set comprises: If the time step type of the sample time step is a cache pruning step, determining a first sample importance score of each sample image unit in the sample image unit sequence in each of the processing blocks using the cache predictor; Determining a second sample importance score for each of the processing blocks based on the first sample importance score; determining a first sample processing block and a second sample processing block from the processing blocks based on the second sample importance score; If the sample processing block is a first sample processing block, generating a first sample image feature corresponding to the sample image unit sequence in the first sample processing block based on the second sample image feature and the third sample image feature; If the processing block is a second processing block, a first sample image feature corresponding to the image unit sequence in the second sample processing block is generated based on the third sample image feature.

18. An image generating device, comprising: An acquisition unit, used to acquire an image unit sequence corresponding to a potential noise image; the image unit sequence includes a plurality of image units; A first inference unit is used for obtaining, if the time step type of the time step is a cache pruning step, a first image feature corresponding to the image unit sequence in each processing block in a feature processing block set, and determining a target image feature of the image unit sequence based on each of the first image features; each of the processing blocks includes at least one first processing block that determines the first image feature corresponding to the image unit sequence in the first processing block based on a second image feature and a third image feature, wherein the second image feature is a historical cache feature of the image unit in the target feature processing block at the previous time step, and the third image feature is an image feature of the image unit generated by a previous processing block of the target feature processing block; a second inference unit, configured to obtain, if the time step type of the time step is a complete inference step, a fourth image feature corresponding to the image unit sequence in each of the processing blocks in the feature processing block set, and determine a target image feature of the image unit sequence based on each of the fourth image features; the fourth image feature is determined based on each of the third image features; A noise determination unit, configured to determine image noise data corresponding to the time step based on the target image feature; A denoising processing unit is used to perform denoising on the potential noise image based on the image noise data to generate a target image corresponding to the time step.

19. An image generating device, comprising: A unitization unit, used to convert a sample potential noise image into a sample image unit sequence by using a unitization module in an image generation model; the sample potential noise image is generated based on the sample image; a parameter determination unit, used to determine the total number of sample time steps and the time step type of each of the sample time steps; A first training unit is used for obtaining, if the time step type of the sample time step is a cache pruning step, a first sample image feature corresponding to the sample image unit sequence in each processing block in a feature processing block set, and determining a sample target image feature of the image unit sequence based on each of the first sample image features; each of the processing blocks includes at least one first processing block that determines the first sample image feature corresponding to the sample image unit sequence in the first processing block based on a second sample image feature and a third sample image feature, wherein the second image feature is a historical cache feature of the image unit in the target feature processing block at the previous time step, and the third image feature is an image feature of the image unit generated by a previous processing block of the target feature processing block; a second training unit, configured to obtain, if the time step type of the sample time step is a complete inference step, fourth sample image features corresponding to the sample image unit sequence in each of the processing blocks in the feature processing block set, and determine a target sample image feature of the sample image unit sequence based on each of the fourth sample image features; The fourth sample image feature is determined based on each of the third sample image features; A sample noise determination unit, configured to determine sample image noise data corresponding to the sample time step based on the sample target image feature; A sample denoising unit, configured to perform denoising processing on the sample potential noise image based on the sample image noise data, and generate a sample generated image corresponding to the sample time step; A time step updating unit, configured to update the sample time step if the number of the sample time step does not reach the total number of the sample time steps, and use the sample generated image as a sample potential noise image, and proceed to execute the step of converting the sample potential noise image into a sample image unit sequence by using a unitization module in the image generation model, until the number of the sample time step reaches the total number of the sample time steps, thereby obtaining a sample generated image; A parameter updating unit is used to determine a first predicted loss value of the sample image and the sample generated image based on a first preset loss function, perform supervised training on the image generation model based on the first predicted loss value, and iteratively update the model parameters of the image generation model until the first predicted loss value converges, thereby obtaining a trained image generation model.

20. An electronic device, comprising: Processor and memory; The memory stores a computer program, wherein the computer program is suitable for being loaded by the processor and executing the steps of the method according to any one of claims 1 to 17.

21. A storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 17.

22. A computer program product comprising: A computer program, when executed by a processor of an electronic device, causes the processor to perform the steps of the method according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • An efficient patch-based method for video denoising

    CN108337402A

  • Image processing method and device, storage medium, terminal and computer program product

    CN118691493A

  • Techniques for content synthesis using denoising diffusion models

    US20230368337A1

  • Multimodal diffusion models

    US20240265505A1