Method and apparatus for media data generation

By analytically deriving coefficients for the consistency function in diffusion models, the training efficiency of consistency trajectory models is enhanced, addressing suboptimal preconditionings and achieving faster media data generation.

WO2026081112A1PCT designated stage Publication Date: 2026-04-23ROBERT BOSCH GMBH +1
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ROBERT BOSCH GMBH
Filing Date
2024-10-16
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing consistency models and consistency trajectory models in diffusion models require suboptimal hand-crafted preconditionings for their consistency functions, leading to inefficiencies in training and generation speed.

Method used

A novel approach to parameterize the consistency function in diffusion models using analytically derived coefficients from the teacher diffusion network, ensuring alignment with the denoiser and adherence to boundary conditions, thereby optimizing the training process.

Benefits of technology

The proposed method accelerates the training of consistency trajectory models by 2-3 times, improving the speed-quality trade-off in media data generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024125176_23042026_PF_FP_ABST
    Figure CN2024125176_23042026_PF_FP_ABST
Patent Text Reader

Abstract

A method for media data generation is disclosed. The method comprises receiving input media data, and generating output media data based on the input media data through one-step or multi-step generation of a consistency trajectory network. The consistency trajectory network is represented by a consistency function and is trained from a teacher diffusion network through consistency distillation. The consistency function maps media data at time t to media data at a time s previous or equal to time t on the trajectory, the time t and the time s range from original time to final time, and the consistency function is parameterized as a sum of the media data at time t multiplied by a first coefficient and a student denoiser function from the media data at time t to the media data at the time s multiplied by a second coefficient, wherein the first coefficient and the second coefficient are determined by the teacher diffusion network.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND APPARATUS FOR MEDIA DATA GENERATIONFIELD

[0001] The present disclosure relates generally to artificial intelligence technology, and more particularly, to generation models for generating media data.BACKGROUND

[0002] Diffusion models are a class of powerful deep generative models, showcasing cutting-edge performance in diverse domains including image synthesis, speech and video generation, controllable image manipulation, density estimation and inverse problem solving, etc. Compared to their generative counterparts like variation auto-encoders (VAEs) and generative adversarial networks (GANs) , diffusion models excel in high-quality generation while circumventing issues of posterior collapse and training instability. Consequently, they serve as the cornerstone of next-generation generative systems like text-to-image and text-to-video synthesis. The primary bottleneck for integrating diffusion models into downstream tasks lies in their slow inference processes, which gradually remove noise from data with hundreds of network evaluations. The sampling process typically involves simulating the probability flow (PF) ordinary differential equation (ODE) backward in time, starting from noise.

[0003] To accelerate diffusion sampling, various training-free samplers have been proposed as specialized solvers of the PF-ODE, yet they still require over 10 steps to generate satisfactory samples due to the inherent discretization errors present in all numerical ODE solvers. Recent advancements in few-step or even single-step generation of diffusion models are concentrated on distillation methods. Particularly, consistency models (CMs) have emerged as a prominent method for diffusion distillation and successfully been applied to various data domains including latent space, audio (including speech) and video. CMs consider training a student network to map arbitrary points on the PF-ODE trajectory to its starting point, thereby enabling one-step generation that directly maps noise to data. A follow-up work named consistency trajectory models (CTMs) extends CMs by changing the mapping destination to encompass not only the starting point but also intermediate ones, facilitating unconstrained backward jumps on the PF-ODE trajectory. This design enhances training flexibility and permits the incorporation of auxiliary losses.

[0004] In both CMs and CTMs, the mapping function (referred to as the consistency function) , that maps arbitrary points on the PF-ODE trajectory to its starting point or intermediate ones, must adhere to certain constraints. For instance, in CMs, there exists a boundary condition dictating that the mapping of a starting point should yield itself. Consequently, the consistency functions are parameterized as a linear combination of the input data and the network output with pre-defined coefficients. This approach ensures that boundary conditions are naturally satisfied without constraining the form or expressiveness of the neural network. However, the coefficients in existing CMs and CTMs are intuitively defined, and may not be superior ones. Therefore, there exists a need to improve the existing CMs and CTMs.SUMMARY

[0005] The following presents a simplified summary of one or more aspects according to the present disclosure in order to provide a basic understanding of such aspects. This summary is not an extensive overview of all contemplated aspects, and is intended to neither identify key or critical elements of all aspects nor delineate the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that is presented later.

[0006] In an aspect of the disclosure, a method for media data generation is provided. The method comprises receiving input media data, and generating output media data based on the input media data through one-step or multi-step generation of a consistency trajectory network. The consistency trajectory network is represented by a consistency function and is trained from a teacher diffusion network through consistency distillation. The consistency function maps media data at time t to media data at a time s previous or equal to time t on the trajectory, the time t and the time s range from original time to final time, and the consistency function is parameterized as a sum of the media data at time t multiplied by a first coefficient and a student denoiser function from the media data at time t to the media data at the time s multiplied by a second coefficient, wherein the first coefficient and the second coefficient are determined by the teacher diffusion network.

[0007] In another aspect of the disclosure, an apparatus for media data generation is provided. The apparatus may comprise a memory and at least one processor coupled to the memory. The at least one processor may be configured to receive input media data, and generate output media data based on the input media data through one-step or multi-step generation of a consistency trajectory network. The consistency trajectory network  is represented by a consistency function and is trained from a teacher diffusion network through consistency distillation. The consistency function maps media data at time t to media data at a time s previous or equal to time t on the trajectory, the time t and the time s range from original time to final time, and the consistency function is parameterized as a sum of the media data at time t multiplied by a first coefficient and a student denoiser function from the media data at time t to the media data at the time s multiplied by a second coefficient, wherein the first coefficient and the second coefficient are determined by the teacher diffusion network.

[0008] In another aspect of the disclosure, a computer readable medium storing computer program codes for media data generation is provided. The computer program codes, when executed by a processor, may cause the processor to receive input media data, and generate output media data based on the input media data through one-step or multi-step generation of a consistency trajectory network. The consistency trajectory network is represented by a consistency function and is trained from a teacher diffusion network through consistency distillation. The consistency function maps media data at time t to media data at a time s previous or equal to time t on the trajectory, the time t and the time s range from original time to final time, and the consistency function is parameterized as a sum of the media data at time t multiplied by a first coefficient and a student denoiser function from the media data at time t to the media data at the time s multiplied by a second coefficient, wherein the first coefficient and the second coefficient are determined by the teacher diffusion network.

[0009] In another aspect of the disclosure, a computer program product for media data generation is disclosed. The computer program product may comprise processor executable computer program codes for receiving input media data, and generating output media data based on the input media data through one-step or multi-step generation of a consistency trajectory network. The consistency trajectory network is represented by a consistency function and is trained from a teacher diffusion network through consistency distillation. The consistency function maps media data at time t to media data at a time s previous or equal to time t on the trajectory, the time t and the time s range from original time to final time, and the consistency function is parameterized as a sum of the media data at time t multiplied by a first coefficient and a student denoiser function from the media data at time t to the media data at the time s multiplied by a second coefficient, wherein the first coefficient and the second coefficient are determined by the teacher diffusion network.

[0010] Other aspects or variations of the disclosure will become apparent by  consideration of the following detailed description and accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The following figures depict various embodiments of the present disclosure for purposes of illustration only. One skilled in the art will readily recognize from the following description that alternative embodiments of the methods and structures disclosed herein may be implemented without departing from the spirit and principles of the disclosure described herein.

[0012] FIG. 1 illustrates a diffusion process of diffusion models in accordance with one aspect of the present disclosure.

[0013] FIG. 2 illustrates parameterization of a consistency function in consistency trajectory models in accordance with one aspect of the present disclosure.

[0014] FIG. 3 illustrates a flow chart of a method for media data generation in accordance with one aspect of the present disclosure.

[0015] FIG. 4 illustrates a block diagram of an apparatus for media data generation in accordance with one aspect of the present disclosure.DETAILED DESCRIPTION

[0016] Before any embodiments of the present disclosure are explained in detail, it is to be understood that the disclosure is not limited in its application to the details of construction and the arrangement of features set forth in the following description. The disclosure is capable of other embodiments and of being practiced or of being carried out in various ways.

[0017] Diffusion models are a type of generative model that may generate data by progressively adding noise to a dataset and then learning to reverse this process for generating data by de-noising. Diffusion models may be trained to generate various media data in real world. For example, diffusion models may generate image, video (each frame of video may also be construed as an image) , audio (including speech) , text, 3-D objects, and other types of media data that can be recognized by humans. All these kinds of data than can be generated by diffusion models are collectively referred to as media data herein.

[0018] FIG. 1 illustrates a diffusion process of diffusion models in accordance with one aspect of the present disclosure. As shown in FIG. 1, the diffusion models may transform a d-dimensional data distribution q0 (x0) such as, the original image x0  into Gaussian noise distribution such as the noise image xT, through a forward stochastic differential equation (SDE) starting from x0~q0 : dxt=f (t) xt dt+g (t) dwt          (1)

[0019] where t∈ [0, T] for some finite horizon is the scalar-valued drift and diffusion term, and is a standard Wiener process. The forward SDE is accompanied by a series of marginal distributions of and f, g are properly designed so that the terminal distribution is approximately a pure Gaussian, i.e.,  An intriguing characteristic of this SDE lies in the presence of the probability flow ordinary differential equation (PF-ODE) :

[0020] whose solution trajectories at time t, when solved backward from time T to time 0, are distributed exactly as qt. The only unknown term is the score function of the marginal density qt and can be learned by denoising score matching (DSM) .

[0021] A prevalent noise schedule is proposed by EDM (Elucidating the design space of diffusion-based generative models. In Advances in Neural Information Processing Systems, 2022) and followed in recent text-to-image generation, video generation, as well as consistency distillation. In this case, the forward transition kernel of the forward SDE (Eq. (1) ) owns a simple form and the terminal distribution Besides, the PF-ODE can be represented by the denoiser function Dφ (xt, t) :

[0022] where the denoiser function is trained to predict x0 given noisy data at any time t, i.e., minimizing for some weighting w (t) . This denoising loss is equivalent to the DSM loss. In EDM, another key insight is to employ preconditioning by parameterizing Dφ as Dφ (x, t) =cskip (t) x+cout (t) Fφ (x, t) , where

[0023] Fφ is a free-form neural network and is the variance of the data distribution.

[0024] More precisely, Dφ (x, t) =csaip (t) x+cout (t) Fφ (cin (t) x, cnoixe (t) ) . Since cin (t) and cnoix (t) take effects inside the network, we absorb them into the definition of Fφ for simplicity.

[0025] Consistency distillation is a prevalent way for accelerating diffusion models adopted in consistency (trajectory) models, in which a student model is trained to traverse backward on the probability flow (PF) ordinary differential equation (ODE) trajectory determined by a teacher model.

[0026] Denote φ as the parameters of the teacher diffusion model, and θ as the parameters of the student model. Given a trajectory with a fixed initial timestep ∈ of a teacher PF-ODE, consistency models (CMs) aim to a consistency function  which maps the point xt at any time t on the trajectory to the initial point x∈. In EDM, the range of timesteps is typically chosen as ∈=0.002, T=80. The consistency function is forced to satisfy the boundary condition fθ (x, ∈) =x, i.e., the mapping of the initial / starting point should yield itself. To ensure unrestricted form and expressiveness of the neural network, fθ is parameterized as

[0027] which naturally satisfies the boundary condition for any free-form network Fθ (xt, t) , since when t=∈, the coefficient of the first term is 1 and the coefficient of the second term is 0. This technique may be called as preconditioning in consistency distillation, aligning with the terminology in EDM. The student network θ can be distilled from the teacher network φ by the training objective:

[0028] where λ (·) is a positive weighting function, d (·, ·) is a distance metric, sg is the (exponential moving average) stop-gradient and Solverφ is any numerical solver for the teacher PF-ODE and the result of Solverφ (xt, t, s) may be xs.

[0029] FIG. 2 illustrates parameterization of a consistency function in consistency trajectory models in accordance with one aspect of the present disclosure. Consistency  trajectory models (CTMs) extend CMs by changing the mapping destination to not only the initial point but also any intermediate ones, enabling unconstrained backward jumps on the PF-ODE. As compared to CMs, the consistency function in CTMs is instead defined as  which maps the point xt at time t on the trajectory 210 to the point xs at any previous time s<t. Generally, the consistency function fθ: (xt, t, s) may be parameterized as a linear combination of the input data and the network output with pre-defined coefficients ɑs, and βs, .

[0030] The boundary condition is fθ (x, t, t) =x, which is forced by the following preconditioning:

[0031] where Dθ (x, t, s) =cskip (t) x+cout (t) Fθ (x, t, s) is the student denoiser function, and Fθ (x, t, s) is a free-form network with an extra timestep input s. The student network is trained by minimizing

[0032] An important property of CTM's precondtioning is that when s→t, the optimal denoiser satisfies Dθ* (x, t, s) →Dφ (xt, t) , i.e. the teacher diffusion denoiser. Consequently, the DSM loss in diffusion models can be incorporated to regularize the training of student network θ, which enhances the sample quality as the number of sampling steps increases, enabling speed-quality trade-off.

[0033] Preconditioning is a vital technique for stabilizing consistency distillation, by linear combining the input data and the network output with pre-defined coefficients as the consistency function. It imposes the boundary condition of consistency functions without restricting the form and expressiveness of the neural network. However, previous preconditionings are hand-crafted and may be suboptimal choices.

[0034] Beyond the hand-crafted preconditionings outlined in Eq. (5) and Eq. (7) , we seek a general paradigm of preconditioning design in consistency distillation. We first analyze their key ingredients and relate them to the discretization of the teacher ODE. Then we derive a generalized ODE form, which can induce a novel family of preconditionings. Finally, we propose a principled way to analytically obtain optimized preconditioning by minimizing the consistency gap.

[0035] The preconditioning in consistency distillation may be analyzed by examining  the form of consistency function fθ (x, t, s) in CTMs, wherein it subsumes CMs as a special case by setting the jumping destination s as the initial timestep ∈. Assume fθis parameterized as the following form of skip connection: fθ (x, t, s) =f (t, s) x+g (t, s) Dθ (x, t, s)     (9)

[0036] where Dθ (x, t, s) =cskip (t) x+cout (t) Fθ (x, t, s) represents the student denoiser function in alignment with EDM, and f (t, s) , g (t, s) are coefficients that linearly combine x and Dθ. This disclosure identifies two essential constraints on the coefficients f and g.

[0037] One of the two constraints is boundary condition. For any free-form network Fθ or denoiser network Dθ, the consistency function fθ must adhere to fθ (x, t, t) =x (in-place jumping retains the original data point) . Therefore, f and g should meet the conditions f (t, t) =1 and g (t, t) =0 for any time t.

[0038] Another one of the two constraints is alignment with the denoiser. Denote the optimal consistency function that precisely follows the teacher PF-ODE trajectory as  and the optimal denoiser as according to Eq. (9) . In CTMs, f and g are properly designed so that the limit  Thus, the student denoiser at s=t, i.e. Dθ (x, t, t) , ideally aligns with the teacher denoiser Dφ. This alignment offers two advantages: (1) Dθ (xt, t, t) acts as a valid diffusion denoiser and is amenable to regularization with the DSM loss. (2) The teacher model Dφ serves as an effective initializer of the student Dθ at s=t, implying that Dθ solely at s<t is suboptimal and requires further optimization.

[0039] Precondionings satisfying these constraints can be derived by discretizing the teacher PF-ODE. Suppose the discretization from time t to time s is expressed as xs=f(t, s) xt+g (t, s) Dφ (xt, t) , then f, g naturally satisfy the two conditions: the discretization from t to t must be xt=xt; as s→t, the discretization error tends to 0, and the optimal student for conducting infinitesimally small jumps is just Dφ (xt, t) . For instance, applying Euler method to the PF-ODE in Eq. (3) yields:

[0040] which exactly matches the preconditioning used in CTMs by replacing Dφ (xt, t) with Dθ (xt, t, s) . Elucidating preconditioning as ODE discretization also closely approximates CMs'choice in Eq. (5) . For t>>, we have t-∈≈t, therefore fθ in Eq. (5) approximately equals the denoiser Dθ. On the other hand, as CTMs' choice in Eq. (7) also indicates fθ≈Dθ. Therefore, CMs’ preconditioning is only distinct from ODE discretization when t is close to ∈, which is not the case in one-step or few-step generation.

[0041] Based on the analyses above, the preconditioning (i.e., the coefficients for parameterization process) can be induced from ODE discretization. According to the dedicated ODE solvers in diffusion models, a generalized representation of the teacher ODE in Eq. (3) may be considered, which can give rise to alternative preconditionings that satisfy the restrictions.

[0042] Firstly, the ODE may be modulated with a continuous function Lt to be transformed into an ODE with respect to Ltxt rather than xt. Leveraging the chain rule of derivatives, we obtain where can be substituted by the original teacher ODE, resulting in

[0043] By changing the time variable from t to λt=-logt, the ODE can be further simplified to

[0044] where we denote and tλ=e-λ is the inverse function of λt. Moreover, Lt can be represented by lt as

[0045] Secondly, instead of using t or λt as the time variable in the ODE (i.e., formulate the ODE as or we can employ a generalized time representation  where St is any positive continuous function. This transformation ensures that η monotonically increases with respect to λ, enabling one-to-one inverse mappings tη, λη. To align with Lt, we express St as where we denote  Using ηt as the new time variable, we have and the  ODE in Eq. (12) is further generalized to

[0046] The final generalized ODE in Eq. (13) is theoretically equivalent to the original teacher PF-ODE in Eq. (3) , albeit with a set of introduced free parameters here ∈ refers to the original time and T refers to the final time. Applying the Euler method leads to different discretizations from Eq. (10) :

[0047] which can be rearranged as

[0048] Hence, the induced preconditioning can be expressed by Eq. (9) with a novel set of coefficients Originating from the Euler discretization of an equivalent teacher ODE, these coefficients adhere to the constraints of boundary condition and alignment with the teacher denoiser under any free parameters thus opening avenues for further optimization. The induced preconditioning can also degenerate to CTM's case under specific selections lt=0, st= -1 for t∈ [∈, T] .

[0049] This disclosure further provides principles for optimizing these coefficients. A range of preconditionings with coefficients f, g from Eq. (15) , derived from the generalized teacher ODE presented in Eq. (13) and governed by the free parameters  are proposed. This disclosure aims to establish guiding principles for discerning the optimal sets of thereby attaining superior preconditioning compared to the original one in Eq. (7) .

[0050] Firstly, considering Rosenbrock-type exponential integrators and their relevance in diffusion models, it is suggested that the parameter lt be chosen to restrict the gradient of Eq. (13) 's right-hand side term with respect to xt. This choice ensures the robustness of the resulting ODE against errors in xt. An analytical solution for ltis derived as follows:

[0051] where d is the data dimensionality, ∥·∥F denotes the Frobenius norm and tr (·) represents the trace of a matrix.

[0052] Secondly, to determine the optimal value of parameter st, we dive deeper into  the relationship between the teacher denoiser Dφ (xt, t) and the student denoiser Dθ (xt, t, s) . To satisfy the second constrain described above, the preconditioning is properly designed to ensure that the optimal student denoiser satisfies  We further explore the scenario where s<t by examining the gap  which we refer to as the consistency gap. Minimizing this gap extends the alignment of Dφ and  to cases where s<t, ensuring that the teacher denoiser also serves as a good trajectory jumper. A bound depicting the asymptotic behavior of the consistency gap can be derived as below.

[0053] Suppose there exists some constant C>0 so that the parameters are bounded by |lt|, |st|≤C, then the optimal student denoiser function  under the preconditioning satisfies

[0054] This conforms to the constraint  when s=t. Moreover, considering s in a local neighborhood of t, by Taylor expansion we have  Therefore, the consistency gap for s∈ (t-δ, t) , when δ is small, is roughly Minimizing this yields an analytic solution for st:

[0055] In this disclosure, the resulting preconditioning may be called as Analytic-Precond, as lt, st are analytically determined by the teacher network φ using Eq. (16) and Eq. (18) . Though lt, st are defined over continuous timesteps, they may be computed on hundreds of discretized ones, while obtaining reasonable estimations of their related terms Lt, St, ηt, as well as ls, ss, Ls, Ss, ηs corresponding to time s. The computation is highly efficient utilizing automatic differentiation in modern deep learning frameworks, requiring less than 1%of the total training time.

[0056] Despite the approximation holding true in local neighborhoods of t, the coefficient in the bound exhibits exponential behavior when In practice, directly applying the preconditioning derived from Eq. (15) may cause training instability, especially on long jumps with large step sizes. In order to improve the training stability, a backward Euler method may be employed for its efficacy in handling stiff equations without step size restrictions. Therefore, this disclosure provides a backward rewriting of Eq. (15) from s to t as where are the original coefficients from Eq. (15) . Rearranging this equation yields giving rise to the backward coefficients  and A comparison between different existing preconditionings and the disclosed preconditioning (termed as Analytic-Precond) used in consistency distillation is listed in the following Table 1, in which bidirectional consistency model (BCM) is also an existing alternative preconditioning to CTM.

[0057] Table 1

[0058] In a specific embodiment, at every time t, the parameters lt and st can be directly computed according to Eq. (16) and Eq. (18) , relying solely on the teacher denoiser model Dφ. The computation of lt involves evaluating which is the trace of a Jacobian matrix. Utilizing Hutchinson's trace estimator, it can be unbiasedly estimated as where v obeys a d-dimensional distribution with zero mean and unit covariance. Thus, only the Jacobian-vector product (JVP) Dφ (xt, t) v is required, achievable in computational cost via automatic differentiation. Once lt is obtained, the function gφ (xt, t) =Dφ (xt, t) - (1-lt) xt is determined. The computation of st involves evaluating which expands as  follows:

[0059] where can also be calculated in time by automatic differentiation.

[0060] For the datasets CIFAR-10 and FFHQ 64×64, the parameters lt and st can be computed across 120 discrete timesteps uniformly distributed in log space, with 4096 samples used to estimate the expectation For the dataset ImageNet 64×64, computations are performed across 160 discretized timesteps following EDM's scheduling using 1024 samples to estimate the expectation The total computation times for datasets CIFAR-10, FFHQ 64×64, and ImageNet 64×64 on 8 NVIDIA A800 GPU cards are approximately 38 minutes, 54 minutes, and 38 minutes, respectively.

[0061] Then, the consistency functions with parameters θ of CTMs based on the disclosed coefficients can be trained following the training procedures of existing CTMs. The teacher models are the pre-trained diffusion models on a corresponding dataset, such as provided by EDM. The network architecture of the student models mirrors that of their respective teacher networks, with the addition of a time-conditioning variable s as input. Training of the student models involves minimizing the consistency loss outlined in Eq. (8) and the denoising score matching loss  For the consistency loss, LPIPS (Learned Perceptual Image Patch Similarity) may be used as the distance metric d (·, ·) , which is also the choice of CMs. t and s in the consistency loss may be chosen from N discretized timesteps determined by EDM's scheduling The Heun sampler in EDM may be employed as the solver in Eq. (8) . The number of sampling steps, determined by the gap between t and s, is restricted to avoid excessive training time. For datasets CIFAR-10 and FFHQ 64×64, we select N=18 and the maximum number of sampling steps as 17 , i.e., not restricting the range of jumping from t to s . For dataset ImageNet 64×64, we set N=40 and the maximum number of sampling steps to 20 , so that the jumping range is at most half of the trajectory length. sg (θ) in Eq. (8) is an exponential moving average stop-gradient  version of θ, updated by sg (θ) =stop-gradient (μsg (θ) + (1-μ) θ)     (20)

[0062] The training of CTMs with the disclosed coefficients may follow the hyperparameters used in EDM, setting σmin=∈=0.002, σmax=T=80.0, σdata = 0.5 and ρ=7.

[0063] FIG. 3 illustrates a flow chart of a method 300 for media data generation in accordance with one aspect of the present disclosure. The method 300 may be performed by an apparatus for media data generation comprising at least one processor and a memory coupled to the at least one processor.

[0064] In block 310, the method 300 may receive input media data for performing the media data generation. The media data may comprise at least one of image data, video data (each frame of video may also be construed as an image) , audio data (including speech) , text data, 3-D object data, and other types of media data that can be recognized by humans. In one embodiment, the input media data may comprise just noise data, i.e., xT as described in connection with FIGs. 1 and / or 2. For example, the noise data may be a noise image or a noise audio. In another embodiment, the input media data may comprise media data with noise, such as, xt-1, xt, xt+1, or xs as described in connection with FIGs. 1 and / or 2. In some other embodiments of image synthesis or controllable image manipulation etc., the input media data may comprise different types of media data, such as, both text data and image data. The input media data means the received media data that will be input into a generative neural network. The input media data may be received (i.e., read) from local memory or from a user via user interfaces.

[0065] In block 320, the method 300 may generate output media data based on the input media data through one-step or multi-step generation of a consistency trajectory network. The consistency trajectory network is based on a type of improved diffusion models and may be trained from a teacher diffusion network through consistency distillation, and thus the consistency trajectory network may also be called as a student network. The consistency trajectory network may be represented by a consistency function, which is also called as mapping function. The consistency function may map media data at time t to media data at a time s previous or equal to time t on an ODE trajectory, where both the time t and the time s may range from the original time to the final time. In this disclosure, the consistency trajectory network may also comprise a consistency network, which maps the media data at time t on the trajectory directly to an initial media data (i.e., single step generation) and thus may be construed as a special case of consistency trajectory network.

[0066] The consistency function may be parameterized as a sum of the media data at  time t multiplied by a first coefficient and a student denoiser function from the media data at time t to the media data at the previous time s multiplied by a second coefficient, such as in Eq. (7) . The parameterization process is also termed as preconditioning. In one example, the student denoiser function Dθ (xt, t, s) may be a linear combination of an input media data and a free-form network Fθ (x, t, s) with a timestep input s. The first coefficient and the second coefficient (such as, f (t, s) and g (t, s) ) may be determined by the teacher diffusion network.

[0067] Considering the constraint of boundary condition, the first coefficient and the second coefficient may be designed as: for any time s equals to time t, the first coefficient equals to 1, and the second coefficient equals to 0.

[0068] Considering the constraint of alignment with teacher denoiser function, the first coefficient and the second coefficient may be designed as: for any time s equals to time t, the student denoiser function aims to align with a teacher denoiser function of the teacher diffusion network at time t. That is, the first coefficient and the second coefficient are properly designed so that the limit Thus, the student denoiser at s=t, i.e. Dθ (x, t, t) , ideally aligns with the teacher denoiser Dφ.

[0069] Generally, in this disclosure, the first coefficient and the second coefficient may be derived from generalized teacher diffusion network ordinary differential equation (such as, represented by Eq. (13) ) and governed by a set of free parameters {lt, st} , where the index time t ranges from original / initial time (∈) to final time (T) .

[0070] In one embodiment, the first coefficient is represented by and the second coefficient is represented by where ls is the parameter lt at time s, Lt and Ls are based on integrations of the parameter lt respectively from final time to time t (i.e., lt) and from final time to time s (i.e., ls) . For example, Lt can be represented by Lt as and accordingly Ls can be represented by ls as where tλ=e-λ. Ss is based on integration of the parameter ss from final time T to time s, ss is the parameter st at time s, and St is based on integration of the parameter st from final time to time t. For example, St can be represented by st as and accordingly Ss can be represented by Ss as ηs is based on  integration of multiplications of Ls and Ss from final time to time s, e.g.,  ηt is based on integration of multiplications of Lt and St from final time T to time t, e.g.,  Thus, the first coefficient and the second coefficient may be represented by a set of free parameters {lt, st} that can be analytically determined by the teacher network. The first coefficient and the second coefficient can by optimized by determining an optimal set of parameters {lt, st} .

[0071] In one embodiment, the parameter lt may be chosen to restrict a gradient of the generalized teacher diffusion network ordinary differential equation with respect to the media data at time t. For example, as derived from Eq. (16) , an optimal solution of the parameter lt may be represented by where xt is media data at time t, Dφ is a teacher denoiser function of the teacher diffusion network,  is an expectation under a distribution of the media data at time t, and d is media data dimensionality.

[0072] In one embodiment, the parameter st may be chosen to minimize a difference (also termed as consistency gap herein, such as represented by Eq. (17) ) between the student denoiser function and a teacher denoiser function of the teacher diffusion network. For example, as derived from Eq. (18) , an optimal solution of the parameter st may be represented by where gφ (xt, t) is defined as Dφ (xt, t) - (1-lt) xt, xt is media data at time t, Dφ is a teacher denoiser function of the teacher diffusion network, λt=-logt, and is an expectation under a distribution of the media data at time t.

[0073] In this disclosure, a design criteria of the preconditioning in consistency distillation is provided. The disclosed novel and principled preconditioning may accelerates the training of CTMs in multi-step generation by 2× to 3×. The disclosed approach connects preconditioning to ODE discretization, and emphasizes the alignment between the consistency function and the denoiser function. Minimizing the consistency gap fosters coordination between the consistency loss and the denoising score-matching loss, thereby facilitating speed-quality trade-offs.

[0074] FIG. 4 illustrates a block diagram of an apparatus 400 for media data generation  in accordance with one aspect of the present disclosure. The apparatus 400 may comprise a memory 410 and at least one processor 420. The processor 420 may be coupled to the memory 410 and configured to perform the method 300 described above with reference to Fig. 3. For example, the processor 420 may be configured to receive input media data, and generate output media data based on the input media data through one-step or multi-step generation of a consistency trajectory network. A consistency function of the consistency trajectory network is parameterized according to the design criteria of the preconditioning in this disclosure, and trained from a teacher diffusion network through consistency distillation. The processor 420 may be a general-purpose processor, a graphic processor, a neural processor, or may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. The memory 410 may store the input data (such as, images) , output data, data generated by processor 420, and / or instructions executed by processor 420.

[0075] The various operations, modules, models and networks described in connection with the disclosure herein may be implemented in hardware, software executed by a processor, firmware, or any combination thereof. According an embodiment of the disclosure, a computer program product for media data generation may comprise processor executable computer program codes for performing the method 300 described above with reference to FIG. 3. According to another embodiment of the disclosure, a computer readable medium may store computer program codes for media data generation. The computer program codes when executed by a processor may cause the processor to perform the method 300 described above with reference to FIG. 3. The computer readable media includes both non-transitory computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. Any connection may be properly termed as a computer-readable medium. Other embodiments and implementations are within the scope of the disclosure.

[0076] The preceding description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the various embodiments. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other embodiments without departing from the scope of the various embodiments. Thus, the claims are not intended to be limited to the embodiments shown herein but is to be accorded the widest scope  consistent with the following claims and the principles and novel features disclosed herein.

Claims

1.A method for media data generation, comprising:receiving input media data; andgenerating output media data based on the input media data through one-step or multi-step generation of a consistency trajectory network, wherein the consistency trajectory network is represented by a consistency function and is trained from a teacher diffusion network through consistency distillation, the consistency function maps media data at time t to media data at a time s previous or equal to time t on the trajectory, the time t and the time s range from original time to final time, and the consistency function is parameterized as a sum of the media data at time t multiplied by a first coefficient and a student denoiser function from the media data at time t to the media data at the time s multiplied by a second coefficient, the first coefficient and the second coefficient are determined by the teacher diffusion network.2.The method of claim 1, wherein the input media data comprises noise data or media data with noise, and wherein the media data comprises at least one of text data, image data, audio data, and video data.3.The method of claim 1, wherein for any time s equals to time t, the first coefficient equals to 1, and the second coefficient equals to 0.4.The method of claim 1, wherein for any previous time s equals to time t, the student denoiser function aims to align with a teacher denoiser function of the teacher diffusion network at time t.5.The method of claim 1, wherein the first coefficient and the second coefficient are derived from generalized teacher diffusion network ordinary differential equation and governed by a set of free parameters {lt, st} , where the index time t ranges from original time to final time.6.The method of claim 5, wherein the parameter lt is chosen to restrict a gradient of the generalized teacher diffusion network ordinary differential equation with respect to the media data at time t; and / or, wherein the parameter st is chosen to minimize a difference between the student denoiser function and a teacher denoiser function of the teacher diffusion network.7.The method of claim 5, wherein the first coefficient is represented by and the second coefficient is represented bywhere ls is the parameter lt at time s, Lt and Ls are based on integrations of the parameter lt respectively from final time to time t and from final time to time s, Ss is based on integration of a parameter ss from final time to time s, ss is the parameter st at time s, ηs is based on integration of multiplications of Ls and Ss from final time to time s, ηt is based on integration of multiplications of Lt and St from final time to time t, and St is based on integration of the parameter st from final time to time t.8.The method of claim 7, wherein the parameter lt is represented by where xt is media data at time t, Dφ is a teacher denoiser function of the teacher diffusion network, is an expectation under a distribution of the media data at time t, and d is media data dimensionality.9.The method of claim 7, wherein the parameter st is represented by where gφ (xt, t) is defined as Dφ (xt, t) - (1-lt) xt, xt is media data at time t, Dφ is a teacher denoiser function of the teacher diffusion network, λt=-log t, andis an expectation under a distribution of the media data at time t.10.An apparatus for media data generation, comprising:a memory; andat least one processor coupled to the memory and configured to perform the method of one of claims 1-9.11.A computer readable medium, storing computer program codes for media data generation, the computer program codes when executed by a processor, causing the processor to perform the method of one of claims 1-9.12.A computer program product for media data generation, comprising: processor executable computer program codes for performing the method of one of claims 1-9.

Citation Information

Patent Citations

  • Network model, method and device for generating video by text

    CN115249062A

  • Diffusion model sampling method and device for image generation

    CN116894778A

  • Video generation method and device based on potential consistency model

    CN118741263A

  • Video generation with latent diffusion models

    US20240169479A1

  • Systems and methods for media content generation

    US20240242427A1