Image sequence generation method based on multi-dimensional denoising probability diffusion model based on copula

By adopting a multi-dimensional denoising diffusion probability model based on Copula in image sequence generation, the problem that univariate Gaussian model is difficult to characterize multi-channel images or image sequence correlation is solved, and more efficient image generation and computing performance is achieved.

CN118229534BActive Publication Date: 2025-05-13YIBIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311143765.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-05
Publication Date
2025-05-13
Estimated Expiration
2043-09-05

AI Technical Summary

Technical Problem

The existing univariate Gaussian diffusion probability model is difficult to effectively characterize the correlation between multi-channel images or image sequences, resulting in poor computational efficiency and fitting performance during diffusion and inverse diffusion.

Method used

The multi-dimensional denoising diffusion probability model (CM-DDPM) based on Copula is used to connect the univariate probability model into a multi-dimensional probability distribution through copula, solving the problem that the correlation between data sources is not effectively utilized.

Benefits of technology

More effective denoising and diffusion of multi-channel images or image sequences is achieved, improving the performance and computing efficiency of the model in image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118229534B_ABST
    Figure CN118229534B_ABST
Patent Text Reader

Abstract

There is a strong correlation between adjacent image frames of a sequence and between channels of a multi-channel image. Aiming at the shortcomings of the Denoising Diffusion Probabilistic Model (DDPM), the method of the present invention proposes a multidimensional denoising diffusion probability model based on copula, called (Copula Multivariate DDPM, CM‑DDPM). The proposed model uses copula to replace the Gaussian probability distribution of a single variable in DDPM to realize a multidimensional denoising diffusion probability model, and solves the shortcomings of the traditional DDPM based on a single variable in characterizing the correlation between data sources. The present invention is suitable for a variety of application scenarios, including video generation, animation production, autonomous driving, human behavior recognition, medical image processing and other image generation fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of image sequence generation, and in particular to an image sequence generation method by converting Copula to a denoising probability diffusion model. Background Art

[0002] Image sequence generation is a task in the field of computer vision and artificial intelligence that involves generating a continuous sequence of images from static images or other forms of data. This task is commonly used in a variety of applications, including video generation, animation production, autonomous driving, human behavior recognition, medical image processing, etc. Common generation techniques include: (1) Generative adversarial network (GAN), (2) Variational Autoencoders (VAE), and (3) Denoising Diffusion Probabilistic Model (DDPM).

[0003] The DDPM diffusion model is an excellent image generation model with good image quality and generation ability. DDPM uses a regression loss function, and the training process is more stable than GAN, and it is not easy to have the problem of gradient disappearance or explosion. However, some DDPMs are univariate diffusion probability models. Because DDPM uses a univariate Gaussian model, the probability model expressions of the forward and reverse intermediate images need to be calculated during the diffusion and reverse diffusion processes, and the univariate method cannot characterize the correlation between multi-channel images or image sequences. For multi-channel images or image sequences, the "independent multivariate Gaussian distribution" (the diagonal of its correlation matrix is) is currently used, that is, the correlation between multi-source data is not considered. Although this method is feasible for the calculation of diffusion and reverse diffusion processes, it cannot achieve optimized fitting performance. At present, DDPM that considers the correlation between data and is applied to multi-image sequence analysis and color image analysis has not been reported. Summary of the invention

[0004] There is a strong correlation between adjacent image frames of a sequence and between channels of a multi-channel image. Aiming at the shortcomings of DDPM, the method of the present invention proposes a multi-dimensional denoising diffusion probability model based on Copula, called (CopulaMultivariate DDPM, CM-DDPM). The proposed model uses copula to replace the Gaussian diffusion of a single variable in DDPM to realize a multi-dimensional denoising diffusion probability model, and solves the shortcomings of the traditional single-variable DDPM in describing the correlation between data sources.

[0005] Due to the complexity of the diffusion process, it is very difficult to directly construct a multidimensional probability distribution DDPM. The subject first independently diffuses the marginal distribution forward and reverse samples it, and then uses copula to connect these univariate probability models into a multidimensional probability distribution DDPM. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Figure 1 Copula-based multidimensional denoising probability diffusion model (CM-DDPM). DETAILED DESCRIPTION

[0007] Assume that the diffusion target image sequence is x t ={x 0t , x 1t ,x kt There are k+1 image sequences in total, where 0≤t≤T represents the diffusion time step, and T is the total time step; the conditional cd variable can be any image, text or voice information, which is used to control the generation of image sequences. Usually x t and cd form a sample pair to train a deep neural network. it The diffusion of is regarded as the change process of the density function of the marginal distribution. Figure 1 Medium θ (x t-1 |x t ) is the distribution of noise images from t to t-1 predicted by the deep neural network. The two processes of DDPM are represented by the following two Gaussian probability density functions:

[0008] Forward noise adding process:

[0009]

[0010] Among them, the t-step noise image is is a number controlled by the diffusion length t; and They represent the mean and variance after t-step noise addition.

[0011] Reverse restoration process:

[0012]

[0013] in, and is the mean and variance when deriving the t-1 step from the t step, expressed as and The diffusion model allows deep neural networks to learn the ability to generate data through the iterative process of forward noise addition and reverse restoration.

[0014] The following is combined with Figure 1The method of the present invention is further described, and the key implementation steps are as follows:

[0015] Step 1: Given the target image sequence x t ={x 0t , x 1t , x kt}, indicating that there are k+1 related images. This sequence is also the target image sequence predicted by CM-DDPM. The forward diffusion process is as follows:

[0016] Step 1.1 Set x t ={x 0t , x 1t , x kt} is input into the DDPM formula (1) to obtain the noise images x of k+1 independent sequence forward processes it Its probability density function q i , where 0≤i≤k.

[0017] Step 1.2: q i and x it They are respectively regarded as the marginal distribution of the copula and the observed data corresponding to its marginal distribution. Then we get the parameter ψ=arg(q i ) and the marginal probability cumulative distribution Q i =CDF i (x it ;ψ);

[0018] Step 1.3 According to q i The generated sequence uses the maximum likelihood algorithm to obtain the parameter Φ of the probability density function c. Using the Gaussian copula connection function c, the multidimensional distribution function h of the forward image sequence is calculated using the following formula:

[0019] h=c(Q1,...,Q k ;Φ)·∏ k q i (x it ;ψ)

[0020] Step 1.4: Use the maximum likelihood method to calculate h and obtain the multidimensional probability distribution model h of the forward diffusion model. Φ and ψ are the parameters of the model.

[0021] Step 2, the reverse restoration process is as follows:

[0022] Step 2.1 Use the formula (2) of the DDPM method to obtain the probability distribution p of k+1 reverse restored images at step t i (x it-1 |x it );

[0023] Step 2.2: i (xit-1 |x it ) and x it -1 is regarded as the marginal distribution of the edge copula and its observed data, and then the cumulative distribution function of these k+1 marginal distributions is calculated using the maximum likelihood method;

[0024] Step 2.3 uses the Gaussian copula connection function c to calculate the multidimensional distribution function h of the inverse restored image sequence using the following formula: θ :

[0025]

[0026] Step 2.4 Use the maximum likelihood method to calculate h θ Get the multidimensional probability distribution model h of the forward diffusion model θ Φ and are model parameters;

[0027] Step 3: The deep network of the CM-DDPM model uses UNet, and the training loss function is expressed as follows:

[0028] loss = λ1×MSE(||Z θ (x t ,cd)-Z||2)+λ2KLD(h φ , h)

[0029] Among them, z θ (x t , cd) is the UNet network, x t ={x 0t , x 1t , x kt ) is an image sequence, cd is a conditional variable, which can be an image, text, or sound data; Z is the noise added to the image sequence; KLD(h φ , h) is the Gaussian copula model h and calculated in step 1.4 and step 2.4 respectively The Kullback-Leibler distance between them; λ1 and λ2 are the weights of the two control loss functions.

[0030] Step 4: Perform inference on the trained UNet. Iterate the trained network T = 1000 times to obtain the final prediction output of the network. The iteration formula is as follows:

[0031] x t-1 =Z θ (x t , cd)

[0032] The above formula, the image x at time t t And condition cd input network, reverse generate image x at time t-1t-1 。

Claims

1. A method for generating image sequences based on a multidimensional denoising probability diffusion model based on copula, characterized in that It includes the following steps: Step 1: Given the target image sequence x t ={x 0t , x 1t , x kt }, indicating that there are k+1 related images. This sequence is also the predicted target image sequence. The forward diffusion process is as follows: Step 1.1 Set x t ={x 0t ,x 1t ,x kt } is input into the DDPM formula (1) to obtain the noise images x of k+1 independent sequence forward processes it Its probability density function q i , where 0≤i≤k Among them, the t-step noise image is is a number controlled by the diffusion length t; and Respectively represent the mean and variance after t-step noise addition; Step 1.2: q i and x it We consider them as the marginal distribution of copula and the observed data corresponding to its marginal distribution, and then get the parameter ψ=agr(q i ) and the marginal probability cumulative distribution Q i =CDF i (x it ;ψ); Step 1.3 According to q i The generated sequence uses the maximum likelihood algorithm to obtain the parameter Φ of the probability density function c, and the Gaussian copula connection function c is used to calculate the multidimensional distribution function h of the forward image sequence using the following formula: h=c(Q1,…,Q k ;F)·P k q i (x it (ψ) Step 1.4: Use the maximum likelihood method to calculate h and obtain the multidimensional probability distribution model h of the forward diffusion model. Φ and ψ are the parameters of the model. Step 2 The reverse restoration process is as follows: Step 2.1 Use the formula (2) of the DDPM method to obtain the probability distribution p of k+1 reverse restored images at step t i (x it-1 |x it ); in, and is the mean and variance when deriving the t-1 step from the t step, expressed as and The diffusion model allows the deep neural network to learn the ability to generate data in the iterative process of continuous forward noise addition and reverse restoration; Step 2.2: i (x it-1 |x it ) and x it-1 Consider the marginal distribution of the edge copula and its observed data and then use the maximum likelihood method to calculate the cumulative distribution function of these k+1 marginal distributions; Step 2.3 uses the Gaussian copula connection function c to calculate the multidimensional distribution function h of the inverse restored image sequence using the following formula: θ : Step 2.4 Use the maximum likelihood method to calculate h θ Get the multidimensional probability distribution model h of the forward diffusion model θ , Φ and are model parameters; Step 3: The deep network of the CM-DDPM model uses UNet, and the training loss function is expressed as follows: loss=λ1×MSE(||Z θ (x t ,cd)-Z||2)+λ2KLD(h φ ,h) Among them, Z θ (x t , cd) is the UNet network, x t ={x 0t , x 1t , x kt } is an image sequence, cd is a conditional variable, which can be an image, text, or sound data; Z is the noise added to the image sequence; KLD(h φ , h) is the Gaussian copula model h and h calculated in step 1.4 and step 2.4 respectively The Kullback-Leibler distance between them; λ1 and λ2 are the weights of the two control loss functions. Step 4: Perform inference on the trained UNet. Iterate the trained network T=1000 times to obtain the final prediction output of the network. The iteration formula is as follows: x t-1 =Z θ (x t ,continued) The above formula, the image x at time t t And condition cd input network, reverse generate t-1 time image x t-1 .

2. The image sequence generation method according to claim 1, characterized in that The Gaussian copula model is used to connect multiple independent denoising diffusion probability models into a multidimensional denoising diffusion probability model.

3. The image sequence generation method according to claim 1, characterized in that The proposed image sequence generation method can be extended and applied to the generation of other correlated data sequences.

Citation Information

Patent Citations

  • Crop income insurance information diffusion actuarial method

    CN112258327A

  • Self-supervised contrast learning method for human body action recognition based on diffusion model

    CN116415152A