Image generation method and system based on grouping Dirichlet diffusion, terminal and storage medium

Through the grouping Dirichlet diffusion method, the problem of insufficient in-group dependency and hierarchical structure capture in the existing diffusion model in multi-channel data processing is solved, and high stability and high-quality image generation effect is achieved.

CN120259477AActive Publication Date: 2025-07-04SHENZHEN MSU-BIT UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510733549.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

In terms of modeling in-group dependencies and inter-group structures, existing diffusion models cannot effectively capture the in-group dependencies and hierarchies of multi-channel data, resulting in poor image data stability.

Method used

The grouped Dirichlet diffusion method is used to divide the image data into independent groups, noise perturbation is performed through the grouped Dirichlet distribution, and noise data is mapped in the logit space. The encoder and decoder are trained using potential embedding vectors, combining long-range dependencies and jump connection features, downsampling and upsampling of image data, and KL divergence optimization training is introduced.

Benefits of technology

Improve the stability and quality of image generation, improve the applicability of the model in a variety of tasks, especially when processing high-dimensional bounded data, and enhance the accuracy of image generation and recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259477A_ABST
    Figure CN120259477A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image generation, and discloses a grouping Dirichlet diffusion-based image generation method and system, a terminal and a storage medium, and the method comprises the steps: inputting image data into a grouping Dirichlet generation model for forward diffusion, and outputting noise data; mapping the noise data to a logit space to obtain a potential embedded vector, and training an encoder and a decoder through the potential embedded vector; inputting the image data into an encoder for down-sampling, and outputting a low-resolution bottleneck feature; and inputting the low-resolution bottleneck features into a decoder for up-sampling, and outputting target image data. According to the method, the data is controlled to be kept in a grouping Dirichlet distribution group in the forward and backward diffusion process, the numerical stability is ensured, and the model is widely applied to various tasks such as image generation, image recovery and structure modeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image generation, and in particular, to an image generation method, system, terminal, and computer-readable storage medium based on grouped Dirichlet diffusion. Background Art

[0002] Image generation is one of the core research directions in the fields of artificial intelligence and computer vision, aiming to automatically generate images with visual authenticity or semantic consistency through algorithms. In recent years, diffusion models have rapidly emerged due to their excellent performance in image clarity, diversity, and training stability, becoming the mainstream method for image generation. Its basic principle is to gradually add noise to the image (forward process) and learn to denoise (reverse process), gradually recovering a realistic image from the noise.

[0003] Current mainstream diffusion generation models are mostly based on Gaussian noise distributions. Although these methods have achieved remarkable results in the field of image generation, they have certain limitations when dealing with high-dimensional bounded data (such as multi-channel images), especially in modeling intra-group dependencies and inter-group structures, and cannot effectively capture the intra-group dependency relationships and hierarchical structures of multi-channel data.

[0004] In addition, although the diffusion of the Beta distribution has a certain range control ability, it lacks the adaptive regulation ability between multiple groups of structures, and its numerical stability is limited.

[0005] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0006] The main objective of the present invention is to provide an image generation method, system, terminal, and computer-readable storage medium based on grouped Dirichlet diffusion, aiming to solve the problem in the existing technology that the existing diffusion models cannot effectively capture the intra-group dependency relationships and hierarchical structures of multi-channel data in modeling intra-group dependencies and inter-group structures, resulting in poor stability of image data.

[0007] To achieve the above objective, the present invention provides an image generation method based on grouped Dirichlet diffusion. The image generation method based on grouped Dirichlet diffusion includes the following steps: Obtain image data, input the image data into a constructed grouped Dirichlet generation model for forward diffusion, and output multiple noise data of the image data; Map all the noise data to the logit space, output a latent embedding vector, and train a constructed initial encoder and initial decoder through the latent embedding vector to obtain a target encoder and a target decoder; Input the image data into the target encoder for downsampling to output low-resolution bottleneck features and skip connection features; Input the low-resolution bottleneck features and the skip connection features into the target decoder for upsampling to output target image data.

[0008] Optionally, in the image generation method based on grouped Dirichlet diffusion, the step of obtaining image data and inputting the image data into the constructed grouped Dirichlet generation model for forward diffusion to output multiple noise data of the image data specifically includes: Obtain the image data input by the user and construct a grouped Dirichlet generation model; Input the image data into the grouped Dirichlet generation model, and the grouped Dirichlet generation model divides the image data to obtain multiple independent groups: ; wherein, represents the number of independent groups, represents the index of the independent group, represents the th independent group, represents the concentration parameter of the th independent group, represents the process of adding noise perturbation to the image data, represents the grouped Dirichlet distribution; Use the concentration parameter of each independent group to perform noise addition and diffusion on each independent group to output noise data at different time steps: ; ; wherein, represents the probability distribution of the noise data at the th time step, represents the probability distribution of the noise data at the th time step, represents the grouped Dirichlet generation model, represents the th time step of the noise data, represents the th time step of the noise data, represents the input data, represents the th time step of the noise intensity, represents the th time step of the noise intensity, represents the global concentration parameter.

[0009] Optionally, in the above-mentioned image generation method based on grouped Dirichlet diffusion, before using the concentration parameter of each independent group to perform noise-added diffusion on each independent group and output noise data at different time steps, it further includes: Using a non-linear strategy to obtain the noise intensity at each time step: ; ; where represents the noise intensity at the -th time step, represents applying a non-linear strategy to , represents the concentration parameter at the -th time step, and both represent constants, represents the normalized time.

[0010] Optionally, in the above-mentioned image generation method based on grouped Dirichlet diffusion, the step of mapping all the noise data to the logit space, outputting a latent embedding vector, and training the constructed initial encoder and initial decoder with the latent embedding vector to obtain a target encoder and a target decoder specifically includes: Mapping all the noise data to the logit space according to the time step encoding of all the noise data to convert all the noise data into high-dimensional vectors: ; where represents the representation of in the logit space, represents the noise data at the -th time step of the -th independent group, represents the noise data at the -th time step of the -th independent group, represents the representation of in the logit space, represents the log-odds ratio of , represents the noise perturbation data of the -th independent group from the -th time step to the -th time step, represents the representation in the logit space of the -th independent group from the -th time step to the -th time step, represents the log odds ratio of represents and the log odds ratio of; All the transformed high-dimensional noise data are subjected to a multi-layer linear transformation to obtain latent embedding vectors; The latent embedding vectors are input into the constructed initial encoder and initial decoder for training to obtain a target encoder and a target decoder.

[0011] Optionally, in the image generation method based on grouped Dirichlet diffusion, wherein the inputting the image data into the target encoder for downsampling to output low-resolution bottleneck features and skip connection features specifically includes: Input the image data into the target encoder, and the initial convolutional layer of the target encoder converts the image data into a base feature map; Through multiple downsampling modules of the target encoder, capture the long-range dependencies of the base feature map and output low-resolution bottleneck features and skip connection features corresponding to each of the downsampling modules.

[0012] Optionally, in the image generation method based on grouped Dirichlet diffusion, wherein the inputting the low-resolution bottleneck features and the skip connection features into the target decoder for upsampling to output target image data specifically includes: Input the low-resolution bottleneck features and all the skip connection features into the target decoder; Multiple upsampling modules of the target decoder perform transposed convolution on the low-resolution bottleneck features to restore the spatial resolution and obtain high-resolution bottleneck features; After fusing the high-resolution bottleneck features with each of the skip connection features, map them to the target data dimension to obtain target image data.

[0013] Optionally, in the image generation method based on grouped Dirichlet diffusion, wherein the inputting the low-resolution bottleneck features and the skip connection features into the target decoder for upsampling to output target image data, and then further includes: Obtain the true distribution and the predicted distribution of the image data in the grouped Dirichlet distribution, and calculate the forward-backward KL divergence and the marginal KL divergence according to the true distribution and the predicted distribution; Calculate the loss function under the grouped Dirichlet distribution according to the forward-backward KL divergence and the marginal KL divergence: ; wherein represents the loss function represents the weight coefficient, represents the true distribution of the image data, represents the predicted distribution of the image data, represents the true distribution in the grouped Dirichlet distribution, represents the predicted distribution in the grouped Dirichlet distribution, represents the number of independent groups, represents the index of the independent group, represents and the true distribution of the image data with significant differences, represents in the grouped Dirichlet distribution and the true distribution of the image data with significant differences, represents the forward - reverse KL divergence, represents the marginal KL divergence; The grouped Dirichlet generation model is optimized and trained using the loss function.

[0014] In addition, to achieve the above object, the present invention also provides an image generation system based on grouped Dirichlet diffusion, wherein the image generation system based on grouped Dirichlet diffusion includes: A noise - adding module, configured to obtain image data, input the image data into the constructed grouped Dirichlet generation model for forward diffusion, and output a plurality of noise data of the image data; A mapping module, configured to map all the noise data to the logit space, output a latent embedding vector, and train the constructed initial encoder and initial decoder through the latent embedding vector to obtain a target encoder and a target decoder; An encoding module, configured to input the image data into the target encoder for downsampling, and output a low - resolution bottleneck feature and a skip connection feature; A decoding module, configured to input the low - resolution bottleneck feature and the skip connection feature into the target decoder for upsampling, and output target image data.

[0015] In addition, to achieve the above object, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and an image generation program based on grouped Dirichlet diffusion stored on the memory and executable on the processor. When the image generation program based on grouped Dirichlet diffusion is executed by the processor, the steps of the above - described image generation method based on grouped Dirichlet diffusion are implemented.

[0016] In addition, to achieve the above object, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an image generation program based on grouped Dirichlet diffusion, and when the image generation program based on grouped Dirichlet diffusion is executed by a processor, the steps of the above-mentioned image generation method based on grouped Dirichlet diffusion are implemented.

[0017] In the present invention, image data is acquired, and the image data is input into a pre-constructed grouped Dirichlet generation model for forward diffusion to output a plurality of noise data of the image data; all the noise data is mapped to the logit space to output a latent embedding vector, and an initial encoder and an initial decoder that have been constructed are trained through the latent embedding vector to obtain a target encoder and a target decoder; the image data is input into the target encoder for downsampling to output a low-resolution bottleneck feature and a skip connection feature; the low-resolution bottleneck feature and the skip connection feature are input into the target decoder for upsampling to output target image data. The present invention controls the data to remain within the distribution family of grouped Dirichlet during the forward and reverse diffusion processes, ensuring numerical stability, enabling the model to be widely applicable to various tasks such as image generation, image restoration, and structure modeling. At the same time, KL divergence is introduced to replace the traditional ELBO divergence (Evidence Lower Bound), improving the stability and generation quality of training. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a flowchart of a preferred embodiment of the image generation method based on grouped Dirichlet diffusion of the present invention; Figure 2 is a structural diagram of the training process of a preferred embodiment of the image generation method based on grouped Dirichlet diffusion of the present invention; Figure 3 is a schematic diagram of adding noise to a first image in a preferred embodiment of the image generation method based on grouped Dirichlet diffusion of the present invention; Figure 4 is a schematic diagram of adding noise to a second image in a preferred embodiment of the image generation method based on grouped Dirichlet diffusion of the present invention; Figure 5 is a schematic diagram of removing noise from a first image in a preferred embodiment of the image generation method based on grouped Dirichlet diffusion of the present invention; Figure 6 is a schematic diagram of removing noise from a second image in a preferred embodiment of the image generation method based on grouped Dirichlet diffusion of the present invention; Figure 7 is a structural diagram of a preferred embodiment of the image generation system based on grouped Dirichlet diffusion of the present invention; Figure 8 is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation Modes

[0019] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the present invention will be further described in detail below with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.

[0020] In view of the fact that traditional methods have certain limitations in dealing with high-dimensional bounded data (such as multi-channel images), especially in modeling intra-group dependencies and inter-group structures, the present invention proposes a novel grouped Dirichlet diffusion model: Grouped Dirichlet Diffusion (GDD). By introducing a multi-group structure modeling mechanism using the Grouped Dirichlet distribution, and by retaining the intra-group probability consistency and adaptively adjusting the inter-group interaction during the diffusion process, high-quality generation of high-dimensional bounded probability data is achieved.

[0021] The image generation method based on grouped Dirichlet diffusion according to a preferred embodiment of the present invention, as Figure 1 shown, the image generation method based on grouped Dirichlet diffusion includes the following steps: Step S10: Obtain image data, input the image data into the constructed grouped Dirichlet generation model for forward diffusion, and output multiple noise data of the image data.

[0022] Among them, in the process of noise scheduling for each group of semantic channels in the image data, a Sigmoid nonlinear strategy is used (referring to a nonlinear mapping method with the Sigmoid function as the core). Specifically: using the nonlinear strategy, obtain the noise intensity at each time step: ; ; Among them, represents the noise intensity at the th time step, represents applying the nonlinear strategy to , represents the concentration parameter at the th time step, and both represent constants, represents the normalized time; through this nonlinear strategy, it can be ensured that the noise intensity of each group of data smoothly decays from the initial value to 0, thus effectively avoiding the problem of boundary mutation in traditional linear scheduling.

[0023] Specifically, obtain the image data input by the user and construct a grouped Dirichlet generative model; input the image data into the grouped Dirichlet generative model, and the grouped Dirichlet generative model divides the image data to obtain multiple independent groups: ; Among them, represents the number of independent groups, represents the index of the independent group, represents the th independent group, represents the th concentration parameter of the independent group, represents the process of adding noise perturbation to the image data, represents the grouped Dirichlet distribution; use the concentration parameter of each independent group to perform noise addition and diffusion on each independent group, and output the noise data at different time steps: ; ; Among them, represents the probability distribution of the noise data at the th time step, represents the probability distribution of the noise data at the th time step, represents the grouped Dirichlet generative model, represents the th time step of the noise data, represents the th time step of the noise data, represents the input data, represents the th time step of the noise intensity, represents the th time step of the noise intensity, represents the global concentration parameter.

[0024] Among them, multiple semantic channels of the image data are represented by the grouped Dirichlet distribution. For example, for RGB (Red - Green - Blue), the image data can be divided into multiple independent groups and gradually perturbed by Dirichlet noise, that is, gradually increasing the noise of each channel data. In this process, the global concentration parameter can be used to effectively control the dispersion degree of the noise, thereby improving the stability of the model during the diffusion process.

[0025] Step S20: Map all the noise data to the logit space (a multi-dimensional space composed of the raw unnormalized scores output by the model), output the latent embedding vectors, and train the constructed initial encoder and initial decoder with the latent embedding vectors to obtain the target encoder and target decoder.

[0026] Among them, the noise data output through the forward diffusion process is used to reconstruct the image data from the noise data, and adaptive denoising is achieved by predicting the Dirichlet parameters. This process can be implemented using the U-Net framework (a symmetric encoder-decoder structure, as Figure 2 shown), where the U-Net framework includes structures such as group normalization, convolutional layers, linear layers, dropout layers, and skip connections. In the mapping stage, after passing through multiple fully connected layers, positional embeddings are performed; residual blocks combined with attention mechanisms are added in the encoder part, and residual blocks combined with skip connections are added in the decoder part.

[0027] Specifically, according to the time step encoding of all the noise data, map all the noise data to the logit space to convert all the noise data into high-dimensional vectors: ; where represents the representation in the logit space, represents the th noise data at the th time step of the th independent group, represents the representation in the logit space, represents the log odds of represents the th noise perturbation data from the th time step to the th time step of the th independent group, represents the representation in the logit space at the time from the th time step to the th time step of the represents and The log odds ratio; all the converted high-dimensional noise data are subjected to multi-layer linear transformation to obtain a latent embedding vector; the latent embedding vector is input into the constructed initial encoder and initial decoder for training to obtain a target encoder and a target decoder.

[0028] Among them, in order to avoid numerical overflow of the probability boundary, all Dirichlet samplings are completed in the logit space; and before denoising, the corresponding noise data are converted into high-dimensional vectors by using the position encoding carried by the noise data. Further, a latent embedding vector is generated through multi-layer linear transformation (where the activation function is SiLU, Sigmoid-Weighted Linear Unit, the S-shaped weighted linear unit), and the generated latent embedding vector can be used to modulate the parameters of the encoder and decoder to ensure that the encoder-decoder can realize the diffusion and reconstruction of grouped Dirichlet of high-dimensional bounded data (such as multi-channel images).

[0029] Step S30: Input the image data into the target encoder for downsampling to output low-resolution bottleneck features and skip connection features.

[0030] Specifically, the image data is input into the target encoder, and the initial convolutional layer of the target encoder converts the image data into a basic feature map; through multiple downsampling modules of the target encoder, the long-range dependencies of the basic feature map are captured, and low-resolution bottleneck features and skip connection features corresponding to each downsampling module are output.

[0031] Among them, in the encoding stage, the initial convolutional layer will convert the input channels (such as RGB channels) into a basic feature map (such as 64 channels), and then in the multi-level downsampling modules, each level of downsampling module contains multiple UNetBlock modules (UNet Building Block, UNet represents the core network architecture, Building Block refers to the basic modular unit that constitutes the network), which are configured in the downsampling mode (such as stride 2 convolution). In these modules, residual connections and self-attention mechanisms are applied to capture long-range dependencies; among them, long-range dependencies refer to the correlation between elements that are far apart in the data. The self-attention mechanism directly models the global relationship between all elements in the sequence and is naturally suitable for capturing long-range dependencies. The residual connection adds the input directly to the output of the network layer through a skip connection to solve the gradient disappearance and network degradation problems in deep networks. The combination of the two can capture long-range dependencies more efficiently.

[0032] Further, the low-resolution bottleneck features output by each level of downsampling module can be transmitted to the decoder through skip connections, and this process can retain the multi-scale details of the image data, thereby improving the accuracy of image generation.

[0033] Step S40: Input the low-resolution bottleneck feature and the skip connection feature into the target decoder for upsampling to output target image data.

[0034] Specifically, input the low-resolution bottleneck feature and all the skip connection features into the target decoder; multiple upsampling modules of the target decoder perform transposed convolution on the low-resolution bottleneck feature to restore the spatial resolution and obtain a high-resolution bottleneck feature; after fusing the high-resolution bottleneck feature with each skip connection feature, map it to the target data dimension to obtain the target image data.

[0035] Among them, for the bottleneck feature and multiple skip connection features input by the target encoder, the multi-level upsampling modules in the target decoder restore the spatial resolution of each noise data through transposed convolution or interpolation upsampling, and then fuse each skip connection feature in the encoder to enhance the detail reconstruction ability.

[0036] Further, after fusion, map the feature map to the target data dimension (such as the RGB three channels) through a convolutional layer. For this process, a Sigmoid activation function can be selected to ensure that the output conforms to the bounded data constraint, such as ensuring that the pixel values of the image are within 0-1, and finally generate the target image data.

[0037] Further, obtain the true distribution and the predicted distribution of the image data in the grouped Dirichlet distribution, and calculate the forward-backward KL divergence and the marginal KL divergence according to the true distribution and the predicted distribution; Calculate the loss function under the grouped Dirichlet distribution according to the forward-backward KL divergence and the marginal KL divergence: ; where represents the loss function, represents the weight coefficient, represents the true distribution of the image data, represents the predicted distribution of the image data, represents the true distribution in the grouped Dirichlet distribution, represents the predicted distribution in the grouped Dirichlet distribution, represents the number of independent groups, represents the index of the independent group, represents related to the true distribution of the image data with significant differences, represents in the grouped Dirichlet distribution related to the true distribution of the image data with significant differences, represents the forward - reverse KL divergence, represents the marginal KL divergence; the grouped Dirichlet generative model is optimized and trained using the loss function.

[0038] Among them, the loss function is trained and optimized using the KL upper bound (KLUB, Kullback - Leibler Upper Bound, KL divergence), which can effectively improve the convergence and stability of the model; for this process, combining the forward - reverse KL divergence and the marginal KL divergence, the loss function of the grouped Dirichlet distribution is calculated, and the weight coefficient of the calculation process can balance the temporal consistency and the generation quality. In this embodiment, the weight coefficient is set to 0.97, and then the grouped Dirichlet generative model is optimized and trained using the loss function.

[0039] Furthermore, in another embodiment, as Figure 3 and Figure 4 shown, it shows the noise addition of the image under the grouped Dirichlet distribution, that is, the degradation process of the image, as Figure 5 and Figure 6 shown, it shows the denoising process of gradually restoring the image details from the chaotic noise of the image under the grouped Dirichlet distribution.

[0040] Among them, on the CIFAR - 10 dataset, the FID metric (Fréchet Inception Distance, a classic metric in the deep learning field for evaluating the quality of images generated by generative models) drops from 16.31 (DDPM, Denoising Diffusion Probabilistic Models) of the traditional diffusion model to 5.13 (GDD), with a 68.5% improvement; the KID (Kernel Inception Distance) score drops to 0.00403, (more than 45% improvement compared to the traditional diffusion model), as shown in Table 1 below: Table 1: Comparative analysis table of FID and KID scores of multiple generative frameworks based on CIFAR - 10 training data

[0041] Among them, the lower the value of the FID score, the better the image generation effect; VAE stands for Variational Autoencoder, a variational autoencoder; GAN stands for Generative Adversarial Network, a generative adversarial network; AutoGAN stands for Automated Generative Adversarial Network, an automated improvement framework based on the generative adversarial network (GAN); DDPM stands for Denoising Diffusion Probabilistic Models, denoising diffusion probabilistic models; PPOGAN stands for Proximal Policy Optimization Generative Adversarial Network, an improved model that combines the proximal policy optimization algorithm in reinforcement learning with the generative adversarial network; LSGM stands for Latent Score-based Generative Model, a latent space score-based generative model; DDIM stands for Denoising Diffusion Implicit Models, denoising diffusion implicit models; ViTGAN stands for Vision Transformer Generative Adversarial Network, an image processing model based on the Transformer architecture; Consistency Models stands for consistency models; EAGAN stands for Efficient Two-stage Evolutionary Architecture Search for GAN, an efficient two-stage evolutionary architecture search generative adversarial network; GENIE stands for Generative Interactive Environments, generative interactive environments; FM stands for Factorization Machine Generative Adversarial Network, a factorization machine generative adversarial network; GLR-GAN stands for Generalized Likelihood Ratio Generative Adversarial Network, a generalized likelihood ratio generative adversarial network.

[0042] Furthermore, as shown in Table 2, GDD reduces the average FID by 30% on datasets such as STL-10 and SVHN: Table 2: Comparison table of FID scores on different datasets

[0043] In the present invention, control data remains within the family of grouped Dirichlet distributions during the forward and reverse diffusion processes, ensuring numerical stability and enabling the model to be widely applicable to various tasks such as image generation, image restoration, and structure modeling. At the same time, KL divergence is introduced to replace the traditional ELBO divergence, increasing the model convergence speed by 50% and further enhancing the training stability and generation quality.

[0044] Furthermore, as Figure 7 shown, based on the above-mentioned grouped Dirichlet diffusion-based image generation method, the present invention also correspondingly provides a grouped Dirichlet diffusion-based image generation system. Among them, the grouped Dirichlet diffusion-based image generation system includes: A noise addition module 51, configured to obtain image data, input the image data into a pre-constructed grouped Dirichlet generation model for forward diffusion, and output multiple noise data of the image data; A mapping module 52, configured to map all the noise data to the logit space, output a latent embedding vector, and train a pre-constructed initial encoder and initial decoder through the latent embedding vector to obtain a target encoder and a target decoder; An encoding module 53, configured to input the image data into the target encoder for downsampling, and output a low-resolution bottleneck feature and a skip connection feature; A decoding module 54, configured to input the low-resolution bottleneck feature and the skip connection feature into the target decoder for upsampling, and output target image data.

[0045] Furthermore, as Figure 8 shown, based on the above-mentioned grouped Dirichlet diffusion-based image generation method and system, the present invention also correspondingly provides a terminal, and the terminal includes a processor 10, a memory 20, and a display 30. Figure 8 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0046] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as the hard disk or memory of the terminal. In some other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk equipped on the terminal, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 20 may also include both the internal storage unit of the terminal and the external storage device. The memory 20 is used to store the application software installed on the terminal and various types of data, such as the program code of the installed terminal, etc. The memory 20 may also be used to temporarily store the data that has been output or will be output. In one embodiment, an image generation program 40 based on grouped Dirichlet diffusion is stored on the memory 20, and the image generation program 40 based on grouped Dirichlet diffusion can be executed by the processor 10, so as to implement the image generation method based on grouped Dirichlet diffusion in the present application.

[0047] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chips, and is used to run the program code stored in the memory 20 or process data, such as executing the image generation method based on grouped Dirichlet diffusion, etc.

[0048] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. The display 30 is used to display the information on the terminal and to display a visual user interface. The components of the terminal communicate with each other through a system bus.

[0049] In one embodiment, when the processor 10 executes the image generation program 40 based on grouped Dirichlet diffusion in the memory 20, the steps of the image generation method based on grouped Dirichlet diffusion as described above are implemented.

[0050] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an image generation program based on grouped Dirichlet diffusion, and when the image generation program based on grouped Dirichlet diffusion is executed by a processor, the steps of the image generation method based on grouped Dirichlet diffusion as described above are implemented.

[0051] In summary, the present invention provides an image generation method and related devices based on grouped Dirichlet diffusion. The method includes: obtaining image data, inputting the image data into a constructed grouped Dirichlet generation model for forward diffusion, and outputting multiple noise data of the image data; mapping all the noise data to the logit space, outputting a latent embedding vector, and training a constructed initial encoder and initial decoder through the latent embedding vector to obtain a target encoder and a target decoder; inputting the image data into the target encoder for downsampling, and outputting a low-resolution bottleneck feature and a skip connection feature; inputting the low-resolution bottleneck feature and the skip connection feature into the target decoder for upsampling, and outputting target image data. The present invention controls the data to remain within the distribution family of grouped Dirichlet during the forward and reverse diffusion processes, ensuring numerical stability, making the model widely applicable to various tasks such as image generation, image restoration, and structure modeling. At the same time, KL divergence is introduced to replace the traditional ELBO divergence, improving the training stability and generation quality.

[0052] It should be noted that in this article, the terms "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or terminal including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or terminal. Without more limitations, the element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or terminal including that element.

[0053] Of course, those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disc, etc.

[0054] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. An image generation method based on grouped Dirichlet diffusion, characterized in that, The image generation method based on grouped Dirichlet diffusion includes: Obtain image data, input the image data into the constructed grouped Dirichlet generation model for forward diffusion, and output multiple noise data of the image data; Map all the noise data to the logit space, output a latent embedding vector, and train the constructed initial encoder and initial decoder through the latent embedding vector to obtain a target encoder and a target decoder; Input the image data into the target encoder for downsampling, and output low-resolution bottleneck features and skip connection features; Input the low-resolution bottleneck features and the skip connection features into the target decoder for upsampling, and output target image data.

2. The image generation method based on grouped Dirichlet diffusion according to claim 1, wherein The obtaining of image data, inputting the image data into the constructed grouped Dirichlet generation model for forward diffusion, and outputting multiple noise data of the image data specifically includes: Obtain the image data input by the user and construct a grouped Dirichlet generation model; Input the image data into the grouped Dirichlet generation model, and the grouped Dirichlet generation model divides the image data to obtain multiple independent groups: ; Among them, represents the number of independent groups, represents the index of the independent group, represents the th independent group, represents the th concentration parameter of the independent group, represents the process of performing noise perturbation on the image data, represents the grouped Dirichlet distribution; Use the concentration parameter of each independent group to perform noise addition and diffusion on each independent group, and output noise data at different time steps: ; ; Among them, represents the probability distribution of the noise data at the th time step, represents the probability distribution of the noise data at the th time step, represents the grouped Dirichlet generative model, represents the noise data at the th time step, represents the noise data at the th time step, represents the input data, represents the noise intensity at the th time step, represents the noise intensity at the th time step, represents the global concentration parameter.

3. The image generation method based on grouped Dirichlet diffusion according to claim 2, wherein Before the step of using the concentration parameter of each independent group to perform noise addition and diffusion on each independent group and outputting noise data at different time steps, it further includes: Use a non-linear strategy to obtain the noise intensity at each time step: ; ; Among them, represents the noise intensity at the th time step, represents applying a non-linear strategy to , represents the concentration parameter at the th time step, and both represent constants, represents the normalized time.

4. The image generation method based on grouped Dirichlet diffusion according to claim 1, wherein The mapping of all the noise data to the logit space, outputting a latent embedding vector, and training the constructed initial encoder and initial decoder through the latent embedding vector to obtain a target encoder and a target decoder specifically includes: According to the time step encoding of all the noise data, map all the noise data to the logit space to convert all the noise data into high-dimensional vectors: ; Among them, represents the representation in the logit space, represents the th noise data at the th time step of the th independent group, represents the th noise data at the th time step of the th independent group, represents the log-odds ratio of represents the th noise perturbation data of the th independent group from the th time step to the th time step, represents the th independent group at the th time step in the logit space from the th time step to the the log-odds ratio of represents and the log-odds ratio of; Perform a multi-layer linear transformation on all the converted high-dimensional noise data to obtain a latent embedding vector; Input the latent embedding vector into the constructed initial encoder and initial decoder for training to obtain a target encoder and a target decoder.

5. The image generation method based on grouped Dirichlet diffusion according to claim 1, wherein The inputting of the image data into the target encoder for downsampling, and outputting low-resolution bottleneck features and skip connection features specifically includes: Input the image data into the target encoder, and the initial convolutional layer of the target encoder converts the image data into a basic feature map; Capture the long-range dependence relationship of the basic feature map through multiple downsampling modules of the target encoder, and output low-resolution bottleneck features and skip connection features corresponding to each downsampling module.

6. The image generation method based on grouped Dirichlet diffusion according to claim 5, wherein The inputting of the low-resolution bottleneck features and the skip connection features into the target decoder for upsampling, and outputting target image data specifically includes: Input the low-resolution bottleneck features and all the skip connection features into the target decoder; Multiple upsampling modules of the target decoder perform transposed convolution on the low-resolution bottleneck features to restore the spatial resolution and obtain high-resolution bottleneck features; After fusing the high-resolution bottleneck features with each of the skip connection features, map them to the target data dimension to obtain the target image data.

7. The image generation method based on grouped Dirichlet diffusion according to claim 1, wherein After inputting the low-resolution bottleneck features and the skip connection features into the target decoder for upsampling to output the target image data, the method further includes: Obtain the true distribution and the predicted distribution of the image data in the grouped Dirichlet distribution, and calculate the forward-backward KL divergence and the marginal KL divergence according to the true distribution and the predicted distribution; Calculate the loss function under the grouped Dirichlet distribution according to the forward-backward KL divergence and the marginal KL divergence; ; Among them, represents the loss function, represents the weight coefficient, represents the true distribution of the image data, represents the predicted distribution of the image data, represents the true distribution in the grouped Dirichlet distribution, represents the predicted distribution in the grouped Dirichlet distribution, represents the number of independent groups, represents the index of the independent group, represents and the true distribution of the image data with significant differences, represents in the grouped Dirichlet distribution and the true distribution of the image data with significant differences, represents the forward - reverse KL divergence, represents the marginal KL divergence; Optimize and train the grouped Dirichlet generative model using the loss function.

8. An image generation system based on grouped Dirichlet diffusion, characterized in that, The image generation system based on grouped Dirichlet diffusion includes: A noise addition module for obtaining image data, inputting the image data into the constructed grouped Dirichlet generative model for forward diffusion, and outputting multiple noise data of the image data; A mapping module for mapping all the noise data to the logit space, outputting a latent embedding vector, and training the constructed initial encoder and initial decoder through the latent embedding vector to obtain a target encoder and a target decoder; An encoding module for inputting the image data into the target encoder for downsampling to output low-resolution bottleneck features and skip connection features; A decoding module for inputting the low-resolution bottleneck features and the skip connection features into the target decoder for upsampling to output the target image data.

9. A terminal, characterized in that, The terminal includes: a memory, a processor, and an image generation program based on grouped Dirichlet diffusion stored on the memory and executable on the processor. When the image generation program based on grouped Dirichlet diffusion is executed by the processor, it implements the steps of the image generation method based on grouped Dirichlet diffusion according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an image generation program based on grouped Dirichlet diffusion. When the image generation program based on grouped Dirichlet diffusion is executed by a processor, it implements the steps of the image generation method based on grouped Dirichlet diffusion according to any one of claims 1-7.

Citation Information

Patent Citations

  • Spoken language understanding method based on Dirichlet variational auto-encoder and related equipment

    CN111724767A

  • Method and device for generating image based on diffusion model, and storage medium

    CN116721179A

  • Generative invisible watermark method based on stable diffusion model

    CN119693214A

  • Diffusion models having continuous scaling through patch-wise image generation

    US20240161327A1