Text-to-Image Generation Based on Diffusion Model and Diffusion Model Training Method, Device and Equipment

The diffusion model with cross-attention mechanisms allows for precise control of image instances by dividing samples into blocks and applying attention-based feature extraction, improving image quality and resolution.

CN119169434BActive Publication Date: 2025-07-15ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411657510.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-19
Publication Date
2025-07-15
Estimated Expiration
2044-11-19

AI Technical Summary

Technical Problem

Existing diffusion models can only control the features of the entire generated image, and cannot accurately control specific instances in the image.

Method used

By dividing the training sample picture into multiple diced pieces, the attention scores of global text encoding and local text encoding for each diced piece are calculated using the interactive attention module, and interact with the cross attention module of the diffusion model to perform diced feature extraction and denoising processing, and finally stitching to generate the image.

Benefits of technology

Accurate control of multiple instances in the image is achieved, while improving the resolution and quality of the generated image, and the generated image has a higher resolution and richer content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119169434B_ABST
    Figure CN119169434B_ABST
Patent Text Reader

Abstract

The present application discloses a text-to-image generation method, a diffusion model training method, a device, and a device based on a diffusion model, including: obtaining a sample image, a bounding box of an instance, a local text description, and a global text description; adding noise through a diffusion process; selecting a training sample image and dividing it into multiple chunks; using the cross-attention module of the diffusion model to perform interactive attention calculation to obtain the attention scores of the local text description / global text description for each chunk; determining that the text description to which the chunk belongs is the local text description of the instance to which it belongs or is empty; inputting the multiple chunks of the training sample image, the text description to which each chunk belongs, the global text description, and the attention scores of the text description to which each chunk belongs for the chunk into the diffusion model feature extraction, denoising and splicing the chunk feature maps, and adjusting the diffusion model parameters. The present application proposes a text-to-image model that can precisely control multiple target instances, generating images with higher quality, richer content, and greater customization.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] The diffusion model is a deep learning model for generating images. It is based on the idea of probability and gradually generates images by continuously "diffusing" and "absorbing" information.

[0003] The training process of the diffusion model usually consists of two stages: the diffusion stage and the generation stage. In the diffusion stage, information is diffused by gradually adding noise, so that the sample image gradually becomes blurred and unclear. In the generation stage, the sample image is input into the model, and the model is used to gradually reduce the noise and gradually restore the details and clarity of the image.

[0004] After the diffusion model training is completed, the model can be used to generate images. The process of generating images is a process of starting from random noise and gradually recovering details.

[0005] Compared with traditional generative models, diffusion models have poor control performance at the instance level, which means that only the characteristics of the entire generated image can be controlled, but specific instances in the image cannot be precisely controlled. Summary of the invention

[0006] The purpose of this application is to provide a diffusion model-based Vincent map and a diffusion model training method, device and equipment to solve the problem that the existing technology can only control the features of the entire generated image but cannot accurately control specific instances in the image.

[0007] In a first aspect, an embodiment of the present application provides a diffusion model training method, the method comprising:

[0008] Obtaining a sample image, annotation boxes of multiple instances in the sample image, local text descriptions for describing each instance, and a global text description for describing the sample image;

[0009] The sample images are denoised through a diffusion process to obtain training sample images at different times;

[0010] Select a training sample image at the current moment, and divide the training sample image into a plurality of blocks;

[0011] Based on the local text description of each instance, the training sample image after the segmentation, and the global text description, the cross attention module of the diffusion model is used to perform interactive attention calculation to obtain the attention score of the local text description / global text description for each segmentation;

[0012] If the segmented block belongs to one of the instances, determining that the text description to which the segmented block belongs is a local text description of the instance to which it belongs, otherwise determining that the text description to which the segmented block belongs is empty;

[0013] Input multiple cutouts of the training sample image, the text description to which each cutout belongs, the global text description, and the attention score of the text description to which each cutout belongs to the cutout into a diffusion model to perform cutout feature extraction to obtain a cutout feature map, denoise the cutout feature map, and splice the denoised cutouts, and adjust the parameters of the diffusion model with the goal of outputting the training sample image at the previous moment.

[0014] In some possible embodiments, the cutout is determined to belong to one of the instances in the following manner:

[0015] According to the position of the cutout and the annotation box of each instance, calculate the intersection over union (IoU) of the cutout and the annotation box of each instance;

[0016] If the calculated maximum IoU is greater than a preset threshold, determine that the cutout belongs to the instance corresponding to the maximum IoU, otherwise determine that the instance does not belong to any instance.

[0017] In some possible embodiments, the input of the diffusion model further includes the position of the cutout and the degree of noise addition to the cutout.

[0018] In some possible embodiments, based on the local text description of each instance, the cutout training sample image, and the global text description, use the cross-attention module of the diffusion model to perform interactive attention calculation to obtain the attention score of the local text description / global text description to each cutout, including:

[0019] Input the local text encoding of any instance, the cutout training sample image, and the global text encoding into the cross-attention module of the diffusion model to perform interactive attention calculation to obtain the first probability matrix corresponding to the local text encoding of the instance, and the elements in the first probability matrix represent the attention score of the local text encoding of the instance to each cutout;

[0020] Input the cutout training sample image and the global text encoding into the cross-attention module of the diffusion model to perform interactive attention calculation to obtain the second probability matrix corresponding to the global text encoding, and the elements in the second probability matrix represent the attention score of the global text encoding to each cutout.

[0021] In some possible embodiments, input multiple cutouts of the training sample image, the text description of each cutout, the global text description, and the attention score of the text description to which each cutout belongs to the cutout into a diffusion model to perform cutout feature extraction to obtain a cutout feature map, including:

[0022] Input multiple cutouts of the training sample image, the text description of each cutout, the global text description, and the attention score of the text description to which each cutout belongs to the cutout into a diffusion model, and use the deep residual network module to perform convolution operations on the input;

[0023] Using the cross-attention layer, calculate the QKV matrix, the K g matrix, and the V g matrix as follows:

[0024]

[0025]

[0026]

[0027]

[0028]

[0029] Based on the calculated QKV matrix, the K g matrix, and the V g matrix, perform a convolution operation to extract a feature map;

[0030] where h is the output of the deep residual network module, and W Q is the parameter of matrix Q, W k is the parameter of matrix K, W V is the parameter of matrix V, MLP represents a fully connected layer, text n represents the text encoding corresponding to the text description to which the chunk belongs, and text g represents the global text encoding corresponding to the global text description.

[0031] In some possible embodiments, based on the calculated QKV matrix, the K g matrix, and the V g matrix, perform a convolution operation to extract a feature map, including calculating:

[0032]

[0033] where Conv represents a convolution operation, represents a hyperparameter, represents the attention score of the chunk in the probability matrix corresponding to the instance to which the chunk belongs, attention represents attention calculation, represents the attention score of the global text description corresponding to the probability matrix for the chunk.

[0034] In some possible embodiments, use the deep residual network module to perform a convolution operation on the input, including calculating:

[0035]

[0036] where Conv represents a convolution operation, MLP represents a fully connected layer, and et Indicates the degree of noise addition.

[0037] In some possible embodiments, denoising the sliced feature map and stitching the denoised slices includes:

[0038] Based on the feature maps of the obtained slices, use the decoder to predict the mean of the current noise distribution;

[0039] Use the slice and the mean of the current noise distribution to predict the denoised slice, and stitch the denoised slices.

[0040] In some possible embodiments, denoising the sliced feature map and stitching the denoised slices includes:

[0041] Input the obtained feature maps of the slices into the feature stitching module of the diffusion model;

[0042] Use the feature stitching module to respectively select local regions from the feature maps of every four adjacent slices to obtain the stitched feature map, and input the stitched feature map into the decoder;

[0043] Use the decoder to predict the mean of the current noise distribution of the stitched feature map based on the stitched feature map;

[0044] Use the predicted mean of the current noise distribution of the stitched feature map to obtain the denoised slices corresponding to the stitched feature map, and stitch the denoised slices corresponding to the stitched feature map.

[0045] In some possible embodiments, using the feature stitching module to respectively select local regions from the feature maps of every four adjacent slices to obtain the stitched feature map includes:

[0046] Use the feature stitching module to respectively select the adjacent local regions of the lower right 1 / 4, lower left 1 / 4, upper right 1 / 4, and upper left 1 / 4 in the feature maps of every four adjacent slices to obtain the stitched feature map.

[0047] In some possible embodiments, obtaining a sample picture, annotation boxes of multiple objects in the sample picture, local text descriptions for describing each object, and a global text description for describing the sample picture includes:

[0048] According to the input annotation information of the user, determine the annotation boxes of multiple instances in the sample picture, local text descriptions for describing each instance, and a global text description for describing the sample picture; and / or

[0049] Perform instance detection on the sample image using an open detection model to obtain the annotation boxes of each instance, and use an image-to-text model to generate local text descriptions for the instances within each annotation box and a global text description for the sample image.

[0050] In a second aspect, an embodiment of the present application provides an image generation method based on a diffusion model, including:

[0051] Obtain the instance annotation boxes and local text descriptions of each instance in the expected image, and encode the local text descriptions to obtain local text encodings;

[0052] Input the instance annotation boxes and the local text encodings of each instance into a diffusion model trained by the method provided in the first aspect above, and use the diffusion model to output the predicted expected image.

[0053] In a third aspect, another embodiment of the present application further provides a diffusion model training device, and the device includes:

[0054] An image information acquisition module, configured to acquire a sample image, annotation boxes of multiple instances in the sample image, local text descriptions for describing each instance, and a global text description for describing the sample image;

[0055] A training sample generation module, configured to add noise to the sample image through a diffusion process to obtain training sample images at different times;

[0056] An image cutting module, configured to select the training sample image at the current time and divide the training sample image into multiple cuts;

[0057] An attention score determination module, configured to perform interactive attention calculation using the cross-attention module of the diffusion model based on the local text descriptions of each instance, the cut training sample image, and the global text description, to obtain the attention scores of the local text description / global text description for each cut;

[0058] A belonging text description determination module, configured to determine that the text description to which the cut belongs is the local text description of the belonging instance if the cut belongs to one of the instances, otherwise determine that the text description to which the cut belongs is empty;

[0059] A model training module, configured to input the multiple cuts of the training sample image, the text description to which each cut belongs, the global text description, and the attention scores of the text description to which each cut belongs for the cut into the diffusion model, perform cut feature extraction to obtain a cut feature map, denoise the cut feature map and splice the denoised cuts, and adjust the parameters of the diffusion model with the output of the training sample image at the previous time as the target.

[0060] Fourthly, an embodiment of the present application further provides an image generation device based on a diffusion model, and the device includes:

[0061] An image information acquisition module, configured to acquire instance annotation frames and local text descriptions of each instance in an expected image, and encode the local text descriptions to obtain local text encodings;

[0062] An image prediction module, configured to input the instance annotation frames and the local text encodings of each instance into a diffusion model trained by the method provided in the first aspect above, and use the diffusion model to output a predicted expected image.

[0063] Fifthly, an embodiment of the present application provides a diffusion model training device, including at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method provided in the first aspect above of the present application.

[0064] Sixthly, an embodiment of the present application provides an image generation device based on a diffusion model, including at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method provided in the second aspect above.

[0065] Seventhly, another embodiment of the present application further provides a computer storage medium, and the computer storage medium stores a computer program, and the computer program is used to enable a computer to execute the method provided in the first aspect above of the embodiments of the present application or execute the method provided in the second aspect above.

[0066] The embodiment of the present application proposes a new image generation method based on a diffusion model, and designs an algorithm that cuts a training sample image into multiple cuts, and allows the diffusion model to generate in parallel, and finally composes a new image, which can keep the generated image at a high resolution. This solves the problem that the prior art can only control the features of the entire generated image, but cannot precisely control specific instances in the image; through the added interactive attention module, calculate the attention scores of the global text encoding and each target text encoding for each patch of the image, and interact with each interactive attention module in the diffusion model unet; when each patch is generated, consider the influence degree of the instance text encoding and the global text encoding on the patch, so that each generated instance conforms to its respective description while also conforming to the description of the global image, and at the same time, the background part also considers the influence of the existence of the instance.

[0067] Other features and advantages of the present application will be described in the subsequent specification, and in part will be obvious from the specification, or can be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written specification, claims, and drawings. Description of the Drawings

[0068] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. Obviously, the following introduced drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0069] Figure 1 Flowchart of the diffusion model training method according to an embodiment of the present application;

[0070] Figure 2 Schematic diagram of the diffusion process according to an embodiment of the present application;

[0071] Figure 3 Schematic diagram of text encoding according to an embodiment of the present application;

[0072] Figure 4 Schematic diagram of the UNet module in the diffusion model according to an embodiment of the present application;

[0073] Figure 5 Schematic diagram of the decoder structure according to an embodiment of the present application;

[0074] Figure 6 Schematic diagram of the chunk splicing process according to an embodiment of the present application;

[0075] Figure 7 Schematic diagram of the diffusion model training process according to an embodiment of the present application;

[0076] Figure 8 Schematic diagram of the diffusion model training device according to an embodiment of the present application;

[0077] Figure 9 Schematic diagram of the text-to-image generation device based on the diffusion model according to an embodiment of the present application;

[0078] Figure 10 Schematic diagram of the diffusion model training device / text-to-image generation device based on the diffusion model according to an embodiment of the present application. Detailed Embodiments

[0079] To further illustrate the technical solutions provided in the embodiments of the present application, the following will be described in detail in conjunction with the accompanying drawings and specific implementation manners. Although the embodiments of the present application provide method operation steps as shown in the following embodiments or drawings, more or fewer operation steps may be included in the method based on routine or non-creative labor. In steps where there is no necessary causal relationship logically, the execution order of these steps is not limited to the execution order provided in the embodiments of the present application. When the method is actually processed or executed by a control device, it can be executed in the method order shown in the embodiments or drawings or executed in parallel.

[0080] In the related art for the task of generating images based on text descriptions, by obtaining multiple sample images of the same target object and obtaining descriptors of the category to which the object belongs, diffusion training is performed on the text-to-image model. However, this method can only target a single target and cannot generate multiple controllable targets in one image at the same time. Therefore, the diffusion model may not meet the requirements at the instance level. At the same time, due to the limitation of the model input size, it is difficult for the diffusion model to train and generate images with higher resolutions.

[0081] The embodiments of the present application mainly propose a text-to-image model that can precisely control multiple target instances, while improving the resolution of the generated images, enabling the generated images to have higher quality, richer content, and be more customized.

[0082] As Figure 1 shown, the embodiments of the present application propose a diffusion model training method, and the method includes:

[0083] Step 101, obtain sample pictures, bounding boxes of multiple instances in the sample pictures, local text descriptions for describing each instance, and a global text description for describing the sample pictures;

[0084] The embodiments of the present application may further encode the local text description and the global text description to obtain local text encoding and global text encoding;

[0085] There are multiple above-mentioned sample pictures, and the multiple sample pictures form a sample picture set, and each sample picture includes multiple instances.

[0086] After obtaining the sample pictures, any one of the following methods or a combination of the two methods can be used to obtain the bounding boxes of multiple instances in the sample pictures, local text descriptions for describing each instance, and a global text description for describing the sample pictures:

[0087] Method 1, determine the bounding boxes of multiple instances in the sample pictures, local text descriptions for describing each instance, and a global text description for describing the sample pictures according to the input annotation information of the user;

[0088] This method uses manual annotation to label the bounding boxes (bboxes) of each target instance in the sample images. At the same time, it describes the attributes, names, and environments of these instances to form local text descriptions. Finally, a global text description is labeled for the entire sample image.

[0089] Method 2: Use an open detection model to perform instance detection on the sample image to obtain the bounding box of each instance. Use an image-to-text model to generate local text descriptions for the instances within each bounding box and a global text description for the sample image.

[0090] The above open detection model and image-to-text model are existing models and will not be described in detail here.

[0091] Step 102: Add noise to the sample image through a diffusion process to obtain training sample images at different times.

[0092] When training a diffusion model with image-text data, as Figure 2 shown, add noise to the sample image \(X_0\) step by step at \(T\) time steps until the sample image \(X_0\) becomes completely noisy. Each time during training, randomly select a time step \(t\), and obtain the training sample image \(X_t\) based on the time step \(t\) and the sample image \(X_0\). t . It should be noted that the training sample images at different times here are images with different noise levels, and the training sample images at different times are denoted as \(X_1\sim X_T\). T .

[0093] Obtaining the training sample image \(X_t\) at any time step \(t\) from the sample image \(X_0\) t is equivalent to sampling from an image distribution with a mean of and a variance of . Assuming there are \(T\) time steps, the representation of the training sample image at any time step \(t\) is as follows:

[0094]

[0095]

[0096] \(N\) represents the normal distribution;

[0097]

[0098]

[0099]

[0100] Among them, represents a hyperparameter, Represents a value between 0 and 1 that follows a normal distribution.

[0101] Through the above formula, the training sample image X at any time step t can be obtained. t , in the related art, during the training process, the diffusion model is used to predict the representation of the training sample image at the previous time step t - 1 as follows:

[0102] First, predict the mean of the distribution to which the noise at time t - 1 belongs :

[0103]

[0104] In the related art, text represents the text encoding of the sample image, represents the diffusion model with parameter .

[0105] Predicting X from X t is equivalent to sampling from a distribution with mean t-1 and variance and variance . The variance can be directly obtained according to t, and is specifically represented by the formula as follows:

[0106]

[0107] .

[0108] Step 103: Select the training sample image at the current moment and divide the training sample image into multiple chunks;

[0109] Step 104: Based on each instance's local text description, the chunked training sample image, and the global text description, use the cross-attention module of the diffusion model to perform interactive attention calculation to obtain the attention scores of the local text description / global text description for each chunk;

[0110] In the embodiment of the present application, an additional interactive attention module is used to calculate the attention scores of the global text encoding and the local text encoding of each instance for each patch of the image. The attention score represents the influence degree of the global text encoding and the local text encoding of each instance on each patch of the image.

[0111] Step 105: If the chunk belongs to one of the instances, determine the text description to which the chunk belongs as the local text description of the belonging instance; otherwise, determine the text description to which the chunk belongs as empty;

[0112] Step 106: Input multiple cut blocks of the training sample image, the text description to which each cut block belongs, the global text description, and the attention score of each cut block's text description for this cut block into the diffusion model to perform cut block feature extraction to obtain a cut block feature map, denoise the cut block feature map, and splice the denoised cut blocks, and adjust the diffusion model parameters with the goal of outputting the training sample image of the previous moment.

[0113] The embodiment of this application proposes a new text-to-image algorithm based on the diffusion model, designs an algorithm that divides the training sample image into multiple cut blocks (patches), and allows the diffusion model to generate in parallel, and finally composes a new image. It can realize instance-level image generation, and this application can keep the generated image at a high resolution. Through the added interactive attention module, calculate the attention scores of the global text encoding and each target text encoding for each patch of the image, and interact with each interactive attention module in the diffusion model's UNet; when each patch is generated, consider the influence degree of the instance's text encoding and the global text encoding on this patch, so that each generated instance conforms to its own description while also conforming to the description of the global image, and at the same time let the background part consider the influence of the instance's existence.

[0114] In some possible embodiments, based on each instance's local text description, the cut-up training sample image, and the global text description, use the cross-attention module of the diffusion model to perform interactive attention calculation to obtain the attention scores of the local text description / global text description for each cut block, including:

[0115] Input the local text encoding of any instance, the cut-up training sample image, and the global text encoding into the cross-attention module of the diffusion model to perform interactive attention calculation to obtain the first probability matrix corresponding to the local text encoding of this instance. The elements in the first probability matrix represent the attention scores of the local text encoding of this instance for each cut block;

[0116] Input the cut-up training sample image and the global text encoding into the cross-attention module of the diffusion model to perform interactive attention calculation to obtain the second probability matrix corresponding to the global text encoding. The elements in the second probability matrix represent the attention scores of the global text encoding for each cut block.

[0117] In some possible embodiments, the input of the diffusion model further includes the position of the cut block and the degree of noise addition to the cut block.

[0118] In the embodiment of this application, a learnable parameter global position embedding is initialized, which is represented as Each block will have this parameter, indicating the block position, and encoding the noise level time step t to obtain time embeddings, which is expressed in the subsequent formula as .

[0119] In the embodiment of the present application, the local text description and the global text description are encoded to obtain the local text encoding and the global text encoding, and the local text encoding and the global text encoding are used as the input of the diffusion model. Figure 3 As shown, the embodiment of the present application uses a clip text encoder to encode the local text description caption corresponding to the instance to obtain local text encoding Text embeddings, and encodes the global text description Global caption to obtain global text encoding Global Text embeddings, and the weights are fixed.

[0120] Assume there are m instances and corresponding m local text descriptions C, and a global text description C g . Then the expressions of local text encoding and global text encoding obtained by clipping text encoder are as follows:

[0121]

[0122] Where E represents the clip text encoder, c1~c m represents the local text description corresponding to m instances, c g Represents global text description, text1~text m Represents m local text encoding text embedding, text g Represents global text embeddings.

[0123] The embodiment of the present application cuts the training sample image into blocks, and inputs the obtained blocks, block positions, noise levels, text descriptions of each block, and global text encoding into the unet (Unity Networking) module in the diffusion model to extract feature maps of the blocks.

[0124] like Figure 4 As shown, the unet module in the embodiment of the present application mainly includes the following structure:

[0125] Convolution module Conv;

[0126] The cross-attention downsampling module cross attention downblock connected to Conv represents a combination of a deep residual network module resnet block, a spatial transformation module spatial transformer block, and a downsampling module Downblock, and has multiple layers;

[0127] The downsampling module Downblock connected to the cross-attention downsampling module cross attention downblock;

[0128] The cross-attention module middle sampling module cross attention midblock connected to the downsampling module Down block represents a combination of a deep residual network module resnet block, a spatial transformation module spatial transformer block, and a deep residual network module resnet block, and has multiple layers;

[0129] The upsampling module upblock connected to the cross-attention module middle sampling module cross attention midblock;

[0130] The cross-attention upsampling module cross attention upblock connected to the upsampling module up block represents a combination of a deep residual network module resnet block, a spatial transformation module spatial transformer block, and an upsampling module up block, and has multiple layers;

[0131] The GSC connected to the cross-attention upsampling module cross attention upblock represents a combination of normalization GroupNorm, activation function silu, and convolution conv.

[0132] Figure 4 The solid lines in it represent residual connections skip connection. Skip connection is an important component in the deep neural network architecture. Its basic idea is to directly connect the input to the output in some layers of the network to allow information to jump between different layers.

[0133] In the embodiments of the present application, there are two implementation manners for denoising the sliced feature map and splicing the denoised slices, where:

[0134] Method 1: Based on the feature maps of the obtained chunks, use the decoder to predict the mean of the current noise distribution; use the chunks and the mean of the current noise distribution to predict the denoised slices, and splice the denoised slices.

[0135] Method 2: Input the obtained feature maps of the chunks into the feature splicing module of the diffusion model; use the feature splicing module to respectively select local regions from the feature maps of every four adjacent chunks to obtain the spliced feature maps, and input the spliced feature maps into the decoder; use the decoder to predict the mean of the current noise distribution of the spliced feature maps based on the spliced feature maps; use the predicted mean of the current noise distribution of the spliced feature maps to obtain the denoised chunks corresponding to the spliced feature maps, and splice the denoised chunks corresponding to the spliced feature maps.

[0136] As Figure 5 shown is the schematic diagram of the decoder structure in the embodiment of the present application, which mainly includes:

[0137] Convolution module Conv;

[0138] Denote the combined module of self-attention module and downsampling self attention Downblock;

[0139] Upsampling module up block;

[0140] Deep residual network module resnet block;

[0141] GSC represents the combination of normalization GroupNorm, activation function silu and convolutional layer conv.

[0142] In order to make the image obtained after splicing multiple chunks patch smoother, the diffusion model in the embodiment of the present application uses the feature splicing feature collage module. For each chunk patch, the corresponding feature map is obtained through the unet module, and the feature splicing module respectively selects local regions from the feature maps of every 4 adjacent chunks to obtain the spliced feature map.

[0143] In some possible embodiments, use the feature splicing module to respectively select local regions from the feature maps of every four adjacent chunks to obtain the spliced feature map, as Figure 6 shown. Specifically, use the feature splicing module to respectively select the lower right 1 / 4, lower left 1 / 4, upper right 1 / 4, and upper left 1 / 4 and adjacent local regions in the feature maps of every four adjacent chunks to obtain the spliced feature map. Input the spliced feature map into the decoder, which can be specifically represented by the formula as follows:

[0144]

[0145] P1, P2, P3, and P4 represent taking the lower - right, lower - left, upper - right, and upper - left local 1 / 4 of the adjacent 4 patches respectively.

[0146] Taking the example that the sample image includes k instances, the following gives the specific process of training the diffusion model of the embodiment of the present application, as Figure 7 shown, mainly including:

[0147] Adding noise to the sample image X0 through the above - mentioned diffusion process to obtain the training sample image X t ;

[0148] Initialize a Unet0 module and a pre - trained decoder VAE decoder;

[0149] Cut the training sample image X t and after being processed by the convolutional layer, input it into the cross - attention block added in the embodiment of the present application;

[0150] For any instance local text encoding text embedding (text1, text2... text k ) and global text encoding global text embedding (text g ) among the k instances, calculate them respectively with the training sample image X t input into the cross - attention block corss attention block to calculate the probability matrix of the local text encoding of this instance. Input the global text encoding global text embedding (text g ) and the training sample image X t input into the cross - attention block corssattention block to calculate, obtain the probability matrix of the global text encoding, and finally obtain k + 1 probability matrices (Probability matrixs) , and the elements in the probability matrix represent the attention scores of the local text encoding / global text encoding of each instance to each patch. It can also be interpreted as the influence degree of each text embedding (text1, text2... text k ) and global text embedding (text g ) on each patch, specifically expressed as follows:

[0151]

[0152] Initialize a learnable parameter , and the obtained by encoding for each time step t.

[0153] Set the total number of time steps to T. Each time during training, randomly select a time step t from {0, T}, and randomly select a sample image X0 from the set of sample images. The training sample image X is obtained in the following way t :

[0154] .

[0155] Before inputting the training sample image X t into the unet module of the diffusion model, further perform image padding on the training sample image to make the size of the image meet the requirements. Cut the padded training sample image into m chunks, and p represents each chunk, which is specifically expressed as follows:

[0156]

[0157] Among them, PAD represents padding, and PATCH represents chunking.

[0158] Input each obtained chunk patch, chunk position , noise addition degree time embeddings , the local text encoding or empty belonging to the instance in the text encoding (text1, text2... text k ), the global text encoding (text g ), and the probability matrix M into the unet module in the diffusion model; it should be noted that if a chunk belongs to one of the instances, determine the text description to which the chunk belongs as the local text encoding of the instance to which it belongs, otherwise determine the text description to which the chunk belongs as empty, then the local text encoding to which the chunk belongs is empty.

[0159] In the embodiment of the present application, the following method is used to determine whether the chunk belongs to one of the instances:

[0160] According to the position of the chunk and the annotation box of each instance, calculate the intersection over union of the chunk and the annotation box of each instance;

[0161] If the calculated maximum intersection over union is greater than the preset threshold, determine that the chunk belongs to the instance corresponding to the maximum intersection over union, otherwise determine that the instance does not belong to any instance.

[0162] During specific calculations, for each patch, calculate the intersection over union (IoU) between the patch and the bounding box (bbox) position of each instance, and find the bbox of the k-th instance with the largest IoU. If the IoU is greater than or equal to 0.5, the text encoding to which the n-th patch belongs is the local text encoding (text embedding) of the k-th instance, which is text k , otherwise it is empty, which can be specifically expressed by the following formula:

[0163]

[0164]

[0165]

[0166] represents the value of the probability matrix of the k-th instance corresponding to the n-th patch.

[0167] As mentioned above, the UNet module includes at least one cross-attention module (cross attention block), specifically the cross attention downblock, cross attention midblock, and cross attention upblock shown above.

[0168] In the embodiments of the present application, the above cross-attention module includes a deep residual network module (resetblock) and a cross-attention layer (cross attention). The UNet module extracts the feature map of this patch, including:

[0169] 1) Perform a convolution operation on the input using the deep residual network module;

[0170] In some possible embodiments, input the multiple patches of the training sample image, the local text encoding of each patch, the global text encoding, and the attention score of the text description to which each patch belongs to the diffusion model. Performing a convolution operation on the input using the deep residual network module includes calculating:

[0171]

[0172] where Conv represents the convolution operation, MLP represents the fully connected layer, and e t represents the degree of noise addition

[0173] 2) Use the cross-attention layer to calculate the QKV matrix and the K g matrix, and the V g matrix in the following manner:

[0174]

[0175]

[0176]

[0177]

[0178]

[0179] Among them, h is the output of the depth residual network module, and W Q is the parameter of matrix Q, and W k is the parameter of matrix K, and W V is the parameter of matrix V. MLP represents a fully connected layer, and text n represents the text encoding corresponding to the text description to which the chunk belongs, and text g represents the global text encoding corresponding to the global text description.

[0180] 3) Based on the calculated QKV matrices, K g matrices, and V g matrices, perform convolution operations to extract feature maps;

[0181] In some possible embodiments, based on the calculated QKV matrices, K g matrices, and V g matrices, perform convolution operations to extract feature maps, including calculating:

[0182]

[0183] Among them, Conv represents convolution operation, represents a hyperparameter, represents the attention score of the chunk in the probability matrix corresponding to the instance to which the chunk belongs. If the chunk does not belong to any instance, then is empty; attention represents attention calculation, represents the attention score of the global text description's probability matrix for the chunk.

[0184] The unet model outputs the latent feature h of the feature map of each patch, which is input into the feature collage module, and the following formula is used for the splicing of the feature maps:

[0185]

[0186] Finally, the spliced feature map is given to the vae decoder, and the mean value of the noise distribution at time t is output after decoding by the vae decoder , using to obtain the denoised cut image at the previous moment, which is specifically represented by the following formula:

[0187]

[0188]

[0189] where D represents the decoding operation,

[0190] Stitch all the cut images to obtain the entire predicted denoised sample image.

[0191] If generating an image from complete noise, repeat the prediction for T time steps, and finally obtain X0 from X T During each prediction process, adjust the diffusion model parameters according to the following loss function L SD Specifically, adjust the parameters of the unet module and the decoder:

[0192]

[0193] where is the mean of the noise distribution at time t determined according to the actual training sample image at the previous moment.

[0194] The embodiment of this application proposes a new text-to-image algorithm based on the diffusion model, and designs an algorithm that cuts the original image into multiple patches, and then generates them in parallel using the diffusion model, and finally combines them into a new image. This can keep the generated image at a high resolution.

[0195] The embodiment of this application filters out the patches included in each instance, and performs cross-attention calculation with the text encoding corresponding to the description, so that the positions and descriptions of multiple instances in the image can be specified during generation, and a conforming image can be generated.

[0196] The embodiment of this application uses an additional cross-attention module to calculate the attention scores of the global text encoding and each target text encoding for each patch of the image, and interacts with each cross-attention module in the diffusion model unet. When each patch is generated, it considers the influence of the instance text encoding and the global text encoding on this patch. While making each generated instance conform to its respective description, it also conforms to the description of the global image, and at the same time makes the background part consider the influence of the existence of the instance.

[0197] Based on the same inventive concept, this application also provides a diffusion model training device, as Figure 8 shown, the device includes:

[0198] The image information acquisition module 801 is used to acquire sample images, annotation boxes of multiple instances in the sample images, local text descriptions for describing each instance, and global text descriptions for describing the sample images;

[0199] The training sample generation module 802 is used to add noise to the sample images through a diffusion process to obtain training sample images at different times;

[0200] The image cutting module 803 is used to select the training sample image at the current time and divide the training sample image into multiple cuts;

[0201] The attention score determination module 804 is used to perform interactive attention calculation using the cross-attention module of the diffusion model based on the local text descriptions of each instance, the cut training sample image, and the global text description, to obtain the attention scores of the local text description / global text description for each cut;

[0202] The belonging text description determination module 805 is used to determine that the text description to which the cut belongs is the local text description of the belonging instance if the cut belongs to one of the instances, otherwise determine that the text description to which the cut belongs is empty;

[0203] The model training module 806 is used to input the multiple cuts of the training sample image, the text description to which each cut belongs, the global text description, and the attention scores of the text description to which each cut belongs for the cut into the diffusion model, perform cut feature extraction to obtain a cut feature map, denoise the cut feature map and splice the denoised cuts, and adjust the parameters of the diffusion model with the output of the training sample image at the previous time as the target.

[0204] In some possible embodiments, the belonging text description determination module determines whether the cut belongs to one of the instances in the following manner:

[0205] According to the position of the cut and the annotation box of each instance, calculate the intersection over union of the cut and the annotation box of each instance;

[0206] If the calculated maximum intersection over union is greater than a preset threshold, determine that the cut belongs to the instance corresponding to the maximum intersection over union, otherwise determine that the cut does not belong to any instance.

[0207] In some possible embodiments, the input of the diffusion model further includes the position of the cut and the degree of noise added to the cut.

[0208] In some possible embodiments, the attention score determination module performs interactive attention calculation using the cross-attention module of the diffusion model based on the local text descriptions of each instance, the cut training sample image, and the global text description, to obtain the attention scores of the local text description / global text description for each cut, including:

[0209] Input the local text encoding of any instance, the segmented training sample image, and the global text encoding into the cross-attention module of the diffusion model for interactive attention calculation to obtain the first probability matrix corresponding to the local text encoding of this instance. The elements in the first probability matrix represent the attention scores of the local text encoding of this instance for each segment.

[0210] Input the segmented training sample image and the global text encoding into the cross-attention module of the diffusion model for interactive attention calculation to obtain the second probability matrix corresponding to this global text encoding. The elements in the second probability matrix represent the attention scores of the global text encoding for each segment.

[0211] In some possible embodiments, the model training module inputs multiple segments of the training sample image, the text descriptions of each segment, the global text description, and the attention scores of the text description to which each segment belongs for this segment into the diffusion model to perform segment feature extraction to obtain a segment feature map, including:

[0212] Input the multiple segments of the training sample image, the text descriptions of each segment, the global text description, and the attention scores of the text description to which each segment belongs for this segment into the diffusion model, and use the deep residual network module to perform convolution operations on the input.

[0213] Use the cross-attention layer to calculate the QKV matrix and the K g matrix, V g matrix as follows:

[0214]

[0215]

[0216]

[0217]

[0218]

[0219] Based on the calculated QKV matrix, K g matrix, V g matrix, perform convolution operations to extract the feature map.

[0220] Among them, h is the output of the deep residual network module, W Q is the parameter of matrix Q, W k is the parameter of matrix K, W V is the parameter of matrix V, MLP represents the fully connected layer, text nIndicates the text encoding corresponding to the text description of the slice, text g Indicates the global text encoding corresponding to the global text description.

[0221] In some possible embodiments, the model training module is based on the calculated QKV matrix, K g Matrix, V g The matrix performs convolution operations to extract feature maps, including calculations:

[0222]

[0223] Among them, Conv represents the convolution operation, represents the hyperparameter, Indicates the attention score of the block in the probability matrix corresponding to the instance to which the block belongs. Attention indicates the attention calculation. Represents the attention score of the probability matrix corresponding to the global text description on the slice.

[0224] In some possible embodiments, the model training module uses the deep residual network module to perform convolution operations on the input, including calculating:

[0225]

[0226] Where Conv represents convolution operation, MLP represents fully connected layer, and e t Indicates the degree of noise addition.

[0227] In some possible embodiments, the model training module denoises the slice feature map and splices the denoised slices, including:

[0228] Based on the obtained feature maps of each block, the decoder is used to predict the mean of the current noise distribution;

[0229] The denoised slices are predicted using the cut blocks and the mean of the current noise distribution, and the denoised slices are spliced.

[0230] In some possible embodiments, the model training module denoises the slice feature map and splices the denoised slices, including:

[0231] The obtained feature maps of each cut block are input into the feature splicing module of the diffusion model;

[0232] Using the feature splicing module, local regions are selected from the feature maps of every four adjacent slices to obtain spliced feature maps, and the spliced feature maps are input into the decoder;

[0233] Predicting a mean value of a current noise distribution of the spliced feature map using the decoder based on the spliced feature map;

[0234] The predicted mean value of the current noise distribution of the spliced feature map is used to obtain the denoised blocks corresponding to the spliced feature map, and the denoised blocks corresponding to the spliced feature map are spliced.

[0235] In some possible embodiments, the model training module uses the feature splicing module to select local areas of feature maps of every four adjacent slices to obtain spliced feature maps, including:

[0236] The feature splicing module is used to select the adjacent local areas located at the lower right 1 / 4, the lower left 1 / 4, the upper right 1 / 4, and the upper left 1 / 4 of the feature maps of the four adjacent blocks respectively for the feature maps of every four adjacent blocks to obtain the spliced feature map.

[0237] In some possible embodiments, the image information acquisition module acquires a sample image, annotation boxes of multiple objects in the sample image, a local text description for describing each object, and a global text description for describing the sample image, including:

[0238] Determine, based on the annotation information input by the user, annotation boxes for multiple instances in the sample image, local text descriptions for describing each instance, and a global text description for describing the sample image; and / or

[0239] An open detection model is used to perform instance detection on the sample image to obtain a labeling box for each instance, and a graph-to-text model is used to generate a local text description for each instance in the labeling box, and a global text description is generated for the sample image.

[0240] Based on the same inventive concept, the present application also provides a Wensheng diagram device based on a diffusion model, such as Figure 9 As shown, the device comprises:

[0241] The picture information acquisition module 901 is used to acquire the instance annotation box and the local text description of each instance in the expected picture, and encode the local text description to obtain the local text code;

[0242] The image prediction module 902 is used to input the instance annotation box and the local text encoding of each instance into the diffusion model trained based on the method provided by the above embodiment, and use the diffusion model to output the predicted expected image.

[0243] After introducing the diffusion model training method, diffusion model-based Vincent diagram method and apparatus according to an exemplary embodiment of the present application, next, a diffusion model training device and a diffusion model-based Vincent diagram device according to another exemplary embodiment of the present application are introduced.

[0244] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to herein as "circuits", "modules", or "systems".

[0245] In some possible implementation manners, the diffusion model training device / based on the diffusion model text-to-image device according to the present application may at least include at least one processor and at least one memory. Among them, the memory stores program code, and when the program code is executed by the processor, the processor executes the steps in the diffusion model training method / based on the diffusion model text-to-image method according to various exemplary implementation manners of the present application described above in this specification.

[0246] The following refers to Figure 10 to describe the diffusion model training device / based on the diffusion model text-to-image device 100 according to this implementation manner of the present application. Figure 10 The shown diffusion model training device / based on the diffusion model text-to-image device 100 is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present application.

[0247] As Figure 10 shown, the diffusion model training device / based on the diffusion model text-to-image device 100 is presented in the form of a general-purpose electronic device. The components of the diffusion model training device / based on the diffusion model text-to-image device 100 may include but are not limited to: the above-mentioned at least one processor 101, the above-mentioned at least one memory 102, and a bus 103 connecting different system components (including the memory 102 and the processor 101).

[0248] The bus 103 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a processor, or a local bus using any bus structure in a variety of bus structures.

[0249] The memory 102 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 1021 and / or a cache memory 1022, and may further include a read-only memory (ROM) 1023.

[0250] The memory 102 may further include a program / utilities 1025 having a set (at least one) of program modules 1024. Such program modules 1024 include but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0251] The diffusion model training device / text-to-image generation device 100 based on the diffusion model can also communicate with one or more external devices 104 (such as keyboards, pointing devices, etc.), and can also communicate with one or more devices that enable users to interact with the diffusion model training device / text-to-image generation device 100 based on the diffusion model, and / or communicate with any device (such as routers, modems, etc.) that enables the diffusion model training device / text-to-image generation device 100 based on the diffusion model to communicate with one or more other electronic devices. Such communication can be carried out through the input / output (I / O) interface 105. Moreover, the diffusion model training device / text-to-image generation device 100 based on the diffusion model can also communicate with one or more networks (such as local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) through the network adapter 106. As shown in the figure, the network adapter 106 communicates with other modules for the diffusion model training device / text-to-image generation device 100 through the bus 103. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the diffusion model training device / text-to-image generation device 100 based on the diffusion model, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0252] In some possible implementation manners, various aspects of the diffusion model training method / text-to-image generation method based on the diffusion model provided in this application can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps in the diffusion model training method / text-to-image generation method according to various exemplary embodiments of this application described above in this specification.

[0253] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0254] The program product for monitoring according to an embodiment of the present application may be a portable compact disc read-only memory (CD-ROM) and include program code, and may run on an electronic device. However, the program product of the present application is not limited thereto. In this document, a readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0255] The readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0256] The program code contained on the readable medium may be transmitted by any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.

[0257] The program code for performing the operations of the present application may be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's electronic device, partially on the user's device, executed as an independent software package, partially on the user's electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In the case of a remote electronic device, the remote electronic device may be connected to the user's electronic device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external electronic device (e.g., connected through the Internet using an Internet service provider).

[0258] It should be noted that although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more units described above may be embodied in one unit. Conversely, the features and functions of one unit described above may be further divided and embodied by multiple units.

[0259] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.

Claims

1. A diffusion model training method, characterized in that, The method includes: Obtaining a sample image, annotation boxes of multiple instances in the sample image, local text descriptions for describing each instance, and a global text description for describing the sample image; Adding noise to the sample image through a diffusion process to obtain training sample images at different times; Selecting the training sample image at the current time and dividing the training sample image into multiple chunks; Based on the local text descriptions of each instance, the chunked training sample image, and the global text description, using the cross-attention module of the diffusion model to perform interactive attention calculation to obtain the attention scores of the local text description / global text description for each chunk; If the chunk belongs to one of the instances, determining the text description to which the chunk belongs as the local text description of the belonging instance, otherwise determining the text description to which the chunk belongs as empty; Inputting the multiple chunks of the training sample image, the text description of each chunk, the global text description, and the attention scores of the text description to which each chunk belongs for the chunk into the diffusion model, and using the deep residual network module to perform convolution operations on the input; The QKV matrix and the K g matrix and the V g matrix are calculated in the following manner using a cross-attention layer: ; Based on the calculated QKV matrix, K g matrix, V g matrix, perform convolution operations to extract feature maps. During the convolution operation process, utilize the attention scores of each chunk belonging to the text description for that chunk and the attention scores of the global text description for each chunk to interact with each cross-attention module in the diffusion model; Among them, h is the output of the depth residual network module, and W Q is the parameter of matrix Q, and W k is the parameter of matrix K, and W V is the parameter of matrix V. MLP represents a fully connected layer, and text n represents the text encoding corresponding to the text description to which the chunk belongs, and text g represents the global text encoding corresponding to the global text description; Denosing the chunk feature map and splicing the denoised chunks, and adjusting the parameters of the diffusion model with the goal of outputting the training sample image at the previous time.

2. The method according to claim 1, wherein Determine whether the chunk belongs to one of the instances in the following way: According to the position of the chunk and the annotation box of each instance, calculate the intersection over union (IoU) of the chunk and the annotation box of each instance; If the calculated maximum IoU is greater than a preset threshold, determine that the chunk belongs to the instance corresponding to the maximum IoU, otherwise determine that the chunk does not belong to any instance.

3. The method according to claim 1, wherein The input of the diffusion model further includes the position of the chunk and the degree of noise added to the chunk.

4. The method according to claim 1, characterized in that, The performing interactive attention calculation based on the local text descriptions of each instance, the chunked training sample image, and the global text description, using the cross-attention module of the diffusion model to obtain the attention scores of the local text description / global text description for each chunk includes: Inputting the local text encoding of any instance, the chunked training sample image, and the global text encoding into the cross-attention module of the diffusion model to perform interactive attention calculation to obtain a first probability matrix corresponding to the local text encoding of the instance, where the elements in the first probability matrix represent the attention scores of the local text encoding of the instance for each chunk; Inputting the chunked training sample image and the global text encoding into the cross-attention module of the diffusion model to perform interactive attention calculation to obtain a second probability matrix corresponding to the global text encoding, where the elements in the second probability matrix represent the attention scores of the global text encoding for each chunk.

5. The method according to claim 1, characterized in that, Based on the calculated QKV matrix, K g matrix, V g matrix, perform a convolution operation to extract a feature map, including calculating: ; Among them, Conv represents the convolution operation, represents a hyperparameter, represents the attention score of the patch in the probability matrix corresponding to the instance to which the patch belongs, and attention represents the attention calculation, represents the attention score of the probability matrix corresponding to the global text description for the patch.

6. The method according to claim 1, characterized in that Using the deep residual network module to perform convolution operations on the input includes calculating: ; where Conv represents the convolution operation, MLP represents the fully connected layer, e t represents the degree of noise addition, and x is the input of the deep residual network module.

7. The method according to claim 1, characterized in that, Denosing the chunk feature map and splicing the denoised chunks includes: Based on the obtained feature maps of each chunk, using a decoder to predict the mean of the current noise distribution; Predicting the denoised slice using the chunk and the mean of the current noise distribution, and splicing the denoised slices.

8. The method according to claim 1, wherein Denosing the chunk feature map and splicing the denoised chunks includes: Inputting the obtained feature maps of each chunk into the feature splicing module of the diffusion model; Using the feature splicing module, for the feature maps of every four adjacent cut blocks, local regions are respectively selected to obtain the spliced feature map, and the spliced feature map is input into the decoder; Using the decoder, based on the spliced feature map, predict the mean value of the current noise distribution of the spliced feature map; Using the predicted mean value of the current noise distribution of the spliced feature map, obtain the denoised cut blocks corresponding to the spliced feature map, and splice the denoised cut blocks corresponding to the spliced feature map.

9. The method according to claim 8, characterized in that Using the feature splicing module to respectively select local regions from the feature maps of every four adjacent cut blocks to obtain the spliced feature map, including: Using the feature splicing module to respectively select the lower-right 1 / 4, lower-left 1 / 4, upper-right 1 / 4, and upper-left 1 / 4 and adjacent local regions of the feature maps of four adjacent cut blocks to obtain the spliced feature map.

10. The method according to claim 1, wherein Obtain a sample picture, annotation boxes of multiple objects in the sample picture, local text descriptions for describing each object, and a global text description for describing the sample picture, including: According to the input annotation information of the user, determine the annotation boxes of multiple instances in the sample picture, local text descriptions for describing each instance, and a global text description for describing the sample picture; and / or Use an open detection model to perform instance detection on the sample picture to obtain the annotation box of each instance, use a text generation model from image to generate local text descriptions for each instance within the annotation box, and generate a global text description for the sample picture.

11. A text-to-image generation method based on a diffusion model, characterized in that, Including: Obtain the instance annotation boxes and local text descriptions of each instance in the expected picture, and encode the local text descriptions to obtain local text encodings; Input the instance annotation boxes and the local text encodings of each instance into the diffusion model trained by any one of the methods in claims 1 to 10, and use the diffusion model to output the predicted expected picture.

12. A diffusion model training device, characterized in that, The device includes: A picture information acquisition module, configured to acquire a sample picture, annotation boxes of multiple instances in the sample picture, local text descriptions for describing each instance, and a global text description for describing the sample picture; A training sample generation module, configured to add noise to the sample picture through a diffusion process to obtain training sample pictures at different times; A picture cutting module, configured to select the training sample picture at the current time and divide the training sample picture into multiple cut blocks; An attention score determination module, configured to perform interactive attention calculation using the cross-attention module of the diffusion model based on the local text descriptions of each instance, the cut training sample picture, and the global text description, to obtain the attention scores of the local text description / global text description for each cut block; A belonging text description determination module, configured to, if the cut block belongs to one of the instances, determine that the text description to which the cut block belongs is the local text description of the belonging instance, otherwise determine that the text description to which the cut block belongs is empty; A model training module, configured to input the multiple cut blocks of the training sample picture, the text description of each cut block, the global text description, and the attention scores of the text description to which each cut block belongs for the cut block into the diffusion model, and use the deep residual network module to perform convolution operations on the input; The QKV matrix and the K g matrix and the V g matrix are calculated in the following manner using a cross-attention layer: ; Based on the calculated QKV matrix, K g Matrix, V g The matrix is convolved to extract feature maps. During the convolution operation, the attention score of each block to the text description and the attention score of each block to the global text description are used to interact with each interactive attention module in the diffusion model. where h is the output of the deep residual network module, and W Q is the parameter of matrix Q, and W k is the parameter of matrix K, and W V is the parameter of matrix V. MLP represents the fully connected layer, and text n represents the text encoding corresponding to the text description to which the chunk belongs, and text g represents the global text encoding corresponding to the global text description; Denoise the sliced feature map and splice the denoised slices, and adjust the diffusion model parameters with the goal of outputting the training sample image of the previous moment.

13. An image generation device based on a diffusion model, characterized in that, The device includes: An image information acquisition module, configured to acquire instance annotation frames and local text descriptions of each instance in the expected image, and encode the local text descriptions to obtain local text encodings; An image prediction module, configured to input the instance annotation frames and the local text encodings of each instance into a diffusion model trained by any one of the methods recited in claims 1 to 10, and use the diffusion model to output a predicted expected image.

14. A diffusion model training device, characterized in that, Comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-10.

15. A text-to-image device based on a diffusion model, characterized in that, Comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to claim 11.

16. A computer storage medium, characterized in that, The computer storage medium stores a computer program, and the computer program is used to cause a computer to execute the method according to any one of claims 1-10, or execute the method according to claim 11.

Citation Information

Patent Citations

  • Fine-tuning and control of diffusion models

    DE102023127111A1

  • KR20220050758A