Training method, generation method and device for visual content generation model of multiple objects

By introducing position coding and cross attention modules into the diffusion model, the problem of matching visual features and text in multi-character literary pictures is solved, and the consistency and control accuracy of visual content in multi-object generation scenes are achieved, and the generated visual content is accurately matched with text description.

CN119784876BActive Publication Date: 2025-07-25BEIJING SHENGSHU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411992015.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-07-25
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

In the literary and artistic task where multiple characters exist, the existing technology lacks effective solutions to how to accurately match the visual characteristics of each character with text prompts.

Method used

By introducing position encoding module and cross attention module into the pretrained diffusion model, local visual and text feature representations of multiple sets of training data pairs are obtained, image position encoding and text position encoding are embedded, feature fusion and denoising are used for cross attention modules, network parameters are iteratively adjusted, and high-quality visual content is generated.

Benefits of technology

The consistency and control accuracy of visual content in the multi-object generation scene is realized, ensuring that the generated visual content of the target object is consistent with the original image attribute information and accurately matches the text description.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119784876B_ABST
    Figure CN119784876B_ABST
Patent Text Reader

Abstract

The present application relates to a training method, a generation method, and a device for a visual content generation model for multiple objects. The training method includes: obtaining, based on training data pairs, a first local visual feature representation and a local text feature representation corresponding to each target object; respectively embedding an image position encoding and a text position encoding in each target object through a position encoding module to obtain a second local visual feature representation and a second local text feature representation; inputting the second local visual feature representations and the second local text feature representations of the respective target objects into a cross-attention module and a diffusion model, so that the diffusion model performs denoising according to the fused features output by the cross-attention module; fixing the network parameters of each layer of the diffusion model, and iteratively adjusting the network parameters of the cross-attention module and the position encoding module to obtain a trained visual content generation model. The solution provided by the present application can ensure the consistency and control accuracy of visual content in a multi-object generation scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a training method, a generation method, and a device for a visual content generation model of multiple objects. Background Art

[0002] With the rapid development of artificial intelligence technology, AIGC (Artificial Intelligence Generated Content) has been widely used in various fields. Among them, diffusion models are increasingly widely used in the field of content generation, and their powerful generation capabilities have had a profound impact on industries such as design, art, and multimedia. Especially in the field of character generation, through text-based prompts, users can generate specific character images according to their imagination. However, character generation not only depends on language description. Especially when there are multiple characters, it is difficult to achieve precise visual feature control only by language.

[0003] Currently, there is still a lack of an effective solution for accurately matching the visual features of each character with text prompts in the text-to-image task with multiple characters. Summary of the Invention

[0004] To solve or partially solve the problems existing in the related technologies, this application provides a training method, a generation method, and a device for a visual content generation model of multiple objects, which can ensure the consistency and control accuracy of visual content in a multiple-object generation scenario for a trained visual content generation model.

[0005] In a first aspect of this application, a training method for a visual content generation model of multiple objects is provided. The visual content generation model includes a position encoding module to be trained, a cross-attention module to be trained, and a pre-trained diffusion model. Among them:

[0006] Obtain multiple groups of training data pairs. Each group of the training data pairs includes a training image and a text description. The training image includes at least two target objects, and the text description contains attribute descriptions of each target object.

[0007] Based on the training data pairs, obtain a first local visual feature representation and a first local text feature representation corresponding to each target object.

[0008] Embed an image position encoding and a text position encoding into the first local visual feature representation and the first local text feature representation corresponding to each target object through the position encoding module, respectively, to obtain a second local visual feature representation and a second local text feature representation.

[0009] Input the second local visual feature representation and the second local text feature representation of each of the target objects into the cross-attention module and the diffusion model, so that the diffusion model denoises according to the fused features output by the cross-attention module and outputs the predicted visual content;

[0010] Fix the network parameters of each layer of the diffusion model, and backpropagate according to the loss function of the diffusion model, and iteratively adjust the network parameters of the cross-attention module and the position encoding module to obtain a trained visual content generation model.

[0011] In some embodiments, obtaining the first local visual feature representation and the first local text feature representation corresponding to each of the target objects based on the training data pair includes:

[0012] Detect and segment each training image, obtain the local image corresponding to each target object in the training image, and generate the corresponding first local visual feature representation;

[0013] Perform semantic recognition on the text description, obtain the local text description corresponding to each target object, and generate the corresponding first local text feature representation.

[0014] In some embodiments, the position encoding module includes an image position encoding network and a text position encoding network;

[0015] Embedding the image position encoding and the text position encoding in the first local visual feature representation and the first local text feature representation corresponding to each of the target objects through the position encoding module includes:

[0016] For each of the target objects, embed the image position encoding in the corresponding first local visual feature representation through the image position encoding network, and embed the text position encoding in the corresponding first local text feature representation through the text position encoding network.

[0017] In some embodiments, the cross-attention module includes a first attention network and a second attention network;

[0018] The first attention network is used to output the corresponding first fused feature according to the second local text feature representation and the sampled features output by the diffusion model;

[0019] The second attention network is used to output the second fused feature according to the second local visual feature representation and the first fused feature, and the second fused feature is used to input the diffusion model for feature sampling.

[0020] In some embodiments, the diffusion model includes N layers of feature sampling layers;

[0021] The diffusion model denoises based on the fused features output by the cross-attention module and outputs predicted visual content, including:

[0022] Performing cross-attention calculation on the second local text feature representation and the output features of the k-th layer of the diffusion model through a first attention network to obtain the k-th first fused feature; where 1 ≤ k ≤ N, and both k and N are natural numbers;

[0023] Performing cross-attention calculation on the first fused feature and the second local visual feature representation through a second attention network to obtain the k-th second fused feature; where the k-th second fused feature is used as an input to the (k + 1)-th layer of the diffusion model to obtain the output features of the (k + 1)-th layer; when k + 1 = N, the diffusion model obtains the output features at the current time step and generates predicted visual content based on the output features at the last time step.

[0024] In some embodiments, the loss function includes at least one of the overall object feature similarity, the key object feature similarity, and the text and image matching degree.

[0025] In some embodiments, each frame of the training image includes at least one target object; the target object is a person, an animal, a plant, or an item; the attribute description includes at least one of appearance features, facial features, clothing features, color features, and orientation features.

[0026] The second aspect of the present application provides a method for generating visual content of multiple objects, which includes:

[0027] Obtaining P frames of original images and prompt texts, where the total number of target objects in each frame of the original images is Y, Y ≥ P, and P ≥ 1;

[0028] Generating corresponding target visual content using the visual content generation model obtained according to the training method of any one of the embodiments in the first aspect above, where the attribute information of the target object in the target visual content is consistent with that in the original image.

[0029] In some embodiments, when P is greater than 1, each frame of the original image contains at least one target object; or

[0030] When P is equal to 1 and the original image contains multiple target objects, segmenting the original image to obtain local images corresponding to each target object; or

[0031] When P is equal to 1 and the original image contains 1 target object, copying to generate multiple of the target objects.

[0032] The second aspect of the present application provides a training device for a visual content generation model of multiple objects, where:

[0033] A training data acquisition module, configured to acquire multiple sets of training data pairs, each set of the training data pairs including a training image and a text description, the training image including at least two target objects, and the text description containing an attribute description of each target object;

[0034] A local feature extraction module, configured to obtain a first local visual feature representation and a first local text feature representation corresponding to each of the target objects based on the training data pairs;

[0035] A position encoding embedding module, configured to respectively embed an image position encoding and a text position encoding in the first local visual feature representation and the first local text feature representation corresponding to each of the target objects through the position encoding module, to obtain a second local visual feature representation and a second local text feature representation;

[0036] A cross-denoising module, configured to input the second local visual feature representations and the second local text feature representations of the target objects into the cross-attention module and the diffusion model, so that the diffusion model performs denoising according to the fused features output by the cross-attention module and outputs predicted visual content;

[0037] A network parameter iteration module, configured to fix the network parameters of each layer of the diffusion model, and perform backpropagation according to the loss function of the diffusion model, and iteratively adjust the network parameters of the cross-attention module and the position encoding module to obtain a trained visual content generation model.

[0038] A fifth aspect of the present application provides a visual content generation device for multiple objects, including:

[0039] A data acquisition module, configured to acquire P-frame original images and a prompt text, the total number of target objects in each frame of the original images being Y, Y≥P, P≥1;

[0040] A visual generation module, configured to generate corresponding target visual content according to the visual content generation model obtained by the training method described in the first aspect, the attribute information of the target objects in the target visual content being consistent with that in the original images.

[0041] A sixth aspect of the present application provides an electronic device, including:

[0042] A processor; and

[0043] A memory, on which executable code is stored, and when the executable code is executed by the processor, the processor is caused to execute the training method of the visual content generation model for multiple objects described in the first aspect or the visual content generation method for multiple objects described in the second aspect.

[0044] The sixth aspect of the present application provides a computer-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor is caused to execute the training method of the multi-object visual content generation model as described in the first aspect or the multi-object visual content generation method as described in the second aspect.

[0045] The technical solutions provided by the present application may include the following beneficial effects:

[0046] The training method of the multi-object visual content generation model of the present application can, based on the image position encoding network and the text position encoding network, achieve one-to-one binding and pairing of the visual features and text features of multiple target objects. Then, through the cross-attention mechanism of the first attention network and the second attention network, the visual features and text features of a single target object are cross-fused, and the natural interaction of multiple target objects in the same scene is realized. While maintaining the relative independence of the respective visual features of the target objects and the consistency of the pairing with the text features, it ensures the high-quality generation of visual content.

[0047] The multi-object visual content generation method of the present application can, based on the original image and the prompt text input by the user, through the trained visual content generation model, make the target visual content predicted and generated by the target object have consistency with the attribute information of the original image. At the same time, make the target visual content displayed by each target object accurately match the indication information in the prompt text, and make the interaction between multiple target objects adapt and match the indication information in the prompt text, so as to achieve the generation of high-quality target visual content and meet the visual creation needs of users.

[0048] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] By describing the exemplary embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features and advantages of the present application will become more obvious. Among them, in the exemplary embodiments of the present application, the same reference numerals generally represent the same components.

[0050] Figure 1 is a schematic diagram of the network structure of a visual content generation model shown in an embodiment of the present application;

[0051] Figure 2 is a schematic flowchart of the training method of the multi-object visual content generation model shown in an embodiment of the present application;

[0052] Figure 3 is another schematic flowchart of the training method of the multi-object visual content generation model shown in an embodiment of the present application;

[0053] Figure 4 is a schematic flowchart of a method for generating visual content of multiple objects shown in an embodiment of the present application;

[0054] Figure 5 is a schematic structural diagram of a training device for a visual content generation model of multiple objects shown in an embodiment of the present application;

[0055] Figure 6 is another schematic structural diagram of a training device for a visual content generation model of multiple objects shown in an embodiment of the present application;

[0056] Figure 7 is a schematic structural diagram of a device for generating visual content of multiple objects shown in an embodiment of the present application;

[0057] Figure 8 is another schematic structural diagram of a device for generating visual content of multiple objects shown in an embodiment of the present application;

[0058] Figure 9 is a schematic structural diagram of an electronic device shown in an embodiment of the present application. Detailed implementation manners

[0059] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.

[0060] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "said", and "the" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0061] It should be understood that although the terms "first", "second", "third", etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, "a plurality" means two or more unless otherwise specifically defined.

[0062] Currently, when there is only one character in an image, the pre-trained diffusion model can fuse with the user's text prompt and visual features to generate new visual content. However, when there are multiple characters in an image, how to accurately match the visual features of each character with the corresponding text description becomes a major problem.

[0063] In view of the above problems, an embodiment of the present application provides a training method for a multi-object visual content generation model, which can ensure the consistency and control accuracy of the visual content in a multi-object generation scenario for the trained visual content generation model. Among them, the visual content can be a static image or a dynamic video.

[0064] The technical solutions of the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0065] Figure 1 It is a schematic diagram of the network structure of a visual content generation model shown in an embodiment of the present application.

[0066] In some embodiments, the visual content generation model of the present application includes a position encoding module to be trained, a cross-attention module to be trained, and a pre-trained diffusion model. Based on the pre-trained diffusion model, by adding a position encoding module and a cross-attention module outside the diffusion model to adjust the generated data of the diffusion model. Specifically, the training method of the present application trains the network parameters of the position encoding module and the cross-attention module, so as to insert additional network parameters into the pre-trained diffusion model to adapt to the new task, rather than directly adjusting the parameters of the diffusion model (the parameters of the diffusion model are on the order of billions or tens of billions), which can not only reduce the number of parameters required for training, but also reduce the demand for computing resources during the training process and speed up the training speed. Through the combination of the position encoding module and the cross-attention module with the diffusion model, the accurate matching of multi-character visual content and text description is realized based on the AIGC technology, ensuring the high-quality expression of the new visual content and not affecting the generality of the diffusion model.

[0067] See Figure 2 , Figure 2 It is a schematic flowchart of the training method of the multi-object visual content generation model shown in an embodiment of the present application.

[0068] In some embodiments, a training method for a multi-object visual content generation model shown in the present application includes:

[0069] S110, obtaining multiple groups of training data pairs, each group of training data pairs including a training image and a text description, the training image including at least two target objects, and the text description including the attribute description of each target object.

[0070] In this application, each pair of training data includes at least one frame of training image or multiple frames of training images. Each frame of training image includes at least one target object, and the cumulative number of target objects in all training images is at least two. For example, a pair of training data includes at least two frames of training images, and each frame of training image includes one target object. For another example, a pair of training data includes one frame of training image, and the single frame of training image includes at least one target object. For yet another example, when a pair of training data has only one frame of training image and there is only one target object in the image, at least two target objects can be obtained by copying the same target object.

[0071] In some embodiments, the target object can be categories such as people, animals, plants, or items, etc.; each target object can be a real image or a virtual image, such as a cartoon image, etc., and there is no limitation here. In a set of training data, the categories of target objects in different training images can be the same or different, and there is no limitation here. In some embodiments, the attribute description of the target object can include at least one of appearance features, physical features, clothing features, color features, action features, and orientation features. Here is only an example for illustration and there is no limitation. Taking the target object as a person as an example, its attribute description can include its physical features, clothing features, action features, etc., so as to clearly and accurately define the visual features of the target object.

[0072] Further, the text description in each pair of training data describes in text form the visual content expected to be generated by each target object. Taking a pair of training data including two frames of training images as an example, one of the training images 1 shows a curly-haired man wearing a red leather jacket, and the other training image 2 shows a human wearing a red metal armor; the text description of the expected generated content can be "a curly-haired man wearing a red leather jacket; a human wearing a red metal armor; Figure 1 and Figure 2 , hugging". Obviously, there are two target objects in this pair of training data; the text description includes the attribute descriptions of these two target objects respectively, and expects the visual content to be generated by these two independent target objects. That is to say, with a paragraph of text description, the visual content that each target object needs to display in the same scene can be comprehensively expressed.

[0073] It can be understood that the visual content of the training image includes the target object and the background image. To improve the training efficiency, the background image in each frame of training image can be set as a base map with a monotonous color, such as a solid-color base map or a base map rendered in the same color system as the background image, so that the system can accurately identify the edge texture of the target object and then quickly extract the image features of the target object.

[0074] S120. Based on the pair of training data, obtain the first local visual feature representation and the first local text feature representation corresponding to each target object.

[0075] In this step, through related technologies, the image features in the training images and the semantic features in the text descriptions can be encoded to form data readable by the visual generation model, such as feature vectors or feature matrices. Specifically, according to the pixel regions corresponding to each target object in the training images, the corresponding first local visual feature representation S = [s(i)], where i = 2, 3, 4..., representing the i-th target object, is encoded through an Image Encoder. According to the text prompts corresponding to each target object in a text description, the corresponding first local text feature representation Y = [y(i)], where i = 2, 3, 4..., representing the i-th target object, is encoded and extracted through a Text Encoder. It can be understood that in S = [s(i)] and Y = [y(i)], the same numerical value i is used to refer to the same target object.

[0076] It can be understood that for the two different modalities of images and texts, through related technologies such as representing them in the same feature space, the first local visual feature representation and the first local text feature representation are semantically aligned to ensure that the visual content generation model can generate content based on the information unified by multiple modalities.

[0077] S130, through the position encoding module, image position encoding and text position encoding are respectively embedded in the first local visual feature representation and the first local text feature representation corresponding to each target object to obtain the second local visual feature representation and the second local text feature representation.

[0078] In this step, the position encoding module is used to embed the image position encoding into the first local visual feature representation to form the second local visual feature representation; and the position encoding module is used to embed the text position encoding into the first local text feature representation to form the second local text feature representation. Among them, the image position encoding and the text position encoding embedded for the same target object form a one-to-one mapping relationship. That is to say, for a group of s(i) and y(i) of the same target object, the corresponding position encodings (id embedding) are respectively embedded through the position encoding module to obtain s(i)' and y(i)', and then the corresponding relationship F = [s(i), y(i)] between the first local visual feature and the corresponding first local text feature of each target object is determined. Correspondingly, after the image position encoding and the text position encoding are embedded for all target objects, the corresponding second local visual feature representation S' = [s(i)'] and the second local text feature representation Y' = [y(i)'] can be obtained. Such a design enables a one-to-one mapping relationship to be generated between the second local visual feature representation and the second local text feature representation of each target object through their respective corresponding image position encoding and text position encoding, so that in the entire training process of the visual generation content model subsequently, based on this mapping relationship, the visual features and text features of each target object can be accurately paired.

[0079] In some embodiments, the position encoding module of the present application includes an image position encoding network and a text position encoding network; among them, for each target object, the image position encoding is embedded into the corresponding first local visual feature representation through the Image Position Embedding network, and the text position encoding is embedded into the corresponding first local text feature representation through the Text Position Embedding network. That is to say, by using independent image position encoding network and text position encoding network for encoding respectively, the training amount of parameters can be reduced, and it is easier to train to obtain a target network with specific functions, and it has better scalability; in addition, by the two encoding networks acting on the position encodings of different modalities simultaneously, the processing efficiency can be improved.

[0080] It can be understood that both the image position encoding network and the text position encoding network in the position encoding module of the present application have network parameters to be trained, that is, the network parameters of the two need to be determined through training.

[0081] S140, input the second local visual feature representation and the second local text feature representation of each target object into the cross-attention module and the diffusion model, so that the diffusion model denoises according to the fused features output by the cross-attention module and outputs the predicted visual content.

[0082] In this step, the diffusion model is a pre-trained general model, and the cross-attention module is the model to be trained. By inputting S’ = [s(i)’] and Y’ = [y(i)’] from the above step into the cross-attention module and the diffusion model respectively, the input data and output data of the two interact and influence each other during the training process. The diffusion model can adjust the output of the data based on the influence of the cross-attention module, so as to obtain the prediction result fine-tuned by the cross-attention module.

[0083] The pre-trained diffusion model has the ability to generate visual content. Based on the input information data, it can generate new images or new videos. That is, the predicted visual content can be an image or a video. The network parameter scale of this pre-trained diffusion model is relatively large, such as in the order of billions or tens of billions. If the parameters of this pre-trained diffusion model are directly re-trained and adjusted for a specific task, it will not only affect the generality of this pre-trained diffusion model, but also require a very large amount of training resources and a long cycle. Taking the U-Net network as an example for the diffusion model, it includes multiple layers of feature sampling layers. Specifically, U-Net is a structure of a downsampling module - upsampling module with skip connections, used to denoise random noise and achieve the generation from noise to video or image. Specifically, by performing denoising processing on random noise for t time steps, that is, both downsampling and upsampling are performed in each time step, the latent space expression of the target video or image is obtained, and then the latent space expression of the target video or image is further mapped to the pixel space through a decoder to obtain the target video or image. Among them, the downsampling module is used to gradually downsample the input noise, extract features and reduce the spatial resolution of the image. The upsampling module gradually restores the spatial resolution of the image through transposed convolution (or inverse convolution) and upsampling operations; through skip connections, the upsampling module can combine the low-level features (such as details like edges and textures) in the downsampling module with high-level features (such as semantic information of the prompt text). The fusion of features helps to reconstruct the details of the image; finally, the target video or target image with the same resolution as the input image is output.

[0084] In an embodiment of the present invention, a pre-trained visual content generation model is obtained through joint training of a diffusion model and a cross-attention module. Without directly adjusting the parameters of the diffusion model, it can not only reduce the number of parameters required for fine-tuning, but also reduce the demand for computing resources during training, thereby accelerating the training speed. At the same time, when the diffusion model performs a specific task, through the trained cross-attention module, feature injection is performed on the vectors of the feature sampling layer of the diffusion model, so that it is possible to efficiently and accurately complete the specific task without large-scale overall training of the diffusion model. It can be seen that through the joint training of the diffusion model and the cross-attention module, a pre-trained visual content generation model is obtained, which realizes the rapid generation of high-quality visual content based on the AIGC technology, including multiple objects and accurately matching the corresponding text prompts, and does not affect the generality of the diffusion model.

[0085] Further, as Figure 1 shown, in some embodiments, the cross-attention module (Cross-Attention) of the present application includes a first attention network (Cross-Attention 1) and a second attention network (Cross-Attention 2); the first attention network is used to output a corresponding first fusion feature according to the second local text feature representation and the sampling features output by the diffusion model; the second attention network is used to output a second fusion feature according to the second local visual feature representation and the first fusion feature, and the second fusion feature is used to input the diffusion model for feature sampling.

[0086] In a specific embodiment, the diffusion model includes N layers of feature sampling layers; cross-attention calculation is performed on the second local text feature representation and the output features of the k-th layer of the diffusion model through the first attention network to obtain the k-th first fusion feature; where 1 ≤ k ≤ N, both k and N are natural numbers; cross-attention calculation is performed on the first fusion feature and the second local visual feature representation through the second attention network to obtain the k-th second fusion feature; where the k-th second fusion feature is used to input the (k + 1)-th layer of the diffusion model to obtain the output features of the (k + 1)-th layer; when k + 1 = N, the diffusion model obtains the output features of the current time step and generates predicted visual content based on the output features of the last time step.

[0087] For ease of understanding, in combination with Figure 1 , the following specific examples are used for illustration. The specific numerical values are only for convenience of understanding and are not limited herein.

[0088] Exemplarily, assume that N is 7 and t is 10, that is, the diffusion model has 7 feature sampling layers, and the diffusion model needs to execute 10 time steps. i is 2, that is, the number of target objects is 2. After the position encoding module generates the second local visual feature representation S’ = [s(1)’, s(2)’] and the second text feature representation Y’ = [y(1)’, y(2)’] for all target objects, they are both input into the first attention network (Cross-Attention 1) and the second attention network (Cross-Attention 2).

[0089] In the 1st time step, S’ and Y’ are input into the 1st feature sampling layer of the diffusion model to obtain the 1st output feature, which is then input into the first attention network. At the same time, Y’ is input into the first attention network to perform cross-attention calculation with the 1st output feature, obtaining the 1st first fusion feature. S’ and the 1st first fusion feature are respectively input into the second attention network to perform cross-attention calculation, obtaining the 1st second fusion feature. Then, the 1st second fusion feature is used as the input of the 2nd feature sampling layer of the diffusion model, and then the 2nd output feature is obtained and input into the first attention network. Through the first attention network, Y’ and the 2nd output feature are used to perform cross-attention calculation, obtaining the 2nd first fusion feature. S’ and the 2nd first fusion feature are respectively input into the second attention network to perform cross-attention calculation, obtaining the 2nd second fusion feature. The 2nd second fusion feature is used as the input of the 3rd feature sampling layer of the diffusion model, and so on in accordance with the foregoing steps until the 7th second fusion feature is obtained, and then the latent space representation X corresponding to the predicted noise image at the 1st time step is obtained. t-1 。

[0090] The same process as above exists in each time step. In the 2nd time step, X t-1 is input into the 1st feature sampling layer of the diffusion model, and in accordance with the same process as above, the latent space representation X at the 2nd time step is obtained. t-2 And so on until the latent space representation X1 corresponding to the predicted noise image at the last time step is obtained, and then the latent space representation of the target video or image can be mapped to the pixel space through the decoder to obtain the target video or image.

[0091] As can be seen from the above, the second local visual feature representation and the second local text feature representation with the position encoding of the unique mapping relationship embedded are fine-tuned by the cross-attention module diffusion model, and cross-calculation is performed in the feature sampling of each layer at each time step of the diffusion model. Using the image position encoding and text position encoding with a unique mapping relationship as the associated markers, the corresponding relationship between the visual features and text features of the same target object is locked throughout the process, realizing the precise pairing of the two, so that the features in the multi-object generation scenario will not be confused, ensuring the consistency and control accuracy of the generated visual content in the text-image.

[0092] S150, fix the network parameters of each layer of the diffusion model, and backpropagate according to the loss function of the diffusion model to iteratively adjust the network parameters of the cross-attention module and the position encoding module to obtain a trained visual content generation model.

[0093] It can be understood that after the diffusion model trains and outputs the predicted visual content, the loss can be calculated according to the loss function, and backpropagation is performed according to the calculation result of the loss function. Among them, the network parameters of each layer of the diffusion model remain unchanged, and the model network parameters of the cross-attention module and the position encoding module are adjusted accordingly. After completing multiple rounds of training, when the loss function converges or reaches the preset training stop condition, the training can be ended to obtain the trained network parameters.

[0094] It can be understood that the first attention network and the second attention network in the above cross-attention module obtain trained network parameters through training. The position encoding module includes an image position encoding network and a text position encoding network, which obtain trained network parameters through training.

[0095] From this example, it can be seen that in this application, by connecting a position encoding module and a cross-attention module outside the diffusion model, it is not necessary to perform large-scale parameter training on the diffusion model, and only a small number of parameters need to be trained, improving the training efficiency and reducing the training cost. At the same time, by embedding the image position encoding and text position encoding of the same target object and introducing an id linkage mechanism, the keywords in the text description can be precisely paired with the visual features of the corresponding target object, reducing the problem of feature confusion in the multi-person generation scenario and improving the visual consistency and control accuracy of the generated content. The text features and visual features of the same target object in the scene are deeply bound and fused under the action of the cross-attention module, and the visual features of different target objects interact naturally under the action of the cross-attention module, ensuring that the trained visual content generation model can generate high-quality visual content.

[0096] See Figure 1 and Figure 3 , this application also provides a training method for a visual content generation model of multiple objects in an embodiment, which includes:

[0097] S210. Obtain multiple sets of training data pairs. Each set of training data pairs includes a training image and a text description. The training image includes at least two target objects, and the text description contains the attribute description of each target object.

[0098] For the introduction of this step, please refer to the above S110, which will not be elaborated here.

[0099] S220. Detect and segment each training image to obtain the local image corresponding to each target object in the training image, and generate the corresponding first local visual feature representation; perform semantic recognition on the text description to obtain the local text description corresponding to each target object, and generate the corresponding first local text feature representation.

[0100] It can be understood that through the object detection model in related technologies, the target objects in each frame of the training image can be accurately detected, and then image segmentation can be performed. It can be understood that a frame of training image includes one or more target objects, and the visual contents of each target image may be the same or different. Through object detection and image segmentation, a frame of training image can be segmented into one or more corresponding local images, and each frame of local image contains a unique target object. If there are multiple target objects in a frame of training image, there may be overlapping regions between the multiple target objects, that is, individual target images will be blocked by other target images. The training method of the present application can perform natural segmentation according to the detected target objects. For the blocked target objects in the segmented local images, the situation of image mutilation is allowed. By not restricting the image of the target object, the diversity of the training data of the visual content generation model is fully expanded, so that the generated effect after training can generate the visual content with the target object blocked.

[0101] To adapt to the application of multi-modal training data, in related technologies, for example, the CLIP model (Contrastive Language-Image Pre-training) is a multi-modal pre-trained neural network. Its core function is to learn the relationship between images and texts from natural language supervision and map them to the same embedding space, so that cross-modal comparison and matching can be performed by calculating the similarity between vectors. The CLIP model includes an Image Encoder and a Text Encoder. The Image Encoder converts the local images corresponding to each target object into local visual feature vectors. The Text Encoder converts the input text description into local text feature vectors. By mapping the feature vectors corresponding to images and texts to a shared embedding space, the similarity between them can be directly compared, thus completing the precise one-to-one pairing of the image features and text features of the same target object. It can be understood that the first local visual feature representation in this embodiment is the local visual feature vector, and the first local text feature representation is the local text feature vector.

[0102] In some embodiments, the first local visual features and the first local text descriptions of each target object are associated and labeled manually or by machine. For example, taking two training images as an example, the corresponding images are respectively labeled as Figure 1 and Figure 2 , and it is known that Figure 1 and Figure 2 respectively contain their respective target objects. Correspondingly, in a text description, the prompt words contain the fields of " Figure 1 " and " Figure 2 ". Such a design realizes the associated mapping of the first local visual features and the first local text descriptions of each target object.

[0103] S230, for each target object, embed the image position encoding in the corresponding first local visual feature representation through the image position encoding network, and embed the text position encoding in the corresponding first local text feature representation through the text position encoding network to obtain the second local visual feature representation and the second local text feature representation.

[0104] In this step, both the image position encoding network and the text position encoding network add id embeddings (image position encoding and text position encoding) for each target object to label the visual features and text features of the same target object with the same tag. That is to say, the image position encoding and text position encoding of the same target object can be uniquely corresponding encodings, for example, they can be the same ID encoding; different target objects use different ID encodings. With such a design, through the id linkage mechanism, specific words in a text description of the same target object can be bound to the corresponding visual information, enabling precise control of visual features according to text prompts.

[0105] For other introductions of this step, refer to the above S230 and will not be elaborated here.

[0106] S240, input the second local visual feature representation and the second local text feature representation of each target object into the cross-attention module and the diffusion model; among them, cross-attention calculation is performed on the second local text feature representation and the output feature of the k-th layer of the diffusion model through the first attention network to obtain the k-th first fusion feature; where 1 ≤ k ≤ N, and both k and N are natural numbers; cross-attention calculation is performed on the first fusion feature and the second local visual feature representation through the second attention network to obtain the k-th second fusion feature; among them, the k-th second fusion feature is used to input into the (k + 1)-th layer of the diffusion model to obtain the output feature of the (k + 1)-th layer; when k + 1 = N, the diffusion model obtains the output feature of the current time step and generates the predicted visual content based on the output feature of the last time step.

[0107] In this step, through the fusion network in the attention mechanism, the mapped and paired second local visual feature representation and second local text feature representation are fused to ensure that the visual features generated according to the text features can be accurately mapped to the corresponding target objects, enabling the target objects predicted by the diffusion model to be consistent with the attribute descriptions in the training images while the generated visual features match the text descriptions.

[0108] For other introductions of this step, refer to the above S140 and will not be elaborated here.

[0109] S250, fix the network parameters of each layer of the diffusion model, and backpropagate according to the loss function of the diffusion model to iteratively adjust the network parameters of the first attention network, the first attention network, the image position encoding network, and the text position encoding network to obtain a trained visual content generation model.

[0110] In some embodiments, the loss function includes at least one of the overall object feature similarity, the key object feature similarity, and the matching degree between the text and the image. Taking the target object as a person as an example, the loss function includes at least one of the person feature similarity (character sim), the face feature similarity (face sim), and the matching degree between the text and the image (CLIP text-Imagesim). When the target object is other categories such as animals, plants, or items, similar settings and calculations can be made with reference to the following concepts.

[0111] Specifically, the person feature similarity is used to measure the similarity between the person body image features in the generated image and the person body image features in the character image. The specific method is to perform human detection on the generated image, and extract the human body area to obtain the human body image (persons_in_image). For each character image, human detection is also performed to extract the human body area image (character_image). If there are two character images, they correspond to character_image_1 and character_image_2. For each character_image, the image similarity is calculated with the persons_in_image image using CLIP, and the maximum value is taken as the sim of that character_image. The sims of the two persons are averaged as the character_sim of this image. The character_sims of multiple images are averaged as the final person feature similarity.

[0112] Specifically, the face feature similarity is used to measure the similarity between the person face image features in the generated image and the face image features in the character image. The specific method is to perform face detection on the generated image, and use the face recognition algorithm insightface to extract the face features (face_embeddings_in_image). For each character image, face feature extraction is also performed to obtain the face_embedding (if there are two character images, they correspond to face_embedding_1 and face_embedding_2). For each face_embedding, the face similarity is calculated with the face_embeddings_in_image, and the maximum value is taken as the sim of that face. The sims of multiple persons are averaged as the face_sim of this image. The face_sims of multiple images are averaged as the final person feature similarity.

[0113] Specifically, the matching degree between the text and the image is used to calculate the CLIP similarity between the generated image and the text used to generate the image.

[0114] When the loss function includes the overall object feature similarity, the key object feature similarity, and the matching degree between the text and the image, calculate the calculation results of each loss function respectively. After multiple rounds of training and loss calculation of the prediction results, when the calculation results of each item reach the preset standard, the training can be ended or the training data can be updated and the training can continue.

[0115] The introduction of this step refers to S150 above and will not be elaborated here.

[0116] From this example, it can be seen that the training method of the multi-object visual content generation model of the present application can, based on the image position encoding network and the text position encoding network, achieve one-to-one binding and pairing of the visual features and text features of multiple target objects, and then through the cross-attention mechanism of the first attention network and the second attention network, cross-fuse the visual features and text features of a single target object, and enable the natural interaction of multiple target objects in the same scene. While maintaining the relative independence of the respective visual features of the target objects and the consistency of the pairing with the text features, it ensures the high-quality generation of visual content.

[0117] See Figure 4 , the present application also provides a multi-object visual content generation method. The generation method of the present application is applicable to the creative scenario of multiple objects by the user using the visual content generation model trained by any of the above embodiments. The multi-object visual content generation method of the present application includes:

[0118] S310, obtain the P-frame original image and the prompt text. The total number of target objects in each frame of the original image is Y, Y≥P, P≥1.

[0119] The visual content generation method of the present application can obtain at least one frame of the original image input by the user. The total number of all target objects in all the original images is Y. Among them:

[0120] When P is greater than 1, each frame of the original image contains at least one target object. When the original image is multiple frames, each frame of the original image can contain one or more target objects, and the respective target objects can be the same or different.

[0121] When P is equal to 1 and the original image contains multiple target objects, segment the original image to obtain the local images corresponding to each target object. The number of original images can be only one frame. If one frame of the original image contains multiple target objects, through object detection and segmentation in the image, the local image of each target object can be obtained.

[0122] When P is equal to 1 and the original image contains 1 target object, copy and generate multiple target objects. When the original image is only 1 frame and only contains 1 target object, at least 2 target objects can be generated by copying the image.

[0123] In some embodiments, the prompt text includes the attribute information of the target object, so that the visual features and text features of each target object can be more accurately locked. Taking the target object in the original image as a person as an example, the prompt text can include the appearance features and clothing features of each target object, so that the visual content generated by subsequent prediction can accurately maintain the corresponding appearance features and clothing features. Further, the prompt text can also include the generation actions, display positions, etc. of multiple target objects expected by the user, so that the visual content generated by subsequent prediction has the expected action features and position features.

[0124] S320. According to the visual content generation model, generate the corresponding target visual content, and the attribute information of the target object in the target visual content is consistent with that in the original image.

[0125] The visual content generation model of the present application can be used to predict and generate a target image or a target video according to the input data of S310, that is, both the target image and the target video have the target visual content, and the target image or the target video is a carrier form of the target visual content.

[0126] Taking the target object as an example of a person, the attribute information of the target object in the target visual content is consistent with that in the original image, that is, the appearance of the same person in the target visual content is the same as his appearance in the original image, and the clothing of the same person in the target visual content is the same as his clothing in the original image. At the same time, multiple people show the action features defined in the prompt text, that is, the interaction of multiple target objects conforms to the instructions in the prompt text input by the user, realizing the natural interaction between multiple target objects to ensure the high-quality display of the target visual content.

[0127] Specifically, the visual content generation model, according to the original image and the prompt text input by the user, after detecting and segmenting the image, generates the first local visual feature representation of each target object through the image encoder based on each frame of the original image, and generates the first local text feature representation of each target object through the text encoder based on the prompt text. The first local visual feature representation and the first local text feature representation are aligned and paired in the same feature space.

[0128] Further, for the same target object, the image position encoding network in the visual generation model embeds the first local visual feature representation with image position encoding to obtain the second local visual feature representation; the text position encoding network embeds the first local text feature representation with image position encoding to obtain the second local text feature representation.

[0129] Input the second local visual feature representation and the second local text feature representation of each target object into the cross-attention module and the diffusion model. Among them, perform cross-attention calculation on the second local text feature representation and the output feature of the k-th layer of the diffusion model through the first attention network to obtain the k-th first fusion feature; where 1 ≤ k ≤ N, and both k and N are natural numbers; perform cross-attention calculation on the first fusion feature and the second visual feature representation through the second attention network to obtain the k-th second fusion feature; where the k-th second fusion feature is used to be input into the (k + 1)-th layer of the diffusion model to obtain the output feature of the (k + 1)-th layer; when k + 1 = N, the diffusion model obtains the output feature of the current time step, and predicts and generates the target visual content according to the output feature of the last time step.

[0130] As can be seen from this example, the method for generating visual content of multiple objects in this application can, based on the original image and the prompt text input by the user, through the trained visual content generation model, make the target objects have consistency in the predicted and generated target visual content and the attribute information of the original image. At the same time, make the target visual content displayed by each target object match the indication information in the prompt text precisely, and make the interaction between multiple target objects adapt and match the indication information in the prompt text, so as to realize the generation of high-quality target visual content and meet the visual creation needs of users.

[0131] Corresponding to the foregoing embodiments of the application function implementation method, the present application also provides a training device for a multi-object visual content generation model, a multi-object visual content generation device, an electronic device, and corresponding embodiments.

[0132] Figure 5 It is a schematic structural diagram of a training device for a multi-object visual content generation model shown in an embodiment of the present application.

[0133] See Figure 5 , the training device for a multi-object visual content generation model of the present application includes a training data acquisition module 510, a local feature extraction module 520, a position encoding embedding module 530, a cross-denoising module 540, and a network parameter iteration module 550. Among them:

[0134] The training data acquisition module 510 is used to acquire multiple groups of training data pairs, each group of training data pairs includes a training image and a text description, the training image includes at least two target objects, and the text description contains the attribute description of each target object.

[0135] The local feature extraction module 520 is used to obtain the first local visual feature representation and the first local text feature representation corresponding to each target object based on the training data pair.

[0136] The position encoding embedding module 530 is used to respectively embed the image position encoding and the text position encoding in the first local visual feature representation and the first local text feature representation corresponding to each target object through the position encoding module, so as to obtain the second local visual feature representation and the second local text feature representation.

[0137] The cross-denoising module 540 is used to input the second local visual feature representation and the second local text feature representation of each target object into the cross-attention module and the diffusion model, so that the diffusion model denoises according to the fused features output by the cross-attention module and outputs the predicted visual content.

[0138] The network parameter iteration module 550 is used to fix the network parameters of each layer of the diffusion model, and backpropagate according to the loss function of the diffusion model to iteratively adjust the network parameters of the cross-attention module and the position encoding module, so as to obtain a trained visual content generation model.

[0139] See Figure 6 In some specific embodiments, the local feature extraction module 520 includes an image feature encoding module 521 and a text feature encoding module 522. The image feature encoding module 521 is used to detect and segment each training image, obtain the local image corresponding to each target object in the training image, and generate the corresponding first local visual feature representation. The text feature encoding module 522 is used to perform semantic recognition on the text description, obtain the local text description corresponding to each target object, and generate the corresponding first local text feature representation.

[0140] In some specific embodiments, the position encoding embedding module 530 includes an image position embedding module 531 and a text position embedding module 532. The image position embedding module 531 is used to embed the image position encoding in the corresponding first local visual feature representation for each target object through the image position encoding network. The text position embedding module 532 is used to embed the text position encoding in the corresponding first local text feature representation through the text position encoding network.

[0141] In some specific embodiments, the cross-denoising module 540 includes a first attention calculation module 541, a second attention calculation module 542, and a denoising prediction module 543. Among them, the first attention calculation module 541 is used to output the corresponding first fused feature through the first attention network according to the second local text feature representation and the sampled features output by the diffusion model. The second attention calculation module 542 is used to output the second fused feature through the second attention network according to the second local visual feature representation and the first fused feature, and the second fused feature is used to input the diffusion model for feature sampling. The denoising prediction module 543 is used to denoise according to the first fused feature and the second fused feature through the diffusion model and output the predicted visual content.

[0142] Specifically, the first attention calculation module 541 performs cross-attention calculation on the second local text feature representation and the output features of the k-th layer of the diffusion model through the first attention network to obtain the k-th first fusion feature; where 1 ≤ k ≤ N, and both k and N are natural numbers; the second attention calculation module 542 performs cross-attention calculation based on the first fusion feature and the second visual feature representation through the second attention network to obtain the k-th second fusion feature; where the denoising prediction module 543 inputs the k-th second fusion feature into the (k + 1)-th layer of the diffusion model to obtain the output features of the (k + 1)-th layer; when k + 1 = N, the diffusion model obtains the output features of the current time step and generates predicted visual content based on the output features of the last time step.

[0143] The training device of the multi-object visual content generation model of the present application can accurately pair and bind the keywords in the text description with the visual features of the target object by introducing the id linkage mechanism, improving the adaptability of the model in multi-person scenarios, reducing the feature confusion problem in multi-person generation scenarios, and enhancing the visual consistency and control accuracy of the generated images; through the cross-attention mechanism and the denoising process of the diffusion model for feature fusion, realizing the precise allocation of visual features in multi-object scenarios, and being able to flexibly control the actions, positions and other attributes of different target objects, improving the control ability of the text-to-image model.

[0144] Figure 7 It is a schematic structural diagram of the multi-object visual content generation device shown in the embodiments of the present application.

[0145] The multi-object visual content generation device in an embodiment of the present application includes a data acquisition module 710 and a visual generation module 720. Among them:

[0146] The data acquisition module 710 is used to acquire P-frame original images and prompt texts, and the total number of target objects in each frame of the original images is Y, where Y ≥ P and P ≥ 1;

[0147] The visual generation module 720 is used to generate corresponding target visual content according to the visual content generation model obtained by the above training method, and the attribute information of the target object in the target visual content is consistent with that in the original image.

[0148] See Figure 8 , in a specific implementation manner, the visual generation module 720 includes a local feature extraction module 520, a position encoding embedding module 530, and a cross-denoising module 540. The structures and functions of the local feature extraction module 520, the position encoding embedding module 530, and the cross-denoising module 540 in the generation device of the present application are the same as those of the corresponding modules in the above training device, and will not be elaborated here.

[0149] The multi-object visual content generation device of the present application can make the generated target visual content include multiple target objects input by the user, and the interaction between the multiple target objects conforms to the semantic indication in the text prompt input by the user.

[0150] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here.

[0151] Figure 9 It is a schematic structural diagram of an electronic device shown in an embodiment of the present application.

[0152] See Figure 9 , the electronic device 1000 includes a memory 1010 and a processor 1020.

[0153] The processor 1020 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0154] The memory 1010 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, the ROM may store static data or instructions required by the processor 1020 or other modules of the computer. The permanent storage device may be a readable and writable storage device. The permanent storage device may be a non-volatile storage device that does not lose the stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device employs a mass storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In some other embodiments, the permanent storage device may be a removable storage device (such as a floppy disk, optical drive). The system memory may be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory may store some or all of the instructions and data required by the processor during operation. In addition, the memory 1010 may include any combination of computer-readable storage media, including various types of semiconductor storage chips (such as DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks may also be used. In some embodiments, the memory 1010 may include a removable storage device that is readable and / or writable, such as a compact disc (CD), read-only digital versatile disc (such as DVD-ROM, dual-layer DVD-ROM), read-only Blu-ray disc, super density disc, flash memory card (such as SD card, min SD card, Micro-SD card, etc.), magnetic floppy disk, etc. The computer-readable storage medium does not include carrier waves and instantaneous electronic signals transmitted wirelessly or by wire.

[0155] Executable code is stored on the memory 1010, and when the executable code is processed by the processor 1020, it may cause the processor 1020 to execute some or all of the methods described above.

[0156] In addition, the method according to the present application may also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing some or all of the steps of the above method according to the present application.

[0157] Alternatively, the present application may also be implemented as a computer-readable storage medium (or non-transitory machine-readable storage medium or machine-readable storage medium), on which executable code (or computer program or computer instruction code) is stored. When the executable code (or computer program or computer instruction code) is executed by a processor of an electronic device (or a server, etc.), it causes the processor to execute some or all of the steps of the above method according to the present application.

[0158] The embodiments of the present application have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A training method for a visual content generation model of multiple objects, characterized in that, The visual content generation model includes a position encoding module to be trained, a cross-attention module to be trained, and a pre-trained diffusion model; where: Obtain multiple sets of training data pairs, each set of the training data pairs including a training image and a text description, the training image including at least two target objects, and the text description containing the attribute description of each target object; Based on the training data pairs, obtain a first local visual feature representation and a first local text feature representation corresponding to each of the target objects; Embed an image position encoding and a text position encoding respectively in the first local visual feature representation and the first local text feature representation corresponding to each of the target objects through the position encoding module, to obtain a second local visual feature representation and a second local text feature representation; Input the second local visual feature representations and the second local text feature representations of the respective target objects into the cross-attention module and the diffusion model, so that the diffusion model performs denoising according to the fused features output by the cross-attention module and outputs a predicted visual content; where, the diffusion model includes N layers of feature sampling layers; the cross-attention module includes a first attention network and a second attention network; the first attention network is used to output a corresponding first fused feature according to the second local text feature representation and the sampled features output by the diffusion model; the second attention network is used to output a second fused feature according to the second local visual feature representation and the first fused feature, and the second fused feature is used to input the diffusion model for feature sampling; where, the diffusion model performs denoising according to the fused features output by the cross-attention module and outputs a predicted visual content, including: performing cross-attention calculation on the second local text feature representation and the output features of the k-th layer of the diffusion model through the first attention network to obtain the k-th first fused feature; where, 1 ≤ k ≤ N, and both k and N are natural numbers; performing cross-attention calculation on the first fused feature and the second local visual feature representation through the second attention network to obtain the k-th second fused feature; where, the k-th second fused feature is used to input the (k + 1)-th layer of the diffusion model to obtain the output features of the (k + 1)-th layer; when k + 1 = N, the diffusion model obtains the output features of the current time step and generates a predicted visual content according to the output features of the last time step; Fix the network parameters of each layer of the diffusion model, and perform backpropagation according to the loss function of the diffusion model, and iteratively adjust the network parameters of the cross-attention module and the position encoding module to obtain a trained visual content generation model.

2. The training method according to claim 1, wherein The obtaining, based on the training data pairs, of a first local visual feature representation and a first local text feature representation corresponding to each of the target objects includes: Perform detection and segmentation on each training image, obtain the local image corresponding to each target object in the training image, and generate a corresponding first local visual feature representation; Perform semantic recognition on the text description, obtain the local text description corresponding to each target object, and generate a corresponding first local text feature representation.

3. The training method according to claim 1, wherein The position encoding module includes an image position encoding network and a text position encoding network; Embedding the image position encoding and the text position encoding respectively into the first local visual feature representation and the first local text feature representation corresponding to each target object through the position encoding module includes: For each target object, embedding the image position encoding into the corresponding first local visual feature representation through the image position encoding network, and embedding the text position encoding into the corresponding first local text feature representation through the text position encoding network.

4. The training method according to claim 1, wherein The loss function includes at least one of the overall object feature similarity, the key object feature similarity, and the text and image matching degree.

5. The training method according to claim 1, wherein Each frame of the training image includes at least one target object; the target object is a person, an animal, a plant or an item; The attribute description includes at least one of appearance features, facial features, clothing features, color features, and orientation features.

6. A visual content generation method for multiple objects, characterized in that, Includes: Obtaining the P-frame original image and the prompt text, the total number of target objects in each frame of the original image is Y, Y≥P, P≥1; The visual content generation model obtained according to any one of the training methods of claims 1 to 5 generates the corresponding target visual content, and the attribute information of the target object in the target visual content is consistent with that in the original image.

7. The visual content generation method according to claim 6, wherein: When P is greater than 1, each frame of the original image contains at least one target object; or When P is equal to 1 and the original image contains multiple target objects, segmenting the original image to obtain local images corresponding to each target object; Or When P is equal to 1 and the original image contains 1 target object, replicating to generate multiple target objects.

8. A training device for a visual content generation model of multiple objects, characterized in that The visual content generation model includes a position encoding module to be trained, a cross-attention module to be trained, and a pre-trained diffusion model; wherein: The training data acquisition module is used to acquire multiple groups of training data pairs, each group of training data pairs includes a training image and a text description, the training image includes at least two target objects, and the text description contains the attribute description of each target object; The local feature extraction module is used to obtain the first local visual feature representation and the first local text feature representation corresponding to each target object based on the training data pair; The position encoding embedding module is used to embed the image position encoding and the text position encoding respectively into the first local visual feature representation and the first local text feature representation corresponding to each target object through the position encoding module, to obtain the second local visual feature representation and the second local text feature representation; A cross-denoising module, which is configured to input the second local visual feature representation and the second local text feature representation of each of the target objects into the cross-attention module and the diffusion model, so that the diffusion model performs denoising based on the fused features output by the cross-attention module and outputs predicted visual content; wherein, the diffusion model includes N layers of feature sampling layers; the cross-attention module includes a first attention network and a second attention network; the first attention network is configured to output a corresponding first fused feature according to the second local text feature representation and the sampled features output by the diffusion model; the second attention network is configured to output a second fused feature according to the second local visual feature representation and the first fused feature, and the second fused feature is used to be input into the diffusion model for feature sampling; wherein, the diffusion model performs denoising based on the fused features output by the cross-attention module and outputs predicted visual content, including: performing cross-attention calculation on the second local text feature representation and the output features of the k-th layer of the diffusion model through the first attention network to obtain the k-th first fused feature; wherein, 1 ≤ k ≤ N, and both k and N are natural numbers; performing cross-attention calculation on the first fused feature and the second local visual feature representation through the second attention network to obtain the k-th second fused feature; wherein, the k-th second fused feature is used to be input into the (k + 1)-th layer of the diffusion model to obtain the output features of the (k + 1)-th layer; when k + 1 = N, the diffusion model obtains the output features of the current time step and generates predicted visual content based on the output features of the last time step; A network parameter iteration module, which is configured to fix the network parameters of each layer of the diffusion model, and perform backpropagation according to the loss function of the diffusion model to iteratively adjust the network parameters of the cross-attention module and the position encoding module, so as to obtain a trained visual content generation model.

9. A visual content generation device for multiple objects, characterized in that, Comprising: A data acquisition module, which is configured to acquire P-frame original images and prompt texts, and the total number of target objects in each frame of the original images is Y, where Y ≥ P and P ≥ 1; A visual generation module, which is configured to generate corresponding target visual content according to the visual content generation model obtained by the training method according to any one of claims 1 to 5, and the attribute information of the target object in the target visual content is consistent with that in the original image.

10. An electronic device, characterized in that, Comprising: A processor; And A memory, on which executable code is stored, and when the executable code is executed by the processor, the processor is caused to execute the training method of the visual content generation model for multiple objects according to any one of claims 1-5 or the visual content generation method for multiple objects according to any one of claims 6-7.

11. A computer-readable storage medium, on which executable code is stored, and when the executable code is executed by a processor of an electronic device, the processor is caused to execute the training method of the visual content generation model for multiple objects according to any one of claims 1-5 or the visual content generation method for multiple objects according to any one of claims 6-7.

Citation Information

Patent Citations

  • Training method of image generation model, image generation method, device and equipment

    CN117351115A

  • Object combination image generation method and device, electronic equipment and storage medium

    CN117808932A