Text-driven visual content generation model training method and device

By introducing conditional loss and unconditional loss in the training process of text-driven visual content generation model, and combining correction weights and power-normal distance optimization loss function, the problem of insufficient matching degree in the text description object combination scenario is solved, and higher matching degree of visual content and text and model generalization ability are achieved.

CN120496096APending Publication Date: 2025-08-15BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510622339.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing text-driven visual content generation model has poor matching between the generated visual content and the text input by the user when generating visual content, especially when facing a combined scene of text description objects.

Method used

By constructing a combination of objects in the training data, combining the introduction of conditional loss and unconditional loss during the diffusion model training process, the first loss and the second loss are used to update the model parameters. The first loss characterizes the difference between the noise added at the target time step and the predicted first noise, the second loss characterizes the difference between the added noise and the unconditional predicted second noise, and optimizes the loss function by correcting the weight and power norm distance, and balancing the general characteristics and conditional guidance generation of the model.

Benefits of technology

The matching degree between the visual content generated by the visual content generation model and the input text is improved, the generalization ability of the model and the rationality and accuracy of the generated visual content are enhanced, the model overfitting text conditions is reduced, and the quality of the generated visual content is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496096A_ABST
    Figure CN120496096A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a text-driven visual content generation model training method and device. According to the main technical scheme, the method comprises the steps of obtaining training data comprising a plurality of training samples; training a diffusion model by using the training data to obtain a visual content generation model; wherein the training comprises the following steps: adding noise of T time steps to a visual content sample to obtain first noise visual content; inputting the first noise visual content into a diffusion model, obtaining first noise which is respectively predicted for the T time steps when the diffusion model takes the text sample as a condition, and obtaining second noise which is respectively predicted for the T time steps when the diffusion model is unconditional; and obtaining a value of a loss function by using the first loss and the second loss, and updating model parameters of the diffusion model by using the value of the loss function. According to the method, the matching degree between the visual content of the visual content generation model and the text input by the user can be improved for the combined scene of the text description object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to a training method and device for a text-driven visual content generation model. Background Art

[0002] With the rapid development of artificial intelligence (AI), AI Generated Content (AIGC) models have made significant progress. Through their ability to interpret and reorganize real-world objects, AIGC models are able to create objects and scenes that do not exist in the physical world.

[0003] Text-driven visual content generation models can generate corresponding visual content based on the text input by users. However, in practical applications, the matching degree between the visual content generated by current visual content generation models and the text input by users still needs to be improved, especially when the text input by users describes a combination of objects, the matching degree between the generated visual content and the text input by users is poor. Summary of the Invention

[0004] In view of this, the present application provides a training method and device for a text-driven visual content generation model, which is used to improve the matching degree between the visual content of the visual content generation model and the text input by the user for the combination scenario of text description objects.

[0005] This application provides the following solutions:

[0006] In a first aspect, a training method for a text-driven visual content generation model is provided, the method comprising: obtaining training data comprising a plurality of training samples, the training samples comprising: a text sample and a visual content sample, the text sample being used to describe a combination of at least two objects, and the visual content sample displaying a combination of at least two objects; training a diffusion model using the training data to obtain a visual content generation model; wherein the training comprises: adding noise of T time steps to the visual content sample to obtain a first noisy visual content, where T is a positive integer; inputting the first noisy visual content into the diffusion model, obtaining the first noise predicted by the diffusion model for the T time steps when the text sample is used as a condition, and obtaining the second noise predicted by the diffusion model for the T time steps when the diffusion model is unconditional; obtaining the value of a loss function using the first loss and the second loss, and updating the model parameters of the diffusion model using the value of the loss function; wherein the first loss represents a first difference between the noise added at a target time step and the first noise predicted for the target time step, and the second loss represents a second difference between the noise added at the target time step and the second noise predicted for the target time step, where the target time step is at least one of the T time steps.

[0007] Optionally, obtaining the value of the loss function using the first loss and the second loss includes: determining the first loss using the corrected weight and the first difference, and determining the second loss using the second difference, wherein the corrected weight is determined based on the second difference; and using a preset weight coefficient to weight the first loss and the second loss to obtain the value of the loss function.

[0008] Optionally, the correction weight is determined in the following manner: based on a preset power value, the power norm distance corresponding to the second difference is determined, and the power norm distance is marked as a stop gradient; and the correction weight is determined using the power norm distance corresponding to the second difference.

[0009] Optionally, the power value is determined based on the data balance of the training data.

[0010] Optionally, the method also includes: obtaining an evaluation data set including multiple evaluation texts, the evaluation texts are used to describe a combination of at least two objects; using the evaluation data set to evaluate the trained visual content generation model to obtain evaluation results, the evaluation including: inputting the second noisy visual content and the evaluation text into the visual content generation model, obtaining the predicted visual content obtained after the visual content generation model denoises the second noisy visual content using the evaluation text as a condition, and obtaining the evaluation results based on the predicted visual content.

[0011] Optionally, obtaining an evaluation data set including multiple evaluation texts includes: extracting a candidate object combination from an object data set, the candidate object combination including at least a first object and a second object, the frequency of occurrence of the first object in the object data set is greater than or equal to a first frequency threshold, the frequency of occurrence of the second object in the object data set is less than or equal to a second frequency threshold, and the first frequency threshold is greater than the second frequency threshold; using the candidate object combination, obtaining evaluation texts to constitute an evaluation data set.

[0012] Optionally, using candidate object combinations to obtain evaluation texts to form an evaluation data set includes: extracting a target object combination from the candidate object combination, where the frequency of occurrence of the first object and the second object in the target object combination in the same text of the object data set meets a preset condition; calling a large language model to generate at least one description text for each target object combination as an evaluation text.

[0013] In a second aspect, a text-driven visual content generation method is provided, the method comprising: obtaining a target text, the target text being used to describe a combination of at least two objects; and using a visual content generation model to generate target visual content corresponding to the target text based on the target text, the visual content generation model being trained using the method of the first aspect, and the target visual content showing a combination of at least two objects.

[0014] In a third aspect, a training device for a text-driven visual content generation model is provided, the device comprising: a first acquisition unit, configured to acquire training data comprising a plurality of training samples, the training samples comprising: a text sample and a visual content sample, the text sample being used to describe a combination of at least two objects, and the visual content sample displaying a combination of at least two objects; a model training unit, configured to train a diffusion model using the training data to obtain a visual content generation model; wherein the training comprises: adding noise of T time steps to the visual content sample to obtain a first noisy visual content, where T is a positive integer; inputting a noisy visual content into the diffusion model, obtaining a first noise predicted by the diffusion model for each of the T time steps when the text sample is used as a condition, and obtaining a second noise predicted by the diffusion model for each of the T time steps when the diffusion model is unconditional; obtaining a value of a loss function using a first loss and a second loss, and updating a model parameter of the diffusion model using the value of the loss function; wherein the first loss represents a first difference between the noise added at a target time step and the first noise predicted for the target time step, and the second loss represents a second difference between the noise added at a target time step and the second noise predicted for the target time step, where the target time step is at least one of the T time steps.

[0015] In a fourth aspect, a text-driven visual content generation device is provided, which includes: a second acquisition unit, configured to acquire a target text, where the target text is used to describe a combination of at least two objects; a content generation unit, configured to use a visual content generation model to generate target visual content corresponding to the target text based on the target text, where the visual content generation model is trained using the device of the third aspect, and the target visual content shows a combination of at least two objects.

[0016] In a fifth aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed, the steps of the method of any one of the first aspect or the second aspect are implemented.

[0017] In a sixth aspect, an electronic device is provided, comprising: one or more processors; and a memory associated with the one or more processors, the memory being used to store program instructions, which, when read and executed by the one or more processors, execute the steps of the method of any one of the first or second aspects above.

[0018] In a seventh aspect, a computer program product is provided, comprising a computer program, which implements the steps of any one of the methods in the first or second aspect when executed.

[0019] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0020] 1) This application constructs training data after combining objects. In the training process based on the diffusion model, not only the conditional loss (i.e., the first loss obtained based on the first difference between the added noise and the first noise predicted when the text sample is used as the condition) is considered, but also the unconditional loss (the second loss obtained based on the second difference between the added noise and the second noise predicted when it is unconditional) is further introduced to avoid the model's excessive reliance on specific vocabulary in the text conditions, which causes the model to ignore the overall semantics and only match the local features of the text, resulting in the generated visual content not matching the object combination expressed in the actual text. The introduction of unconditional loss can make the model better analogous to humans, learning both the global features and underlying distribution of visual content distribution to improve the rationality of visual content generation, and learning the local features of object combinations under specific conditions. By balancing the "general feature distribution" and "condition-guided generation", the overall matching degree between the generated visual content and the input text is improved.

[0021] 2) In this application, the first loss can constrain the model's process of generating visual content with text as a condition to improve the accuracy of the generated visual content. The second loss can enable the model to better understand the global features and general distribution of the visual content and improve the rationality of the visual content. The first loss is obtained by correcting the first difference using the unconditional second difference, which can better balance the rationality of visual content generation and dependence on text conditions. It can not only ensure the rationality of the generated visual content, but also make the generated visual content consistent with the input text. In addition, it can also enhance the generalization ability of the visual content generation model. Moreover, the correction weight is determined based on the second difference, and the degree of influence of the first difference and the second difference on the model can be balanced based on the model's unconditional performance, that is, the model's "general feature distribution" and "condition-guided generation" capabilities can be adaptively balanced.

[0022] 3) When determining the power norm distance corresponding to the second difference, the present application marks the power norm distance as a stop gradient to stop the back propagation of the gradient. On the one hand, it can improve the stability of the model, and on the other hand, it can save computing resources and improve the training efficiency of the model.

[0023] 4) This application determines the power value used for the power norm distance by the balance of the training data, which can adaptively adjust the sensitivity of the distance metric to different samples or features during the training of the diffusion model, thereby improving the model's attention to the features of long-tail objects and further improving the consistency between the visual content generated by the model and the input text.

[0024] 5) This application evaluates the visual content generation model by constructing an evaluation dataset including multiple evaluation texts. On the one hand, the evaluation results can measure the effectiveness of the visual content generation model. On the other hand, the evaluation results can be further used to optimize the visual content generation model, thereby further improving the quality of the generated visual content and improving the matching degree between the generated visual content and the input text.

[0025] 6) This application extracts a first object with a higher frequency of occurrence and a second object with a lower frequency of occurrence from an object dataset to form a candidate object combination, and uses these candidate object combinations to determine an evaluation dataset. This evaluation dataset can focus on the tail objects in the object combination that pose a challenge to the model, thereby truly reflecting the effect of the model in generating visual content in real-world scenes.

[0026] 7) This application can extract the target object combination with the smallest frequency from the candidate object combination by determining the frequency of the first object and the second object in the candidate object combination appearing in the same text of the object data set, thereby increasing attention to the tail combination in the object combination to solve the problem of low evaluation accuracy of the evaluation data set due to uneven data distribution.

[0027] Of course, any product implementing the present application does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0029] Figure 1 is a system architecture diagram applicable to the embodiments of the present application;

[0030] Figure 2 A flowchart of a training method for a text-driven visual content generation model provided in an embodiment of the present application;

[0031] Figure 3 A schematic diagram of generating an image using a visual content generation model provided in an embodiment of the present application;

[0032] Figure 4 A schematic diagram of another method of generating an image using a visual content generation model provided in an embodiment of the present application;

[0033] Figure 5 A schematic diagram of another method of generating an image using a visual content generation model provided in an embodiment of the present application;

[0034] Figure 6 A schematic diagram of another method of generating an image using a visual content generation model provided in an embodiment of the present application;

[0035] Figure 7 A flowchart of a text-driven visual content generation method provided in an embodiment of the present application;

[0036] Figure 8 A schematic block diagram of a training device for a text-driven visual content generation model provided in an embodiment of the present application;

[0037] Figure 9 A schematic block diagram of a text-driven visual content generation device provided in an embodiment of the present application;

[0038] Figure 10 A schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0039] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0040] The terms used in the embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "an", "the" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.

[0041] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.

[0042] The word "if," as used herein, may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.

[0043] First, let’s explain the terms used in this application:

[0044] Objects, also known as "concepts," are an important foundation for models to understand and generate visual content. They can refer to things, scenes, and so on. For example, things refer to specific physical objects, such as "cat," "car," and "house," while scenes refer to overall environments or situations, such as "beach," "forest," and "city streets." Objects can further include the attributes and actions of things. Attributes can include colors (such as "red" and "blue"), shapes (such as "round" and "square"), and sizes (such as "big" and "small"). Actions can include actions such as "run," "jump," and "fly."

[0045] Visual content: refers to various information perceived through the visual senses, such as images (including static visual content such as photographs, paintings, and illustrations), videos (dynamic visual content composed of a series of continuous image frames that can record and display dynamic scenes and events), animation (creating imaginative and creative visual content through dynamic drawing and design of virtual characters or objects), graphic design (visual content that achieves specific visual communication purposes through the combination and creative design of elements such as text, images, and colors), virtual reality (VR) and augmented reality (AR) content (using special equipment to immerse users in a virtual environment or superimpose virtual elements on the real world to form visual content that users can interact with).

[0046] Diffusion model: By gradually adding noise to the data, and then learning the ability to recover the original data from the noise, its core includes the forward diffusion process, the reverse generation process and the learning mechanism based on the neural network.

[0047] Forward Diffusion Process: Sampling the initial data point x0 from the real data distribution, and then in each time step t, according to a certain probability distribution q(x t |x t-1 ) adds noise to the data, gradually transforming the data point x0 into a noise distribution q(x T ), usually this noise process is modeled as Gaussian noise, whose variance gradually increases with time step t, and finally x T Obey a simple Gaussian distribution.

[0048] Reverse Generation Process: Take a known prior distribution (usually a Gaussian distribution) as the starting point of the generation process, such as x T Then, by learning an inverse diffusion process p θ (x t-1 |x t), gradually recovering the approximation of the original data from the noise. This process is a Markov chain, where θ is the parameter of the model, which needs to be learned through training. At each time step t, the model is based on the current noise data x t Predict the data x at the previous time step t-1 The distribution of , and then sample from this distribution to get x t-1 , repeat this process until a sample x0 close to the true data distribution is generated.

[0049] Neural network-based learning mechanism: Neural networks (such as U-Net, etc.) are usually used to parameterize the inverse diffusion process p θ (x t-1 |x t During training, the noise at the target time step is sampled from the forward diffusion process, and the noise actually predicted for the target time step is sampled from the backward generation process. The neural network parameters θ are updated by minimizing the difference between the predicted noise and the actual added noise (i.e., the noise at the target time step sampled from the forward diffusion process) (e.g., a loss function such as mean squared error). After a large amount of training data and repeated steps, the model gradually learns how to accurately recover the original data from the noise, and is able to generate new samples that are similar to the training data.

[0050] In order to facilitate the understanding of this application, the system architecture on which this application is based is first described. Figure 1 This is a system architecture diagram applicable to the embodiments of the present application, such as Figure 1 As shown in , the system architecture may include: a user device, a visual content generation device located on the server side, and a device for training a visual content generation model.

[0051] The user equipment and the server can communicate with each other. The user equipment and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in this application.

[0052] User devices may include, but are not limited to, smart mobile terminals, smart home devices, wearable devices, and personal computers (PCs). Smart mobile devices may include mobile phones, tablets, laptops, PDAs (Personal Digital Assistants), and internet-connected cars. Smart home devices may include smart TVs and smart refrigerators. Wearable devices may include smart watches, smart glasses, virtual reality devices, augmented reality devices, and mixed reality devices (i.e., devices that support both virtual reality and augmented reality).

[0053] A server can be a standalone server, a server cluster, or even a cloud server. A cloud server, also known as a cloud computing server or cloud host, is a hosting product within the cloud computing service ecosystem. It addresses the management difficulties and limited scalability of traditional physical hosting and virtual private server (VPS) services.

[0054] Before performing a visual content generation task, the apparatus for training a visual content generation model may train a diffusion model using the method provided in an embodiment of the present application to obtain a trained visual content generation model.

[0055] A user can input target text through a user device. The user device then sends the target text along with a visual content generation request over the network to a server-side visual content generation device. The visual content generation device uses a trained visual content generation model to generate target visual content corresponding to the target text and returns this target visual content to the user device over the network. The user device then displays the received target visual content to the user.

[0056] Apart from Figure 1 In addition to the shown architecture, a computer terminal device with strong computing power can also use the method provided in the embodiment of the present application to train the visual content generation model and / or generate target visual content corresponding to the target text.

[0057] It should be understood that Figure 1 The number of user devices and server ends in the embodiment is only for illustration. Any number of user devices and server ends may be provided according to implementation requirements.

[0058] It should be noted that the limitations such as "first" and "second" involved in the present disclosure do not have restrictions on size, order and quantity, but are only used to distinguish them in name. For example, "first noise visual content" and "second noise visual content" are used to distinguish two types of noise visual content, and "first loss" and "second loss" are used to distinguish two types of losses, and so on.

[0059] Traditionally, training of text-driven visual content generation models typically uses only conditional loss, which constructs a loss function based on the difference between the noise added at the target time step and the noise predicted by the diffusion model using the text sample as a condition for the target time step. This approach causes the model to over-rely on specific vocabulary in the text, leading to overfitting of the model to the text conditions, which can lead to severe semantic drift, including object loss (for example, when multiple objects are described in the text, some of them are lost in the generated visual content), attribute leakage (for example, the attributes of an object described in the text are mistakenly "leaked" to other objects in the generated visual content), and object entanglement (for example, in the generated visual content, objects described in the text are mistakenly generated as other objects, or multiple objects are mistakenly merged or distorted, resulting in abnormal spatial relationships or morphology).

[0060] In view of this, this application provides a new approach to training visual content generation models. Figure 2 A flowchart of a training method for a text-driven visual content generation model provided in an embodiment of the present application. The method can be performed by Figure 1 The user equipment in the system shown in FIG. Figure 2 As shown in , the method may include the following steps:

[0061] Step 201: Acquire training data including a plurality of training samples, where the training samples include: text samples and visual content samples, where the text samples are used to describe a combination of at least two objects, and the visual content samples display a combination of at least two objects.

[0062] Step 202: Use the training data to train the diffusion model to obtain a visual content generation model.

[0063] The training includes:

[0064] Add T time steps of noise to the visual content sample to obtain a first noisy visual content, where T is a positive integer; input the first noisy visual content into a diffusion model to obtain the first noise predicted by the diffusion model for T time steps when the text sample is used as a condition, and obtain the second noise predicted by the diffusion model for T time steps when the diffusion model is unconditional; use the first loss and the second loss to obtain the value of the loss function, and use the value of the loss function to update the model parameters of the diffusion model; wherein the first loss represents a first difference between the noise added at the target time step and the first noise predicted for the target time step, and the second loss represents a second difference between the noise added at the target time step and the second noise predicted for the target time step, and the target time step is at least one of the T time steps.

[0065] As can be seen from the above process, the present application constructs training data after combining objects. In the training process based on the diffusion model, not only the conditional loss is considered (i.e., the first loss obtained based on the first difference between the added noise and the first noise predicted when the text sample is used as a condition), but also the unconditional loss is further introduced (the second loss obtained based on the second difference between the added noise and the second noise predicted when it is unconditional) to avoid the model's excessive reliance on specific vocabulary in the text conditions, which causes the model to ignore the overall semantics and only match the local features of the text, resulting in the generated visual content not matching the object combination expressed in the actual text. The introduction of unconditional loss can enable the model to better resemble humans, learning both the global features and underlying distribution of visual content distribution to improve the rationality of visual content generation, and learning the local features of object combinations under specific conditions. By balancing the "general feature distribution" and "condition-guided generation", the overall matching between the generated visual content and the input text is improved.

[0066] The following describes in detail the steps in the above process and the effects that can be further produced in conjunction with the embodiments.

[0067] First, the above step 201, namely "obtaining training data including multiple training samples, the training samples including: text samples and visual content samples, the text samples are used to describe the combination of at least two objects, and the visual content samples show the combination of at least two objects" is described in detail in conjunction with the embodiment.

[0068] In an embodiment of the present application, a text sample is used to describe a combination of at least two objects, such as "a red bird flying in the sky", "glowing mushrooms in the forest at night", "a sunny park with green grass, colorful flowers and white benches", etc.; a visual content sample shows a combination of at least two objects, thereby showing an effect that conforms to the text description.

[0069] Next, the above step 202, namely "using training data to train a diffusion model to obtain a visual content generation model", is described in detail with reference to an embodiment.

[0070] During the forward diffusion process, for each visual content sample in the training sample, T time steps of noise are added to obtain a first noisy visual content, where T is a positive integer. In an embodiment of the present application, the noise added at the target time step can be sampled, for example, as , where the target time step can be one or more of the T time steps. The target time step can be pre-set or randomly selected.

[0071] In one embodiment, the noise added to the visual content sample can be sampled from a standard normal distribution. The time step can be sampled from a uniform distribution (e.g., a 0-1 uniform distribution) or a non-uniform distribution (e.g., a Gaussian distribution).

[0072] Add T time steps of noise to the visual content sample to obtain the first noisy visual content, where for each time step t, the visual content x obtained after adding noise ∈ t It can be expressed in the following form:

[0073]

[0074] Among them, x0 represents the visual content sample; represents the noise factor associated with the time step number t, and As t increases monotonically or decreases monotonically. Alternatively, It can be obtained from the noise planner. Noise is added to each time step t, and after reaching time step T, the first noise visual content x is obtained. T .

[0075] In the reverse generation process, the first noise visual content is input into the diffusion model, and the first noise predicted by the diffusion model for T time steps when the text sample is used as a condition is obtained. The second noise predicted by the diffusion model for T time steps is also obtained when the diffusion model is unconditioned. The diffusion model can be implemented using a neural network such as U-Net.

[0076] Alternatively, if θ is used to represent the diffusion model, then for the t-th time step, the first noise predicted by the model can be expressed as ∈ θ (x t ,y,t), y represents the text sample; the second noise can be expressed as Indicates a null condition.

[0077] Then, the first loss and the second loss are used to obtain the value of the loss function, and the value of the loss function is used to update the model parameters of the diffusion model. The first loss represents the first difference between the noise added at the target time step and the first noise predicted for the target time step, and the second loss represents the second difference between the noise added at the target time step and the second noise predicted for the target time step. In other words, the first loss can be regarded as the "conditional loss" obtained when there is text as a condition, and the second loss can be regarded as the "unconditional loss" obtained when there is no condition as a guide. Under unconditional guidance, the model can learn the global features and underlying distribution of visual content distribution (such as color, shape, texture, etc.), thereby improving the rationality of visual content generation.

[0078] As one of the feasible ways, the first loss can be determined by using the first difference, the second loss can be determined by using the second difference, and then the first loss and the second loss can be weighted using a preset weight coefficient to obtain the value of the loss function.

[0079] As another more preferred implementation method, the first loss can be determined using the corrected weight and the first difference, the second loss can be determined using the second difference, and then the first loss and the second loss can be weighted using a preset weight coefficient to obtain the value of the loss function.

[0080] For example, the first loss can be expressed as follows:

[0081] L * =D||∈―∈ θ (x t ,y,t)|| 2

[0082] Among them, L * represents the first loss; D represents the modified weight; ∈―∈ θ (x t ,y,t) represents the first difference.

[0083] The second loss can be expressed as follows:

[0084]

[0085] Among them, L u Indicates the second loss; Indicates the second difference.

[0086] In one embodiment, the modified weight D can be determined based on the second difference. That is, the influence of the first and second differences on the model is balanced based on the model's unconditional performance, adaptively balancing the model's "general feature distribution" and "condition-guided generation" capabilities.

[0087] In a preferred embodiment, the above-mentioned correction weight can be determined as follows: based on a preset power value, the power norm distance corresponding to the second difference is determined, and the power norm distance is marked as stopping the gradient calculation, so that during the backpropagation process, the power norm distance is not involved in the gradient calculation. Then, the power norm distance corresponding to the second difference is used to determine the correction weight. It can be expressed as follows:

[0088]

[0089] Here, γ represents the power value, a preset hyperparameter determined based on the data balance of the training data. The more balanced the training data, the smaller the value of γ, and the more unbalanced the training data, the larger the value of γ. γ can be selected from empirical or experimental values based on the data balance of the training data. Preferably, 0 < γ < 10. For example, γ can be set to 0.8 to reduce color shift in the generated visual content. sg indicates that the gradient calculation of D is stopped, that is, the gradient propagation of D is stopped during the backpropagation process. This can improve the stability of the model on the one hand, and save computing resources on the other hand, improving the training efficiency of the model on the other.

[0090] The value of the loss function can be expressed in the following form:

[0091] L=λL * +(1―λ)L u

[0092] Here, L represents the loss function, which is essentially the Imbalance in Balance Loss (IMBA Loss); λ represents the preset weight coefficient. By setting λ, the diffusion model can be trained in the absence of some conditional information, so that the diffusion model will not overly rely on the text samples used as conditions during training, thereby enhancing the generalization ability of the trained visual content generation model. At the same time, it can also enable the visual content generation model to learn more common features and patterns, improving its adaptability to different conditions. In addition, the loss function obtained by combining the conditional first loss and the unconditional second loss can improve the degree of match between the visual content generated by the trained visual content generation model and the input text.

[0093] Preferably, 0<λ<1, and λ=0.9 can be set to make the generalization ability of the visual content generation model better.

[0094] In one embodiment, after obtaining the loss function, the gradient can be calculated and the model parameters can be updated using a method such as gradient descent during backpropagation until the training termination condition is met. The training termination condition may include, for example, the loss function value being less than or equal to a preset loss function threshold, or the number of iterations reaching a preset threshold. This approach allows the diffusion model to gradually learn the distribution and characteristics of the training data during training, thereby enabling the trained visual content generation model to generate high-quality visual content.

[0095] After training completes and a visual content generation model is obtained, the visual content generation model can be evaluated using an evaluation dataset. In one embodiment, an evaluation dataset comprising multiple evaluation texts can be obtained, wherein the evaluation texts are used to describe a combination of at least two objects. The trained visual content generation model is evaluated using the evaluation dataset to obtain an evaluation result. The evaluation can include: inputting the second noisy visual content and the evaluation texts into the visual content generation model, obtaining predicted visual content obtained by denoising the second noisy visual content using the evaluation texts as a condition, and obtaining an evaluation result based on the predicted visual content.

[0096] The second noise visual content can be random noise visual content or preset noise visual content. The evaluation results can include a similarity score (Contrastive Language-Image Pretraining Score, CLIP Score) between the predicted visual content and the evaluation text, a visual question answering (VQA) generation success rate, and the color, shape, texture, non-spatial semantic information, and spatial semantic information of the predicted visual content.

[0097] In some embodiments, after obtaining the predicted visual content, the similarity between the feature vectors of the predicted visual content and the review text can be calculated as a CLIP Score. For example, the predicted visual content and the review text are mapped to a shared semantic embedding space using the CLIP model, respectively obtaining a feature vector for the predicted visual content and a feature vector for the review text. The CLIP Score is then used to calculate the similarity between the feature vectors of the predicted visual content and the feature vectors of the review text. The CLIP Score measures the degree of semantic alignment between the predicted visual content and the review text.

[0098] In some embodiments, after obtaining the predicted visual content, questions about the predicted visual content can be asked, and predicted answers can be obtained based on the information of the predicted visual content and the questions. The proportion of consistent predicted answers and true answers is used as the VQA generation success rate to evaluate the performance of the visual content generation model.

[0099] In some embodiments, after obtaining the predicted visual content, the performance of the visual content generation model can be evaluated by analyzing the color, shape, texture, non-spatial, spatial, etc. of the predicted visual content. Among them, color is used to measure the accuracy, consistency and rationality of the color in the predicted visual content, such as whether the color conforms to the inherent properties of the real object, whether the distribution is natural and harmonious, etc. Shape is used to evaluate the correctness of the geometric structure and contour of the object in the generated image, such as whether the contour is complete and without distortion, whether the details of the structure are accurate, whether the spatial logic of the shape is reasonable, etc. Texture is used to measure the authenticity of the surface details and texture of the object, such as the fineness of the surface texture, the restoration of the material characteristics, the degree of fit with the shape, etc. Non-spatial is used to evaluate semantic attributes that do not depend on spatial position or structure, focusing on the inherent characteristics or abstract attributes of the object, such as category and semantic correctness, non-spatial relationships, etc. Spatial is used to evaluate the position and arrangement of objects in space, as well as the spatial relationship and depth level between objects. These five indicators can be evaluated manually by scoring, or they can be automatically evaluated by combining computer vision algorithms to extract features.

[0100] In some preferred embodiments, when obtaining an evaluation data set including multiple evaluation texts, a candidate object combination can be extracted from the object data set, where the candidate object combination includes at least a first object and a second object, the frequency of occurrence of the first object in the object data set is greater than or equal to a first frequency threshold, the frequency of occurrence of the second object in the object data set is less than or equal to a second frequency threshold, and the first frequency threshold is greater than the second frequency threshold; using the candidate object combination, the evaluation text is obtained to constitute the evaluation data set.

[0101] An object dataset may include multiple texts, each describing a combination of at least two objects. However, the texts included in the object dataset may be large in size and unevenly distributed. To reduce data complexity and address the uneven data distribution in the evaluation dataset, the evaluation dataset focuses not only on the head objects but also on the tail objects. Head objects and tail objects can be extracted from the object dataset. Then, the same number of representative objects are selected from the head objects and the tail objects, respectively, to obtain a first object and a second object. The first object and the second object are then combined to obtain a candidate object combination that includes both the head object and the tail object.

[0102] By combining candidate objects, we obtain evaluation texts to form an evaluation dataset, and then use this evaluation dataset to evaluate the visual content generation model, which can obtain more accurate evaluation results. The evaluation results can be further used to optimize the visual content generation model, improve the object combination ability of the visual content generation model, improve the quality of the generated visual content, and improve the matching degree between the generated visual content and the input text.

[0103] The frequency of the first object appearing in the object dataset is greater than the frequency of the second object appearing in the object dataset. For example, the ratio of the frequency of the first object appearing in the object dataset to the frequency of the second object appearing in the object dataset can be 100.

[0104] In one embodiment, using candidate object combinations to obtain evaluation texts to form an evaluation data set includes: extracting a target object combination from the candidate object combination, wherein the frequency of occurrence of the first object and the second object in the target object combination in the same text of the object data set meets a preset condition; and calling a large language model to generate at least one description text for each target object combination as the evaluation text.

[0105] The preset conditions may include: the first object and the second object are ranked in ascending order of frequency of occurrence in the same text of the object dataset, where k is a preset positive integer. This allows the evaluation dataset to focus more on unusual object combinations. For example, if the first object is "wings" and the second object is "car," the combination of "wings" and "car" is an unusual object combination; if the first object is "coffee" and the second object is "child," the combination of "coffee" and "child" is an unusual object combination, and so on.

[0106] The above-mentioned preset condition may also include: the frequency of the first object and the second object appearing in the same text of the object data set is less than a preset third frequency threshold. In addition, other conditions may also be used.

[0107] Then, the large language model is called to generate at least one description text as evaluation text for each target object combination, thereby obtaining an evaluation dataset.

[0108] For example, using the words "wings" and "car," the large language model could generate descriptions like "a car with colorful feather wings, driving towards the clouds" or "transparent wings on the top of the car." For example, using the words "coffee" and "child," the large language model could generate descriptions like "a child sitting at a table with a cup of coffee on it" or "a child walking down the street holding a cup of coffee," and so on.

[0109] If the objects in the evaluation text belong to an unconventional or uncommon object combination, and the quality of the visual content generated by the visual content generation model based on the evaluation text is high, then for the case where the objects in the evaluation text belong to a regular object combination, the quality of the visual content generated by the visual content generation model will be higher. Therefore, by evaluating the visual content generation model through unconventional object combinations, the visual content generation model can have better visual content generation capabilities.

[0110] After experiments, the visual content generation model trained by the above method provided in the embodiment of the present application (referred to as visual content generation model 1) was compared with the visual content generation model trained by the traditional method (referred to as visual content generation model 2): on the LC-Mis evaluation dataset, the CLIP Score was improved from 0.3045 to 0.3121, and the VQA generation success rate was improved from 46.21% to 62.89%; on the T2I-CompBench evaluation dataset, the color was improved from 0.5812 to 0.7067, the shape was improved from 0.4307 to 0.5151, the texture was improved from 0.6188 to 0.6861, the non-spatial was improved from 0.3041 to 0.3071, and the spatial was improved from 0.1966 to 0.2518; on the Inert-CompBench evaluation dataset, the CLIP The score increased from 0.3194 to 0.3229, and the VQA generation success rate increased from 44% to 57%.

[0111] Select several evaluation texts and compare the visual content generation model 1 with the visual content generation model 2. Figures 3 to 6 , Figure 3 A schematic diagram of generating an image using a visual content generation model provided in an embodiment of the present application is provided. Figure 4 This is another schematic diagram of generating an image using a visual content generation model provided in an embodiment of the present application. Figure 5 This is another schematic diagram of generating an image using a visual content generation model provided in an embodiment of the present application. Figure 6 This is another schematic diagram of generating an image using a visual content generation model provided in an embodiment of the present application. Figures 3 to 5 (a) is the image generated by the visual content generation model 2. Figures 3 to 5 (b) is the image generated by the visual content generation model 1. Figure 6 (a), (c), and (e) are images generated by the visual content generation model 2. Figure 6 (b), (d), and (f) are images generated by the visual content generation model 1:

[0112] Evaluation text 1: Feathers and fur intermingling on a beaver's sleek body. The images generated by visual content generation model 1 and visual content generation model 2 are as follows: Figure 3 As shown in (b) and (a), it can be seen that the visual content generation model 1 significantly improves the problem of object missing.

[0113] Evaluation text 2: A ball is on a square table and a cube is on a round table. The images generated by the visual content generation model 1 and the visual content generation model 2 are as follows: Figure 4 As shown in (b) and (a), it can be seen that the visual content generation model 1 significantly improves the problem of attribute leakage.

[0114] Evaluation text 3: A pair of shoes is running on the road. The images generated by the visual content generation model 1 and the visual content generation model 2 are as follows: Figure 5 As shown in (b) and (a), it can be seen that the visual content generation model 1 significantly improves the problem of object entanglement.

[0115] Evaluation text 4: A small bathroom with a small white toilet next to a brown sink. This evaluation text 4 comes from the LC-Mis evaluation dataset. The images generated by the visual content generation model 1 and the visual content generation model 2 are as follows: Figure 6 As shown in (b) and (a), it is obvious that the visual images generated by the visual content generation model 1 are more consistent with the evaluation text.

[0116] Evaluation text 5: A circular dining table and a triangular table runner. This evaluation text comes from the T2I-CompBench evaluation dataset. The images generated by the visual content generation model 1 and the visual content generation model 2 are as follows: Figure 6 As shown in (d) and (c), it is obvious that the visual images generated by the visual content generation model 1 are more consistent with the evaluation text.

[0117] Evaluation text 6: A sliding door painted with giant onigiri designs. This evaluation text comes from the Inert-CompBench evaluation dataset. The images generated by the visual content generation model 1 and the visual content generation model 2 are as follows: Figure 6 As shown in (f) and (e), it is obvious that the visual images generated by the visual content generation model 1 are more consistent with the evaluation text.

[0118] Figure 7 This is a flowchart of a text-driven visual content generation method provided in an embodiment of the present application. The method can be performed by Figure 1 The user equipment in the system shown in FIG. Figure 7 As shown in , the method may include the following steps:

[0119] Step 701: Acquire target text, where the target text is used to describe a combination of at least two objects.

[0120] Step 702: Generate target visual content corresponding to the target text based on the target text using a visual content generation model, wherein the visual content generation model is trained using the aforementioned text-driven visual content generation model training method, and the target visual content displays a combination of at least two objects.

[0121] Specifically, the noisy visual content and the target text can be input into the visual content generation model. The visual content generation model takes the target text as a condition and performs denoising processing on the noisy visual content for T time steps to obtain the target visual content.

[0122] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0123] According to another embodiment, a training apparatus for a text-driven visual content generation model is provided. Figure 8 A schematic block diagram of a training device for a text-driven visual content generation model provided in an embodiment of the present application, wherein the device is provided at Figure 1 User equipment in the architecture shown. Figure 8 As shown, the device 800 mainly includes: a first acquisition unit 801 and a model training unit 802. The main functions of each component unit are as follows:

[0124] The first acquisition unit 801 is configured to acquire training data including a plurality of training samples, where the training samples include: text samples and visual content samples, where the text samples are used to describe a combination of at least two objects, and the visual content samples show a combination of at least two objects.

[0125] The model training unit 802 is configured to train a diffusion model using training data to obtain a visual content generation model; wherein the training includes: adding T time steps of noise to the visual content sample to obtain a first noisy visual content, where T is a positive integer; inputting a noisy visual content into the diffusion model, obtaining the first noise predicted by the diffusion model for T time steps when the diffusion model takes the text sample as a condition, and obtaining the second noise predicted by the diffusion model for T time steps when the diffusion model is unconditional; obtaining the value of the loss function using the first loss and the second loss, and updating the model parameters of the diffusion model using the value of the loss function; wherein the first loss represents a first difference between the noise added at the target time step and the first noise predicted for the target time step, and the second loss represents a second difference between the noise added at the target time step and the second noise predicted for the target time step, and the target time step is at least one of the T time steps.

[0126] As one of the feasible ways, when the model training unit 802 obtains the value of the loss function using the first loss and the second loss, it can be specifically configured as follows: determining the first loss using the correction weight and the first difference, and determining the second loss using the second difference, wherein the correction weight is determined based on the second difference; and weighting the first loss and the second loss using a preset weight coefficient to obtain the value of the loss function.

[0127] As one of the feasible methods, the correction weight is determined in the following manner: based on a preset power value, the power norm distance corresponding to the second difference is determined, and the power norm distance is marked as a stop gradient; and the correction weight is determined using the power norm distance corresponding to the second difference.

[0128] As one of the possible implementations, the power value is determined based on the data balance of the training data.

[0129] Furthermore, the device 800 also includes: an evaluation unit 803, which is configured to obtain an evaluation data set including multiple evaluation texts, and the evaluation texts are used to describe a combination of at least two objects; use the evaluation data set to evaluate the trained visual content generation model to obtain an evaluation result, and the evaluation includes: inputting the second noisy visual content and the evaluation text into the visual content generation model, obtaining the predicted visual content obtained after the visual content generation model denoises the second noisy visual content using the evaluation text as a condition, and obtaining the evaluation result based on the predicted visual content.

[0130] As one of the possible implementation methods, when obtaining an evaluation data set including multiple evaluation texts, the evaluation unit 803 can be specifically configured as follows: extracting a candidate object combination from the object data set, the candidate object combination including at least a first object and a second object, the frequency of occurrence of the first object in the object data set is greater than or equal to a first frequency threshold, the frequency of occurrence of the second object in the object data set is less than or equal to a second frequency threshold, and the first frequency threshold is greater than the second frequency threshold; using the candidate object combination, obtaining the evaluation text to constitute the evaluation data set.

[0131] As one of the feasible ways, when the evaluation unit 803 uses the candidate object combination to obtain the evaluation text to form the evaluation data set, it can be specifically configured as follows: extracting the target object combination from the candidate object combination, and the frequency of the first object and the second object in the target object combination appearing in the same text of the object data set meets the preset conditions; calling the large language model to generate at least one description text for each target object combination as the evaluation text.

[0132] According to another embodiment, a text-driven visual content generation apparatus is provided. Figure 9 A schematic block diagram of a text-driven visual content generation device provided in an embodiment of the present application, wherein the device is provided at Figure 1 User equipment in the architecture shown. Figure 9 As shown, the apparatus 900 includes: a second acquisition unit 901 and a content generation unit 902. The main functions of each component unit are as follows:

[0133] The second acquiring unit 901 is configured to acquire a target text, where the target text is used to describe a combination of at least two objects.

[0134] The content generation unit 902 is configured to use a visual content generation model to generate target visual content corresponding to the target text based on the target text. The visual content generation model is trained using the training method of the aforementioned text-driven visual content generation model, and the target visual content displays a combination of the at least two objects.

[0135] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or device embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0136] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0137] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.

[0138] And an electronic device comprising:

[0139] one or more processors; and

[0140] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the aforementioned method embodiments.

[0141] The present application also provides a computer program product, comprising a computer program, which implements the steps of any one of the methods described in the aforementioned method embodiments when executed by a processor.

[0142] in, Figure 10The schematic block diagram of an electronic device provided in an embodiment of the present application may include a processor 1010, a video display adapter 1011, a disk drive 1012, an input / output interface 1013, a network interface 1014, and a memory 1020. The processor 1010, the video display adapter 1011, the disk drive 1012, the input / output interface 1013, the network interface 1014, and the memory 1020 may be communicatively connected via a communication bus 1030. The input / output interface 1013 may also be referred to as an I / O interface 1013.

[0143] Among them, the processor 1010 can be implemented by a general-purpose CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., to execute relevant programs to implement the technical solutions provided in this application.

[0144] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store an operating system 1021 for controlling the operation of the electronic device 1000, and a basic input and output system (BIOS) 1022 for controlling the low-level operations of the electronic device 1000. In addition, a web browser 1023, a data storage management system 1024, a text-driven visual content generation model training device 800 and a text-driven visual content generation device 900, etc. can also be stored. The above-mentioned text-driven visual content generation model training device 800 and text-driven visual content generation device 900 can be the application program that specifically implements the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided by the present application is implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0145] The input / output interface 1013 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.

[0146] The network interface 1014 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WIFI, Bluetooth, etc.).

[0147] The bus 1030 comprises a pathway for transmitting information between the various components of the device (eg, the processor 1010 , the video display adapter 1011 , the disk drive 1012 , the input / output interface 1013 , the network interface 1014 , and the memory 1020 ).

[0148] It should be noted that although the above device only shows the processor 1010, the video display adapter 1011, the disk drive 1012, the input / output interface 1013, the network interface 1014, the memory 1020, the bus 1030, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may also include only the components necessary to implement the solution of the present application, and does not necessarily include all the components shown in the figure.

[0149] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer program product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.

[0150] The above is a detailed introduction to the technical solutions provided by this application. Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the contents of this specification should not be understood as limiting this application.

Claims

1. A training method for a text-driven visual content generation model, characterized in that: The method comprises: Acquire training data comprising a plurality of training samples, the training samples comprising: a text sample and a visual content sample, the text sample being used to describe a combination of at least two objects, and the visual content sample showing the combination of the at least two objects; The diffusion model is trained using the training data to obtain the visual content generation model; wherein the training includes: Adding T time steps of noise to the visual content sample to obtain a first noisy visual content, where T is a positive integer; Inputting the first noise visual content into the diffusion model, obtaining first noises predicted by the diffusion model for the T time steps when the diffusion model takes the text sample as a condition, and obtaining second noises predicted by the diffusion model for the T time steps when the diffusion model takes no condition; A value of a loss function is obtained using a first loss and a second loss, and model parameters of the diffusion model are updated using the value of the loss function; wherein the first loss represents a first difference between the noise added at a target time step and the first noise predicted for the target time step, and the second loss represents a second difference between the noise added at the target time step and the second noise predicted for the target time step, and the target time step is at least one of the T time steps.

2. The method according to claim 1, characterized in that The value of the loss function obtained by using the first loss and the second loss includes: determining the first loss using a revised weight and the first difference, and determining the second loss using the second difference, wherein the revised weight is determined based on the second difference; The first loss and the second loss are weighted using a preset weight coefficient to obtain a value of the loss function.

3. The method according to claim 2, characterized in that The correction weight is determined in the following manner: Determining a power norm distance corresponding to the second difference based on a preset power value, wherein the power norm distance is marked as stopping gradient calculation; The correction weight is determined using the power norm distance corresponding to the second difference.

4. The method according to claim 3, characterized in that The power value is determined based on the data balance of the training data.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Acquire an evaluation data set including a plurality of evaluation texts, wherein the evaluation texts are used to describe a combination of at least two objects; The visual content generation model obtained through training is evaluated using the evaluation data set to obtain an evaluation result, wherein the evaluation includes: inputting the second noisy visual content and the evaluation text into the visual content generation model, obtaining the predicted visual content obtained after the visual content generation model denoises the second noisy visual content using the evaluation text as a condition, and obtaining the evaluation result based on the predicted visual content.

6. The method according to claim 5, characterized in that The step of obtaining an evaluation data set including a plurality of evaluation texts includes: Extracting a candidate object combination from an object dataset, the candidate object combination comprising at least a first object and a second object, wherein a frequency of occurrence of the first object in the object dataset is greater than or equal to a first frequency threshold, a frequency of occurrence of the second object in the object dataset is less than or equal to a second frequency threshold, and the first frequency threshold is greater than the second frequency threshold; The candidate objects are combined to obtain evaluation texts to form the evaluation data set.

7. The method according to claim 6, characterized in that The step of utilizing the candidate object combination to obtain the evaluation text to form the evaluation data set includes: Extracting a target object combination from the candidate object combination, wherein the frequencies of occurrence of a first object and a second object in the target object combination in the same text of the object dataset meet a preset condition; The large language model is called to generate at least one description text for each of the target object combinations as the evaluation text.

8. A text-driven visual content generation method, characterized in that: The method comprises: Acquire a target text, where the target text is used to describe a combination of at least two objects; A visual content generation model is used to generate target visual content corresponding to the target text based on the target text, wherein the visual content generation model is trained using the method described in any one of claims 1 to 7, and the target visual content displays a combination of the at least two objects.

9. A training device for a text-driven visual content generation model, characterized in that: The device comprises: A first acquisition unit is configured to acquire training data comprising a plurality of training samples, wherein the training samples include: a text sample and a visual content sample, wherein the text sample is used to describe a combination of at least two objects, and the visual content sample shows the combination of the at least two objects; A model training unit is configured to train a diffusion model using the training data to obtain the visual content generation model; wherein the training includes: adding T time steps of noise to the visual content sample to obtain first noisy visual content, where T is a positive integer; inputting the first noisy visual content into the diffusion model to obtain the first noise predicted by the diffusion model for the T time steps when the diffusion model takes the text sample as a condition, and obtaining the second noise predicted by the diffusion model for the T time steps when the diffusion model is unconditional; using the first loss and the second loss to obtain the value of the loss function, and using the value of the loss function to update the model parameters of the diffusion model; wherein the first loss represents a first difference between the noise added at the target time step and the first noise predicted for the target time step, and the second loss represents a second difference between the noise added at the target time step and the second noise predicted for the target time step, and the target time step is at least one of the T time steps.

10. A text-driven visual content generation device, characterized in that: The device comprises: a second acquiring unit, configured to acquire a target text, where the target text is used to describe a combination of at least two objects; The content generation unit is configured to generate target visual content corresponding to the target text based on the target text using a visual content generation model, wherein the visual content generation model is trained using the device described in claim 9, and the target visual content displays a combination of the at least two objects.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed, the steps of the method according to any one of claims 1 to 8 are implemented.

12. An electronic device, characterized in that: include: one or more processors; as well as A memory associated with the one or more processors, the memory being configured to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method according to any one of claims 1 to 8.

13. A computer program product comprising a computer program, characterized in that When the computer program is executed, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Minkowski distance-based center loss expansion method

    CN115730179A

  • Beautifying material setting method and device, equipment and storage medium

    CN116366762A

  • Figure graph and model training method and device, electronic equipment and storage medium

    CN118155023A

  • Multi-object visual content generation model training method, generation method and device

    CN119784876A

  • Image generation model training method and device, electronic equipment and storage medium

    CN119919755A