Image generation method and device, electronic equipment and computer readable storage medium
By adjusting the parameters of the diffusion model, it can identify and learn the independent location area and attribute characteristics of each object in the training sample image, solving the problem of attribute leakage in the generation of multiple types of objects and improving the fidelity of image generation.
Patent Information
- Application Number
- CN202510765942.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-19
AI Technical Summary
Existing diffusion models are prone to attribute leakage when generating images containing multiple types of objects, resulting in a decrease in the fidelity of the generated images.
By obtaining text-image paired samples, the target diffusion model is used to generate independent position area information of each target object in the training sample image, and the model parameters are adjusted through loss calculation, so that the model learns to find the independent attribute features in the actual independent position area of each object to avoid attribute confusion.
The fidelity of the generated images is improved, ensuring that each object generates visual information using its own independent attribute features within its generation area, avoiding attribute leakage.
Smart Images

Figure CN120672908A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an image generation method, device, electronic device, and computer-readable storage medium. Background Art
[0002] The text-to-image task is a cross-modal generation task whose goal is to automatically generate visual images that match the semantic content based on text descriptions.
[0003] Currently, text-to-image generation models, such as the diffusion model, perform well in generating images containing a single type of object. For example, when the diffusion model is asked to generate an image containing a "Chow Chow", the diffusion model can accurately generate an image containing a "Chow Chow".
[0004] However, when the diffusion model is asked to generate multiple categories of objects, if the training samples are insufficient, the diffusion model cannot separate the independence between multiple categories of objects from a small number of training samples, which may lead to attribute leakage (i.e., attribute confusion or feature error migration) in the generated image. Among them, attribute leakage refers to the diffusion model incorrectly associating unspecified attributes in the text with the generated object, or confusing objects with different attributes. The essence is the failure of cross-modal alignment, resulting in inaccurate semantic control. For example, if there are only a small number of husky and corgi training samples in the training samples, the diffusion model may generate a deformed dog with "the coat color of a husky (black and white) mixed with the short legs of a corgi", or confuse the tiger's stripes with the body shape of a domestic cat. This will affect the fidelity of the generated image. Summary of the Invention
[0005] In view of this, the purpose of this application is to provide an image generation method, device, electronic device and computer-readable storage medium, so that when a diffusion model generates an image containing multiple types of objects based on text prompts, it can prevent attribute leakage problems and improve the fidelity of the generated image.
[0006] In a first aspect, an embodiment of the present application provides an image generation method, comprising:
[0007] Acquire a text-image pairing sample; wherein the text-image pairing sample includes a training sample image and a matching content text prompt, and the training sample image contains at least two target objects;
[0008] Inputting the content text prompt into a target diffusion model to generate first position region information for representing an independent position region of each target object in the training sample image determined by the target diffusion model;
[0009] Acquire second position area information for representing an actual independent position area of each target object in the training sample image, and perform loss calculation on the first position area information and the second position area information to obtain a first loss value;
[0010] Adjusting parameters of the target diffusion model using the first loss value so that the target diffusion model learns to find the actual independent position region of each target object in the training sample image, and learns the independent attribute features of each target object in the actual independent position region of each target object;
[0011] When an image containing at least two target objects is generated using the target diffusion model after parameter adjustment, visual information of each target object is generated using the independent attribute features of each target object in the generation area of each target object.
[0012] In a second aspect, an embodiment of the present application further provides an image generating device, comprising:
[0013] A first acquisition module is configured to acquire a text-image pairing sample, wherein the text-image pairing sample includes a training sample image and a matching content text prompt, and the training sample image contains at least two target objects;
[0014] A first generating module is configured to input the content text prompt into a target diffusion model to generate first position region information representing an independent position region of each target object in the training sample image determined by the target diffusion model;
[0015] a first calculation module, configured to obtain second position region information representing an actual independent position region of each target object in the training sample image, and perform loss calculation on the first position region information and the second position region information to obtain a first loss value;
[0016] an adjustment module, configured to adjust parameters of the target diffusion model using the first loss value, so that the target diffusion model learns to find the actual independent position region of each target object in the training sample image, and learns the independent attribute features of each target object in the actual independent position region of each target object;
[0017] The second generation module is used to generate visual information of each target object respectively within the generation area of each target object by using the independent attribute features of each target object when generating an image containing at least two target objects by using the target diffusion model after parameter adjustment.
[0018] In a third aspect, an embodiment of the present application further provides an electronic device comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate through the bus, and when the machine-readable instructions are executed by the processor, the steps of any possible implementation method of the first aspect above are performed.
[0019] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any possible implementation method of the above-mentioned first aspect are executed.
[0020] Embodiments of the present application provide an image generation method, apparatus, electronic device, and computer-readable storage medium. Considering that existing target diffusion models perform well when generating images containing a single target object based on text prompts, but are prone to attribute leakage when generating images containing different types of target objects, in this embodiment, parameters of the existing target diffusion model are further adjusted to prevent attribute leakage when the parameter-adjusted target diffusion model generates images containing at least two types of target objects based on text prompts.
[0021] In this embodiment, during the process of further parameter adjustment of the target diffusion model, the target diffusion model is inputted into the target diffusion model so that the target diffusion model determines first position region information representing the independent position region of each target object in the training sample image. Second position region information representing the actual independent position region of each target object in the training sample image is obtained. A first loss value is calculated between the first position region information and the second position region information, and the target diffusion model parameters are adjusted using the first loss value. This enables the target diffusion model to learn to find the actual independent position region of each target object in the training sample image and learn the independent attribute features of each target object within the actual independent position region of each target object, thereby avoiding learning attribute features that do not belong to the target object itself. In other words, when learning the attribute features of target object A, the attribute features of target object B are avoided from being mixed in. Finally, when generating an image containing at least two target objects using the parameter-adjusted target diffusion model, the independent attribute features of each target object are used within each target object's respective generation region to generate visual information of each target object, thereby avoiding attribute leakage and improving the fidelity of the generated image.
[0022] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 A flowchart of an image generation method provided in an embodiment of the present application is shown;
[0025] Figure 2 A schematic diagram showing a process of generating a training sample image provided by an embodiment of the present application is shown;
[0026] Figure 3 A schematic diagram showing a process of calculating a reference loss and a localized refinement loss provided in an embodiment of the present application is shown;
[0027] Figure 4 A schematic diagram illustrating a process of calculating a priori loss and localized refinement loss provided in an embodiment of the present application is shown;
[0028] Figure 5 A schematic structural diagram of an image generating device provided in an embodiment of the present application is shown;
[0029] Figure 6 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of illustration and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowcharts can be implemented out of sequence, and steps without logical context can be reversed or implemented simultaneously. In addition, those skilled in the art, under the guidance of the contents of this application, can add one or more other operations to the flowchart, or remove one or more operations from the flowchart.
[0031] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.
[0032] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.
[0033] Currently, text-to-image generation models, such as the diffusion model, perform well in generating images containing a single type of object. For example, when the diffusion model is asked to generate an image containing a "Chow Chow", the diffusion model can accurately generate an image containing a "Chow Chow".
[0034] However, when the diffusion model is asked to generate multiple categories of objects, if the training samples are insufficient, the diffusion model cannot separate the independence between multiple categories of objects from a small number of training samples, which may lead to attribute leakage (i.e., attribute confusion or feature error migration) in the generated image. Attribute leakage refers to the fact that the diffusion model mistakenly associates unspecified attributes in the text with the generated object, or confuses objects with different attributes. The essence is the failure of cross-modal alignment, which leads to inaccurate semantic control. For example, if there are only a small number of husky and corgi training samples in the training samples, the diffusion model may generate a deformed dog with "a mixture of husky's coat color (black and white) and corgi's short legs", or confuse tiger's stripes with the body shape of a domestic cat, which will affect the fidelity of the generated image.
[0035] Based on this, the embodiments of the present application provide an image generation method, device, electronic device and computer-readable storage medium. By further adjusting the parameters of the existing target diffusion model, the target diffusion model learns to find the actual independent position area of each target object in the training sample image, and learns the independent attribute characteristics of each target object in the actual independent position area of each target object, thereby avoiding the problem of attribute leakage and improving the fidelity of the generated image when the target diffusion model generates an image containing at least two target objects.
[0036] In one embodiment of the present application, an image generation method can be run on a terminal device or a server. The terminal device can be a local terminal device. When the image generation method is run on a server, the image generation method can be implemented and executed based on a cloud interaction system, wherein the cloud interaction system includes a server and a client device (i.e., a terminal device).
[0037] To facilitate understanding of this embodiment, an image generation method, device, electronic device, and computer-readable storage medium provided in an embodiment of the present application are described in detail below.
[0038] like Figure 1 As shown, Figure 1 A flowchart of an image generation method provided in an embodiment of the present application is shown, wherein the image generation method includes the following steps S101-S105:
[0039] S101: Acquire text-image pairing samples; wherein the text-image pairing samples include training sample images and matching content text prompts, and the training sample images contain at least two target objects.
[0040] S102: Inputting the content text prompt into the target diffusion model to generate first position region information for representing the independent position region of each target object determined by the target diffusion model in the training sample image.
[0041] S103: Obtain second position region information for representing the actual independent position region of each target object in the training sample image, perform loss calculation on the first position region information and the second position region information, and obtain a first loss value.
[0042] S104: Parameters of the target diffusion model are adjusted using the first loss value, so that the target diffusion model learns to find the actual independent position region of each target object in the training sample image, and learns the independent attribute features of each target object in the actual independent position region of each target object.
[0043] S105: When the target diffusion model after parameter adjustment is used to generate an image containing at least two target objects, visual information of each target object is generated in each generation area of each target object using independent attribute features of each target object.
[0044] In step S101 , the text-image pairing sample includes: a training sample image and a content text prompt that matches the training sample image.
[0045] Among them, the training sample image contains at least two target objects, and the target object refers to the entity contained in the training sample image. For example, the target object can be an animal (dog, cat, tiger, etc.), a plant (tree, flower, etc.), a person (man, woman, old man, child, etc.), an object (table, bag, lamp, etc.), etc.
[0046] For example, different types of target objects include: cats and dogs, animals and plants, border collies and Chow Chows, and red and yellow roses. The same type of target object includes: two border collies in different positions (e.g., one standing and the other lying) and two yellow roses of different sizes.
[0047] The content text hints that match the training sample images are used to describe the content contained in the training sample images.
[0048] In step S102, the target diffusion model is specifically an existing Stable Diffusion model. The text prompt content is input into the target diffusion model, which generates first location region information representing the independent location region of each target object in the training sample image. The first location region information is determined (estimated, interpreted) by the target diffusion model and is not necessarily the actual independent location region.
[0049] In step S103 , the second position region information may specifically be an independent segmentation mask of each target object in the training sample image. That is, the second position region information is used to represent the actual independent position region of each target object in the training sample image.
[0050] A localized refinement loss is calculated for the first position area information and the second position area information to obtain a first loss value. The first loss value is intended to enable the target diffusion model to focus on the actual independent position area where each target object is located during the process of learning the independent attribute features of each target object.
[0051] In step S104, the first loss value is used to adjust the parameters of the target diffusion model so that the target diffusion model learns to find the actual independent position area of each target object in the training sample image, and learns the independent attribute characteristics of each target object in the actual independent position area of each target object.
[0052] For example, when the two target objects contained in the training sample image are a Chow Chow and a Border Collie, the target diffusion model is allowed to learn the independent attribute characteristics of the Chow Chow itself within the actual independent position area of the Chow Chow in the training sample image; and to learn the independent attribute characteristics of the Border Collie within the actual independent position area of the Border Collie in the training sample image.
[0053] In step S105 , after adjusting the parameters of the existing target diffusion model, the obtained target diffusion model is made more suitable for outputting an image containing at least two target objects.
[0054] Specifically, when generating an image containing at least two target objects using the parameter-adjusted target diffusion model, visual information for each target object is generated using the independent attribute features of each target object within its respective generation region. In other words, for each target object, the independent attribute features of other target objects are prohibited from being received within the generation region of that target object.
[0055] In a possible implementation, the content text prompt includes first text prompt information and second text prompt information; the first text prompt information is used to describe each target object contained in the training sample image, and the second text prompt information includes a placeholder for each target object, and the placeholder is used to represent the initial attribute characteristics of the target object;
[0056] For example, when the training sample image includes a border collie dog and a chow chow dog, the first text prompt included in the text prompt content matching the training sample image is “A photo of a border collie dog and chow chow dog”.
[0057] The second text prompt includes the first text prompt and a placeholder for each target object. The placeholder is used to prompt the target diffusion model with the initial attribute characteristics of each target object, so that the target diffusion model can more quickly understand each target object. The initial attribute characteristics can specifically be the category characteristics of the target object.
[0058] In this embodiment, the target diffusion model can understand some target objects, while it cannot quickly understand other target objects. For example, if the target object is a certain game character, the target diffusion model may not be able to quickly understand the game character. In this case, the target diffusion model can be informed of whether the game character is a human or an animal through the placeholder of each target object in the second text prompt, thereby allowing the target diffusion model to understand the game character more quickly.
[0059] Specifically, the placeholder for each target object in the second text prompt is located before the name of each target object in the first text prompt. The placeholder can specifically be a category feature of the target object. For example, when the first text prompt is "A photo of a border collie dog and a chow chow dog," the second text prompt is "A photo of a [dog] border collie dog and a [dog] chow chow dog," where the first "[dog]" is a placeholder for the target object "border collie dog"; the second "[dog]" is a placeholder for the target object "chow chow dog."
[0060] When executing step S102, the following steps may be specifically performed:
[0061] The first text prompt, the second text prompt and the noise image are input into the target diffusion model. The target diffusion model generates a semantically matched first image according to the guidance of the first text prompt and the second text prompt, and generates independent position area information of each target object in the noise image according to the guidance, and uses the independent position area information as the first position area information; wherein the noise image is an image obtained by adding noise to the training sample image.
[0062] In this embodiment, the target diffusion model includes at least a text encoder, a cross-attention layer, and a self-attention layer.
[0063] The first text prompt and the second text prompt are input into a text encoder in the target diffusion model, and the first text prompt is converted into a first embedding vector and the second text prompt is converted into a second embedding vector by the text encoder.
[0064] The noise image can specifically be a pure noise image obtained by adding noise to the training sample image. The noise image is input into the target diffusion model. The denoising network in the target diffusion model (the self-attention layer is the key layer in the denoising network) will gradually remove the noise in the noise image based on the guidance of the first embedding vector and the second embedding vector, thereby outputting the denoised first image.
[0065] Furthermore, the cross-attention layer in the target diffusion model finds the independent location region of each target object in the noisy image based on the guidance of the first embedding vector and the second embedding vector, generates independent location region information of each target object in the noisy image, and uses the independent location region information as the first location region information. The first location region information exists in the form of a high-dimensional matrix vector.
[0066] In a possible implementation, before performing step S104 of adjusting parameters of the target diffusion model using the first loss value, the method may further be performed according to the following steps:
[0067] A standard diffusion loss is calculated for the first image and the noise image to obtain a second loss value.
[0068] In this embodiment, a standard diffusion loss (Latent Diffusion Model, LDM) is calculated on the first image and the noise image to obtain a second loss value, which is intended to minimize the difference between the noise predicted by the target diffusion model and the actual noise added to the training sample image.
[0069] When performing step S104 to adjust the parameters of the target diffusion model using the first loss value so that the target diffusion model learns to find the actual independent position region of each target object in the training sample image, and learns the independent attribute features of each target object in the actual independent position region of each target object, the following steps may be specifically performed:
[0070] Adjusting parameters of the target diffusion model using the first loss value and the second loss value so that the target diffusion model learns to find the actual independent position region of each target object in the training sample image based on the first loss value, and learns the independent attribute features of each target object within the actual independent position region of each target object based on the second loss value;
[0071] And iteratively optimize the placeholder using the first loss value and the second loss value.
[0072] In this embodiment, in the process of further adjusting the parameters of the target diffusion model using the first loss value and the second loss value, the parameters in the text encoder, the cross attention layer, and the self-attention layer in the target diffusion model are specifically adjusted.
[0073] In this embodiment, the target diffusion model is prompted with the initial attribute characteristics of each target object by the placeholder for each target object in the second text prompt. In subsequent parameter adjustments, the placeholder for each target object is iteratively optimized to enable the target diffusion model to better understand the independent attribute characteristics of each target object. Therefore, in the subsequent image generation process, when the target diffusion model generates target objects through the text prompt, it can clearly understand the independent attribute characteristics of each target object.
[0074] In one possible implementation, when adjusting the parameters of a target diffusion model using only a small amount of training data, overfitting can occur. Overfitting refers to the target diffusion model's over-reliance on noise or specific details in the training data, resulting in excellent performance on the training set but poor generalization (significantly reduced performance on new data or new objects). Furthermore, current training sample images typically contain only a single target object, making it difficult for the target diffusion model to learn to combine multiple target objects in the same scene, thus limiting the diversity and flexibility of the generated images.
[0075] Based on this, in this embodiment, before executing step S101, multiple training sample images can be generated according to the following steps S1001-S1004 to alleviate the problem of overfitting:
[0076] S1001: For each target object, obtain multiple reference images containing only the target object, and generate multiple prior images containing only the target object using a target diffusion model.
[0077] S1002: For each reference image, segment the target object contained in the reference image to obtain a reference segmented image of the target object; and for each prior image, segment the target object contained in the prior image to obtain a prior segmented image of the target object.
[0078] S1003: Recombining reference segmented images of different target objects to form an enhanced reference image set including multiple enhanced reference images, and recombining prior segmented images of different target objects to form an enhanced prior image set including multiple enhanced prior images.
[0079] S1004: Use the enhanced reference image and the enhanced prior image as training sample images.
[0080] In step S1001 , two target objects, namely, Chow Chow and Border Collie, are used as examples for illustration.
[0081] For the target object "Chow Chow", 3-5 images containing only Chow Chow are obtained and used as reference images of the Chow Chow. The shape, position, and size of the Chow Chow contained in different reference images may be different.
[0082] For the target object "Border Collie", 3-5 images containing only Border Collie are obtained and used as reference images of the Border Collie. The shape, position, and size of the Border Collie contained in different reference images can be different.
[0083] These reference images demonstrate the specific visual appearance of each target object that the user wants the target diffusion model to learn.
[0084] Since existing target diffusion models perform well in generating images containing only one type of target object, Figure 2 As shown, for each target object, by inputting text prompts into the target diffusion model, the target diffusion model generates multiple prior images that only contain the target object.
[0085] For example, taking the target object "Chow Chow" as an example, by inputting "Chow Chow" into the target diffusion model, the target diffusion model generates dozens of images containing only "Chow Chow", and the generated images are used as the prior images of the Chow Chow. The shape, position, and size of the Chow Chow contained in different prior images can be different.
[0086] The prior image generated by the target diffusion model provides prior knowledge about the general characteristics of the target object for subsequent further optimization of the existing target diffusion model, helping the target diffusion model to better capture the posture, shape and viewpoint changes of each target object.
[0087] In step S1002 , the target objects contained in each reference image and each prior image may be segmented using a semantic segmentation model (Grounded SAM).
[0088] In step S1003, illustratively, when the target objects are a border collie and a chow chow, the reference segmented image of the border collie and the reference segmented image of the chow chow are pasted into a solid color background image, thereby obtaining an enhanced reference image containing the two target objects.
[0089] Furthermore, the prior segmentation image of the Border Collie and the prior segmentation image of the Chow Chow are pasted into a solid color background image, thereby obtaining an enhanced prior image containing the two target objects.
[0090] In step S1004, the generated enhanced reference image and enhanced prior image are both used as training sample images.
[0091] By using the above method, a training sample image containing at least two target objects is generated, which enables the target diffusion model to learn to combine at least two target objects in the same scene.
[0092] In a possible implementation, when executing step S1002, the following steps S10021-S10024 may be specifically performed:
[0093] S10021: Inputting a reference image and a third text prompt for instructing the semantic segmentation model to segment a target object contained in the reference image into the semantic segmentation model to obtain a first segmentation mask of the target object in the reference image.
[0094] S10022: Extracting the region where the target object is located and covered by the first segmentation mask from the reference image, and using the extracted region as a reference segmentation image of the target object.
[0095] S10023: Inputting the prior image and a fourth text prompt for instructing the semantic segmentation model to segment the target object contained in the prior image into the semantic segmentation model to obtain a second segmentation mask of the target object in the prior image.
[0096] S10024: Extracting the region where the target object is located and covered by the second segmentation mask from the prior image, and using the segmented region as the prior segmentation image of the target object.
[0097] In step S10021, if Figure 2 As shown, the reference image corresponding to each target object and the third text prompt corresponding to each reference image (such as "Border Collie", "Chow Chow") are input into the semantic segmentation model (Grounded SAM), and the semantic segmentation model outputs the first segmentation mask of the target object in each reference image.
[0098] In step S10023, if Figure 2 As shown, the prior image corresponding to each target object and the fourth text prompt corresponding to each prior image (such as "Border Collie", "Chow Chow") are input into the semantic segmentation model, and the second segmentation mask of the target object in each prior image is output through the semantic segmentation model.
[0099] In a possible implementation, when executing step S1003, the following steps S10031-S10033 may be specifically performed:
[0100] S10031: Perform data enhancement processing on each reference segmentation image of each target object to obtain multiple reference segmentation enhanced images corresponding to each reference segmentation image; and perform data enhancement processing on each prior segmentation image of each target object to obtain multiple prior segmentation enhanced images corresponding to each prior segmentation image.
[0101] S10032: Randomly combine the reference segmentation enhancement images corresponding to different target objects to obtain multiple first combination methods, paste the reference segmentation enhancement images corresponding to the different target objects contained in each first combination method onto a background image, and generate an enhanced reference image corresponding to each first combination method; and randomly combine the priori segmentation enhancement images corresponding to different target objects to obtain multiple second combination methods, paste the reference segmentation enhancement images corresponding to the different target objects contained in each second combination method onto a background image, and generate an enhanced priori image corresponding to each second combination method.
[0102] S10033: forming an enhanced reference image set based on each generated enhanced reference image, and forming an enhanced priori image set based on each generated enhanced priori image.
[0103] In step S10031, if Figure 2 As shown, each reference segmentation image of each target object is randomly translated and scaled to obtain multiple reference segmentation enhanced images corresponding to each reference segmentation image. Also, each prior segmentation image of each target object is randomly translated and scaled to obtain multiple prior segmentation enhanced images corresponding to each prior segmentation image.
[0104] In step S10032, each reference segmentation-enhanced image included in the first combination is a reference segmentation-enhanced image of a different target object. Similarly, each a priori segmentation-enhanced image included in the second combination is a priori segmentation-enhanced image of a different target object. The background image is a solid color image, specifically a white background image. In this process, some overlap between target objects is allowed to simulate a more realistic scene.
[0105] In step S10033, based on the enhanced reference images of each target object, an enhanced reference dataset D is formed. ref ; Based on the prior segmentation images of each target object, an enhanced prior dataset D is formed prior .
[0106] Enhanced reference dataset D ref The enhanced reference image and enhanced prior dataset D in priorThe prior segmented images in the target diffusion model are used as training sample images for training. Therefore, the training sample images in the text image pairing samples in step S101 are respectively from the enhanced reference dataset D ref and enhanced prior dataset D prior .
[0107] In this embodiment, the diversity of training data is increased by separating the target objects in the image from the background and recombining them into a synthetic image. Specifically, a semantic segmentation model (such as GroundedSAM) is first used to automatically extract the segmentation mask of the target object. Then, an enhanced image is created by randomly translating and scaling these segmented target objects and pasting them onto a solid background image. This method is applied not only to reference images, but also to prior images generated by the target diffusion model, thereby further expanding the diversity of the dataset. Concept fusion reduces the risk of overfitting by enriching the different combinations in the training dataset and enhances the ability of the target diffusion model to learn single and multiple target objects.
[0108] In a possible implementation, when performing the step of calculating the standard diffusion loss on the first image and the noise image to obtain the second loss value, the following steps may be specifically performed:
[0109] Using the standard diffusion loss function, the standard diffusion loss is calculated for each text-image paired sample included in the current training round, corresponding to the first image and the noise image, to obtain the second loss value corresponding to the current training round.
[0110] In the current training round, multiple text-image paired samples need to be obtained in step S101. In fact, multiple text-image paired samples need to be obtained in each training round. In each training round (including the current training round), the training sample images contained in the multiple text-image paired samples obtained are respectively from the enhanced reference dataset D ref and enhanced prior dataset D prior .
[0111] In this embodiment, when performing the step of using the standard diffusion loss function to perform standard diffusion loss calculation on the first image and the noise image corresponding to each text image pairing sample included in the current training round to obtain the second loss value corresponding to the current training round, specifically:
[0112] like Figure 3 As shown, for the current training round from the enhanced reference dataset D refThe training sample image (i.e., enhanced reference image) is used to calculate the standard diffusion loss based on the enhanced reference data set for the first image and the noise image corresponding to all the enhanced reference images in the current training round, and the reference loss L corresponding to all the enhanced reference images in the current training round is obtained. ref :
[0113]
[0114] Where θ is a learnable parameter in the target diffusion model; p1 is the first and second text prompts corresponding to the enhanced reference image (i.e., the training sample image); T(P1) is the first and second embedding vectors corresponding to the enhanced reference image; z1=ε(x1) is the latent space representation of the noise image corresponding to the enhanced reference image (i.e., the training sample image) obtained by the VAE encoder ε; x1 represents the enhanced reference image; t is the time step uniformly sampled from {1...T}; T represents the maximum value in the denoising step; ∈ is the noise sampled from the standard overall distribution N(0,I); z t is the noise-added latent variable obtained by adding noise ∈ to z according to the predetermined noise schedule; ∈ θ (z t ,t,T(p1)) is the denoising network (U-Net) under the given noise latent variable z t , time step t, the first embedding vector, and the second embedding vector T(P1); E[·] represents the expected value.
[0115] like Figure 4 As shown, the current training round comes from the enhanced prior dataset D prior The training sample image (enhanced prior image) is used to calculate the standard diffusion loss based on the enhanced prior data set for the first image and the noise image corresponding to all the enhanced prior images in the current training round, and the prior loss L corresponding to all the enhanced prior images in the current training round is obtained. prior :
[0116]
[0117] Where p2 is the first and second text prompts corresponding to the enhanced prior image (i.e., the training sample image); T(P2) is the first embedding vector and the second embedding vector corresponding to the enhanced prior image; z2 = ε(x2) is the latent space representation of the noise image corresponding to the enhanced prior image (i.e., the training sample image) obtained by the VAE encoder ε; x2 represents the enhanced prior image; ∈ θ (z t ,t,T(p2)) is the denoising network (U-Net) under the given noise latent variable z t, time step t and the first embedding vector and the second embedding vector T(P2).
[0118] After obtaining the reference loss L corresponding to all enhanced reference images in the current training round ref and the prior loss L corresponding to all enhanced prior images in the current training round prior Then, the second loss value L2 between the first image and the noise image corresponding to the current training round is calculated by the following formula:
[0119] L2=L ref +μL prior
[0120] Among them, μ is a predefined hyperparameter (scaling factor) used to balance the weight of the prior loss, and its value range is 0 to 1.
[0121] In one possible implementation, the first position region information includes a weight for each position in the noise image as the location of the target object; the second position region information is an independent segmentation mask for each target object in the training sample image; when performing step S103 to calculate the loss on the first position region information and the second position region information to obtain the first loss value, the following steps S1031-S1032 may be specifically performed:
[0122] S1031: For each position in the noise image, using a log-likelihood loss function, calculate a log-likelihood loss value between the weight of the position in the first position region information and the mask of the position in the independent segmentation mask.
[0123] S1032: Calculate the sum of the log-likelihood loss values corresponding to each position, and determine a first loss value corresponding to the text-image paired sample based on the sum.
[0124] In step S1031, the independent segmentation mask can be a segmentation mask of the target object obtained by semantically segmenting the training sample image using a semantic segmentation model (Grounded SAM). The independent segmentation mask can represent the actual independent position region of each target object in the noisy image (and in the training sample image). The first position region information represents the independent position region of each target object in the noisy image predicted (assessed, understood) by the target diffusion model.
[0125] In this embodiment, the first position region information and the independent segmentation mask are compared element by element through the log-likelihood loss function, and the log-likelihood loss value Loss between the weight of each position in the first position region information and the mask of the corresponding position in the independent segmentation mask is calculated. h,w :
[0126] Lossh,w =M Ci [h,w]·logA Ci [h,w]+(1-M Ci [h,w])·log(1-M Ci [h,w])
[0127] Where [h, w] represents each position in the first position area information; Ci represents the target object; M Ci A represents an independent segmentation mask of the target object; Ci Indicates the first location area information of the target object.
[0128] For the position where the mask is 1: when M Ci When [h,w]=1, the log-likelihood loss function encourages A Ci [h,w] is close to 1. This means that the position in the first position area information should have a higher weight. For the position where the mask is 0: When M Ci When [h,w]=0, the log-likelihood loss function encourages A Ci [h,w] is close to 0. This means that the position should have a lower weight in the first position area information.
[0129] In this embodiment, the log-likelihood loss is used to measure the difference between the first position region information and the independent segmentation mask. The mathematical form of the log-likelihood loss function is:
[0130] Loss h,w =logA Ci [h,w] If M Ci [h,w]=1
[0131] Loss h,w =log(1-A Ci [h,w]) if M Ci [h,w]=0
[0132] For the position where the mask is 1: When A Ci When [h,w] is close to 1, logA Ci The larger the value of [h,w], the smaller the loss. For the position where the mask is 0: when A Ci When [h,w] is close to 0, log(1-A Ci [h,w]) is larger, the loss is smaller.
[0133] Through element-wise comparison and log-likelihood loss, the target diffusion model is forced to adjust the distribution of the first position region information so that it is highly aligned with the independent segmentation mask. Specifically:
[0134] Where the mask is 1, the weight of the first position region information is increased, ensuring that the target diffusion model pays attention to these areas. Where the mask is 0, the weight of the first position region information is decreased, ensuring that the target diffusion model ignores these areas.
[0135] Independent segmentation mask M Ci This provides the true spatial location information of the target object in the training sample image. By optimizing the log-likelihood loss function, the target diffusion model learns to focus on the area indicated by the independent segmentation mask instead of distributing it uniformly across the entire training sample image.
[0136] By ensuring that the first position region information of different target objects does not overlap, the target diffusion model can clearly distinguish different target objects. For example, when generating an image of two dogs, the first position region information of each dog is concentrated in its own independent position region to avoid feature mixing. Specifically, the first position region information of each target object is forced to be concentrated at the location of its own independent segmentation mask. The overlap between the first position region information of different target objects is minimized, thus preventing attribute leakage.
[0137] For example, Figure 3 As shown, the location of the Chow Chow is in the upper right corner of the training sample image, so that the target diffusion model learns the attribute characteristics of the Chow Chow in the upper right corner area, and does not learn the attribute characteristics of the Border Collie in the Border Collie area.
[0138] In step S1032, the first loss value corresponding to each text-image paired sample is equal to:
[0139]
[0140] The resolution of the image corresponding to the first location area information is N×N.
[0141] In a possible implementation, when the step of performing parameter adjustment on the target diffusion model and iterative optimization on the placeholder using the first loss value and the second loss value is executed, the following steps may be specifically performed:
[0142] S1041: After obtaining the first loss value corresponding to each text-image paired sample included in the current training round, calculate the expectation of the first loss value corresponding to each text-image paired sample included in the current training round to obtain the expected value.
[0143] S1042: Using the second loss value and expected value corresponding to the current training round, adjust the parameters of the target diffusion model for the current training round and perform iterative optimization on the placeholder for the current training round.
[0144] In step S1041, the expected value L is calculated by the following formula:local :
[0145]
[0146] The expected value L local as a localized refinement loss.
[0147] In step S1042, the total loss is calculated using the following formula:
[0148] L total =L ref +μL prior +γL local
[0149] Among them, L total represents the total loss; γ is a predefined hyperparameter (scaling factor) used to balance the weight of the localized refinement loss.
[0150] Utilize the total loss L total The parameters of the target diffusion model are adjusted for the current training round, and the placeholders are iteratively optimized for the current training round.
[0151] Specifically, according to the calculated total loss L total Update the learnable parameters in the text encoder, cross attention layer, and self-attention layer in the target diffusion model using a gradient descent optimization algorithm, such as the AdamW optimizer. In each training iteration, calculate the total loss L total Gradients of the learnable parameters in the text encoder, cross-attention layer, and self-attention layer are calculated. The learnable parameters are updated using the calculated gradients according to the optimizer's update rule. This process is repeated until the target diffusion model converges or the preset number of training steps is reached.
[0152] Through the above method, the method proposed in this embodiment can effectively adjust the existing target diffusion model so that when generating an image containing at least two target objects (especially target objects of similar categories), it can not only accurately reproduce the visual features of each target object, but also effectively avoid attribute confusion and leakage between target objects, while maintaining good image generation quality and realism.
[0153] In one possible implementation, when executing step S105 and using the parameter-adjusted target diffusion model to generate an image containing at least two target objects, generating visual information of each target object using independent attribute features of each target object within each target object's respective generation area may be performed in the following manner:
[0154] When an image containing at least two target objects is generated using the target diffusion model after parameter adjustment, a fifth text prompt for indicating the generation of an image containing at least two target objects is input into the target diffusion model after parameter adjustment. Based on the guidance of the fifth text prompt, the target expansion model generates visual information for each target object in each generation area of each target object using the learned independent attribute features of each target object.
[0155] In this embodiment, after adjusting the parameters of the existing target diffusion model, the obtained target diffusion model is more suitable for outputting images containing multiple target objects.
[0156] Based on the same technical concept, the embodiment of the present application also provides an image generating device corresponding to the above-mentioned image generating method. Since the principle of solving the problem by the image generating device in the embodiment of the present application is similar to the above-mentioned image generating method in the embodiment of the present application, the implementation of the image generating device can refer to the implementation of the above-mentioned image generating method, and the repeated parts will not be repeated.
[0157] like Figure 5 As shown, Figure 5 The following is a schematic diagram of the structure of an image generation device provided in an embodiment of the present application, wherein the image generation device includes:
[0158] The first acquisition module 501 is configured to acquire a text-image pairing sample, wherein the text-image pairing sample includes a training sample image and a matching content text prompt, and the training sample image contains at least two target objects;
[0159] A first generating module 502 is configured to input the content text prompt into a target diffusion model to generate first position region information representing an independent position region of each target object in the training sample image determined by the target diffusion model;
[0160] A first calculation module 503 is configured to obtain second position region information representing an actual independent position region of each target object in the training sample image, and perform loss calculation on the first position region information and the second position region information to obtain a first loss value;
[0161] an adjustment module 504, configured to adjust parameters of the target diffusion model using the first loss value, so that the target diffusion model learns to find the actual independent location region of each target object in the training sample image, and learns the independent attribute features of each target object in the actual independent location region of each target object;
[0162] The second generation module 505 is used to generate visual information of each target object respectively within the generation area of each target object by using the independent attribute features of each target object when generating an image containing at least two target objects by using the target diffusion model after parameter adjustment.
[0163] In one possible implementation, the content text prompt includes first text prompt information and second text prompt information; the first text prompt information is used to describe each target object contained in the training sample image, and the second text prompt information includes a placeholder for each target object, where the placeholder is used to represent the initial attribute characteristics of the target object; when the first generation module 502 is used to input the content text prompt into the target diffusion model and generate first position area information for representing the independent position area of each target object in the training sample image determined by the target diffusion model, it is specifically used to:
[0164] The first text prompt, the second text prompt and the noise image are input into a target diffusion model. The target diffusion model generates a semantically matched first image according to the guidance of the first text prompt and the second text prompt, and generates independent position area information of each target object in the noise image according to the guidance, so as to use the independent position area information as the first position area information; wherein the noise image is an image obtained by adding noise to the training sample image.
[0165] In a possible implementation, the device further includes:
[0166] a second calculation module, configured to perform standard diffusion loss calculation on the first image and the noise image to obtain a second loss value before the adjustment module 504 adjusts the parameters of the target diffusion model using the first loss value;
[0167] When the adjustment module 504 is used to adjust the parameters of the target diffusion model using the first loss value so that the target diffusion model learns to find the actual independent position region of each target object in the training sample image, and learns the independent attribute features of each target object in the actual independent position region of each target object, it is specifically used to:
[0168] Adjusting parameters of the target diffusion model using the first loss value and the second loss value, so that the target diffusion model learns to find the actual independent position region of each target object in the training sample image based on the first loss value, and learns the independent attribute features of each target object within the actual independent position region of each target object based on the second loss value;
[0169] And using the first loss value and the second loss value, iteratively optimize the placeholder.
[0170] In one possible implementation, the first position region information includes a weight for each position in the noise image as the position of each target object; the second position region information is an independent segmentation mask for each target object in the training sample image; and the first calculation module 503, when performing loss calculation on the first position region information and the second position region information to obtain a first loss value, is specifically configured to:
[0171] For each position in the noise image, using a log-likelihood loss function, calculate a log-likelihood loss value between the weight of the position in the first position region information and the mask of the position in the independent segmentation mask;
[0172] The sum of the log-likelihood loss values corresponding to each position is calculated, and a first loss value corresponding to the text-image paired sample is determined based on the sum.
[0173] In a possible implementation, when the second calculation module is used to perform standard diffusion loss calculation on the first image and the noise image to obtain the second loss value, it is specifically used to:
[0174] Using the standard diffusion loss function, perform standard diffusion loss calculation on the first image and the noise image corresponding to each text-image pair sample included in the current training round to obtain the second loss value corresponding to the current training round;
[0175] When the adjustment module 504 is used to adjust the parameters of the target diffusion model and iteratively optimize the placeholder using the first loss value and the second loss value, it is specifically used to:
[0176] After obtaining the first loss value corresponding to each text-image paired sample included in the current training round, calculating the expectation of the first loss value corresponding to each of the text-image paired samples included in the current training round to obtain an expected value;
[0177] Utilizing the second loss value and the expected value corresponding to the current training round, adjusting parameters of the target diffusion model for the current training round and iteratively optimizing the placeholder for the current training round.
[0178] In a possible implementation, the device further includes:
[0179] A second acquisition module is configured to acquire, for each target object, a plurality of reference images containing only the target object before the first acquisition module 501 acquires the text-image paired samples, and generate a plurality of prior images containing only the target object using the target diffusion model;
[0180] a segmentation module configured to segment, for each reference image, a target object contained in the reference image to obtain a reference segmented image of the target object; and to segment, for each prior image, a target object contained in the prior image to obtain a prior segmented image of the target object;
[0181] a combining module, configured to recombine the reference segmented images of different target objects to form an enhanced reference image set comprising a plurality of enhanced reference images, and to recombine the a priori segmented images of different target objects to form an enhanced a priori image set comprising a plurality of enhanced a priori images;
[0182] A determination module is configured to use the enhanced reference image and the enhanced prior image as the training sample images.
[0183] In a possible implementation, the segmentation module is configured to segment, for each reference image, the target object contained in the reference image to obtain a reference segmented image of the target object; and, for each prior image, segment the target object contained in the prior image to obtain a prior segmented image of the target object, specifically configured to:
[0184] Inputting the reference image and a third text prompt for instructing the semantic segmentation model to segment the target object contained in the reference image into the semantic segmentation model to obtain a first segmentation mask of the target object in the reference image;
[0185] extracting a region where the target object is located and covered by the first segmentation mask from the reference image, and using the extracted region as a reference segmentation image of the target object;
[0186] Inputting the prior image and a fourth text prompt for instructing the semantic segmentation model to segment the target object contained in the prior image into the semantic segmentation model to obtain a second segmentation mask of the target object in the prior image;
[0187] The region where the target object is located and covered by the second segmentation mask is extracted from the prior image, and the extracted region is used as the prior segmentation image of the target object.
[0188] In a possible implementation, when the second generating module 505 is configured to generate an image containing at least two target objects using the target diffusion model after parameter adjustment, generating visual information of each target object using independent attribute features of each target object within a generation region of each target object, specifically:
[0189] When an image containing at least two of the target objects is generated using a target diffusion model whose parameters have been adjusted, a fifth text prompt for indicating the generation of an image containing at least two of the target objects is input into the target diffusion model whose parameters have been adjusted. Based on the guidance of the fifth text prompt, the target expansion model generates visual information for each of the target objects in each generation area of each of the target objects using the learned independent attribute features of each of the target objects.
[0190] In the above-mentioned image generation device provided by the embodiment of the present application, during the process of further parameter adjustment of the target diffusion model, the target diffusion model is inputted into the target diffusion model so that the target diffusion model determines first position region information representing the independent position region of each target object in the training sample image. Second position region information representing the actual independent position region of each target object in the training sample image is obtained. A first loss value is calculated between the first position region information and the second position region information, and the first loss value is used to adjust the parameters of the target diffusion model. This enables the target diffusion model to learn to find the actual independent position region of each target object in the training sample image and learn the independent attribute features of each target object within the actual independent position region of each target object, thereby avoiding learning attribute features that do not belong to the target object itself, that is, avoiding confusing the attribute features of target object B when learning the attribute features of target object A. Finally, when generating an image containing at least two target objects using the parameter-adjusted target diffusion model, the independent attribute features of each target object are used within each target object's respective generation region to generate visual information of each target object, thereby avoiding attribute leakage and improving the fidelity of the generated image.
[0191] Based on the same technical concept, the embodiment of the present application also provides an electronic device corresponding to the above-mentioned image generation method. Since the principle of solving the problem by the electronic device in the embodiment of the present application is similar to the above-mentioned image generation method in the embodiment of the present application, the implementation of the electronic device can refer to the implementation of the above-mentioned image generation method, and the repeated parts will not be repeated.
[0192] Figure 6 This is a structural diagram of an electronic device 600 provided in an embodiment of the present application, including: a processor 601, a memory 602, and a bus 603. The memory 602 stores machine-readable instructions executable by the processor 601. When the electronic device runs an image generation method such as that in the embodiment, the processor 601 communicates with the memory 602 via the bus 603, and the processor 601 executes the machine-readable instructions. When the processor 601 executes the machine-readable instructions, the following steps are implemented, specifically:
[0193] Acquire a text-image pairing sample; wherein the text-image pairing sample includes a training sample image and a matching content text prompt, and the training sample image contains at least two target objects;
[0194] Inputting the content text prompt into a target diffusion model to generate first position region information for representing an independent position region of each target object in the training sample image determined by the target diffusion model;
[0195] Acquire second position area information for representing an actual independent position area of each target object in the training sample image, and perform loss calculation on the first position area information and the second position area information to obtain a first loss value;
[0196] Adjusting parameters of the target diffusion model using the first loss value so that the target diffusion model learns to find the actual independent position region of each target object in the training sample image, and learns the independent attribute features of each target object in the actual independent position region of each target object;
[0197] When an image containing at least two target objects is generated using the target diffusion model after parameter adjustment, visual information of each target object is generated using the independent attribute features of each target object in the generation area of each target object.
[0198] In one possible implementation, the content text prompt includes first text prompt information and second text prompt information; the first text prompt information is used to describe each target object contained in the training sample image, and the second text prompt information includes a placeholder for each target object, where the placeholder is used to represent an initial attribute feature of the target object; when inputting the content text prompt into a target diffusion model to generate first location area information representing an independent location area of each target object in the training sample image determined by the target diffusion model, the processor 601 is used to:
[0199] The first text prompt, the second text prompt and the noise image are input into a target diffusion model. The target diffusion model generates a semantically matched first image according to the guidance of the first text prompt and the second text prompt, and generates independent position area information of each target object in the noise image according to the guidance, so as to use the independent position area information as the first position area information; wherein the noise image is an image obtained by adding noise to the training sample image.
[0200] In a possible implementation, before adjusting parameters of the target diffusion model using the first loss value, the processor 601 is further configured to:
[0201] performing a standard diffusion loss calculation on the first image and the noise image to obtain a second loss value;
[0202] When adjusting the parameters of the target diffusion model using the first loss value so that the target diffusion model learns to find the actual independent position region of each target object in the training sample image and learns the independent attribute features of each target object within the actual independent position region of each target object, the processor 601 is configured to:
[0203] Adjusting parameters of the target diffusion model using the first loss value and the second loss value, so that the target diffusion model learns to find the actual independent position region of each target object in the training sample image based on the first loss value, and learns the independent attribute features of each target object within the actual independent position region of each target object based on the second loss value;
[0204] And using the first loss value and the second loss value, iteratively optimize the placeholder.
[0205] In a possible implementation, the first position region information includes a weight for each position in the noise image as a location of each target object; the second position region information is an independent segmentation mask for each target object in the training sample image; and when performing loss calculation on the first position region information and the second position region to obtain a first loss value, the processor 601 is configured to:
[0206] For each position in the noise image, using a log-likelihood loss function, calculate a log-likelihood loss value between the weight of the position in the first position region information and the mask of the position in the independent segmentation mask;
[0207] The sum of the log-likelihood loss values corresponding to each position is calculated, and a first loss value corresponding to the text-image paired sample is determined based on the sum.
[0208] In a possible implementation, when performing the standard diffusion loss calculation on the first image and the noise image to obtain the second loss value, the processor 601 is configured to:
[0209] Using the standard diffusion loss function, perform standard diffusion loss calculation on the first image and the noise image corresponding to each text-image pair sample included in the current training round to obtain the second loss value corresponding to the current training round;
[0210] When adjusting parameters of the target diffusion model and iteratively optimizing the placeholder using the first loss value and the second loss value, the processor 601 is configured to:
[0211] After obtaining the first loss value corresponding to each text-image paired sample included in the current training round, calculating the expectation of the first loss value corresponding to each of the text-image paired samples included in the current training round to obtain an expected value;
[0212] Utilizing the second loss value and the expected value corresponding to the current training round, adjusting parameters of the target diffusion model for the current training round and iteratively optimizing the placeholder for the current training round.
[0213] In a possible implementation, before obtaining the text-image pairing sample, the processor 601 is further configured to:
[0214] For each target object, a plurality of reference images containing only the target object are acquired, and a plurality of prior images containing only the target object are generated using the target diffusion model;
[0215] For each of the reference images, segmenting the target object contained in the reference image to obtain a reference segmented image of the target object; and for each of the prior images, segmenting the target object contained in the prior image to obtain a prior segmented image of the target object;
[0216] Recombining the reference segmented images of different target objects to form an enhanced reference image set including a plurality of enhanced reference images, and recombining the a priori segmented images of different target objects to form an enhanced a priori image set including a plurality of enhanced a priori images;
[0217] The enhanced reference image and the enhanced prior image are used as the training sample images.
[0218] In a possible implementation, for each of the reference images, segmenting the target object contained in the reference image to obtain a reference segmented image of the target object; and for each of the prior images, segmenting the target object contained in the prior image to obtain a priori segmented image of the target object, the processor 601 is configured to:
[0219] Inputting the reference image and a third text prompt for instructing the semantic segmentation model to segment the target object contained in the reference image into the semantic segmentation model to obtain a first segmentation mask of the target object in the reference image;
[0220] extracting a region where the target object is located and covered by the first segmentation mask from the reference image, and using the extracted region as a reference segmentation image of the target object;
[0221] Inputting the prior image and a fourth text prompt for instructing the semantic segmentation model to segment the target object contained in the prior image into the semantic segmentation model to obtain a second segmentation mask of the target object in the prior image;
[0222] The region where the target object is located and covered by the second segmentation mask is extracted from the prior image, and the extracted region is used as the prior segmentation image of the target object.
[0223] In a possible implementation, when generating an image containing at least two target objects using the target diffusion model after parameter adjustment, the processor 601 is configured to generate visual information of each target object using independent attribute features of each target object within a generation area of each target object:
[0224] When an image containing at least two of the target objects is generated using a target diffusion model whose parameters have been adjusted, a fifth text prompt for indicating the generation of an image containing at least two of the target objects is input into the target diffusion model whose parameters have been adjusted. Based on the guidance of the fifth text prompt, the target expansion model generates visual information for each of the target objects in each generation area of each of the target objects using the learned independent attribute features of each of the target objects.
[0225] In the electronic device provided by this embodiment, during further parameter adjustment of the target diffusion model, the target diffusion model determines first position region information representing the independent position region of each target object in the training sample image by inputting the information into the target diffusion model. Second position region information representing the actual independent position region of each target object in the training sample image is obtained. A first loss value is calculated between the first position region information and the second position region information, and the target diffusion model parameters are adjusted using the first loss value. This enables the target diffusion model to learn to find the actual independent position region of each target object in the training sample image and learn the independent attribute features of each target object within the actual independent position region of each target object, thereby avoiding learning attribute features that do not belong to the target object itself. Specifically, when learning the attribute features of target object A, the attribute features of target object B are avoided. Finally, when generating an image containing at least two target objects using the parameter-adjusted target diffusion model, the independent attribute features of each target object are used within each target object's respective generation region to generate visual information of each target object, thereby avoiding attribute leakage and improving the fidelity of the generated image.
[0226] Based on the same technical concept, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. The computer program is executed when a processor is run, and the processor performs the following steps:
[0227] Acquire a text-image pairing sample; wherein the text-image pairing sample includes a training sample image and a matching content text prompt, and the training sample image contains at least two target objects;
[0228] Inputting the content text prompt into a target diffusion model to generate first position region information for representing an independent position region of each target object in the training sample image determined by the target diffusion model;
[0229] Acquire second position area information for representing an actual independent position area of each target object in the training sample image, and perform loss calculation on the first position area information and the second position area information to obtain a first loss value;
[0230] Adjusting parameters of the target diffusion model using the first loss value so that the target diffusion model learns to find the actual independent position region of each target object in the training sample image, and learns the independent attribute features of each target object in the actual independent position region of each target object;
[0231] When an image containing at least two target objects is generated using the target diffusion model after parameter adjustment, visual information of each target object is generated using the independent attribute features of each target object in the generation area of each target object.
[0232] In one possible implementation, the content text prompt includes first text prompt information and second text prompt information; the first text prompt information is used to describe each target object contained in the training sample image, and the second text prompt information includes a placeholder for each target object, where the placeholder is used to represent an initial attribute feature of the target object; when the processor is used to input the content text prompt into a target diffusion model to generate first position area information representing an independent position area of each target object in the training sample image determined by the target diffusion model, it is specifically used to:
[0233] The first text prompt, the second text prompt and the noise image are input into a target diffusion model. The target diffusion model generates a semantically matched first image according to the guidance of the first text prompt and the second text prompt, and generates independent position area information of each target object in the noise image according to the guidance, so as to use the independent position area information as the first position area information; wherein the noise image is an image obtained by adding noise to the training sample image.
[0234] In a possible implementation, the processor is further configured to:
[0235] Before adjusting parameters of the target diffusion model using the first loss value, performing standard diffusion loss calculation on the first image and the noise image to obtain a second loss value;
[0236] When the processor is used to adjust the parameters of the target diffusion model using the first loss value so that the target diffusion model learns to find the actual independent position area of each target object in the training sample image, and learns the independent attribute features of each target object in the actual independent position area of each target object, the processor is specifically used to:
[0237] Adjusting parameters of the target diffusion model using the first loss value and the second loss value, so that the target diffusion model learns to find the actual independent position region of each target object in the training sample image based on the first loss value, and learns the independent attribute features of each target object within the actual independent position region of each target object based on the second loss value;
[0238] And using the first loss value and the second loss value, iteratively optimize the placeholder.
[0239] In one possible implementation, the first position region information includes a weight for each position in the noise image as a location of each target object; the second position region information is an independent segmentation mask for each target object in the training sample image; and when the processor is configured to perform loss calculation on the first position region information and the second position region to obtain a first loss value, is specifically configured to:
[0240] For each position in the noise image, using a log-likelihood loss function, calculate a log-likelihood loss value between the weight of the position in the first position region information and the mask of the position in the independent segmentation mask;
[0241] The sum of the log-likelihood loss values corresponding to each position is calculated, and a first loss value corresponding to the text-image paired sample is determined based on the sum.
[0242] In a possible implementation, when the processor is used to perform standard diffusion loss calculation on the first image and the noise image to obtain the second loss value, it is specifically used to:
[0243] Using the standard diffusion loss function, perform standard diffusion loss calculation on the first image and the noise image corresponding to each text-image pair sample included in the current training round to obtain the second loss value corresponding to the current training round;
[0244] When the processor is used to adjust parameters of the target diffusion model and iteratively optimize the placeholder using the first loss value and the second loss value, the processor is specifically used to:
[0245] After obtaining the first loss value corresponding to each text-image paired sample included in the current training round, calculating the expectation of the first loss value corresponding to each of the text-image paired samples included in the current training round to obtain an expected value;
[0246] Utilizing the second loss value and the expected value corresponding to the current training round, adjusting parameters of the target diffusion model for the current training round and iteratively optimizing the placeholder for the current training round.
[0247] In a possible implementation, before obtaining the text-image pairing sample, the processor is further configured to:
[0248] For each target object, a plurality of reference images containing only the target object are acquired, and a plurality of prior images containing only the target object are generated using the target diffusion model;
[0249] For each of the reference images, segmenting the target object contained in the reference image to obtain a reference segmented image of the target object; and for each of the prior images, segmenting the target object contained in the prior image to obtain a prior segmented image of the target object;
[0250] Recombining the reference segmented images of different target objects to form an enhanced reference image set including a plurality of enhanced reference images, and recombining the a priori segmented images of different target objects to form an enhanced a priori image set including a plurality of enhanced a priori images;
[0251] The enhanced reference image and the enhanced prior image are used as the training sample images.
[0252] In a possible implementation, the processor, when used to segment, for each reference image, the target object contained in the reference image to obtain a reference segmented image of the target object; and, for each prior image, segment the target object contained in the prior image to obtain a priori segmented image of the target object, is specifically used to:
[0253] Inputting the reference image and a third text prompt for instructing the semantic segmentation model to segment the target object contained in the reference image into the semantic segmentation model to obtain a first segmentation mask of the target object in the reference image;
[0254] extracting a region where the target object is located and covered by the first segmentation mask from the reference image, and using the extracted region as a reference segmentation image of the target object;
[0255] Inputting the prior image and a fourth text prompt for instructing the semantic segmentation model to segment the target object contained in the prior image into the semantic segmentation model to obtain a second segmentation mask of the target object in the prior image;
[0256] The region where the target object is located and covered by the second segmentation mask is extracted from the prior image, and the extracted region is used as the prior segmentation image of the target object.
[0257] In one possible implementation, when the processor is configured to generate an image containing at least two target objects using the target diffusion model after parameter adjustment, generating visual information of each target object using independent attribute features of each target object within a generation area of each target object, the processor is specifically configured to:
[0258] When an image containing at least two of the target objects is generated using a target diffusion model whose parameters have been adjusted, a fifth text prompt for indicating the generation of an image containing at least two of the target objects is input into the target diffusion model whose parameters have been adjusted. Based on the guidance of the fifth text prompt, the target expansion model generates visual information for each of the target objects in each generation area of each of the target objects using the learned independent attribute features of each of the target objects.
[0259] Through the computer-readable storage medium provided in the embodiments of the present application, during the process of further parameter adjustment of the target diffusion model, the target diffusion model is inputted into the target diffusion model so that the target diffusion model determines first position region information representing the independent position region of each target object in the training sample image. Second position region information representing the actual independent position region of each target object in the training sample image is obtained. A first loss value is calculated between the first position region information and the second position region information, and the first loss value is used to adjust the parameters of the target diffusion model. This enables the target diffusion model to learn to find the actual independent position region of each target object in the training sample image, and learn the independent attribute features of each target object within the actual independent position region of each target object, thereby avoiding learning attribute features that do not belong to the target object itself, that is, when learning the attribute features of target object A, avoiding mixing the attribute features of target object B. Finally, when generating an image containing at least two target objects using the parameter-adjusted target diffusion model, the independent attribute features of each target object are used within each target object's respective generation region to generate visual information of each target object, thereby avoiding the problem of attribute leakage and improving the fidelity of the generated image.
[0260] In an embodiment of the present application, the computer-readable storage medium can also execute other machine-readable instructions when run by the processor to execute the image generation method as described in other embodiments. For the specific steps and principles of the image generation method, please refer to the description of the method side embodiment, which will not be repeated here.
[0261] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. The system embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. There may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some communication interface, indirect coupling or communication connection of the system or unit, which may be electrical, mechanical or other forms.
[0262] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0263] In addition, each functional unit in the embodiments provided in the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0264] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0265] It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures. In addition, the terms "first," "second," "third," etc. are used only to distinguish and describe, and should not be understood as indicating or implying relative importance. Finally, it should be noted that the above-described embodiments are merely specific implementation methods of this application, used to illustrate the technical solutions of this application, rather than to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the above-described embodiments, a person of ordinary skill in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed in this application, or make equivalent substitutions for some of the technical features therein; and such modifications, changes, or substitutions do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of this application. They should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. An image generation method, characterized in that: include: Acquire a text-image pairing sample; wherein the text-image pairing sample includes a training sample image and a matching content text prompt, and the training sample image contains at least two target objects; Inputting the content text prompt into a target diffusion model to generate first position region information for representing an independent position region of each target object in the training sample image determined by the target diffusion model; Acquire second position area information for representing an actual independent position area of each target object in the training sample image, and perform loss calculation on the first position area information and the second position area information to obtain a first loss value; Adjusting parameters of the target diffusion model using the first loss value so that the target diffusion model learns to find the actual independent position region of each target object in the training sample image, and learns the independent attribute features of each target object in the actual independent position region of each target object; When an image containing at least two target objects is generated using the target diffusion model after parameter adjustment, visual information of each target object is generated using the independent attribute features of each target object in the generation area of each target object.
2. The method according to claim 1, characterized in that The content text prompt includes first text prompt information and second text prompt information; the first text prompt information is used to describe each target object contained in the training sample image, and the second text prompt information contains a placeholder for each target object, and the placeholder is used to represent the initial attribute characteristics of the target object; The step of inputting the content text prompt into a target diffusion model to generate first position region information representing an independent position region of each target object in the training sample image determined by the target diffusion model includes: The first text prompt, the second text prompt and the noise image are input into a target diffusion model. The target diffusion model generates a semantically matched first image according to the guidance of the first text prompt and the second text prompt, and generates independent position area information of each target object in the noise image according to the guidance, so as to use the independent position area information as the first position area information; wherein the noise image is an image obtained by adding noise to the training sample image.
3. The method according to claim 2, characterized in that Before adjusting the parameters of the target diffusion model using the first loss value, the method further includes: performing a standard diffusion loss calculation on the first image and the noise image to obtain a second loss value; The step of adjusting parameters of the target diffusion model using the first loss value so that the target diffusion model learns to find the actual independent position region of each target object in the training sample image, and learns the independent attribute features of each target object in the actual independent position region of each target object, including: Adjusting parameters of the target diffusion model using the first loss value and the second loss value, so that the target diffusion model learns to find the actual independent position region of each target object in the training sample image based on the first loss value, and learns the independent attribute features of each target object within the actual independent position region of each target object based on the second loss value; And using the first loss value and the second loss value, iteratively optimize the placeholder.
4. The method according to claim 3, characterized in that The first position area information includes a weight for each position in the noise image as the position of each target object; the second position area information is an independent segmentation mask for each target object in the training sample image; The performing loss calculation on the first location area information and the second location area to obtain a first loss value includes: For each position in the noise image, using a log-likelihood loss function, calculate a log-likelihood loss value between the weight of the position in the first position region information and the mask of the position in the independent segmentation mask; The sum of the log-likelihood loss values corresponding to each position is calculated, and a first loss value corresponding to the text-image paired sample is determined based on the sum.
5. The method according to claim 4, characterized in that: The performing standard diffusion loss calculation on the first image and the noise image to obtain a second loss value includes: Using the standard diffusion loss function, perform standard diffusion loss calculation on the first image and the noise image corresponding to each text-image pair sample included in the current training round to obtain the second loss value corresponding to the current training round; The using the first loss value and the second loss value to adjust parameters of the target diffusion model and iteratively optimize the placeholder includes: After obtaining the first loss value corresponding to each text-image paired sample included in the current training round, calculating the expectation of the first loss value corresponding to each of the text-image paired samples included in the current training round to obtain an expected value; Utilizing the second loss value and the expected value corresponding to the current training round, adjusting parameters of the target diffusion model for the current training round and iteratively optimizing the placeholder for the current training round.
6. The method according to claim 1, characterized in that Before obtaining the text-image paired samples, the method further includes: For each target object, obtaining a plurality of reference images containing only the target object, and generating a plurality of prior images containing only the target object using the target diffusion model; For each of the reference images, segmenting the target object contained in the reference image to obtain a reference segmented image of the target object; and for each of the prior images, segmenting the target object contained in the prior image to obtain a prior segmented image of the target object; Recombining the reference segmented images of different target objects to form an enhanced reference image set including a plurality of enhanced reference images, and recombining the a priori segmented images of different target objects to form an enhanced a priori image set including a plurality of enhanced a priori images; The enhanced reference image and the enhanced prior image are used as the training sample images.
7. The method according to claim 6, characterized in that for each of the reference images, segmenting the target object contained in the reference image to obtain a reference segmented image of the target object; And for each of the prior images, segmenting the target object contained in the prior image to obtain a priori segmented image of the target object, including: Inputting the reference image and a third text prompt for instructing the semantic segmentation model to segment the target object contained in the reference image into the semantic segmentation model to obtain a first segmentation mask of the target object in the reference image; extracting a region where the target object is located and covered by the first segmentation mask from the reference image, and using the extracted region as a reference segmentation image of the target object; Inputting the prior image and a fourth text prompt for instructing the semantic segmentation model to segment the target object contained in the prior image into the semantic segmentation model to obtain a second segmentation mask of the target object in the prior image; The region where the target object is located and covered by the second segmentation mask is extracted from the prior image, and the extracted region is used as the prior segmentation image of the target object.
8. The method according to claim 1, characterized in that: When the target diffusion model after parameter adjustment is used to generate an image containing at least two target objects, generating visual information of each target object by using the independent attribute features of each target object in the generation area of each target object, including: When an image containing at least two of the target objects is generated using a target diffusion model whose parameters have been adjusted, a fifth text prompt for indicating the generation of an image containing at least two of the target objects is input into the target diffusion model whose parameters have been adjusted. Based on the guidance of the fifth text prompt, the target expansion model generates visual information for each of the target objects in each generation area of each of the target objects using the learned independent attribute features of each of the target objects.
9. An image generating device, characterized in that: include: A first acquisition module is configured to acquire a text-image pairing sample, wherein the text-image pairing sample includes a training sample image and a matching content text prompt, and the training sample image contains at least two target objects; A first generating module is configured to input the content text prompt into a target diffusion model to generate first position region information representing an independent position region of each target object in the training sample image determined by the target diffusion model; a first calculation module, configured to obtain second position region information representing an actual independent position region of each target object in the training sample image, and perform loss calculation on the first position region information and the second position region information to obtain a first loss value; an adjustment module, configured to adjust parameters of the target diffusion model using the first loss value, so that the target diffusion model learns to find the actual independent position region of each target object in the training sample image, and learns the independent attribute features of each target object in the actual independent position region of each target object; The second generation module is used to generate visual information of each target object respectively within the generation area of each target object by using the independent attribute features of each target object when generating an image containing at least two target objects by using the target diffusion model after parameter adjustment.
10. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 8 are performed.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the method according to any one of claims 1 to 8.