Image generation method and device and electronic equipment
The diffusion model method uses mask images and cross-attention activation maps to refine parameter adjustments, addressing the lack of spatial control in existing image generation methods, achieving precise alignment of object positions and directions in generated images.
Patent Information
- Application Number
- CN202510161856.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-07-15
AI Technical Summary
In the prior art, when controlling the spatial distribution of an image through text description, the control information is limited, which makes the generated image difficult to meet user needs, especially in the shape and extension direction of the object.
By acquiring the target text and drawing the image, the diffusion model is used to generate mask images and cross attention activation maps, and the model parameters are adjusted based on these images generation loss values, so that the latent spatial characteristics output by the diffusion model match the expected spatial distribution area, graph position and extension direction, thereby generating the target image.
Accurate control of image content is achieved, ensuring that the shape, position and extension direction of objects in the generated image are consistent with the user's drawing graphics, providing rich control information to meet the user's image generation needs.
Smart Images

Figure CN120318347A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to an image generation method, apparatus, and electronic device. Background Art
[0002] Diffusion models can generate images based on text input by users. In related technologies, control methods such as text descriptions, bounding boxes, or region masks are used to control the spatial distribution of objects in the image, for example, position, shape, extension direction, etc. However, the control information provided by these control methods is limited, resulting in poor accuracy in controlling the spatial distribution and making it difficult for the generated images to meet user requirements. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide an image generation method, apparatus, and electronic device to accurately control the spatial distribution such as the shape, position, and extension direction of the image content and meet the user's image generation requirements.
[0004] In a first aspect, an embodiment of the present invention provides an image generation method, the method comprising: obtaining a target text and a drawing image; wherein, the drawing image includes at least one drawing graphic, and the drawing graphic is set with a text marker corresponding to the target text; inputting the target text into a preset diffusion model to output a first latent space feature; generating a mask image corresponding to the drawing image and a cross-attention activation map; wherein, the mask image is used to indicate: the spatial distribution area of the image content corresponding to the text marker; the cross-attention activation map includes: the degree of correlation between the feature vectors at each position in the first latent space feature and the text marker; generating a loss value based on the mask image and the cross-attention activation map; wherein, the loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches at least one of the following: the spatial distribution area indicated by the mask image, the position of the drawing graphic in the drawing image, the extension direction of the drawing graphic in the drawing image; adjusting the model parameters of the diffusion model based on the loss value, outputting a second latent space feature through the adjusted diffusion model, and generating a target image based on the second latent space feature.
[0005] In a second aspect, an embodiment of the present invention provides an image generation device. The device includes: a data acquisition module configured to acquire a target text and a drawn image; wherein, the drawn image includes at least one drawn graphic, and the drawn graphic is provided with a text mark corresponding to the target text; a feature output module configured to input the target text into a preset diffusion model and output a first latent space feature; an intermediate generation module configured to generate a mask image and a cross-attention activation map corresponding to the drawn image; wherein, the mask image is used to indicate: the spatial distribution area of the image content corresponding to the text mark; the cross-attention activation map includes: the degree of correlation between the feature vectors at each position in the first latent space feature and the text mark; a loss generation module configured to generate a loss value based on the mask image and the cross-attention activation map; wherein, the loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches at least one of the following: the spatial distribution area indicated by the mask image, the position of the drawn graphic in the drawn image, the extension direction of the drawn graphic in the drawn image; an image generation module configured to adjust the model parameters of the diffusion model based on the loss value, output a second latent space feature through the adjusted diffusion model, and generate a target image based on the second latent space feature.
[0006] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the above image generation method.
[0007] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the above image generation method.
[0008] The embodiments of the present invention bring the following beneficial effects:
[0009] The above-mentioned image generation method, device and electronic device obtain target text and a drawn image; wherein, the drawn image includes at least one drawn graphic, and the drawn graphic is set with a text mark corresponding to the target text; input the target text into a preset diffusion model to output first latent space features; generate a mask image and a cross-attention activation map corresponding to the drawn image; wherein, the mask image is used to indicate: the spatial distribution area of the image content corresponding to the text mark; the cross-attention activation map includes: the degree of correlation between the feature vectors at each position in the first latent space features and the text mark; generate a loss value based on the mask image and the cross-attention activation map; wherein, the loss value is used to: adjust the model parameters of the diffusion model so that the second latent space features output by the adjusted diffusion model match at least one of the following: the spatial distribution area indicated by the mask image, the position of the drawn graphic in the drawn image, the extension direction of the drawn graphic in the drawn image; adjust the model parameters of the diffusion model based on the loss value, output second latent space features through the adjusted diffusion model, and generate a target image based on the second latent space features.
[0010] In this method, after inputting the target text and the drawn image, a loss value is generated through the mask image and the cross-attention activation function, and the model parameters of the diffusion model are adjusted through the loss value, so that the latent space vectors output by the diffusion model have a high degree of matching with the drawn image. This method can provide rich control information through the drawn image and precisely control the spatial distribution such as the shape, position, and extension direction of the image content, meeting the user's image generation requirements.
[0011] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings.
[0012] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings
[0013] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0014] Figure 1 It is a flowchart of an image generation method provided by an embodiment of the present invention;
[0015] Figure 2Schematic diagram for aligning the directions of the cross-attention activation maps provided by the embodiments of the present invention;
[0016] Figure 3 Schematic diagram for aligning by image moments provided by the embodiments of the present invention;
[0017] Figure 4 Schematic diagram for expanding the spatial distribution region provided by the embodiments of the present invention;
[0018] Figure 5 Schematic diagram for the expansion of the spatial distribution region with time steps provided by the embodiments of the present invention;
[0019] Figure 6 Schematic diagram for the target image of the expanded spatial distribution region provided by the embodiments of the present invention;
[0020] Figure 7 Schematic diagram for the processing process of the diffusion model at multiple time steps provided by the embodiments of the present invention;
[0021] Figure 8 Schematic diagram for the target image of aligning by cross-attention and expanding the spatial distribution region provided by the embodiments of the present invention;
[0022] Figure 9 Schematic diagram for an image generation device provided by the embodiments of the present invention;
[0023] Figure 10 Schematic diagram for an electronic device provided by the embodiments of the present invention. Detailed implementation manners
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0025] For ease of understanding, first, the terms related to the embodiments of the present invention are explained.
[0026] 1. Doodle: In this application, a doodle refers to simple lines or marks randomly drawn by a user. The doodle guides the process of image generation as control information. Usually, the doodle contains the approximate positions and shapes of the objects in the image to be generated.
[0027] 2. Diffusion process: The diffusion process is a process of simulating data, such as image data, gradually being covered by noise until noise data is generated. Then, starting from the noise data, the noise is gradually removed to finally form a clear image.
[0028] 3. Attention Mechanism: The attention mechanism enables the diffusion model to focus on specific parts of the input data. During the process of generating an image from text, the attention mechanism helps the model focus on the key information in the text description and match this key information with specific parts of the image.
[0029] 4. Cross-Attention: That is, Cross-Attention. Cross-Attention is a special attention mechanism that allows the diffusion model to focus on relevant parts of another type of data when processing one type of data; for example, cross-attention allows the diffusion model to focus on relevant parts of image data when processing text data.
[0030] 5. Loss Function: The loss function is a mathematical function that measures the difference between the model's predicted results and the actual results. During the training process, the principle is to minimize the function value of the loss function to adjust the model parameters and improve the model's performance.
[0031] 6. Moments: In image processing, moments are statistics that describe the distribution characteristics of an image. For example, the first-order moment represents the centroid, and the second-order moment represents the extension direction, etc. In this application, through the method of moment alignment, it is ensured that the extension direction of the object in the generated image is consistent with the extension direction of the scribble line, and the position of the object is consistent with the position of the scribble line.
[0032] 7. Alignment: During the image generation process, alignment refers to adjusting the generated image content to match the input scribble, specifically including that the extension direction, position, and shape of the image content respectively match the extension direction, position, and shape of the scribble line.
[0033] 8. Binarization: Converting the pixel values of the scribble into two pixel values, such as 0 and 1, to distinguish the scribble area from the non-scribble area.
[0034] 9. Self-Attention Map: In a neural network, the self-attention map represents the mutual relationship between different parts of an image. In this application, the self-attention map is used to identify and expand the scribble area.
[0035] 10. Kullback-Leibler Divergence: The Kullback-Leibler divergence can measure the difference between two probability distributions and is used to evaluate the similarity between the scribble area and the surrounding area of the scribble area to determine the expansion method of the scribble area.
[0036] 11. Centroid: Also known as the first-order moment, it is used to determine the geometric center of an image or an image area and is used to align the position of the object in the generated image.
[0037] 12. Central Moment: A high-order moment that describes the shape and extension direction of an image area and is used to capture the directionality and distribution characteristics of an object.
[0038] 13. Direction alignment: The process of ensuring that the extension direction of the object in the generated image is consistent with the extension direction of the scribble line.
[0039] 14. Cross-attention activation map: A map indicating the relationship between the scribble and the content of the generated image, used to adjust and optimize the image generation process.
[0040] 15. Focal loss: A loss function used to solve the problem of class imbalance, which is used in this application to align the scribble area with the content of the generated image.
[0041] 16. Gradient descent: An optimization algorithm that iteratively adjusts the model parameters to minimize the loss value of the loss function, used for parameter optimization of the model.
[0042] 17. Backpropagation: In neural network training, it is used to calculate the gradient of the loss value of the loss function with respect to the network parameters for parameter update.
[0043] In the related art, the spatial distribution of the image can be controlled through text description. However, due to the lack of explicit spatial information in the text description itself, the matching degree between the image generated by the diffusion model and the text description is very low. The spatial distribution of the image can also be controlled by combining text with a bounding box or text with a region mask. Among them, the bounding box can provide the position of the object in the image, but it is difficult to provide the shape and extension direction of the object, resulting in the shape and extension direction of the object in the generated image not matching the user's expectations. Compared with the bounding box, the region mask can provide more detailed control information for the spatial distribution, but it requires a high annotation cost and is still difficult to control the extension direction of the object.
[0044] Based on this, the embodiments of the present invention provide an image generation method, device, and electronic device, which can be applied to generate various types of images.
[0045] See Figure 1 An image generation method as shown, the method includes the following steps:
[0046] Step S102, obtain the target text and the drawn image; wherein, the drawn image includes at least one drawn graphic, and the drawn graphic is set with a text mark corresponding to the target text;
[0047] The target text is a text description used to describe the target image to be generated; the target text may include descriptions of the object in the target image, as well as the color, shape, posture, clothing, position, etc. of the object. The object may be a person, an animal, a plant, a still life, a natural scene, etc.
[0048] The drawn image can be generated by the user's drawing or scribbling. The drawn figures in the drawn image can include one or more, and the drawn figures can be points, lines, closed figures, etc. The text label can specifically be the object name. For example, clouds, bridges, rivers, etc. The text label corresponding to the drawn figure represents the object that the user expects to draw at the position where the drawn figure is located, and the shape and extension direction of the drawn figure represent the shape and extension direction of the object that the user expects.
[0049] The text label can be all or part of the aforementioned target text. For example, if the target text is "a cat on the beach", the text label is "cat"; the text label can also be text related to the target text. For example, if the target text is "a cat on the beach", the text label is "orange cat".
[0050] Step S104, input the target text into a preset diffusion model to output the first latent space feature;
[0051] The diffusion model can specifically be a denoising diffusion probabilistic model (abbreviated as DDPM), a classifier-free guidance diffusion model (such as the GLIDE model), a distilled diffusion model, etc.
[0052] In the diffusion model, multi-time-step diffusion processing needs to be performed. In the first time step, the target text and noise data are jointly input into the diffusion model to output the first latent space feature corresponding to the first time step. The noise data can specifically be Gaussian noise.
[0053] In each subsequent time step, the second latent space feature output in the previous time step is the first latent space feature of the current time step. In this embodiment, the process of adjusting the model parameters in the following steps S106, S108, and S110 requires multiple time steps to loop and iterate until the second latent space feature is output in the last time step, thereby generating the target image.
[0054] Step S106, generate a mask image and a cross-attention activation map corresponding to the drawn image; where the mask image is used to indicate: the spatial distribution area of the image content corresponding to the text label; the cross-attention activation map includes: the degree of correlation between the feature vectors at each position in the first latent space feature and the text label;
[0055] The area occupied by the drawn figure in the drawn image corresponds to the spatial distribution area in the mask image; the spatial distribution area can be the same size and shape as the area occupied by the drawn figure; the spatial distribution area can also be an area after a certain expansion of the area occupied by the drawn figure. In the mask image, different pixel values can be used to distinguish the spatial distribution area and the area outside the spatial distribution area; for example, the pixel value within the spatial distribution area is 1, and the pixel value of the area outside the spatial distribution area is 0.
[0056] In the drawn image, the drawn graphic is set with corresponding text labels. Based on the correspondence between the drawn graphic and the text labels, the spatial distribution area in the masked image also has a correspondence with the text labels. Therefore, the spatial distribution area in the masked image is also set with corresponding text labels, and the spatial distribution area is used to display the image content corresponding to the text labels. For example, image content such as the sea, dolphins, tourists, etc.
[0057] In different time steps, the masked image can remain unchanged. At this time, the spatial distribution area in the masked image is the area occupied by the drawn graphic in the drawn image. In addition, the masked image can also be continuously updated as the time step changes. For example, the spatial distribution area continuously expands as the time step changes, etc., so that the shape and size of the spatial distribution area better match the image content corresponding to the text labels.
[0058] The above cross-attention activation map is usually generated based on the aforementioned first latent space features and text labels. Each pixel position in the cross-attention activation map stores a weight value, and this weight value is used to indicate the degree of correlation between the feature vector corresponding to this pixel position and the text label. When the weight value is higher, it indicates that the correlation between this pixel position and the text label is higher; when the weight value is lower, it indicates that the correlation between this pixel position and the text label is lower.
[0059] Step S108, generate a loss value based on the masked image and the cross-attention activation map. Among them, the loss value is used to: adjust the model parameters of the diffusion model so that the second latent space features output by the adjusted diffusion model match at least one of the following: the spatial distribution area indicated by the masked image, the position of the drawn graphic in the drawn image, the extension direction of the drawn graphic in the drawn image.
[0060] In subsequent time steps except the first time step, before generating the second latent space features of this time step, it is necessary to adjust the model parameters of the diffusion model. By adjusting the model parameters of the diffusion model, the second latent space features output by the diffusion model are indirectly adjusted.
[0061] The diffusion model includes a Unet network, and the model parameters of the diffusion model mainly include the neural network parameters of the Unet network, such as weight parameters, bias parameters, activation function parameters, etc.
[0062] Since the masked image indicates the spatial distribution area corresponding to the text labels, and the cross-attention activation map indicates the degree of correlation between the feature vectors at each position of the first latent space features and the text labels. In this embodiment, in multiple time steps, a loss value is generated through the masked image and the cross-attention activation map, and then the model parameters of the diffusion model are adjusted based on the loss value, so that the matching degree between the second latent space features output by the adjusted diffusion model and the drawn image can be getting higher and higher.
[0063] Specifically, after adjusting the model parameters through the loss value, the second latent space features output by the adjusted diffusion model can be made to match the spatial distribution area indicated by the mask image. Also, since the spatial distribution area of the mask image corresponds to the area of the drawn figure in the drawn image, the spatial distribution of each image content in the target image is made to match the spatial distribution of the drawn figure in the drawn image.
[0064] In addition, after adjusting the model parameters through the loss value, the second latent space features output by the adjusted diffusion model can be made to match the position of the drawn figure in the drawn image and the extension direction of the drawn figure in the drawn image. Thus, the position and extension direction of each image content in the target image are respectively made to match the position and extension direction of the drawn figure in the drawn image.
[0065] Step S110: Adjust the model parameters of the diffusion model based on the loss value, output the second latent space features through the adjusted diffusion model, and generate the target image based on the second latent space features.
[0066] In the last time step, after adjusting the model parameters of the diffusion model, the diffusion model outputs the second latent space features; and converts the second latent space features into the target image.
[0067] The target image contains the image content corresponding to the foregoing text markers, and this image content matches the text description in the foregoing target text; at the same time, the spatial distribution of the image content in the target image matches the spatial distribution of the drawn figure in the foregoing drawn image, and the position and extension direction of the graphic content respectively match the position and extension direction of the drawn figure in the drawn image.
[0068] For the above image generation method, obtain the target text and the drawn image; wherein, the drawn image includes at least one drawn figure, and the drawn figure is provided with a text marker corresponding to the target text; input the target text into a preset diffusion model to output the first latent space features; generate a mask image and a cross-attention activation map corresponding to the drawn image; wherein, the mask image is used to indicate: the spatial distribution area of the image content corresponding to the text marker; the cross-attention activation map includes: the degree of correlation between the feature vectors at each position in the first latent space features and the text marker; generate a loss value based on the mask image and the cross-attention activation map; wherein, the loss value is used to: adjust the model parameters of the diffusion model so that the second latent space features output by the adjusted diffusion model match at least one of the following: the spatial distribution area indicated by the mask image, the position of the drawn figure in the drawn image, the extension direction of the drawn figure in the drawn image; adjust the model parameters of the diffusion model based on the loss value, output the second latent space features through the adjusted diffusion model, and generate the target image based on the second latent space features.
[0069] In this method, after inputting the target text and drawing an image, a loss value is generated through a mask image and a cross-attention activation function, and the model parameters of the diffusion model are adjusted by the loss value, so that the latent space vector output by the diffusion model has a high degree of matching with the drawn image. This method can provide rich control information by drawing an image, and accurately control the spatial distribution such as the shape, position, and extension direction of the image content, meeting the user's image generation requirements.
[0070] In a specific implementation manner, a focal loss value is generated based on a mask image and a cross-attention activation map; wherein, the focal loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the spatial distribution area indicated by the mask image.
[0071] The spatial distribution area indicated by the mask image can comprehensively affect the position, size, shape, extension direction, etc. of the image content in the target graph. Specifically, the focal loss value is used to continuously adjust the cross-attention activation map in multiple time steps, so that the cross-attention activation map matches the spatial distribution area indicated by the mask image, and further makes the cross-attention activation map match the drawn graph in the drawn image. The focal loss value can make the adjusted diffusion model focus on the spatial distribution area indicated by the mask image and suppress the importance of areas outside the spatial distribution area.
[0072] Specifically, the focal loss function can be generated in the following manner.
[0073] Based on the cross-attention activation map, an adjustment factor for the focal loss is generated; specifically, the cross-attention activation map can be input into a specified function to output the adjustment factor. For example, a linear mapping function, an activation function, etc.; or a preset arithmetic formula is used to perform operations on the cross-attention activation map to obtain the adjustment factor.
[0074] In one example, the adjustment factor of the focal loss is wherein, is the cross-attention activation map, σ is the Sigmoid function, which is used to map the cross-attention activation map to the interval (0, 1), expressed as a binary probability; β is a parameter used to adjust the size of the adjustment factor;
[0075] Adjust the first weight of the spatial distribution area indicated by the mask image and the second weight of the area outside the spatial distribution area through the adjustment factor of focal loss; the adjustment factor of focal loss is used to reduce the weight of the pixel values that are easy to classify in the cross-attention activation map and increase the weight of the pixel values that are difficult to classify. Among them, the pixel values that are easy to classify are the pixel values whose Sigmoid function values are close to 0 or close to 1, and the weight of the pixel values that are difficult to classify is the pixel values whose Sigmoid function values are close to 0.5. Since the positive and negative probabilities are equal, it is difficult for the model to make a judgment; the above parameters can adjust the intensity of the adjustment of the aforementioned first weight and second weight.
[0076] The mask image is represented as Ms, the pixel values within the spatial distribution area are 1, and the pixel values outside the spatial distribution area are 0. The process of adjusting the first weight of the spatial distribution area indicated by the mask image and the second weight of the area outside the spatial distribution area can be expressed as where a is a parameter used to balance the loss contributions within and outside the spatial distribution area.
[0077] When α is small, reduce the penalty for incorrect predictions within the spatial distribution area, making the model pay more attention to the pixels within the spatial distribution area; when α is large, increase the penalty for incorrect predictions outside the spatial distribution area, suppressing the attention outside the spatial distribution area.
[0078] Then, determine the cross-entropy loss value between the mask image and the cross-attention activation map; based on the adjusted first weight, second weight, and cross-entropy loss value, generate the focal loss value. The cross-entropy loss value is used to measure the difference between the cross-attention activation map and the mask image.
[0079] The focal loss value can be expressed by the following formula:
[0080]
[0081] where represents the focal loss value, S is the set of drawn lines, s represents a drawn line; |S| is the number of drawn lines; C(s) is the set of text tokens; c represents a text token; represents the cross-entropy loss value.
[0082] In actual implementation, the diffusion model first generates a cross-attention activation map This activation map represents the correlation between text tokens and the second latent space features. For each pixel in the activation map, calculate its Sigmoid function value, denoted as The Sigmoid function is used to map the pixel values in the activation map to the interval (0, 1). For pixel values that are difficult to classify, i.e., pixel values where the Sigmoid function value is close to 0.5, the weight is increased by an adjustment factor According to the mask image Ms and the parameter α, the weights inside and outside the spatial distribution region are calculated, and the loss function is adjusted to focus on the pixels within the spatial distribution region. Then, the binary cross-entropy loss LBCE between the cross-attention activation map and the mask image is calculated. Combining all the above parts, the total focal loss is calculated
[0083] During the diffusion process of the diffusion model, the model parameters are continuously adjusted to minimize the focal loss The cross-attention activation map is adjusted by adjusting the model parameters to better align with the mask image.
[0084] The above focal loss value aims to optimize the cross-attention activation map to better align with the spatial distribution region in the mask image. The focal loss value helps the diffusion model focus its attention on the spatial distribution region of the mask image and suppress the activations outside the spatial distribution region.
[0085] In a specific implementation, the first specified image moment of the drawn image and the second specified image moment of the cross-attention activation map are determined; based on the first specified image moment and the second specified image moment, an image moment loss value is generated; wherein, the image moment loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position and / or extension direction of the drawn figure in the drawn image.
[0086] The image moment is a statistic that describes the distribution characteristics of an image. The image moment includes multiple orders. Among them, the first-order moment represents the centroid of the image, and the second-order moment represents the extension direction of the image. In this embodiment, by aligning the image moments, the direction and position of the image content in the generated target image are respectively consistent with the direction and position of the drawn image in the drawn image.
[0087] The aforementioned first specified image moment may include the first-order moment or the second-order moment of the drawn image, and the second specified image moment includes the first-order moment or the second-order moment of the cross-attention activation map; correspondingly, when the first specified image moment may include the first-order moment of the drawn image and the second specified image moment includes the first-order moment of the cross-attention activation map, the aforementioned image moment loss value may make the second latent space feature output by the adjusted diffusion model match the position of the drawn figure in the drawn image.
[0088] When the first specified image moment can include the second moment of the drawn image and the second specified image moment includes the second moment of the cross-attention activation map, the aforementioned image moment loss value can make the second latent space feature output by the adjusted diffusion model match the extension direction of the drawn figure in the drawn image.
[0089] Specifically, determine the first moment of the drawn image and the first moment of the cross-attention activation map; determine the position loss value based on the distance between the first moment of the drawn image and the first moment of the cross-attention activation map; use the position loss value as the image moment loss value; where the position loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position of the drawn figure in the drawn image.
[0090] The first moment is the centroid of the image. The first moment of the drawn image is the centroid of the drawn image; the first moment of the cross-attention activation map is the centroid of the cross-attention activation map. The centroid represents the geometric center of the image. Centroid can be represented by the following formula:
[0091]
[0092] where, m 10 , m 00 , m 01 , m 00 are represented as m pq , m pq can be obtained through the following formula:
[0093]
[0094] where, I(x, y) represents the pixel value of image I at the (x, y) coordinate; when image I is the drawn image, the centroid of the drawn image, that is, the first moment, can be obtained through the above formula; when image I is the cross-attention activation map, the centroid of the cross-attention activation map, that is, the first moment, can be obtained through the above formula.
[0095] The above position loss value can be represented by the following formula:
[0096]
[0097] where, is the position loss value; C(s) is the set of text tokens; c represents a text token; (x c , y c ) is the centroid of the cross-attention activation map, and (x s , y s ) is the centroid of the drawn image.
[0098] By using the above-mentioned position loss value, the difference between the centroid of the drawn image and the centroid of the cross-attention activation map can be compared, and the deviation between the position of the image content in the generated target image and the position of the user's drawn figure can be quantified; by calculating and aligning the first-order image moments, it can be ensured that the position of the image content in the generated target image is consistent with the position of the user's drawn figure.
[0099] In another way, determine the second moment of the drawn image and the second moment of the cross-attention activation map; based on the distance between the second moment of the drawn image and the second moment of the cross-attention activation map, determine the direction loss value; use the direction loss value as the image moment loss value; wherein, the direction loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the extension direction of the drawn figure in the drawn image.
[0100] In addition to position alignment, the alignment of the extension direction of the image content also needs to be considered, which can be achieved by aligning the direction moments in the image moments: the direction moments are the above-mentioned second moments, such as m 20 , m 02 and m 11 etc., and these second moments describe the shape and direction of the image region.
[0101] The above-mentioned direction loss value can be expressed by the following formula:
[0102]
[0103] Wherein, is the direction loss value, C(s) is the set of text tokens, c represents a text token; θ c is the direction moment of the cross-attention activation map, θ s is the direction moment of the drawn image; the direction moment of the cross-attention activation map or the direction moment of the drawn image can both be expressed by the following formula:
[0104]
[0105] Wherein, when calculating the direction moment of the cross-attention activation map, is the centroid of the cross-attention activation map, when calculating the direction moment of the drawn image, is the centroid of the drawn image. m 00 , m 11 , m 01 , m 00 is expressed as m pq , and can be calculated by the above formula of m pq .
[0106] Through the direction loss value, the direction moments of the drawn image and the cross-attention activation map are compared. The direction loss value can quantify the deviation between the extension direction of the image content in the generated target image and the extension direction of the user's drawn figure. By calculating and aligning the second-order image moments, it can be ensured that the extension direction of the image content in the generated target image is consistent with the extension direction of the user's drawn figure.
[0107] In order to make the extension direction and position of the image content in the generated target image consistent with the extension direction and position of the drawn figure input by the user respectively, the loss value needs to include the aforementioned direction loss value and position loss value. Based on this, the loss value can be generated in the following manner.
[0108] Based on the distance between the first-order moment of the drawn image and the first-order moment of the cross-attention activation map, the position loss value is determined; wherein, the position loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position of the drawn figure in the drawn image; based on the distance between the second-order moment of the drawn image and the second-order moment of the cross-attention activation map, the direction loss value is determined; wherein, the direction loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the extension direction of the drawn figure in the drawn image; the weighted sum of the direction loss value and the position loss value is determined as the image moment loss value.
[0109] The direction loss value and the position loss value can be obtained by calculation in the foregoing embodiments. The image moment loss value can be expressed by the following formula:
[0110] L moment = λ1L centroid + λ2L central
[0111] wherein, L moment is the image moment loss value, L centroid is the position loss value, L central is the direction loss value; λ1, λ2 are weighting parameters, and the weights of the direction loss value and the position loss value can be adjusted.
[0112] Since the image moment is an image feature with translation, rotation, and scale invariance, it can be used for image recognition. In this embodiment, the spatial distribution of the cross-attention activation map is regarded as the image moment, and these image moments are used to align the position of the image content in the generated target image with the position of the drawn figure in the drawn image, and to align the extension direction of the image content in the target image with the extension direction of the drawn figure in the drawn image.
[0113] In Figure 2 the example of Figure 1In it, the red arrow indicates that the extension direction of the image area is from the lower left to the upper right. In the mask image, the extension direction of the image area is from the upper left to the lower right. Through the adjustment of the above direction loss value, the direction alignment is achieved. In the cross-attention activation Figure 2 the extension direction is adjusted to the upper left to lower right direction to match the mask image.
[0114] In Figure 3 the example of, the drawn image includes three drawn graphics, and the corresponding text labels are cat, butterfly, and meadow respectively. For the drawn graphic with the text label cat, the extension direction of this drawn graphic is from the lower left to the upper right. If the model parameters are not updated using the above image moment loss value, the extension direction of the cat in the generated image content is from the upper left to the lower right, which does not match the extension direction of the drawn graphic. In the corresponding cross-attention activation map a, the extension direction of the activation area corresponding to this drawn graphic does not match the mask image.
[0115] If the model parameters are updated using the above image moment loss value, the extension direction of the cat in the generated image content is from the lower left to the upper right, which matches the extension direction of the drawn graphic. In the corresponding cross-attention activation map b, the extension direction of the activation area corresponding to this drawn graphic matches the mask image.
[0116] Furthermore, based on the mask image and the cross-attention activation map, a focal loss value is generated; wherein, the focal loss value is used for: adjusting the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the spatial distribution area indicated by the mask image; based on the first specified image moment of the drawn image and the second specified image moment of the cross-attention activation map, an image moment loss value is generated; wherein, the image moment loss value is used for: adjusting the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position and / or extension direction of the drawn graphic in the drawn image; the sum of the focal loss value and the image moment loss value is determined as the final loss value.
[0117] Both the focal loss value and the image moment loss value can be calculated through the foregoing embodiments. The final loss value can be expressed by the following formula:
[0118] L cross = L focal + L moment
[0119] Wherein, L cross is the final loss value, L focal is the focal loss value, and L moment is the image moment loss value.
[0120] The focal loss value is used to enhance the attention of the diffusion model to the drawn image, ensuring the spatial consistency between the generated target image and the drawn image. By adjusting the attention distribution of the diffusion model to the input drawn image during the image generation process, the alignment degree, that is, the matching degree, between the second latent space features and the drawn image is continuously improved.
[0121] The image moment loss value is used to further improve the control precision of the extension direction and position of the image content in the target image. By adjusting the position and direction of the image content, it is ensured that the generated image content not only matches the drawn figure in shape, but also remains consistent with the drawn figure in terms of direction and position.
[0122] During the operation of the diffusion model, the drawn image is used to guide the adjustment of the second latent space features. At each time step of the reverse diffusion, the second latent space features are adjusted by aligning according to the focal loss value and the image moment loss value. The gradient of the loss value is calculated through the backpropagation algorithm, and the model parameters are updated. The update of the model parameters aims to minimize the loss value, ensuring that the generated image content reflects both the spatial distribution of the drawn figure and the extension direction of the drawn figure in the drawn image. Although the drawn image does not directly become a part of the second latent space features, it indirectly affects the iterative inference process of the second latent space features at multiple time steps. Through the iterative inference at multiple time steps, the diffusion model can generate a target image that is aligned with the user's drawn image in terms of space and direction.
[0123] In the above method, combining the focal loss value and the image moment loss value can help the diffusion model improve the image generation effect in multiple dimensions, including spatial alignment, directionality, and overall consistency. The user can precisely control the spatial distribution of the image content in the generated image through the drawn image, meeting the user's image generation requirements.
[0124] In a specific implementation, considering that the drawn figure may be relatively sparse or thin, resulting in the influence on the control precision of the drawn figure on the target image, this embodiment can also control the mask image to be continuously updated along with the time steps.
[0125] Specifically, generate the self-attention map of the first latent space features; wherein, the self-attention map is used to indicate the correlation degree between the feature vectors at different positions in the first latent space features; the self-attention map represents the association between the feature vectors at each position within the first latent space features.
[0126] Determine a target image region corresponding to the spatial distribution region indicated by the mask image from the self-attention map; based on the target image region, determine an extended anchor point from the mask image; wherein, the extended anchor point is used to indicate: the extension direction and / or extension distance of the extension of the spatial distribution region in the mask image; in actual implementation, the extended anchor point is located outside the spatial distribution region, and when the spatial distribution region is extended, the extended anchor point needs to be included in the extended spatial distribution region. Therefore, the position of the extended anchor point determines the extension direction and extension distance of the spatial distribution region.
[0127] The pixel positions in the self-attention map contain weight values, and the weight values indicate the association between the feature vector corresponding to the pixel position and the feature vectors corresponding to other upward positions. Therefore, based on the weight values of each pixel position in the target image region, pixel positions with a strong association with the target image region can be screened from outside the spatial distribution region as extended anchor points.
[0128] Specifically, determine the average distribution of the self-attention values in the target image region; based on the average distribution and the self-attention map, determine an extended anchor point from the mask image. The average distribution can be represented by the following formula:
[0129]
[0130] where, represents the average distribution, |S| is the number of pixels in the target image region, and A self [i, j] is the pixel value at the pixel position [i, j].
[0131] Furthermore, obtain the self-attention values corresponding to the positions outside the spatial distribution region in the mask image from the self-attention map; based on the self-attention values corresponding to the positions outside the spatial distribution region and the average distribution, determine the similarity between the positions outside the spatial distribution region and the edge positions of the spatial distribution region; based on the similarity, determine a specified number of extended anchor points from the positions outside the spatial distribution region.
[0132] The edge positions of the spatial distribution region are denoted as Bs, and assume the extended anchor point is (x, y); specifically, the Kullback-Leibler divergence algorithm can be used to calculate the similarity, and the Kullback-Leibler divergence algorithm can be used to measure the difference between the positions outside the spatial distribution region and the edge positions of the spatial distribution region.
[0133] First, the distance D s between the edge position B s (x, y):
[0134]
[0135] Among them, D KL represents the Kullback-Leibler divergence algorithm function; the distance D s (x, y) enables the adjusted diffusion model to determine which regions outside the spatial distribution area are closest to the semantic content within the spatial distribution area, and these regions should be extended into the spatial distribution area.
[0136] After calculating D s (x, y), the D s (x, y) corresponding to each pixel outside the spatial distribution area can be sorted, and then a specified number of pixels with the smallest D s (x, y) values are selected as the expansion anchor points.
[0137] Based on the expansion anchor points, the spatial distribution area in the mask image is expanded to obtain the processed mask image. Among them, the expanded spatial distribution area can be represented by the following formula:
[0138]
[0139] Among them, N(B s ) represents the neighborhood of the aforementioned edge position B s , S ′ is the expanded spatial distribution area. argmin represents taking the minimum value of the independent variable D s (x, y), S represents the spatial distribution area before expansion, and k is the number of expansion anchor points.
[0140] In the Figure 4 example, in mask image 1, the pixel positions of the white box form the spatial distribution area. After expanding the spatial distribution area, the pixel positions of the red box in mask image 2 are expanded into the spatial distribution area. In the Figure 5 example, four mask images are shown, and the spatial distribution area in each mask image expands with the time step.
[0141] In the Figure 6 example, the drawn image includes five drawn graphics. Among them, the text label of one drawn graphic is "ship", the text label of one drawn graphic is "ocean", and the text labels of the other three drawn graphics are "dolphins"; if the spatial distribution area is not expanded, the diffusion model may not fully capture the three drawn graphics and may not fully understand that the three drawn graphics are all dolphins. Then, in the target image, only one dolphin is generated among the three drawn graphics labeled as dolphins, and the other two drawn graphics generate waves.
[0142] After expanding the spatial distribution area, the diffusion model can fully capture the three drawn figures and fully understand that all three drawn figures are dolphins. Then, in the target image, among the three drawn figures marked as dolphins, three dolphins are correspondingly generated, which is more in line with the user's intention.
[0143] In the above method, at the early stage of denoising, a rough mask image is created according to the drawn image, and then refined in a more detailed denoising stage, continuously expanding the spatial distribution area in the mask image, thereby improving the alignment between the mask image and the drawn image and improving the control accuracy of the drawn image over the image generation process.
[0144] Figure 7 Shows the processing process of the diffusion model at multiple time steps; where Unet represents the network in the diffusion model, r represents the r-th time step, and the time steps decrease sequentially until the 0-th time step.
[0145] First, the target text and Gaussian noise are input into the diffusion module, and the first latent space feature r-1 is output; the mask image corresponding to the drawn image and the first latent space feature are input into the cross-attention alignment module, and the cross-attention alignment module implements the following steps: generating a loss value based on the mask image and the cross-attention activation map; where the loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches at least one of the following: the spatial distribution area indicated by the mask image, the position of the drawn figure in the drawn image, and the extension direction of the drawn figure in the drawn image; adjusting the model parameters of the diffusion model based on the loss value.
[0146] That is, the cross-attention module is used to adjust the model parameters of the diffusion model through the loss value, and then in the next time step r-2, the first latent space feature r-1 is further iteratively processed through the diffusion model with adjusted parameters to output the second latent space feature r-2; at the same time, in the time step r-2, the mask image is processed for expanding the spatial distribution area to obtain the processed mask image; the latent space feature r-2 and the processed mask image are input into the cross-attention alignment module to continue adjusting the parameters of the diffusion model.
[0147] And so on, at time step 0, the diffusion model outputs the second latent space feature 0, and the target image is obtained by converting the second latent space feature 0.
[0148] In Figure 8 In the example, the drawn image includes two drawn figures, with the text labels being plant and cat respectively; compared with the target image 1 without using cross-attention alignment and the target image 2 without expanding the spatial distribution area, the positions, directions, etc. of the plant and the cat in the target image 3 are more in line with the drawn figures in the drawn image, which is more in line with the user's intention.
[0149] The expansion process of the spatial distribution region is gradually realized during the diffusion process of multiple time steps in the diffusion model, thereby improving the shape of the image content and enhancing visual coherence. As the time step progresses, the spatial distribution region gradually expands, helping to better define the outline of the image content in the target image. The aforementioned cross-attention activation map can ensure that the generated image content objects not only match the drawn graphics in position but also are consistent with the drawn graphic scribbles in the extension direction. This alignment helps capture the direction information of the drawn graphics, making the generated image content conform to the user's intention.
[0150] By continuously expanding the spatial distribution region, more abundant information can be provided for the cross-attention activation map, making it easier to perform accurate direction alignment. At the same time, due to the expansion of the spatial distribution region, the cross-attention alignment can more precisely adjust the extension direction of the generated image content, and then feedback to the expansion process of the spatial distribution region, guiding the direction and degree of the next expansion of the spatial distribution region. This method enables the generated image content to more accurately reflect the user's drawn graphics and be more visually coherent.
[0151] Generally speaking, the expansion process of the spatial distribution region and the cross-attention alignment do not operate in isolation but complement each other. The expansion process of the spatial distribution region can provide more spatial information, and the cross-attention activation map uses this spatial information to adjust the direction of the image content. The combination of the two improves the spatial control and consistency of the image content.
[0152] For the image generation method provided in this embodiment, the user simply provides the drawn graphics as visual cues. By expanding the spatial distribution region during the diffusion model processing and cross-attention alignment, the user's intention can be captured more accurately, and the spatial control of image generation can be achieved more precisely. Among them, by iteratively expanding the spatial distribution region, the coverage range of the spatial distribution region can be enhanced, thereby providing more abundant and accurate spatial information; using image moments to align the direction of the objects in the generated image with the direction of the scribbles to ensure the consistency of the generated image and the scribble cues in terms of spatial layout and object direction.
[0153] The image generation method of this embodiment does not require model training, avoiding the expensive training process. Through the optimization strategy and loss function design, fast deployment and flexible adaptation to new conditional inputs can be achieved. At the same time, in this embodiment, by the user inputting the drawn graphics, the image generation can be guided through scribbles, providing an intuitive and user-friendly way to guide image generation. At the same time, fine-grained adjustment of the object details in the generated image can be realized, thus achieving more accurate, efficient, and user-intention-compliant text-to-image generation.
[0154] Corresponding to the above method embodiment, refer to Figure 9Schematic diagram of an image generation device, the device comprising:
[0155] A data acquisition module 90, configured to acquire a target text and a drawing image; wherein, the drawing image includes at least one drawing graphic, and the drawing graphic is provided with a text marker corresponding to the target text;
[0156] A feature output module 92, configured to input the target text into a preset diffusion model and output a first latent space feature;
[0157] An intermediate generation module 94, configured to generate a mask image and a cross-attention activation map corresponding to the drawing image; wherein, the mask image is used to indicate: the spatial distribution area of the image content corresponding to the text marker; the cross-attention activation map includes: the correlation degree between the feature vectors at each position in the first latent space feature and the text marker;
[0158] A loss generation module 96, configured to generate a loss value based on the mask image and the cross-attention activation map; wherein, the loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches at least one of the following: the spatial distribution area indicated by the mask image, the position of the drawing graphic in the drawing image, the extension direction of the drawing graphic in the drawing image;
[0159] An image generation module 98, configured to adjust the model parameters of the diffusion model based on the loss value, output the second latent space feature through the adjusted diffusion model, and generate a target image based on the second latent space feature.
[0160] The above image generation device acquires a target text and a drawing image; wherein, the drawing image includes at least one drawing graphic, and the drawing graphic is provided with a text marker corresponding to the target text; inputs the target text into a preset diffusion model and outputs a first latent space feature; generates a mask image and a cross-attention activation map corresponding to the drawing image; wherein, the mask image is used to indicate: the spatial distribution area of the image content corresponding to the text marker; the cross-attention activation map includes: the correlation degree between the feature vectors at each position in the first latent space feature and the text marker; generates a loss value based on the mask image and the cross-attention activation map; wherein, the loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches at least one of the following: the spatial distribution area indicated by the mask image, the position of the drawing graphic in the drawing image, the extension direction of the drawing graphic in the drawing image; adjusts the model parameters of the diffusion model based on the loss value, outputs the second latent space feature through the adjusted diffusion model, and generates a target image based on the second latent space feature.
[0161] In this method, after inputting the target text and drawing an image, a loss value is generated through a mask image and a cross-attention activation function, and the model parameters of the diffusion model are adjusted through the loss value, so that the latent space vector output by the diffusion model has a high degree of matching with the drawn image. This method can provide rich control information by drawing an image and precisely control the spatial distribution such as the shape, position, and extension direction of the image content, meeting the user's image generation requirements.
[0162] The above loss generation module is used to: generate a focal loss value based on the mask image and the cross-attention activation map; wherein, the focal loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the spatial distribution area indicated by the mask image.
[0163] The above loss generation module is used to: generate an adjustment factor for the focal loss based on the cross-attention activation map; adjust the first weight of the spatial distribution area indicated by the mask image and the second weight of the area outside the spatial distribution area through the adjustment factor; determine the cross-entropy loss value between the mask image and the cross-attention activation map; generate a focal loss value based on the adjusted first weight, second weight, and cross-entropy loss value.
[0164] The above loss generation module is used to: determine the first specified image moment of the drawn image and the second specified image moment of the cross-attention activation map; generate an image moment loss value based on the first specified image moment and the second specified image moment; wherein, the image moment loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position and / or extension direction of the drawn figure in the drawn image.
[0165] The above loss generation module is used to: determine the first-order moment of the drawn image and the first-order moment of the cross-attention activation map; determine a position loss value based on the distance between the first-order moment of the drawn image and the first-order moment of the cross-attention activation map; use the position loss value as the image moment loss value; wherein, the position loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position of the drawn figure in the drawn image.
[0166] The above loss generation module is used to: determine the second-order moment of the drawn image and the second-order moment of the cross-attention activation map; determine a direction loss value based on the distance between the second-order moment of the drawn image and the second-order moment of the cross-attention activation map; use the direction loss value as the image moment loss value; wherein, the direction loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the extension direction of the drawn figure in the drawn image.
[0167] The above loss generation module is used to: determine a position loss value based on the distance between the first moment of the drawn image and the first moment of the cross-attention activation map; wherein, the position loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position of the drawn figure in the drawn image; determine an orientation loss value based on the distance between the second moment of the drawn image and the second moment of the cross-attention activation map; wherein, the orientation loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the extension direction of the drawn figure in the drawn image; determine the weighted sum of the orientation loss value and the position loss value as the image moment loss value.
[0168] The above loss generation module is used to: generate a focal loss value based on the mask image and the cross-attention activation map; wherein, the focal loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the spatial distribution area indicated by the mask image; generate an image moment loss value based on the first specified image moment of the drawn image and the second specified image moment of the cross-attention activation map; wherein, the image moment loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position and / or extension direction of the drawn figure in the drawn image; determine the sum of the focal loss value and the image moment loss value as the final loss value.
[0169] The above device further includes a region expansion module, which is used to: generate a self-attention map of the first latent space feature; wherein, the self-attention map is used to indicate the correlation degree between the feature vectors at different positions in the first latent space feature; determine a target image region corresponding to the spatial distribution area indicated by the mask image from the self-attention map; determine an expansion anchor point from the mask image based on the target image region; wherein, the expansion anchor point is used to indicate the expansion direction and / or expansion distance for expanding the spatial distribution area in the mask image; perform an expansion process on the spatial distribution area in the mask image based on the expansion anchor point to obtain a processed mask image.
[0170] The above region expansion module is used to: determine the average distribution of the self-attention values within the target image region; determine the expansion anchor point from the mask image based on the average distribution and the self-attention map.
[0171] The above region expansion module is used to: obtain the self-attention values corresponding to the positions outside the spatial distribution area in the mask image from the self-attention map; determine the similarity between the positions outside the spatial distribution area and the edge positions of the spatial distribution area based on the self-attention values corresponding to the positions outside the spatial distribution area and the average distribution; determine a specified number of expansion anchor points from the positions outside the spatial distribution area based on the similarity.
[0172] This embodiment also provides an electronic device, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the above image generation method. The electronic device can be a server or a terminal device.
[0173] See Figure 10 As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores computer-executable instructions that can be executed by the processor 100, and the processor 100 executes the computer-executable instructions to implement the above image generation method.
[0174] Furthermore, Figure 10 the electronic device shown also includes a bus 102 and a communication interface 103. The processor 100, the communication interface 103, and the memory 101 are connected through the bus 102.
[0175] Among them, the memory 101 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 103 (which can be wired or wireless), a communication connection is established between this system network element and at least one other network element. The Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 102 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 10 only a bidirectional arrow is used in
[0176] The processor 100 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method may be completed by the integrated logic circuit of the hardware in the processor 100 or the instructions in the form of software. The above-mentioned processor 100 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention may be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and combines its hardware to complete the steps of the method in the foregoing embodiments.
[0177] The processor in the above electronic device can implement the following operations in the above image generation method by executing computer-executable instructions:
[0178] An image generation method, which obtains target text and draws an image; wherein, drawing the image includes at least one drawing graphic, and the drawing graphic is set with a text marker corresponding to the target text; inputting the target text into a preset diffusion model to output a first latent space feature; generating a mask image and a cross-attention activation map corresponding to the drawn image; wherein, the mask image is used to indicate: the spatial distribution area of the image content corresponding to the text marker; the cross-attention activation map includes: the degree of correlation between the feature vectors at each position in the first latent space feature and the text marker; generating a loss value based on the mask image and the cross-attention activation map; wherein, the loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches at least one of the following: the spatial distribution area indicated by the mask image, the position of the drawing graphic in the drawn image, the extension direction of the drawing graphic in the drawn image; adjusting the model parameters of the diffusion model based on the loss value, outputting a second latent space feature through the adjusted diffusion model, and generating a target image based on the second latent space feature.
[0179] Generating a focal loss value based on the mask image and the cross-attention activation map; wherein, the focal loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the spatial distribution area indicated by the mask image.
[0180] Generating an adjustment factor for the focal loss based on the cross-attention activation map; adjusting, through the adjustment factor, the first weight of the spatial distribution area indicated by the mask image and the second weight of the area outside the spatial distribution area; determining the cross-entropy loss value between the mask image and the cross-attention activation map; generating a focal loss value based on the adjusted first weight, second weight, and cross-entropy loss value.
[0181] Determining the first specified image moment of the drawn image and the second specified image moment of the cross-attention activation map; generating an image moment loss value based on the first specified image moment and the second specified image moment; wherein, the image moment loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position and / or extension direction of the drawing graphic in the drawn image.
[0182] Determining the first-order moment of the drawn image and the first-order moment of the cross-attention activation map; determining a position loss value based on the distance between the first-order moment of the drawn image and the first-order moment of the cross-attention activation map; taking the position loss value as the image moment loss value; wherein, the position loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position of the drawing graphic in the drawn image.
[0183] Determine the second-order moment of the drawn image and the second-order moment of the cross-attention activation map; determine the direction loss value based on the distance between the second-order moment of the drawn image and the second-order moment of the cross-attention activation map; use the direction loss value as the image moment loss value; wherein, the direction loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the extension direction of the drawn figure in the drawn image.
[0184] Determine the position loss value based on the distance between the first-order moment of the drawn image and the first-order moment of the cross-attention activation map; wherein, the position loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position of the drawn figure in the drawn image; determine the direction loss value based on the distance between the second-order moment of the drawn image and the second-order moment of the cross-attention activation map; wherein, the direction loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the extension direction of the drawn figure in the drawn image; determine the weighted sum of the direction loss value and the position loss value as the image moment loss value.
[0185] Generate a focal loss value based on the mask image and the cross-attention activation map; wherein, the focal loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the spatial distribution area indicated by the mask image; generate an image moment loss value based on the first specified image moment of the drawn image and the second specified image moment of the cross-attention activation map; wherein, the image moment loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position and / or extension direction of the drawn figure in the drawn image; determine the sum of the focal loss value and the image moment loss value as the final loss value.
[0186] Generate a self-attention map of the first latent space feature; wherein, the self-attention map is used to indicate: the degree of correlation between the feature vectors at different positions in the first latent space feature; determine the target image area corresponding to the spatial distribution area indicated by the mask image from the self-attention map; determine an extended anchor point from the mask image based on the target image area; wherein, the extended anchor point is used to indicate: the extension direction and / or extension distance of the spatial distribution area in the mask image for extension; perform an extension process on the spatial distribution area in the mask image based on the extended anchor point to obtain a processed mask image.
[0187] Determine the average distribution of the self-attention values within the target image area; determine the extended anchor point from the mask image based on the average distribution and the self-attention map.
[0188] Obtain the self-attention values corresponding to the positions outside the spatially distributed region in the masked image from the self-attention map; determine the similarity between the positions outside the spatially distributed region and the edge positions of the spatially distributed region based on the self-attention values corresponding to the positions outside the spatially distributed region and the average distribution; determine a specified number of extended anchor points from the positions outside the spatially distributed region based on the similarity.
[0189] In this method, after inputting the target text and the drawn image, a loss value is generated through the masked image and the cross-attention activation function, and the model parameters of the diffusion model are adjusted through the loss value, so that the latent space vector output by the diffusion model has a high degree of matching with the drawn image. This method can provide rich control information through the drawn image and precisely control the spatial distribution such as the shape, position, and extension direction of the image content, meeting the user's image generation requirements.
[0190] This embodiment also provides a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the above image generation method.
[0191] The computer-executable instructions stored in the above computer-readable storage medium can, by executing the computer-executable instructions, implement the following operations in the above image generation method:
[0192] An image generation method includes obtaining a target text and a drawn image; wherein, the drawn image includes at least one drawn graphic, and the drawn graphic is set with a text marker corresponding to the target text; inputting the target text into a preset diffusion model to output a first latent space feature; generating a masked image and a cross-attention activation map corresponding to the drawn image; wherein, the masked image is used to indicate the spatially distributed region of the image content corresponding to the text marker; the cross-attention activation map includes the correlation degree between the feature vectors at each position in the first latent space feature and the text marker; generating a loss value based on the masked image and the cross-attention activation map; wherein, the loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches at least one of the following: the spatially distributed region indicated by the masked image, the position of the drawn graphic in the drawn image, the extension direction of the drawn graphic in the drawn image; adjusting the model parameters of the diffusion model based on the loss value, outputting a second latent space feature through the adjusted diffusion model, and generating a target image based on the second latent space feature.
[0193] Generate a focal loss value based on the masked image and the cross-attention activation map; wherein, the focal loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the spatially distributed region indicated by the masked image.
[0194] Generate an adjustment factor for the focal loss based on the cross-attention activation map; adjust the first weight of the spatial distribution area indicated by the mask image and the second weight of the area outside the spatial distribution area through the adjustment factor; determine the cross-entropy loss value between the mask image and the cross-attention activation map; generate a focal loss value based on the adjusted first weight, second weight, and cross-entropy loss value.
[0195] Determine the first specified image moment of the drawn image and the second specified image moment of the cross-attention activation map; generate an image moment loss value based on the first specified image moment and the second specified image moment; wherein, the image moment loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position and / or extension direction of the drawn figure in the drawn image.
[0196] Determine the first-order moment of the drawn image and the first-order moment of the cross-attention activation map; determine a position loss value based on the distance between the first-order moment of the drawn image and the first-order moment of the cross-attention activation map; use the position loss value as the image moment loss value; wherein, the position loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position of the drawn figure in the drawn image.
[0197] Determine the second-order moment of the drawn image and the second-order moment of the cross-attention activation map; determine an orientation loss value based on the distance between the second-order moment of the drawn image and the second-order moment of the cross-attention activation map; use the orientation loss value as the image moment loss value; wherein, the orientation loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the extension direction of the drawn figure in the drawn image.
[0198] Determine a position loss value based on the distance between the first-order moment of the drawn image and the first-order moment of the cross-attention activation map; wherein, the position loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position of the drawn figure in the drawn image; determine an orientation loss value based on the distance between the second-order moment of the drawn image and the second-order moment of the cross-attention activation map; wherein, the orientation loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the extension direction of the drawn figure in the drawn image; determine the weighted sum of the orientation loss value and the position loss value as the image moment loss value.
[0199] Generate a focal loss value based on a masked image and a cross-attention activation map; wherein, the focal loss value is used to: adjust the model parameters of the diffusion model so that the second latent space features output by the adjusted diffusion model match the spatial distribution area indicated by the masked image; Generate an image moment loss value based on the first specified image moment of the drawn image and the second specified image moment of the cross-attention activation map; wherein, the image moment loss value is used to: adjust the model parameters of the diffusion model so that the second latent space features output by the adjusted diffusion model match the position and / or extension direction of the drawn figure in the drawn image; Determine the sum of the focal loss value and the image moment loss value as the final loss value.
[0200] Generate a self-attention map of the first latent space features; wherein, the self-attention map is used to indicate the degree of correlation between feature vectors at different positions in the first latent space features; Determine the target image area corresponding to the spatial distribution area indicated by the masked image from the self-attention map; Based on the target image area, determine an extended anchor point from the masked image; wherein, the extended anchor point is used to indicate the extension direction and / or extension distance of the expansion of the spatial distribution area in the masked image; Based on the extended anchor point, perform an expansion process on the spatial distribution area in the masked image to obtain a processed masked image.
[0201] Determine the average distribution of the self-attention values within the target image area; Based on the average distribution and the self-attention map, determine an extended anchor point from the masked image.
[0202] Obtain the self-attention values corresponding to the positions outside the spatial distribution area in the masked image from the self-attention map; Based on the self-attention values corresponding to the positions outside the spatial distribution area and the average distribution, determine the similarity between the positions outside the spatial distribution area and the edge positions of the spatial distribution area; Based on the similarity, determine a specified number of extended anchor points from the positions outside the spatial distribution area.
[0203] In this method, after inputting the target text and the drawn image, a loss value is generated through the masked image and the cross-attention activation function, and the model parameters of the diffusion model are adjusted through the loss value, so that the latent space vector output by the diffusion model has a high degree of matching with the drawn image. This method can provide rich control information through the drawn image and accurately control the spatial distribution such as the shape, position, and extension direction of the image content, meeting the user's image generation requirements.
[0204] The computer program product of the image generation method, device, and electronic device provided by the embodiments of the present invention includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments and will not be elaborated here.
[0205] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0206] In addition, in the description of the embodiments of the present invention, unless otherwise clearly defined and limited, the terms "install", "connect", and "couple" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0207] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0208] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0209] Finally, it should be noted that the above embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. An image generation method, characterized in that, The method includes: Obtaining a target text and a drawn image; wherein, the drawn image includes at least one drawn graphic, and the drawn graphic is provided with a text marker corresponding to the target text; Inputting the target text into a preset diffusion model to output a first latent space feature; Generating a mask image and a cross-attention activation map corresponding to the drawn image; wherein, the mask image is used to indicate: the spatial distribution area of the image content corresponding to the text marker; the cross-attention activation map includes: the correlation degree between the feature vectors at each position in the first latent space feature and the text marker; Generating a loss value based on the mask image and the cross-attention activation map; wherein, the loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches at least one of the following: the spatial distribution area indicated by the mask image, the position of the drawn graphic in the drawn image, the extension direction of the drawn graphic in the drawn image; Adjusting the model parameters of the diffusion model based on the loss value, outputting the second latent space feature through the adjusted diffusion model, and generating a target image based on the second latent space feature.
2. The method according to claim 1, characterized in that The step of generating a loss value based on the mask image and the cross-attention activation map includes: Generating a focal loss value based on the mask image and the cross-attention activation map; wherein, the focal loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the spatial distribution area indicated by the mask image.
3. The method according to claim 2, wherein The step of generating a focal loss value based on the mask image and the cross-attention activation map includes: Generating an adjustment factor for the focal loss based on the cross-attention activation map; Adjusting a first weight of the spatial distribution area indicated by the mask image and a second weight of the area outside the spatial distribution area through the adjustment factor; Determining the cross-entropy loss value between the mask image and the cross-attention activation map; Generating a focal loss value based on the adjusted first weight, the second weight, and the cross-entropy loss value.
4. The method according to claim 1, characterized in that, The step of generating a loss value based on the mask image and the cross-attention activation map includes: Determining a first specified image moment of the drawn image and a second specified image moment of the cross-attention activation map; Generating an image moment loss value based on the first specified image moment and the second specified image moment; wherein, the image moment loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position and / or extension direction of the drawn graphic in the drawn image.
5. The method according to claim 4, wherein The step of determining a first specified image moment of the drawn image and a second specified image moment of the cross-attention activation map includes: determining a first-order moment of the drawn image and a first-order moment of the cross-attention activation map; The step of generating an image moment loss value based on the first specified image moment and the second specified image moment includes: determining a position loss value based on the distance between the first moment of the drawn image and the first moment of the cross-attention activation map; and using the position loss value as the image moment loss value. Wherein, the position loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position of the drawn figure in the drawn image.
6. The method according to claim 4, wherein The step of determining the first specified image moment of the drawn image and the second specified image moment of the cross-attention activation map includes: determining the second moment of the drawn image and the second moment of the cross-attention activation map. The step of generating an image moment loss value based on the first specified image moment and the second specified image moment includes: determining an orientation loss value based on the distance between the second moment of the drawn image and the second moment of the cross-attention activation map; and using the orientation loss value as the image moment loss value. Wherein, the orientation loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the extension direction of the drawn figure in the drawn image.
7. The method according to claim 4, wherein The step of generating an image moment loss value based on the first specified image moment and the second specified image moment includes: Determining a position loss value based on the distance between the first moment of the drawn image and the first moment of the cross-attention activation map; wherein, the position loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position of the drawn figure in the drawn image; Determining an orientation loss value based on the distance between the second moment of the drawn image and the second moment of the cross-attention activation map; wherein, the orientation loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the extension direction of the drawn figure in the drawn image; Determining the weighted sum of the orientation loss value and the position loss value as the image moment loss value.
8. The method according to claim 1, characterized in that, The step of generating a loss value based on the mask image and the cross-attention activation map includes: Generating a focal loss value based on the mask image and the cross-attention activation map; wherein, the focal loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the spatial distribution area indicated by the mask image; Generating an image moment loss value based on the first specified image moment of the drawn image and the second specified image moment of the cross-attention activation map; wherein, the image moment loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches the position and / or extension direction of the drawn figure in the drawn image; Determining the sum of the focal loss value and the image moment loss value as the final loss value.
9. The method according to claim 1, characterized in that, After the step of generating the mask image corresponding to the drawn image, the method further includes: Generate a self-attention map of the first latent space feature; wherein, the self-attention map is used to indicate the correlation degree between feature vectors at different positions in the first latent space feature; Determine a target image region corresponding to the spatial distribution region indicated by the mask image from the self-attention map; Based on the target image region, determine an expansion anchor point from the mask image; wherein, the expansion anchor point is used to indicate the expansion direction and / or expansion distance of the expansion of the spatial distribution region in the mask image; Based on the expansion anchor point, perform an expansion process on the spatial distribution region in the mask image to obtain the processed mask image.
10. The method according to claim 9, wherein The step of determining an expansion anchor point from the mask image based on the target image region includes: Determine the average distribution of self-attention values within the target image region; Based on the average distribution and the self-attention map, determine an expansion anchor point from the mask image.
11. The method according to claim 10, wherein The step of determining an expansion anchor point from the mask image based on the average distribution and the self-attention map includes: Obtain the self-attention values corresponding to the positions outside the spatial distribution region in the mask image from the self-attention map; Based on the self-attention values corresponding to the positions outside the spatial distribution region and the average distribution, determine the similarity between the positions outside the spatial distribution region and the edge positions of the spatial distribution region; Based on the similarity, determine a specified number of expansion anchor points from the positions outside the spatial distribution region.
12. An image generation device, characterized in that, The device includes: A data acquisition module, configured to acquire a target text and a drawing image; wherein, the drawing image includes at least one drawing graphic, and the drawing graphic is provided with a text mark corresponding to the target text; A feature output module, configured to input the target text into a preset diffusion model and output a first latent space feature; An intermediate generation module, configured to generate a mask image and a cross-attention activation map corresponding to the drawing image; wherein, the mask image is used to indicate the spatial distribution region of the image content corresponding to the text mark; the cross-attention activation map includes the correlation degree between the feature vectors at each position in the first latent space feature and the text mark; A loss generation module, configured to generate a loss value based on the mask image and the cross-attention activation map; wherein, the loss value is used to: adjust the model parameters of the diffusion model so that the second latent space feature output by the adjusted diffusion model matches at least one of the following: the spatial distribution region indicated by the mask image, the position of the drawing graphic in the drawing image, the extension direction of the drawing graphic in the drawing image; An image generation module, configured to adjust the model parameters of the diffusion model based on the loss value, output the second latent space feature through the adjusted diffusion model, and generate a target image based on the second latent space feature.
13. An electronic device, characterized in that, It includes a processor and a memory, the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the image generation method according to any one of claims 1-11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the image generation method according to any one of claims 1-11.