A method, device, equipment and medium for generating images from text

By optimizing bounding box position and potential representation, the deviation problem of diffusion model in the target quantity control is solved, and precise control of the target quantity in the generated image and image quality improvement is achieved.

CN119722875BActive Publication Date: 2025-05-13ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510244910.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-13
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

Diffusion models such as Stable Diffusion have deviations in fine control of the number of targets, making it difficult to ensure that the number of targets in the generated image is exactly in line with expectations.

Method used

The initial image is generated by responding to the text prompt word, and the corresponding target object is generated based on the number of targets in the prompt word. If the number of targets in the initial image does not match, optimize the position of the candidate bounding box and adjust the potential representation to match the peak of the cross attention score graph to ensure that the number of target objects matches the expected one.

Benefits of technology

Accurate control of the number of targets in the generated image is achieved, the accuracy and quality of image generation is improved, and the high consistency and controllability of the target image is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119722875B_ABST
    Figure CN119722875B_ABST
Patent Text Reader

Abstract

The present application relates to the field of image processing technology, and discloses a method, device, equipment and medium for generating images from text, wherein the method comprises: in response to the acquired text prompt words containing target objects, generating an initial image corresponding to the text prompt words; wherein the text prompt words include the target number of target objects, and the initial image includes the generated number of target objects; when the target number is not equal to the generated number, optimizing the position of the candidate bounding box in the initial image generation process to generate an optimized bounding box; optimizing the current potential representation in the initial image generation process based on the optimized bounding box to obtain the target potential representation; updating the initial image based on the target potential representation to generate the target image, so that the generated number of target objects in the target image is equal to the target number. The technical solution provided by the present application can accurately control the number of targets in the generated image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method, device, equipment and medium for generating images from text. Background Art

[0002] Diffusion models such as Stable Diffusion generate high-quality images by gradually adding noise to the image and then denoising it in reverse. They are widely used in fields such as art creation, medical imaging, and film and television production. Their advantages include generating clear and fidelity images and flexible conditional generation, but they still face some challenges, especially in the fine control of the number of targets. Since the generation process relies on gradual denoising, the model has deviations in accurately controlling the number of generated objects, making it difficult to ensure that the number of targets in the generated image is exactly as expected. Summary of the invention

[0003] The present application provides a method, device, equipment and medium for generating images from text, which achieves the technical effect of accurately controlling the number of targets in the generated image.

[0004] In order to achieve the above objectives, the main technical solutions adopted in this application include:

[0005] In a first aspect, an embodiment of the present application provides a method for generating an image from text, the method comprising:

[0006] In response to the acquired text prompt word containing the target object, an initial image corresponding to the text prompt word is generated; wherein the text prompt word includes the target number of the target object, and the initial image includes the generated number of the target object;

[0007] When the target number is not equal to the generated number, optimizing the position of the candidate bounding box in the initial image generation process to generate an optimized bounding box;

[0008] Optimizing the current potential representation in the initial image generation process based on the optimized bounding box so that the peak of the cross-attention score map in the initial image generation process matches the optimized bounding box to obtain a target potential representation;

[0009] The initial image is updated based on the target latent representation to generate a target image so that the generated number of the target objects in the target image is equal to the target number.

[0010] A text-generated image method provided in this embodiment generates an initial image by responding to a text prompt word, and generates a corresponding target object according to the target number in the prompt word. However, since the initial image may not fully meet the target number, the subsequent bounding box position optimization further adjusts the distribution of the target object to ensure that the generated image meets the expected number. Then, by optimizing the potential representation and matching it with the peak of the cross-attention score map, the position and number of the target objects are further accurately controlled. Finally, by updating the optimized potential representation, the target image is generated, and the deviation in the initial generation process is corrected, thereby ensuring the accuracy of the image quality and the number of target objects. The accuracy and quality of image generation are effectively improved, ensuring the high consistency and controllability of the target image.

[0011] In one embodiment, the step of optimizing the position of the candidate bounding box in the initial image generation process to generate the optimized bounding box includes:

[0012] Extracting a cross-attention score map during the generation of the initial image, and performing a binarization conversion on the cross-attention score map to obtain a mask image corresponding to the text prompt word;

[0013] Generating a plurality of candidate bounding boxes corresponding to the target object in the mask image;

[0014] Construct bounding box position optimization function;

[0015] The position of the candidate bounding box is optimized using the bounding box position optimization function to generate the optimized bounding box.

[0016] This embodiment accurately determines the area related to the target object by extracting the cross-attention score map and performing a binary conversion, thereby generating a clear mask image. Then, multiple candidate bounding boxes are generated to provide different positioning references for the target object. A bounding box position optimization function is constructed to further improve the positioning accuracy of the candidate bounding box. Finally, the position of the candidate bounding box is adjusted through the optimization function so that the final generated bounding box is more accurately aligned with the target area, thereby effectively improving accuracy and robustness.

[0017] In one embodiment, constructing a bounding box position optimization function includes:

[0018] Determine the complete intersection-over-union ratio between each of the candidate bounding boxes;

[0019] Determine a single Gaussian kernel corresponding to the candidate bounding box;

[0020] Accumulating all the single Gaussian kernels at corresponding positions of the mask image to obtain a target Gaussian kernel;

[0021] Determining a distance loss function between the candidate bounding box and the peak area in the cross-attention score map based on the target Gaussian kernel and the score of the peak area in the cross-attention score map;

[0022] The bounding box position optimization function is determined based on the complete intersection-over-union ratio and the distance loss function.

[0023] This embodiment not only considers the overlap of candidate bounding boxes when calculating the complete intersection-union ratio, but also takes into account the relationship between position, shape and scale, thereby reducing redundant candidate bounding boxes and improving detection accuracy. Each candidate bounding box generates a Gaussian kernel that reflects the influence range of the target, and forms a target Gaussian kernel by accumulation, smoothing the target area and reducing background interference. Then, the distance loss function is calculated using the target Gaussian kernel and the peak area in the cross-attention score map to quantify the degree of match between the candidate bounding box and the target area. Finally, combined with the complete intersection-union ratio and the distance loss function, the position of the candidate bounding box is optimized, the alignment accuracy of the candidate bounding box and the target area is improved, and the candidate bounding box regression is more accurate.

[0024] In one embodiment, determining the single Gaussian kernel corresponding to the candidate bounding box includes:

[0025] Determining an initial side length of the candidate bounding box according to the area of ​​the target region in the mask image and the number of targets;

[0026] Determining a standard deviation based on the scale factor and the initial side length;

[0027] For any candidate bounding box, the weight of each pixel in the mask image is determined according to the standard deviation and the center coordinates of the candidate bounding box, so as to obtain a single Gaussian kernel corresponding to the any candidate bounding box.

[0028] This embodiment accurately generates candidate bounding boxes and determines the initial side length of the candidate bounding box based on the area and number of the target area, so that it matches the actual size of the target area, thereby improving detection accuracy. The scale factor is combined with the initial side length to calculate the standard deviation, and then adjust the smoothness of the Gaussian kernel to avoid over-smoothing or fitting. Using the standard deviation and the center coordinates of the candidate bounding box, an appropriate weight is assigned to each pixel in the mask image, so that the closer the pixel to the center of the bounding box, the greater the weight, and the farther the pixel from the center, the smaller the weight, thereby smoothly weighting the target area and reducing noise interference. Overall, this embodiment can more accurately generate a bounding box that matches the target area.

[0029] In one embodiment, the optimizing the current potential representation in the initial image generation process based on the optimized bounding box so that the peak of the cross-attention score map in the initial image generation process matches the optimized bounding box to obtain the target potential representation includes:

[0030] constructing a target object position mask according to the position of the optimized bounding box;

[0031] Based on the target object position mask, constructing a potential representation optimization function using a weighted cross entropy method;

[0032] The current latent representation in the initial image generation process is optimized according to the latent representation optimization function to obtain the target latent representation.

[0033] This embodiment generates a target object position mask through a bounding box, clearly marks the position of the target area in the image, and provides precise guidance for the subsequent optimization process. Next, a potential representation optimization function is constructed based on weighted cross entropy to ensure that the pixels in the target area have higher weights, thereby promoting the optimization of the potential representation in the target area and making the generated image more accurately present the target features. Finally, by optimizing the initial potential representation, it gradually approaches the target potential representation, and finally generates an image with more refined and clear details in the target area. This embodiment can not only accurately locate the target area, but also effectively optimize the potential representation of the image and improve the quality of image generation.

[0034] In one embodiment, the optimization of the current potential representation in the initial image generation process based on the optimized bounding box is performed as follows:

[0035] Iteratively performing optimization of a position of a candidate bounding box in the process of generating the initial image and optimizing a current potential representation in the process of generating the initial image to obtain an optimized potential representation;

[0036] The optimized latent representation is used to replace the current latent representation until a preset number of cross-optimization steps is met to obtain the target latent representation.

[0037] This embodiment iteratively optimizes the position of the candidate bounding box and the current potential representation, so that the image generation process gradually converges to a more accurate state. Optimizing the position of the bounding box can better capture the target area and improve the quality of local details, while optimizing the potential representation ensures the accurate expression of image features. By replacing the current potential representation with the optimized potential representation and continuously iterating until the preset number of cross-optimization steps is met, the target potential representation is finally obtained. This embodiment continuously adjusts the potential representation to ensure that the quality of the generated image meets expectations and can more accurately reflect the characteristics and details of the target area, thereby achieving high-precision image generation effects.

[0038] In one embodiment, a Stable Diffusion model is used to generate an initial image corresponding to the text prompt word.

[0039] In a second aspect, an embodiment of the present application provides a text-generated image device, the device comprising:

[0040] An initial image generating unit, configured to generate an initial image corresponding to the text prompt word in response to the acquired text prompt word containing the target object; wherein the text prompt word includes the target number of the target object, and the initial image includes the generated number of the target object;

[0041] a bounding box optimization unit, configured to optimize the positions of the candidate bounding boxes in the initial image generation process to generate an optimized bounding box when the target number is not equal to the generated number;

[0042] A potential representation optimization unit, configured to optimize the current potential representation in the initial image generation process based on the optimized bounding box, so that the peak of the cross-attention score map in the initial image generation process matches the optimized bounding box, and obtain a target potential representation;

[0043] The target image generating unit is used to update the initial image based on the target potential representation to generate a target image so that the generated number of the target objects in the target image is equal to the target number.

[0044] In a third aspect, an embodiment of the present application provides a computer device, including:

[0045] A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the above-mentioned method of generating images from text by executing the computer instructions.

[0046] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to enable a computer to execute the above-mentioned method of generating an image from text. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0048] Figure 1 A schematic diagram of a specific scenario application provided by an embodiment of the present application;

[0049] Figure 2 A flowchart of a method for generating an image from text provided in an embodiment of the present application;

[0050] Figure 3 A flowchart of step S3 provided in an embodiment of the present application;

[0051] Figure 4 A flowchart of step S35 provided in an embodiment of the present application;

[0052] Figure 5 A flowchart of step S353 provided in an embodiment of the present application;

[0053] Figure 6 A flowchart of step S5 provided in an embodiment of the present application;

[0054] Figure 7 A flowchart for optimizing the current potential representation provided by an embodiment of the present application;

[0055] Figure 8 A block diagram of a text-to-image device provided in an embodiment of the present application;

[0056] Fig. 9 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0058] In recent years, image generation technology has made significant progress in the field of computer vision. The main goal of image generation models is to generate realistic visual images based on input information. These technologies have a wide range of applications, covering many fields such as computer graphics, virtual reality, augmented reality, and automated design. Among these technologies, Generative Adversarial Networks (GANs) once dominated. However, with the introduction and development of diffusion models, this situation has changed. Diffusion models have quickly become a research hotspot in the field of image generation due to their advantages in generating high-quality images and their unique theoretical foundation.

[0059] The basic concept of the diffusion model is derived from the diffusion process in statistical physics. Its working principle includes two main stages: the forward diffusion process and the reverse generation process. In the forward diffusion stage, the model gradually transforms the image data into noise by gradually adding Gaussian noise to the image. This process is controlled by a defined noise scheduling mechanism. The reverse generation stage gradually removes the noise by processing the samples sampled from the noise, and finally restores a clear image. This process relies on a parameterized neural network to predict the noise intensity in each step. This generation method of the diffusion model not only ensures the generation of high-fidelity images, but also provides higher flexibility and controllability in the image generation process.

[0060] Among them, Stable Diffusion, as an important implementation of the diffusion model, has received particular attention. It has greatly expanded the application scope of image generation through efficient computing methods and flexible conditional generation. Compared with traditional diffusion models, Stable Diffusion uses latent space and optimizes noise scheduling, effectively reducing the demand for computing resources, allowing it to run on ordinary consumer-grade hardware. This feature enables Stable Diffusion to be widely used in many fields such as artistic creation, medical image generation, film and television production, and promotes the improvement of creative efficiency and quality in various industries.

[0061] However, despite the remarkable achievements of diffusion models, especially Stable Diffusion, it still faces some challenges in practical applications. In particular, the diffusion model has not yet achieved the desired effect in terms of fine control of the number of objects in image generation. Since its generation process relies on step-by-step denoising, although it can well preserve the overall style and details of the image, the performance of the model still has deviations when it comes to precisely controlling the number of generated objects, and the number of objects in the generated image often deviates from the expected number of objects.

[0062] Based on this, in order to solve the above technical problems, please refer to Figure 1, in a specific scenario example:

[0063] Step S1001: Input the prompt word "A photo of apples." into the Stable Diffusion model to generate an initial image (I). Use the Judge Model to analyze the number of target apples generated in the initial image (I) and compare it with the target number specified in the prompt word. If the target number matches, the process ends; if the target number does not match, the cross-optimization process is entered. It can be seen from the figure that the number of apples in the initial image (I) is 7, which does not match the number 5 in the text prompt word, so the cross-optimization process is entered. In this process, the cross-attention score map in the initial image generation process is extracted from the U-Net, and the cross-attention score map is binarized to obtain the mask image M corresponding to the text prompt word, which is used for bounding box optimization in subsequent steps.

[0064] Step S1002: Optimize the position of the bounding box using the bounding box position optimization function L1, where the bounding box position optimization function L1 consists of two parts:

[0065] CIOU: Calculate the overlap between bounding boxes to avoid excessive overlap of bounding boxes.

[0066] L CA : Calculate bounding box and cross attention map The distance loss of the peak area ensures that the bounding box is aligned with the main feature area in the image.

[0067] The position of the bounding box is adjusted by the gradient optimization method to move it toward the peak area of ​​the cross-attention score map to generate an optimized bounding box.

[0068] Step S1003: Use the optimized bounding box to construct the target position mask B. The latent representation Z is optimized by the latent representation optimization function L2. t The potential representation optimization function L2 encourages higher cross attention scores inside the bounding box and lower cross attention scores outside the bounding box. t , and obtain the optimized potential representation , so that the peak area of ​​the cross attention map appears in the optimized bounding box.

[0069] Step S1004: At a given cross optimization step number Topt The inner iteration performs steps S1002 and S1003. After each denoising step, the optimized latent representation is used. Replace the current potential representation Z t , and re-enter the StableDiffusion denoising process. In each optimization process, steps S1002 and S1003 are no longer repeated to ensure that the bounding box and the potential representation are continuously optimized in the correct direction of the target quantity during the cross-optimization process. After all the cross-optimization steps, the final potential representation Z is obtained t+1 , the number of targets in the final generated target image I* is consistent with the text prompt words, and the image quality is not affected.

[0070] According to an embodiment of the present application, an embodiment of a method for generating an image from text is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0071] In this embodiment, a method for generating an image from text is provided. Figure 2 A flowchart of a method for generating an image from text provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the process includes the following steps:

[0072] Step S1, in response to the acquired text prompt word containing the target object, generating an initial image corresponding to the text prompt word; wherein the text prompt word includes the target number of the target object, and the initial image includes the generated number of the target object.

[0073] Specifically, the Stable Diffusion model is used to generate the initial image. The Stable Diffusion model is a deep learning text-to-image model based on diffusion technology, which uses the latent diffusion model (LDM) to generate high-quality images. It is mainly used to generate detailed images conditioned on text descriptions, but can also be applied to other tasks, such as image repair (inpainting), image expansion (outpainting), and image-to-image conversion based on text prompts. In this embodiment, the Stable Diffusion model receives a text prompt containing a target object, and the model can generate an initial image containing a prompt description based on the text prompt. The initial image contains the number of target objects to be generated. For example, the text prompt mentions "an image of 5 apples", and Stable Diffusion will generate an image containing 5 apples based on the prompt. The model will adjust the number of target objects in the image based on the quantity information.

[0074] Step S3, when the target number is not equal to the generated number, the position of the candidate bounding box in the initial image generation process is optimized to generate an optimized bounding box.

[0075] Specifically, when the Stable Diffusion model is used to generate the initial image, the input text prompt contains the target number of the target object. If the generated number of target objects in the generated initial image is not equal to the target number in the text prompt, the position of the candidate bounding box needs to be optimized to ensure that the generated image is consistent with the target number in the text prompt. The judgment model can be used to extract the target number of the target object from the input text prompt. This can be achieved through natural language processing technology, such as using a pre-trained language model (such as BERT or GPT) to parse the text prompt and extract the target number.

[0076] Step S5, optimizing the current potential representation in the initial image generation process based on the optimized bounding box so that the peak of the cross-attention score map in the initial image generation process matches the optimized bounding box to obtain the target potential representation.

[0077] Specifically, when using the Stable Diffusion model to generate images, optimizing the latent representation in the initial image generation process based on the optimized bounding box is a key step to ensure the accurate position and number of target objects. Specifically, the target object position mask is constructed through the optimized bounding box, and the pixels inside the bounding box are set to 1 and the pixels outside are set to 0. Next, based on the target object position mask and the cross-attention score map, a latent representation optimization function is constructed to optimize the current latent representation until the loss function value converges. After each cross-optimization step, the optimized latent representation replaces the current latent representation, and the denoising process is re-executed until the predetermined number of optimization steps is reached. Finally, after multiple cross-optimizations, the target latent representation is obtained.

[0078] Step S7, updating the initial image based on the target latent representation, and generating a target image, so that the number of generated target objects in the target image is equal to the target number.

[0079] Specifically, by updating the initial image based on the target latent representation and generating the target image, it is ensured that the generated image meets the user's expectations. Specifically, this is achieved by performing the denoising process on the target latent representation as the new latent representation. The resulting target image will contain the same number of target objects as the target, and these target objects will appear at the locations specified by the optimized bounding box.

[0080] A text-generated image method provided in this embodiment generates an initial image by responding to a text prompt word, and generates a corresponding target object according to the target number in the prompt word. However, since the initial image may not fully meet the target number, the subsequent bounding box position optimization further adjusts the distribution of the target object to ensure that the generated image meets the expected number. Then, by optimizing the potential representation and matching it with the peak of the cross-attention score map, the position and number of the target objects are further accurately controlled. Finally, by updating the optimized potential representation, the target image is generated, and the deviation in the initial generation process is corrected, thereby ensuring the accuracy of the image quality and the number of target objects. The accuracy and quality of image generation are effectively improved, ensuring the high consistency and controllability of the target image.

[0081] Figure 3 The flowchart of step S3 provided in the embodiment of the present application may include the following steps:

[0082] Step S31, extracting the cross-attention score map in the process of generating the initial image, and binarizing the cross-attention score map to obtain a mask image corresponding to the text prompt word.

[0083] Specifically, during the generation process of the Stable Diffusion model, the cross-attention score map can be extracted from the U-Net structure. Specifically, each time step t of the U-Net generates a cross-attention score map, which shows the interaction between the text prompt and the different parts of the image generation. Through these cross-attention score maps, we can understand in which areas the model focuses more attention. In order to fuse the cross-attention score maps (extracted at multiple time steps) into a representative image, it is necessary to average the cross-attention score maps of all time steps to obtain an h×w global cross-attention score map A. avg .

[0084] The cross attention score map A avg The purpose of converting to a binary image is to extract the areas that the model pays special attention to through a threshold segmentation method. These areas usually correspond to the important features mentioned in the text prompt words, or the parts that should be focused on in image generation. First, select a suitable threshold T. This threshold is used to determine which areas have a high cross-attention score and are worth retaining. For the cross-attention score map A avg For each pixel (x,y) in A avg If (x,y) is greater than the threshold, the pixel is set to white (indicating that the area is considered to be a key area by the model); otherwise, it is set to black (indicating that the area is not important). The generated binary mask image M is used as the mask image to represent the most important part of the image, that is, the area related to the text prompt.

[0085] Step S33: Generate multiple candidate bounding boxes corresponding to the target object in the mask image.

[0086] Specifically, a bounding box is usually used to identify the location of the region of interest in the image. For the mask image, it is the cross-attention score mentioned above. Figure 2 The valued image contains the areas that the model considers important, and the goal is to extract multiple candidate bounding boxes from these areas to locate different objects in the image. In order to generate these candidate bounding boxes, several methods can be used, including anchor-based, edge detection and contour extraction methods, and deep learning models (such as YOLO, SSD, etc.).

[0087] Step S35, constructing a bounding box position optimization function.

[0088] Specifically, the candidate bounding box is accurately generated and the initial side length is determined according to the area and number of the target area to match the actual size of the target area. At the same time, the standard deviation is calculated by combining the scale factor and the initial side length, and the standard deviation is used to adjust the smoothness of the Gaussian kernel to avoid over-smoothing or fitting. On this basis, according to the center coordinates of the candidate bounding box, an appropriate weight is assigned to each pixel in the mask image, so that the pixel weight closer to the center of the bounding box is larger, and the pixel weight far from the center is smaller, thereby smoothly weighting the target area and reducing noise interference. Furthermore, the complete intersection-over-union ratio and target Gaussian kernel are introduced to effectively improve the accuracy of the candidate bounding box position. When calculating CIOU, not only the overlap of the candidate bounding boxes is considered, but also the relationship between position, shape and scale is taken into account, thereby reducing redundant candidate bounding boxes and improving detection accuracy. Each candidate bounding box generates a Gaussian kernel to reflect the influence range of the target, and the target Gaussian kernel is formed by accumulation to smooth the target area and reduce background interference. Subsequently, the distance loss function is calculated by combining the target Gaussian kernel and the peak area in the cross-attention score map to quantify the matching degree between the candidate bounding box and the target area. Finally, by combining CIOU and distance loss functions, a bounding box position optimization function is constructed to optimize the position of the candidate bounding box, improve its alignment accuracy with the target area, and make the candidate bounding box regression more accurate.

[0089] Step S37, optimizing the position of the candidate bounding box using the bounding box position optimization function to generate an optimized bounding box.

[0090] Specifically, the position of the candidate bounding box is optimized using the bounding box position optimization function within a given number of cross optimization steps, and the position of the candidate bounding box is continuously updated to move in the direction in which the bounding box position optimization function decreases until the loss function value converges or reaches a predetermined number of iterations, thereby obtaining the optimized bounding box.

[0091] This embodiment accurately determines the area related to the target object by extracting the cross-attention score map and performing a binary conversion, thereby generating a clear mask image. Then, multiple candidate bounding boxes are generated to provide different positioning references for the target object. A bounding box position optimization function is constructed to further improve the positioning accuracy of the candidate bounding box. Finally, the position of the candidate bounding box is adjusted through the optimization function so that the final generated bounding box is more accurately aligned with the target area, thereby effectively improving accuracy and robustness.

[0092] Figure 4 The flowchart of step S35 provided in the embodiment of the present application may include the following steps:

[0093] Step S351, determining the complete intersection-over-union ratio between any two candidate bounding boxes.

[0094] Specifically, in order to ensure that the generated bounding boxes do not overlap with each other and can accurately cover the target object, it is necessary to calculate the complete intersection over union (IOU) between the bounding boxes. First, the intersection over union (IOU) is calculated, which is a common indicator to measure the degree of overlap between two bounding boxes:

[0095]

[0096] Among them, B p and B g Represents two different candidate bounding boxes, |B p ∩B g | is the intersection area of ​​two candidate bounding boxes, |B p ∪B g | is the union area of ​​the two candidate bounding boxes.

[0097] Then the Complete Intersection over Union (CIOU) is calculated to ensure that the bounding boxes do not overlap too much:

[0098]

[0099] Wherein, N is the number of targets in the text prompt word. For example, if the text prompt word is “A photo of fiveapples.”, then N is 5.

[0100] Step S353: determine a single Gaussian kernel corresponding to the candidate bounding box.

[0101] In order to more accurately control the position and shape of each candidate bounding box during the optimization process and ensure that they can better cover the target object and avoid overlapping, the Gaussian kernel is used to generate a smooth weight distribution for each candidate bounding box. This distribution reflects the distance between each pixel and the center of the candidate bounding box, so that the candidate bounding box can align the target object more accurately and maintain an appropriate distance during the optimization process, avoiding excessive overlap and preventing multiple candidate bounding boxes from interfering with each other, thereby improving the optimization effect.

[0102] Specifically, the initial side length of the candidate bounding box is first determined by calculating the area and number of the target area to match the actual size of the target area. Then, the standard deviation is calculated by combining the scale factor with the initial side length, and the smoothness of the Gaussian kernel is adjusted accordingly to avoid over-smoothing or overfitting. Finally, according to the standard deviation and the center coordinates of the candidate bounding box, an appropriate weight is assigned to each pixel in the mask image, thereby generating a single Gaussian kernel corresponding to each candidate bounding box.

[0103] Step S355: All single Gaussian kernels are accumulated at corresponding positions of the mask image to obtain a target Gaussian kernel.

[0104] Since each candidate bounding box has a corresponding single Gaussian kernel. These candidate bounding boxes may overlap or be distributed in different areas. Therefore, directly using a single Gaussian kernel to represent the target area may not be sufficient to accurately represent the complete position and shape of the target object. By superimposing the single Gaussian kernels of all candidate bounding boxes, a comprehensive target Gaussian kernel can be formed, which can reflect the total effect of all bounding boxes and provide a more accurate description of the target position and shape.

[0105] Specifically, each individual single Gaussian kernel will be superimposed on the corresponding pixel point according to its position in the mask image, and finally the target Gaussian kernel is obtained. If multiple single Gaussian kernels cover the same position, the weight of the position is the sum of the weights of all single Gaussian kernels at the position. This means that the area where multiple candidate bounding boxes overlap will have a higher weight, thereby better reflecting the integrated features of the target area. If some single Gaussian kernels do not overlap at the same position, the weights of these positions are contributed by only one single Gaussian kernel. In this way, the area not affected by other candidate bounding boxes can still reflect the influence of the corresponding bounding box. Finally, the superimposed target Gaussian kernel K forms a more accurate weight distribution by integrating the influence of all candidate bounding boxes. This target Gaussian kernel can more accurately represent the position and shape of the target area and effectively avoid the limitations of a single bounding box.

[0106] Step S357, based on the target Gaussian kernel and the score of the peak area in the cross-attention score map, determine the distance loss function between the candidate bounding box and the peak area in the cross-attention score map.

[0107] Specifically, the corresponding scores can be calculated for the cross-attention score map:

[0108]

[0109]

[0110] Where Q is the query vector; K is the key vector; QK T K is the transpose of the query vector Q and the key vector K T The result is a score matrix that represents the correlation between each query vector and each key vector; is the dimension of the key vector K, which is used to prevent the score from being too large; the Softmax function is used to convert the score matrix into a probability distribution so that each score is between 0 and 1 and the sum of all scores is 1; in the Stable Diffusion model, and denote the query vector and key vector in the cross-attention calculation, respectively. These vectors are obtained by multiplying the current latent representation Z in U-Net. t Obtained by nonlinear mapping.

[0111] Find areas with higher scores by threshold segmentation or local maximum detection method, and extract the scores of these peak areas , used for subsequent distance loss calculation.

[0112] Distance loss function L CA Used to measure the distance between the candidate bounding box and the peak area in the cross-attention score map:

[0113]

[0114] Among them, K is the target Gaussian kernel.

[0115] Step S359: Determine a bounding box position optimization function based on the complete intersection-over-union ratio and the distance loss function.

[0116] Specifically, the bounding box position optimization function L1 is:

[0117]

[0118] Among them, β is a hyperparameter used to control the distance loss function L CA By optimizing this bounding box position optimization function L1, the model not only considers the overlap (CIOU) between the candidate bounding box and the true target object, but also considers whether the candidate bounding box is close to the high correlation area in the cross-attention score map. This comprehensive loss function helps to achieve more accurate bounding box positioning during training and ensures that the candidate bounding box accurately matches the target area.

[0119] This embodiment not only considers the overlap of candidate bounding boxes when calculating the complete intersection-union ratio, but also takes into account the relationship between position, shape and scale, thereby reducing redundant candidate bounding boxes and improving detection accuracy. Each candidate bounding box generates a Gaussian kernel that reflects the influence range of the target, and forms a target Gaussian kernel by accumulation, smoothing the target area and reducing background interference. Then, the distance loss function is calculated using the target Gaussian kernel and the peak area in the cross-attention score map to quantify the degree of match between the candidate bounding box and the target area. Finally, combined with the complete intersection-union ratio and the distance loss function, the position of the candidate bounding box is optimized, the alignment accuracy of the candidate bounding box and the target area is improved, and the candidate bounding box regression is more accurate.

[0120] Figure 5 The flowchart of step S353 provided in the embodiment of the present application may include the following steps:

[0121] Step S3531, determining the initial side length of the candidate bounding box according to the area of ​​the target region and the number of targets in the mask image.

[0122] Specifically, in order to effectively select the target object in the image, each candidate bounding box needs to have a suitable initial side length. Usually, the total area A of the target area is calculated using the binary mask image M. Specifically, the number of white pixels in the mask image can be calculated, and then the total area of ​​the target area can be obtained by multiplying the area of ​​each pixel. The number of targets N is extracted from the input text prompt word, and the total area A of the target area needs to be evenly distributed to the N target objects. The area of ​​each target area is A / N. Assuming that each target area is an approximate square, the side length l of each candidate bounding box is bb It can be calculated by the total area A of the target area and the number of targets N:

[0123]

[0124] This formula shows that the side length l of each candidate bounding box bb It is the square root of the ratio of the total area of ​​the target area A to the number of targets N. The calculated l bb It is the initial side length value, which is used to provide a suitable size for each candidate bounding box.

[0125] Step S3533, determine the standard deviation based on the scale factor and the initial side length.

[0126] Specifically, the standard deviation is an important parameter in the Gaussian distribution, which controls the "diffusion" of the Gaussian kernel. The larger the standard deviation, the wider the influence range of the kernel; the smaller the standard deviation, the narrower the influence range of the kernel. Here, the scale factor s and the initial side length l are used. bb Calculate the standard deviation σ: σ=s×l bb ,Usually, the scale factor s is set to 0.5 to ensure that the influence range of the Gaussian kernel is moderate.

[0127] Step S3535 : For any candidate bounding box, the weight of each pixel in the mask image is determined according to the standard deviation and the center coordinates of the candidate bounding box, and a single Gaussian kernel corresponding to any candidate bounding box is obtained.

[0128] Specifically, for each candidate bounding box, generate its corresponding single Gaussian kernel k:

[0129]

[0130] Among them, (x, y) is the coordinate of each pixel in the mask image, (μ x ,μ y ) are the center coordinates of the candidate bounding box.

[0131] This embodiment accurately generates candidate bounding boxes and determines the initial side length of the candidate bounding box based on the area and number of the target area, so that it matches the actual size of the target area, thereby improving detection accuracy. The scale factor is combined with the initial side length to calculate the standard deviation, and then adjust the smoothness of the Gaussian kernel to avoid over-smoothing or fitting. Using the standard deviation and the center coordinates of the candidate bounding box, an appropriate weight is assigned to each pixel in the mask image, so that the closer the pixel to the center of the bounding box, the greater the weight, and the farther the pixel from the center, the smaller the weight, thereby smoothly weighting the target area and reducing noise interference. Overall, this embodiment can more accurately generate a bounding box that matches the target area.

[0132] Figure 6 The flowchart of step S5 provided in the embodiment of the present application may include the following steps:

[0133] Step S51, constructing a target object position mask according to the optimized position of the bounding box.

[0134] Specifically, according to the position of the optimized bounding box, the pixel values ​​inside the bounding box are set to 1, and the pixel values ​​outside the bounding box are set to 0, forming a binary target object position mask B. This target object position mask B is used to identify the area of ​​the target object (1 indicates the position of the target object, and 0 indicates the non-target area).

[0135] Step S53: construct a potential representation optimization function based on the target object position mask by using a weighted cross entropy method.

[0136] Specifically, the latent representation optimization function L2 is:

[0137]

[0138] Where i is the index of each pixel position, w i is a weighting coefficient, when B i When w is 1, i is 10, when B i When w is 0, i is 1. In the form of cross entropy, L2 constrains the current potential representation, pushing the cross attention score of the current potential representation to increase in the target area (inside the bounding box) and decrease in the non-target area (outside the bounding box). This optimization process can make the current potential representation Z t The cross attention score map is more concentrated in the target area.

[0139] Step S55, optimizing the current latent representation in the initial image generation process according to the latent representation optimization function to obtain a target latent representation.

[0140] Specifically, the current potential representation in the initial image generation process is optimized according to the potential representation optimization function within a given cross-optimization step number, and the current potential representation is continuously updated to move in the direction where the potential representation optimization function decreases until the loss function value converges or reaches a predetermined number of iterations, and the target potential representation is obtained. This will make the peak area of ​​the cross-attention score map of the target area appear at the position specified by the bounding box.

[0141] This embodiment generates a target object position mask through a bounding box, clearly marks the position of the target area in the image, and provides precise guidance for the subsequent optimization process. Next, a potential representation optimization function is constructed based on weighted cross entropy to ensure that the pixels in the target area have higher weights, thereby promoting the optimization of the potential representation in the target area and making the generated image more accurately present the target features. Finally, by optimizing the initial potential representation, it gradually approaches the target potential representation, and finally generates an image with more refined and clear details in the target area. This embodiment can not only accurately locate the target area, but also effectively optimize the potential representation of the image and improve the quality of image generation.

[0142] Figure 7 A flowchart for optimizing the current potential representation provided in an embodiment of the present application may include the following steps:

[0143] Step S571, iteratively performing optimization of the position of the candidate bounding box in the initial image generation process and optimizing the current potential representation in the initial image generation process to obtain an optimized potential representation;

[0144] Step S573: Use the optimized latent representation to replace the current latent representation until the preset number of cross-optimization steps is met to obtain the target latent representation.

[0145] Specifically, at a given cross optimization step number T opt Optimize the position of the candidate bounding box. This can be achieved by calculating the bounding box position optimization function L1 and using the gradient descent method. Specifically, calculate the gradient of L1 with respect to each bounding box position. Update the position of the bounding box so that it moves in the direction where the loss function value decreases. Repeat the above steps until the loss function value converges or reaches a predetermined number of iterations.

[0146] At a given cross optimization step number T opt Within, the current potential representation is optimized. This can be achieved by calculating the potential representation optimization function L2 and using the gradient descent method. Specifically, L2 is calculated with respect to the current potential representation Z t Update the current potential representation Z t, so that it moves in the direction of decreasing the loss function value. Repeat the above steps until the loss function value converges or reaches the predetermined number of iterations.

[0147] At each cross optimization step T opt At the end, the optimized latent representation is used to replace the current latent representation, and the U-Net is re-entered to perform the denoising process again, but step S571 is no longer executed. That is, at each cross optimization step number T opt Within, position optimization and latent representation optimization are performed. In each cross optimization step T opt At the end, replace the latent representation and re-enter the U-Net to perform the denoising process. opt At the end, the optimized latent representation is used to replace the current latent representation to obtain the target latent representation.

[0148] This embodiment iteratively optimizes the position of the candidate bounding box and the current potential representation, so that the image generation process gradually converges to a more accurate state. Optimizing the position of the bounding box can better capture the target area and improve the quality of local details, while optimizing the potential representation ensures the accurate expression of image features. By replacing the current potential representation with the optimized potential representation and continuously iterating until the preset number of cross-optimization steps is met, the target potential representation is finally obtained. This embodiment continuously adjusts the potential representation to ensure that the quality of the generated image meets expectations and can more accurately reflect the characteristics and details of the target area, thereby achieving high-precision image generation effects.

[0149] Accordingly, please refer to Figure 8 A block diagram of a text-generated image device provided in an embodiment of the present application, the device comprising:

[0150] The initial image generating unit 101 is used to generate an initial image corresponding to the text prompt word in response to the acquired text prompt word containing the target object; wherein the text prompt word includes the target number of the target object, and the initial image includes the generated number of the target object;

[0151] The bounding box optimization unit 103 is used to optimize the position of the candidate bounding box in the initial image generation process and generate an optimized bounding box when the target number is not equal to the generated number;

[0152] A potential representation optimization unit 105 is used to optimize the current potential representation in the initial image generation process based on the optimized bounding box, so that the peak of the cross-attention score map in the initial image generation process matches the optimized bounding box to obtain a target potential representation;

[0153] The target image generating unit 107 is used to update the initial image based on the target potential representation and generate the target image so that the generated number of target objects in the target image is equal to the target number.

[0154] In some optional implementations, optimizing the position of the candidate bounding box in the initial image generation process to generate an optimized bounding box includes:

[0155] Extract the cross-attention score map in the initial image generation process, and convert the cross-attention score map into a binary value to obtain the mask image corresponding to the text prompt word;

[0156] Generate multiple candidate bounding boxes corresponding to the target object in the mask image;

[0157] Construct bounding box position optimization function;

[0158] The position of the candidate bounding box is optimized using the bounding box position optimization function to generate an optimized bounding box.

[0159] In some optional implementations, constructing a bounding box position optimization function includes:

[0160] Determine the complete intersection-over-union ratio between each pair of candidate bounding boxes;

[0161] Determine a single Gaussian kernel corresponding to the candidate bounding box;

[0162] Accumulate all single Gaussian kernels at the corresponding positions of the mask image to obtain the target Gaussian kernel;

[0163] Determine the distance loss function between the candidate bounding box and the peak area in the criss-cross attention score map based on the target Gaussian kernel and the score of the peak area in the criss-cross attention score map;

[0164] Based on the complete intersection-over-union and distance loss functions, the bounding box position optimization function is determined.

[0165] In some optional implementations, determining a single Gaussian kernel corresponding to a candidate bounding box includes:

[0166] Determine the initial side length of the candidate bounding box based on the area of ​​the target region and the number of targets in the mask image;

[0167] Determine the standard deviation based on the scale factor and the initial side length;

[0168] For any candidate bounding box, the weight of each pixel in the mask image is determined according to the standard deviation and the center coordinates of the candidate bounding box, and a single Gaussian kernel corresponding to any candidate bounding box is obtained.

[0169] In some optional implementations, the latent representation optimization unit 105 includes:

[0170] Construct a target object position mask based on the position of the optimized bounding box;

[0171] Based on the target object position mask, a potential representation optimization function is constructed using weighted cross entropy.

[0172] The current latent representation in the initial image generation process is optimized according to the latent representation optimization function to obtain the target latent representation.

[0173] In some optional implementations, the optimization of the current potential representation during the initial image generation process based on the optimized bounding box is performed as follows:

[0174] Iteratively performing optimization of the position of the candidate bounding box in the initial image generation process and optimizing the current potential representation in the initial image generation process to obtain an optimized potential representation;

[0175] The optimized potential representation is used to replace the current potential representation until the preset number of cross-optimization steps is met to obtain the target potential representation.

[0176] In some optional implementations, a Stable Diffusion model is used to generate an initial image corresponding to the text prompt word.

[0177] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0178] A text generation image device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0179] See also Fig. 9 , Fig. 9 A schematic diagram of the structure of a computer device provided in an embodiment of the present application is shown in FIG. Fig. 9As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Fig. 9 A processor 10 is taken as an example.

[0180] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0181] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0182] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0183] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0184] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0185] The embodiment of the present application also provides a computer-readable storage medium. The above method according to the embodiment of the present application can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0186] The methods, devices or units described in the above embodiments may be implemented by a computer chip or entity, or by a product having a certain function. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0187] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0188] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods or devices. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0189] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices, and apparatuses according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0190] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0191] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0192] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0193] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0194] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

[0195] Although the embodiments of the present application have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for generating an image from text, characterized in that: The method comprises: In response to the acquired text prompt word containing the target object, an initial image corresponding to the text prompt word is generated; wherein the text prompt word includes the target number of the target object, and the initial image includes the generated number of the target object; When the target number is not equal to the generated number, optimizing the position of the candidate bounding box in the initial image generation process to generate an optimized bounding box; The step of optimizing the position of the candidate bounding box in the process of generating the initial image to generate an optimized bounding box includes: Extracting a cross-attention score map during the generation of the initial image, and performing a binarization conversion on the cross-attention score map to obtain a mask image corresponding to the text prompt word; Generating a plurality of candidate bounding boxes corresponding to the target object in the mask image; Construct bounding box position optimization function; Optimizing the position of the candidate bounding box using the bounding box position optimization function to generate the optimized bounding box; Optimizing the current potential representation in the initial image generation process based on the optimized bounding box so that the peak of the cross-attention score map in the initial image generation process matches the optimized bounding box to obtain a target potential representation; The initial image is updated based on the target latent representation to generate a target image so that the generated number of the target objects in the target image is equal to the target number.

2. The method according to claim 1, characterized in that The constructing of the bounding box position optimization function comprises: Determine the complete intersection-over-union ratio between each of the candidate bounding boxes; Determine a single Gaussian kernel corresponding to the candidate bounding box; Accumulating all the single Gaussian kernels at corresponding positions of the mask image to obtain a target Gaussian kernel; Determining a distance loss function between the candidate bounding box and the peak area in the cross-attention score map based on the target Gaussian kernel and the score of the peak area in the cross-attention score map; The bounding box position optimization function is determined based on the complete intersection-over-union ratio and the distance loss function.

3. The method according to claim 2, characterized in that The determining a single Gaussian kernel corresponding to the candidate bounding box includes: Determining an initial side length of the candidate bounding box according to the area of ​​the target region in the mask image and the number of targets; Determining a standard deviation based on the scale factor and the initial side length; For any candidate bounding box, the weight of each pixel in the mask image is determined according to the standard deviation and the center coordinates of the candidate bounding box, so as to obtain a single Gaussian kernel corresponding to the any candidate bounding box.

4. The method according to claim 1, characterized in that The step of optimizing the current potential representation in the initial image generation process based on the optimized bounding box so that the peak of the cross-attention score map in the initial image generation process matches the optimized bounding box to obtain the target potential representation includes: constructing a target object position mask according to the position of the optimized bounding box; Based on the target object position mask, constructing a potential representation optimization function using a weighted cross entropy method; The current latent representation in the initial image generation process is optimized according to the latent representation optimization function to obtain the target latent representation.

5. The method according to claim 1 or 4, characterized in that: The execution process of optimizing the current potential representation in the initial image generation process based on the optimized bounding box is as follows: Iteratively performing optimization of a position of a candidate bounding box in the process of generating the initial image and optimizing a current potential representation in the process of generating the initial image to obtain an optimized potential representation; The optimized latent representation is used to replace the current latent representation until a preset number of cross-optimization steps is met to obtain the target latent representation.

6. The method according to claim 1, characterized in that The Stable Diffusion model is used to generate the initial image corresponding to the text prompt word.

7. A text-generated image device, characterized in that: The device comprises: An initial image generating unit, configured to generate an initial image corresponding to the text prompt word in response to the acquired text prompt word containing the target object; wherein the text prompt word includes the target number of the target object, and the initial image includes the generated number of the target object; A bounding box optimization unit is used to optimize the position of the candidate bounding box in the initial image generation process to generate an optimized bounding box when the target number is not equal to the generated number; wherein the optimization of the position of the candidate bounding box in the initial image generation process to generate the optimized bounding box includes: extracting the cross-attention score map in the initial image generation process, and binarizing the cross-attention score map to obtain a mask image corresponding to the text prompt word; generating multiple candidate bounding boxes corresponding to the target object in the mask image; constructing a bounding box position optimization function; optimizing the position of the candidate bounding box using the bounding box position optimization function to generate the optimized bounding box; A potential representation optimization unit, configured to optimize the current potential representation in the initial image generation process based on the optimized bounding box, so that the peak of the cross-attention score map in the initial image generation process matches the optimized bounding box, and obtain a target potential representation; The target image generating unit is used to update the initial image based on the target potential representation to generate a target image so that the generated number of the target objects in the target image is equal to the target number.

8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the text-to-image method according to any one of claims 1 to 6 by executing the computer instructions.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the text-to-image generation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Visual question and answer data enhancement method and device, equipment and storage medium

    CN119128118A

  • Layered contrast anti-fact learning method and system oriented to visual question and answer model

    CN119166795A