An image generation method, apparatus, system, and storage medium

By vectorizing images and text and training models, the problem of objects not appearing or being poorly configured in text-to-image models has been solved, achieving both accuracy and aesthetics in image generation.

CN118470374BActive Publication Date: 2026-07-24GUILIN UNIV OF ELECTRONIC TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUILIN UNIV OF ELECTRONIC TECH
Filing Date
2024-04-03
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing text-to-image diffusion models suffer from problems such as objects not appearing or being poorly configured when generating images containing multiple object categories, resulting in unsatisfactory generation results and low accuracy.

Method used

The original latent representation vector is obtained by mapping the image to be processed. The original prompt text is then vectorized and analyzed. A training model is built and trained to generate an image generation model. This model is then used to predict the prompt text to be generated, thereby achieving accurate capture and mapping of key attributes in the text description.

Benefits of technology

It improves the attribute correspondence between generated images and text, enhances the accuracy of image generation, and ensures that multiple object categories appear accurately and are configured reasonably in the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118470374B_ABST
    Figure CN118470374B_ABST
Patent Text Reader

Abstract

The application provides an image generation method, device and system and a storage medium, and belongs to the technical field of image generation. The method comprises the following steps: performing mapping processing on a to-be-processed image to obtain an original latent representation vector; performing vectorization analysis on an original prompt word text to obtain an original text vector group; training a training model by using the original prompt word text, the original latent representation vector and the original text vector group to obtain an image generation model; and predicting a to-be-generated prompt word text by using the image generation model to obtain an image generation result. The application realizes accurate capture and accurate mapping of key attributes in a text description, improves the attribute correspondence relationship between a generated image and the text, improves the accuracy of image generation, and has important practical application value in the field of image generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates primarily to the field of image generation technology, specifically to an image generation method, apparatus, system, and storage medium. Background Technology

[0002] Despite significant advancements in recent text-to-image diffusion models, which have achieved increasingly higher levels of image quality, resolution, realism, and diversity, a major consistency issue remains between text prompts and generated image content. When multiple object categories, particularly those that don't frequently coexist in the real world, are included in the text prompt, the model loses its combinatorial ability; objects may not appear in the image, or their placement may be aesthetically unappealing. These issues contribute to suboptimal image generation, poor accuracy, and weak relevance between the generated image and the text. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide an image generation method, apparatus, system and storage medium that address the shortcomings of the prior art.

[0004] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: An image generation method, comprising the following steps:

[0005] Import multiple images to be processed and the original prompt text corresponding to each image to be processed;

[0006] Each of the images to be processed is mapped to obtain the original latent representation vector corresponding to each of the images to be processed.

[0007] Each of the original prompt words is vectorized and analyzed to obtain the original text vector group corresponding to each of the original prompt words.

[0008] A training model is constructed by training the model using all the original prompt word texts, all the original latent representation vectors, and all the original text vector groups to obtain an image generation model.

[0009] Import the text to be generated as a prompt word, and use the image generation model to predict the text to be generated, thereby obtaining the image generation result.

[0010] Another technical solution of the present invention to solve the above-mentioned technical problems is as follows: An image generation device, comprising:

[0011] The import module is used to import multiple images to be processed and the original prompt text corresponding to each of the images to be processed.

[0012] The mapping processing module is used to perform mapping processing on each of the images to be processed to obtain the original latent representation vectors corresponding to each of the images to be processed.

[0013] The vectorization analysis module is used to perform vectorization analysis on each of the original prompt word texts to obtain the original text vector group corresponding to each of the prompt word texts;

[0014] The training module is used to build a training model, which is trained by using all the original prompt word texts, all the original latent representation vectors, and all the original text vector groups to obtain an image generation model;

[0015] The import module is also used to import the text of the prompt words to be generated;

[0016] The image generation result acquisition module is used to predict the text of the prompt word to be generated through the image generation model to obtain the image generation result.

[0017] Based on the above-described image generation method, the present invention also provides an image generation system.

[0018] Another technical solution of the present invention to solve the above-mentioned technical problems is as follows: an image generation system, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the image generation method described above is implemented.

[0019] Based on the above-described image generation method, the present invention also provides a computer-readable storage medium.

[0020] Another technical solution of the present invention to solve the above-mentioned technical problems is as follows: a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the image generation method as described above.

[0021] The beneficial effects of this invention are as follows: by mapping the image to be processed to obtain the original latent representation vector, by vectorizing the original prompt text to obtain the original text vector group, by training the training model with the original prompt text, the original latent representation vector, and the original text vector group to obtain the image generation model, and by predicting the prompt text to be generated by the image generation model to obtain the image generation result, the invention achieves accurate capture and mapping of key attributes in the text description, improves the attribute correspondence between the generated image and the text, and improves the accuracy of image generation, thus having important practical application value in the field of image generation. Attached Figure Description

[0022] Figure 1This is a schematic flowchart of an image generation method provided in an embodiment of the present invention;

[0023] Figure 2 This is a block diagram of an image generation device provided in an embodiment of the present invention. Detailed Implementation

[0024] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0025] Figure 1 This is a flowchart illustrating an image generation method provided in an embodiment of the present invention.

[0026] like Figure 1 As shown, an image generation method includes the following steps:

[0027] Import multiple images to be processed and the original prompt text corresponding to each image to be processed;

[0028] Each of the images to be processed is mapped to obtain the original latent representation vector corresponding to each of the images to be processed.

[0029] Each of the original prompt words is vectorized and analyzed to obtain the original text vector group corresponding to each of the original prompt words.

[0030] A training model is constructed by training the model using all the original prompt word texts, all the original latent representation vectors, and all the original text vector groups to obtain an image generation model.

[0031] Import the text to be generated as a prompt word, and use the image generation model to predict the text to be generated, thereby obtaining the image generation result.

[0032] It should be understood that, through the optimized model (i.e., the image generation model), an image containing multiple target objects (i.e., the image generation result) is generated based on the new input text description (i.e., the prompt text to be generated).

[0033] Specifically, the system provides a text description containing multiple target objects (i.e., the text to be generated as prompt words) as input, uses an optimized diffusion model (i.e., an image generation model) to generate an image based on the input text description (i.e., the text to be generated as prompt words), and outputs the generated image (i.e., the image generation result) as output, which is then displayed to the user or processed further.

[0034] In the above embodiments, the original latent representation vector is obtained by mapping the image to be processed, and the original text vector group is obtained by vectorization analysis of the original prompt text. The image generation model is obtained by training the training model with the original prompt text, the original latent representation vector, and the original text vector group. The image generation result is obtained by predicting the prompt text to be generated by the image generation model. This achieves accurate capture and mapping of key attributes in the text description, improves the attribute correspondence between the generated image and the text, and improves the accuracy of image generation. It has important practical application value in the field of image generation.

[0035] Optionally, as an embodiment of the present invention, the process of performing mapping processing on each of the images to be processed to obtain the original latent representation vectors corresponding to each of the images to be processed includes:

[0036] Each of the images to be processed is mapped using a pre-constructed variational autoencoder to obtain the original latent representation vector corresponding to each image to be processed.

[0037] It should be understood that the image (i.e. the image to be processed) is mapped from the pixel space to the latent space.

[0038] Specifically, define an image (i.e., the image to be processed). The trained variational autoencoder (i.e., the pre-built variational autoencoder) is used to map the image (i.e., the image to be processed) from the pixel space to the latent space, to obtain the implicit representation of the image z = ε(x), where z is the latent representation vector (i.e., the original latent representation vector).

[0039] In the above embodiments, the original latent representation vector is obtained by mapping the image to be processed through a pre-constructed variational autoencoder, which realizes the accurate capture and mapping of key attributes in the text description and improves the attribute correspondence between the generated image and the text.

[0040] Optionally, as an embodiment of the present invention, the process of performing vectorization analysis on each of the original prompt word texts to obtain the original text vector group corresponding to each of the prompt word texts includes:

[0041] Each of the original prompt words is converted into a text embedding vector corresponding to each of the original prompt words using the CLIP text encoder.

[0042] The original prompt texts are preprocessed using Python tools to obtain multiple token vectors corresponding to each original prompt text. The original text vector group includes the text embedding vector and the multiple token vectors.

[0043] It should be understood that the input text description (i.e., the original prompt text) is transformed into a text embedding (i.e., a text embedding vector) and a series of tokens (i.e., a token vector).

[0044] Specifically, the CLIP text encoder is used to convert the cue word description (i.e., the original cue word text) into a text embedding. (i.e., text embedding vector), where y is the text cue describing the image (i.e., the original cue text); θ is the neural network parameter.

[0045] Specifically, the original prompt texts are preprocessed using Python tools, as follows:

[0046] First, the input text description (i.e., the original prompt text) (usually a sentence) is preprocessed, including removing stop words (such as the, as, is, which, etc.), punctuation marks, and performing stemming (e.g., converting cats to cat, fisher to fish, etc.) or lemmatization (e.g., converting ate to eat, etc.) for tokenization in subsequent steps. Then, a series of tokens ei (i.e., token vectors) are generated using a part-of-speech tagger (POS Tagger, which can be the NLTK package). in It is the length of the embedded text.

[0047] In the above embodiments, vectorization analysis is performed on each original prompt word text to obtain the original text vector group, which improves the attribute correspondence between the generated image and the text, improves the accuracy of image generation, and has important practical application value in the field of image generation.

[0048] Optionally, as an embodiment of the present invention, the training model includes a semantic alignment model and a diffusion model.

[0049] The process of constructing and training the model, which involves training the model using all the original prompt word texts, all the original latent representation vectors, and all the original text vector groups to obtain the image generation model, includes:

[0050] The semantic alignment model is used to extract the mapping map of each token vector corresponding to each original prompt word text, so as to obtain multiple binary segmentation mapping maps corresponding to each original prompt word text. The multiple binary segmentation mapping maps corresponding to each original prompt word text are then combined to obtain a set of binary segmentation mapping maps corresponding to each original prompt word text.

[0051] The diffusion model is used to add preset noise to each of the original latent representation vectors to obtain the target latent representation vectors corresponding to each of the original latent representation vectors.

[0052] The loss function is analyzed for all the original prompt words, all the text embedding vectors, all the token vectors, all the original latent representation vectors, all the binary segmentation map sets, and all the target latent representation vectors to obtain the target loss function;

[0053] The parameters of the trained model are updated according to the target loss function to obtain the image generation model.

[0054] It should be understood that the binary segmentation map is extracted, and random noise is added during the forward process.

[0055] Specifically, an image-based semantic alignment model (Grounded SAM) is used to extract all binary segmentation maps M corresponding to noun tokens in the image from the cue (i.e., token vector). i This can be achieved by calculating the similarity or distance between image blocks and tokens.

[0056] Specifically, in the forward process of the diffusion model, noise ∈ (i.e., preset noise) is successively added to the latent vector z (i.e., the original latent representation vector) in the latent space over time, resulting in a series of latent space vectors z1,…,z t ,…,Z T (i.e., the target latent representation vector).

[0057] In the above embodiments, the image generation model is obtained by training the training model with all original prompt word texts, all original latent representation vectors, and all original text vector groups. This achieves accurate capture and mapping of key attributes in the text description, improves the attribute correspondence between the generated output and the text, and has important practical application value.

[0058] Optionally, as an embodiment of the present invention, the process of analyzing the loss function of all the original prompt word texts, all the text embedding vectors, all the token vectors, all the original latent representation vectors, all the binary segmentation map sets, and all the target latent representation vectors to obtain the target loss function includes:

[0059] The first loss function is obtained by calculating the loss function of all the text embedding vectors, all the original latent representation vectors, and all the target latent representation vectors using the first equation. The first equation is:

[0060]

[0061] in, Let ε(x) be the first loss function, ε(x) be the original latent representation vector corresponding to the x-th image to be processed, y be the original prompt text of the y-th word, t be the preset time step, and ∈ be random noise. θ () is the noise function, z t (x) is the target latent representation vector corresponding to the x-th image to be processed. Let be the text embedding vector corresponding to the y-th original prompt word. For mathematical expectation, The square of the L2 norm;

[0062] The original cross-attention graph is obtained by calculating the original cross-attention graph corresponding to each of the target latent representation vectors and the token vectors corresponding to each of the original prompt word texts using the second formula. The second formula is:

[0063]

[0064] in,

[0065] Among them, A i Let H be the original cross-attention graph corresponding to the i-th token vector, where H is the total number of cross-attention points. The noise image weights are the values ​​corresponding to the h-th cross-attention event in the x-th image to be processed. Let d be the token weight corresponding to the h-th cross-attention of the i-th token vector. k Let K be the dimension, and T be the transpose. and All of these are weight matrices for the h-th cross-attention. For a flat function, z t (x) is the target latent representation vector corresponding to the x-th image to be processed, e i Let i be the i-th token vector;

[0066] The second loss function is obtained by calculating the loss function of all the original cross-attention maps and all the binary segmentation mapping maps using the third equation. The third equation is:

[0067]

[0068] in,

[0069] in, Let A be the second loss function, where N is the total number of token vectors, and A is the second loss function. (i,u) This is the original cross-attention map corresponding to the u-th region in the i-th token vector. To preset the variable latent representation resolution, For the prediction space region corresponding to the i-th token vector, Ai M is the original cross-attention graph corresponding to the i-th token vector. i Let i be the set of binary segmentation mapping graphs corresponding to the i-th token vector. Pixel-wise multiplication;

[0070] Based on all the token vectors, all the original cross-attention maps are divided into multiple cross-attention maps to be processed corresponding to each of the original cue words, multiple adjective cross-attention maps corresponding to each of the original cue words, and multiple noun cross-attention maps corresponding to each of the original cue words;

[0071] Count the total number of all the original prompt words to obtain the total number of original prompt words.

[0072] By combining multiple adjective cross-attention maps corresponding to each of the original prompt word texts and multiple noun cross-attention maps corresponding to each of the original prompt word texts, a target cross-attention map set corresponding to each of the original prompt word texts is obtained.

[0073] The fourth equation is used to calculate the loss function for the total number of original prompt words, all target cross-attention maps, all adjective cross-attention maps, and all noun cross-attention maps, thus obtaining the third loss function. The fourth equation is:

[0074]

[0075] in,

[0076] in,

[0077] in, The third loss function is denoted by k, where k is the total number of original prompt words, and P(S) is the total number of prompt words. j Let A be the set of target cross-attention graphs corresponding to the j-th original cue word text. mj For the cross-attention graph of the m-th adjective corresponding to the j-th original cue word text, A nj This is the cross-attention graph of the nth noun corresponding to the jth original cue word text, where dist() is the distance metric, and D... KL (||) represents the divergence, and pixels represent the number of pixels.

[0078] Each of the original prompt word texts is combined to obtain a set of cross-attention graphs to be processed corresponding to each of the original prompt word texts.

[0079] The fourth loss function is obtained by calculating the loss function of the total number of original prompt word texts, all sets of cross-attention graphs to be processed, all sets of target cross-attention graphs, all cross-attention graphs to be processed, all adjective cross-attention graphs, and all noun cross-attention graphs using the fifth formula. The fifth formula is:

[0080]

[0081] in,

[0082] in,

[0083] in, For the fourth loss function, P(S) j U(S) represents the target cross-attention graph set corresponding to the j-th original cue word text. j Let A be the set of cross-attention graphs to be processed corresponding to the j-th original cue word text. mj For the cross-attention graph of the m-th adjective corresponding to the j-th original cue word text, A vj For the v-th cross-attention graph corresponding to the j-th original cue word text, A nj This is the cross-attention graph of the nth noun corresponding to the jth original cue word text, where dist() is the distance metric, and D... KL (||) represents the divergence, and pixels represent the number of pixels.

[0084] The target loss function is obtained by calculating the first loss function, the second loss function, the third loss function, and the fourth loss function using the sixth equation. The sixth equation is:

[0085]

[0086] in, Let be the target loss function. Let λ be the first loss function and λ be the scaling factor. Let L be the second loss function of the l-th layer, and L be the preset total number of layers. For the third loss function, This is the fourth loss function.

[0087] It should be understood that optimization is performed using a loss function (i.e., the target loss function).

[0088] Specifically, using the loss function (i.e., the first loss function) denoising function ∈ in the reverse process of training and optimizing the diffusion model θ , where ε and Keep frozen, the formula is as follows:

[0089]

[0090] Where ε(x) is the image embedding vector (i.e., the original latent representation vector); y is the text cue describing the image (i.e., the original cue text); ∈ is random noise following a Gaussian distribution; and t is the time step. It is a noise function, z t It is a noisy embedding vector, obtained by adding random noise ∈ to the embedding vector of the image.

[0091] It should be understood that, due to The function is simply optimized to predict noise, and the latent function of the image is reconstructed through denoising, thus embedding each token. and potentially noisy image z t There is no explicit optimization. This results in a poor understanding of token levels in the LDM.

[0092] Just use Training a diffusion model often results in the activation of cross-attention maps for different instance tokens failing to focus on corresponding instances appearing in the images during training, leading to poor ability to combine multi-class instances during inference. To better achieve multi-class instance combination, a training constraint is added to supervise the activation region of the cross-attention map, as shown in the following formula:

[0093]

[0094]

[0095]

[0096]

[0097] Where K i Q represents the weighted token embedding (i.e., token weight). i Let K represent the weighted noisy image embedding (i.e., the weights of the noisy image), and d_k represent the dimension of K. h∈{1,…,H} represents each head of the multi-head cross-attention mechanism. This represents a function that can potentially flatten a two-dimensional image into a one-dimensional image (i.e., a flattening function). and The weight matrix that needs to be learned is shown in the following formula:

[0098]

[0099] Where A (i,u) ∈R represents A iThe region at point u in the cross-attention map space (i.e., the original cross-attention map), This indicates a variable latent resolution (the latent resolution is different for each layer). The predicted spatial region is equal to the pixel-wise multiplication of the cross-attention map and the binary segmentation map; N is the number of tokens. in This is pixel-wise multiplication.

[0100] Specifically, consider a pair of nouns and their modifiers. It is expected that the cross-attention maps of the modifiers largely overlap with those of the nouns, while largely disjoint with those corresponding to other nouns and modifiers. To encourage the denoising process to conform to the spatial relationships between these attention maps, a loss function is designed that operates on the cross-attention maps.

[0101] First, the formula for minimizing the distance between all modifier pairs and their corresponding entity nouns (m, n) (maximizing overlap) is as follows:

[0102]

[0103] Where k represents k noun modifiers {S1, S2, ..., S...} k},P(S j ) represents the j-th set S j All symbol pairs (m,n) between the noun root n and its modifier m. Using {A1,A2,…,A…} N} represents the attention map of N markers in the prompt, dist(A m A n ) represents attention map A m and A n Distance metric between them.

[0104] A loss function was also constructed, comparing modifier and entity noun pairs with the remaining words in the prompt that are grammatically irrelevant to these pairs. In other words, this loss is defined between words within the (modifier, entity noun) set and words outside the set, as follows:

[0105]

[0106] Where U(S) j ) indicates that S is excluded from the complete set of words. j A set of words that do not match the given words. v It is an attention map corresponding to a given unrelated word v. A iA j It is a normalized attention map.

[0107] It should be understood that the formula for calculating the total loss function (i.e., the target loss function) is as follows:

[0108]

[0109] λ is the scaling factor, which is a hyperparameter.

[0110] In the above embodiments, the target loss function is obtained by analyzing the loss function of all original prompt text, all text embedding vectors, all token vectors, all original latent representation vectors, all binary segmentation mapping graph sets, and all target latent representation vectors. This achieves accurate capture and mapping of key attributes in the text description, improves the attribute correspondence between the generated image and the text, and improves the accuracy of image generation. It has important practical application value in the field of image generation.

[0111] Alternatively, as another embodiment of the present invention, the steps of the present invention are as follows:

[0112] This invention maps images from pixel space to latent space; transforms input text descriptions into text embeddings and a series of tokens; extracts binary segmentation maps, progressively adding random noise during the forward pass; and performs a backward pass using a diffusion model. Optimization is achieved using a loss function; the optimized model then generates images containing multiple target objects based on new input text descriptions. This invention utilizes a diffusion model linguistic binding technique based on attention map alignment to accurately capture and map key attributes in text descriptions, improving the attribute correspondence between the generated output and the text, and possesses significant practical application value.

[0113] Alternatively, as another embodiment of the present invention, the steps of the present invention are as follows:

[0114] S1: Map the image from pixel space to latent space;

[0115] S2: Transform the input text description into a text embedding and a series of tokens;

[0116] S3: Extract the binary segmentation map and add random noise during the forward process;

[0117] S4: Optimize using a loss function;

[0118] S5: Using the optimized model, generate an image containing multiple target objects based on the new input text description.

[0119] Figure 2This is a block diagram of an image generation device provided in an embodiment of the present invention.

[0120] Alternatively, as another embodiment of the present invention, such as Figure 2 As shown, an image generation apparatus includes:

[0121] The import module is used to import multiple images to be processed and the original prompt text corresponding to each of the images to be processed.

[0122] The mapping processing module is used to perform mapping processing on each of the images to be processed to obtain the original latent representation vectors corresponding to each of the images to be processed.

[0123] The vectorization analysis module is used to perform vectorization analysis on each of the original prompt word texts to obtain the original text vector group corresponding to each of the prompt word texts;

[0124] The training module is used to build a training model, which is trained by using all the original prompt word texts, all the original latent representation vectors, and all the original text vector groups to obtain an image generation model;

[0125] The import module is also used to import the text of the prompt words to be generated;

[0126] The image generation result acquisition module is used to predict the text of the prompt word to be generated through the image generation model to obtain the image generation result.

[0127] Optionally, as an embodiment of the present invention, the mapping processing module is specifically used for:

[0128] Each of the images to be processed is mapped using a pre-constructed variational autoencoder to obtain the original latent representation vector corresponding to each image to be processed.

[0129] Optionally, as an embodiment of the present invention, the vectorization analysis module is specifically used for:

[0130] Each of the original prompt words is converted into a text embedding vector corresponding to each of the original prompt words using the CLIP text encoder.

[0131] The original prompt texts are preprocessed using Python tools to obtain multiple token vectors corresponding to each original prompt text. The original text vector group includes the text embedding vector and the multiple token vectors.

[0132] Optionally, another embodiment of the present invention provides an image generation system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the image generation method described above. This system can be a computer or similar system.

[0133] Optionally, another embodiment of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the image generation method described above.

[0134] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0135] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0136] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0137] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention, depending on actual needs.

[0138] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0139] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0140] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An image generation method, characterized in that, Includes the following steps: Import multiple images to be processed and the original prompt text corresponding to each image to be processed; Each of the images to be processed is mapped to obtain the original latent representation vector corresponding to each of the images to be processed. Each of the original prompt words is vectorized and analyzed to obtain the original text vector group corresponding to each of the original prompt words. A training model is constructed by training the model using all the original prompt word texts, all the original latent representation vectors, and all the original text vector groups to obtain an image generation model. Import the text to be generated as a prompt word, and use the image generation model to predict the text to be generated to obtain the image generation result; The training model includes a semantic alignment model and a diffusion model. The process of constructing and training the model, which involves training the model using all the original prompt word texts, all the original latent representation vectors, and all the original text vector groups to obtain the image generation model, includes: The semantic alignment model is used to extract the mapping map of each token vector corresponding to each original prompt word text, so as to obtain multiple binary segmentation mapping maps corresponding to each original prompt word text. The multiple binary segmentation mapping maps corresponding to each original prompt word text are then combined to obtain a set of binary segmentation mapping maps corresponding to each original prompt word text. The diffusion model is used to add preset noise to each of the original latent representation vectors to obtain the target latent representation vectors corresponding to each of the original latent representation vectors. The loss function is analyzed for all the original prompt words, all the text embedding vectors, all the token vectors, all the original latent representation vectors, all the binary segmentation map sets, and all the target latent representation vectors to obtain the target loss function; The parameters of the trained model are updated according to the target loss function to obtain the image generation model; The process of analyzing the loss function of all the original prompt word texts, all the text embedding vectors, all the token vectors, all the original latent representation vectors, all the binary segmentation map sets, and all the target latent representation vectors to obtain the target loss function includes: The first loss function is obtained by calculating the loss function of all the text embedding vectors, all the original latent representation vectors, and all the target latent representation vectors using the first equation. The first equation is: , in, For the first loss function, For the first The original latent representation vector corresponding to each image to be processed. For the first The original prompt text, To preset the time step, It is random noise. For noise function, For the first The target latent representation vector corresponding to each image to be processed For the first The text embedding vector corresponding to each original prompt word text. For mathematical expectation, The square of the L2 norm; The original cross-attention graph is obtained by calculating the original cross-attention graph corresponding to each of the target latent representation vectors and the token vectors corresponding to each of the original prompt word texts using the second formula. The second formula is: , in, , , in, For the first The original cross-attention graph corresponding to each token vector. The total number of cross-attention points, For the first The first image to be processed The weights of the noisy image corresponding to each cross-attention point. For the first The token vector of the first The token weights corresponding to each cross-attention. For the dimension of K, For transpose, and All are the first A weight matrix for cross-attention. It is a flat function. For the first The target latent representation vector corresponding to each image to be processed For the first One token vector; The second loss function is obtained by calculating the loss function of all the original cross-attention maps and all the binary segmentation mapping maps using the third equation. The third equation is: , in, , in, For the second loss function, The total number of token vectors, For the first The th token vector in the th token vector The original cross-attention diagram corresponding to each region To preset the variable latent representation resolution, For the first The prediction space region corresponding to each token vector. For the first The original cross-attention graph corresponding to each token vector. For the first The set of binary segmentation mapping graphs corresponding to each token vector. Pixel-wise multiplication; Based on all the token vectors, all the original cross-attention maps are divided into multiple cross-attention maps to be processed corresponding to each of the original cue words, multiple adjective cross-attention maps corresponding to each of the original cue words, and multiple noun cross-attention maps corresponding to each of the original cue words; Count the total number of all the original prompt words to obtain the total number of original prompt words. By combining multiple adjective cross-attention maps corresponding to each of the original prompt word texts and multiple noun cross-attention maps corresponding to each of the original prompt word texts, a target cross-attention map set corresponding to each of the original prompt word texts is obtained. The fourth equation is used to calculate the loss function for the total number of original prompt words, all target cross-attention maps, all adjective cross-attention maps, and all noun cross-attention maps, thus obtaining the third loss function. The fourth equation is: , in, , in, , , in, For the third loss function, This represents the total number of original prompt words. For the first A set of target cross-attention maps corresponding to each original cue word text. For the first The first original prompt text corresponding to the first A cross-attention diagram of adjectives. For the first The first original prompt text corresponding to the first A diagram showing the intersection of nouns. For distance measurement, For divergence, For pixels; Each of the original prompt word texts is combined to obtain a set of cross-attention graphs to be processed corresponding to each of the original prompt word texts. The fourth loss function is obtained by calculating the loss function of the total number of original prompt word texts, all sets of cross-attention graphs to be processed, all sets of target cross-attention graphs, all cross-attention graphs to be processed, all adjective cross-attention graphs, and all noun cross-attention graphs using the fifth formula. The fifth formula is: , in, , , in, , , , , in, This is the fourth loss function. For the first A set of target cross-attention maps corresponding to each original cue word text. For the first The set of cross-attention graphs to be processed corresponding to the original cue word texts. For the first The first original prompt text corresponding to the first A cross-attention diagram of adjectives. For the first The first original prompt text corresponding to the first One cross-attention diagram to be processed For the first The first original prompt text corresponding to the first A diagram showing the intersection of nouns. For distance measurement, For divergence, For pixels; The target loss function is obtained by calculating the first loss function, the second loss function, the third loss function, and the fourth loss function using the sixth equation. The sixth equation is: , in, Let be the target loss function. For the first loss function, As a scaling factor, For the first The second loss function of the layer, To preset the total number of floors, For the third loss function, This is the fourth loss function.

2. The image generation method according to claim 1, characterized in that, The process of mapping each of the images to be processed to obtain the original latent representation vector corresponding to each of the images to be processed includes: Each of the images to be processed is mapped using a pre-constructed variational autoencoder to obtain the original latent representation vector corresponding to each image to be processed.

3. The image generation method according to claim 1, characterized in that, The process of performing vectorization analysis on each of the original prompt word texts to obtain the original text vector group corresponding to each of the prompt word texts includes: Each of the original prompt words is converted into a text embedding vector corresponding to each of the original prompt words using the CLIP text encoder. The original prompt texts are preprocessed using Python tools to obtain multiple token vectors corresponding to each original prompt text. The original text vector group includes the text embedding vector and the multiple token vectors.

4. An image generation apparatus, characterized in that, include: The import module is used to import multiple images to be processed and the original prompt text corresponding to each of the images to be processed. The mapping processing module is used to perform mapping processing on each of the images to be processed to obtain the original latent representation vectors corresponding to each of the images to be processed. The vectorization analysis module is used to perform vectorization analysis on each of the original prompt word texts to obtain the original text vector group corresponding to each of the prompt word texts; The training module is used to build a training model, which is trained by using all the original prompt word texts, all the original latent representation vectors, and all the original text vector groups to obtain an image generation model; The import module is also used to import the text of the prompt words to be generated; The image generation result acquisition module is used to predict the text of the prompt word to be generated through the image generation model to obtain the image generation result; The training model includes a semantic alignment model and a diffusion model. The training module is specifically used for: The semantic alignment model is used to extract the mapping map of each token vector corresponding to each original prompt word text, so as to obtain multiple binary segmentation mapping maps corresponding to each original prompt word text. The multiple binary segmentation mapping maps corresponding to each original prompt word text are then combined to obtain a set of binary segmentation mapping maps corresponding to each original prompt word text. The diffusion model is used to add preset noise to each of the original latent representation vectors to obtain the target latent representation vectors corresponding to each of the original latent representation vectors. The loss function is analyzed for all the original prompt words, all the text embedding vectors, all the token vectors, all the original latent representation vectors, all the binary segmentation map sets, and all the target latent representation vectors to obtain the target loss function; The parameters of the trained model are updated according to the target loss function to obtain the image generation model; In the training module, the process of analyzing the loss function for all the original prompt word texts, all the text embedding vectors, all the token vectors, all the original latent representation vectors, all the binary segmentation map sets, and all the target latent representation vectors to obtain the target loss function includes: The first loss function is obtained by calculating the loss function of all the text embedding vectors, all the original latent representation vectors, and all the target latent representation vectors using the first equation. The first equation is: , in, For the first loss function, For the first The original latent representation vector corresponding to each image to be processed. For the first The original prompt text, To preset the time step, It is random noise. For noise function, For the first The target latent representation vector corresponding to each image to be processed For the first The text embedding vector corresponding to each original prompt word text. For mathematical expectation, The square of the L2 norm; The original cross-attention graph is obtained by calculating the original cross-attention graph corresponding to each of the target latent representation vectors and the token vectors corresponding to each of the original prompt word texts using the second formula. The second formula is: , in, , , in, For the first The original cross-attention graph corresponding to each token vector. The total number of cross-attention points, For the first The first image to be processed The weights of the noisy image corresponding to each cross-attention point. For the first The token vector of the first The token weights corresponding to each cross-attention. For the dimension of K, For transpose, and All are the first A weight matrix for cross-attention. It is a flat function. For the first The target latent representation vector corresponding to each image to be processed For the first One token vector; The second loss function is obtained by calculating the loss function of all the original cross-attention maps and all the binary segmentation mapping maps using the third equation. The third equation is: , in, , in, For the second loss function, The total number of token vectors, For the first The th token vector in the th token vector The original cross-attention diagram corresponding to each region To preset the variable latent representation resolution, For the first The prediction space region corresponding to each token vector. For the first The original cross-attention graph corresponding to each token vector. For the first The set of binary segmentation mapping graphs corresponding to each token vector. Pixel-wise multiplication; Based on all the token vectors, all the original cross-attention maps are divided into multiple cross-attention maps to be processed corresponding to each of the original cue words, multiple adjective cross-attention maps corresponding to each of the original cue words, and multiple noun cross-attention maps corresponding to each of the original cue words; Count the total number of all the original prompt words to obtain the total number of original prompt words. By combining multiple adjective cross-attention maps corresponding to each of the original prompt word texts and multiple noun cross-attention maps corresponding to each of the original prompt word texts, a target cross-attention map set corresponding to each of the original prompt word texts is obtained. The fourth equation is used to calculate the loss function for the total number of original prompt words, all target cross-attention maps, all adjective cross-attention maps, and all noun cross-attention maps, thus obtaining the third loss function. The fourth equation is: , in, , in, , , in, For the third loss function, This represents the total number of original prompt words. For the first A set of target cross-attention maps corresponding to each original cue word text. For the first The first original prompt text corresponding to the first A cross-attention diagram of adjectives. For the first The first original prompt text corresponding to the first A diagram showing the intersection of nouns. For distance measurement, For divergence, For pixels; Each of the original prompt word texts is combined to obtain a set of cross-attention graphs to be processed corresponding to each of the original prompt word texts. The fourth loss function is obtained by calculating the loss function of the total number of original prompt word texts, all sets of cross-attention graphs to be processed, all sets of target cross-attention graphs, all cross-attention graphs to be processed, all adjective cross-attention graphs, and all noun cross-attention graphs using the fifth formula. The fifth formula is: , in, , , in, , , , , in, This is the fourth loss function. For the first A set of target cross-attention maps corresponding to each original cue word text. For the first The set of cross-attention graphs to be processed corresponding to the original cue word texts. For the first The first original prompt text corresponding to the first A cross-attention diagram of adjectives. For the first The first original prompt text corresponding to the first One cross-attention diagram to be processed For the first The first original prompt text corresponding to the first A diagram showing the intersection of nouns. For distance measurement, For divergence, For pixels; The target loss function is obtained by calculating the first loss function, the second loss function, the third loss function, and the fourth loss function using the sixth equation. The sixth equation is: , in, Let be the target loss function. For the first loss function, As a scaling factor, For the first The second loss function of the layer, To preset the total number of floors, For the third loss function, This is the fourth loss function.

5. The image generation apparatus according to claim 4, characterized in that, The mapping processing module is specifically used for: Each of the images to be processed is mapped using a pre-constructed variational autoencoder to obtain the original latent representation vector corresponding to each image to be processed.

6. The image generation apparatus according to claim 4, characterized in that, The vectorization analysis module is specifically used for: Each of the original prompt words is converted into a text embedding vector corresponding to each of the original prompt words using the CLIP text encoder. The original prompt texts are preprocessed using Python tools to obtain multiple token vectors corresponding to each original prompt text. The original text vector group includes the text embedding vector and the multiple token vectors.

7. An image generation system, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the image generation method as described in any one of claims 1 to 3.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the image generation method as described in any one of claims 1 to 3.