Image generation method and device of embedded image prompt adapter, equipment and storage medium

CN120612385APending Publication Date: 2025-09-09BEIJING XIAOBING YUEDONG TECHNOLOGY CO LTD
View PDF 0 Cites -1 Cited by

Patent Information

Application Number
CN202510584512.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-09-09

Smart Images

  • Figure CN120612385A_ABST
    Figure CN120612385A_ABST
Patent Text Reader

Abstract

The invention provides an image generation method and device of an embedded image prompt adapter, equipment and a storage medium, and relates to the technical field of image processing. The method comprises the following steps: acquiring image prompt information and text prompt information; inputting the image prompt information and the text prompt information into a trained image generation model for processing, and generating a target image; wherein the target image is represented as an image conforming to text description in the text prompt information; the image generation model is obtained by training a neural network by using an image prompt sample, a text prompt sample and decoupling cross-attention output; the decoupling cross-attention output is obtained by performing attention decoupling processing on the image prompt sample and the text prompt sample based on an image prompt adapter. According to the embodiment of the invention, the image prompt adapter is embedded, the target image with relatively high image quality is generated through the decoupling cross-attention output determined by the image prompt adapter, and the calculation complexity can be remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to an image generation method, device, equipment and storage medium embedded with an image prompt adapter. Background Art

[0002] In recent years, text-based image diffusion models have demonstrated remarkable performance in generating high-fidelity images. However, relying solely on textual cues to generate desired images remains challenging, primarily due to the complex cuing engineering involved. As an alternative to textual cues, image cues have garnered attention due to their "a picture is worth a thousand words" nature. However, in practical applications, implementing image cues presents numerous technical challenges, including but not limited to effective encoding of cues, cross-modal information consistency, and the diversity and accuracy of generated images.

[0003] In the related art, users can write text prompts to generate images using powerful text-to-image diffusion models. DALL-E2 is the first attempt to support image prompts. This diffusion model is conditioned on image embeddings rather than text embeddings and requires a prior model to achieve text-to-image capabilities. However, most existing text-to-image diffusion models generate images conditioned on text. For example, the popular SD model is based on text features extracted from a frozen CLIP text encoder.

[0004] In related technologies, generating the desired image using only textual cues typically involves complex cuing engineering. When textual descriptions contain multiple stylistic elements, the model struggles to effectively integrate them, resulting in a mixed or unclear style of the generated image. This limits the user's ability to accurately generate images from simple textual descriptions. While existing methods that directly fine-tune from pre-trained models (such as Stable unCLIP) are effective, they require significant computational resources and are incompatible with other base models, textual cues, and structural control.

[0005] Methods that replace text encoders with image encoders (such as the prior model of DALL-E2) only support image cues, which prevents users from using both text and image cues to enrich and refine generated content. In addition, fine-tuning the image encoder alone is usually not enough to guarantee image quality and may lead to generalization issues. Existing adapters (such as ControlNet and T2I-Adapter) have difficulty matching the performance of fine-tuned image cues models or models trained from scratch. The main reason is that image features cannot be effectively embedded in pre-trained models. Most methods usually input the concatenated features into a frozen cross-attention layer, which prevents the diffusion model from capturing fine-grained information from image cues, resulting in generated images with poor details and quality.

[0006] Therefore, how to use text prompt information to generate high-quality target images is a technical problem that needs to be solved urgently. Summary of the Invention

[0007] The present invention provides an image generation method, apparatus, device and storage medium embedded with an image prompt adapter, which are used to solve the technical defects of poor image quality and high computational complexity in the prior art, realize the embedding of an image prompt adapter, and generate a target image with higher image quality through the decoupled cross-attention output determined by the image prompt adapter, while significantly reducing the computational complexity.

[0008] In a first aspect, the present invention provides an image generation method for an embedded image prompt adapter, comprising the following steps: Get image prompt information and text prompt information; Inputting the image prompt information and the text prompt information into a trained image generation model for processing to generate a target image; Among them, the target image is represented as an image that conforms to the text description in the text prompt information; the image generation model is obtained by training a neural network using image prompt samples, text prompt samples and decoupled cross-attention outputs; the decoupled cross-attention output is obtained by performing attention decoupling processing on the image prompt samples and the text prompt samples based on an image prompt adapter.

[0009] Preferably, according to the image generation method embedded in an image prompt adapter provided by the present invention, the decoupled cross-attention output is obtained by performing attention decoupling processing on the image prompt sample and the text prompt sample based on the image prompt adapter, comprising: Performing computational processing on the text prompt sample to obtain a text cross-attention output corresponding to the text prompt sample; Performing computational processing on the image prompt sample to obtain an image cross-attention output corresponding to the image prompt sample; Based on the image prompt adapter, attention decoupling processing is performed on the text cross-attention output and the image cross-attention output to obtain the decoupled cross-attention output.

[0010] Preferably, according to the image generation method of the embedded image prompt adapter provided by the present invention, the image generation model comprises at least a text encoder and a cross attention layer; The calculating and processing the text prompt sample to obtain a text cross-attention output corresponding to the text prompt sample includes: Using the text encoder to perform text encoding processing on the text prompt sample to obtain text features; The text features are input into the cross attention layer for calculation and processing to obtain the text cross attention output.

[0011] Preferably, according to the image generation method of the embedded image prompt adapter provided by the present invention, the image generation model comprises at least an image encoder, a linear layer, and a normalization layer; The performing computational processing on the image prompt sample to obtain an image cross attention output corresponding to the image prompt sample includes: Performing image encoding processing on the image prompt sample by using the image encoder to obtain initial image features; Inputting the initial image features into the linear layer for linear transformation processing to obtain corresponding linear features; wherein the linear features are determined by multiplying each initial image feature input into the linear layer by an image weight matrix and adding a bias term; Inputting the linear features into the normalization layer for normalization processing to obtain target image features; The target image features are input into the cross attention layer for calculation and processing to obtain the image cross attention output.

[0012] Preferably, according to an image generation method embedded in an image prompt adapter provided by the present invention, the image prompt adapter performs attention decoupling processing on the text cross-attention output and the image cross-attention output, and the formula for obtaining the decoupled cross-attention output is as follows: Where, Denoted as disentangled cross-attention output, Represented as text cross attention output, Represented as image crisscross attention output.

[0013] Preferably, according to the image generation method embedded in the image prompt adapter provided by the present invention, the image prompt information and the text prompt information are input into a trained image generation model for processing to generate a target image, including: Inputting the image prompt information and the text prompt information into a trained image generation model for processing to generate an initial image; When the initial image is denoised, the initial image is optimized based on the decoupled cross-attention output to generate the target image.

[0014] In a second aspect, the present invention further provides an image generating device embedded in an image prompting adapter, comprising: An information acquisition module is used to acquire image prompt information and text prompt information; A target image generation module is used to input the image prompt information and the text prompt information into a trained image generation model for processing to generate a target image; wherein, the target image is represented as an image that conforms to the text description in the text prompt information; the image generation model is obtained by training a neural network using image prompt samples, text prompt samples and decoupled cross-attention outputs; the decoupled cross-attention outputs are obtained by performing attention decoupling processing on the image prompt samples and the text prompt samples based on an image prompt adapter.

[0015] In a third aspect, the present invention further provides an electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the image generation method of the embedded image prompt adapter as described in any one of the above is implemented.

[0016] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image generation method of the embedded image prompt adapter as described in any one of the above.

[0017] In a fifth aspect, the present invention further provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the image generation method of the embedded image prompt adapter as described in any one of the above is implemented.

[0018] The present invention provides an image generation method, device, equipment and storage medium embedded in an image prompt adapter. The method obtains image prompt information and text prompt information; inputs the image prompt information and the text prompt information into a trained image generation model for processing to generate a target image; wherein the target image is characterized as an image that conforms to the text description in the text prompt information; the image generation model is obtained by training a neural network using image prompt samples, text prompt samples and decoupled cross-attention outputs; and the decoupled cross-attention outputs are obtained by performing attention decoupling processing on the image prompt samples and the text prompt samples based on the image prompt adapter. The method is used to solve the technical defects of poor image quality and high computational complexity in the prior art, realize the embedding of an image prompt adapter, and generate a target image with higher image quality through the decoupled cross-attention output determined by the image prompt adapter, while significantly reducing computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1 This is one of the flow charts of the image generation method of the embedded image prompt adapter provided by the present invention.

[0021] Figure 2 This is the second schematic diagram of the image generation method of the embedded image prompt adapter provided by the present invention.

[0022] Figure 3 It is a structural schematic diagram of the image generating device embedded in the image prompt adapter provided by the present invention.

[0023] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0024] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0025] The following combination Figures 1-4 The present invention describes an image generation method, apparatus, device and storage medium embedded in an image prompt adapter, which is used to address the technical defects of poor image quality and high computational complexity in the prior art, realize the embedded image prompt adapter, and generate a target image with higher image quality through the decoupled cross-attention output determined by the image prompt adapter, while significantly reducing the computational complexity.

[0026] Figure 1 This is one of the flow charts of an image generation method embedded in an image prompt adapter provided by the present invention, such as Figure 1 As shown, the method may include but is not limited to steps S100 to S200: S100, obtaining image prompt information and text prompt information; S200, input the image prompt information and the text prompt information into a trained image generation model for processing to generate a target image; wherein, the target image is represented as an image that conforms to the text description in the text prompt information; the image generation model is obtained by training a neural network using image prompt samples, text prompt samples and decoupled cross-attention outputs; the decoupled cross-attention outputs are obtained by performing attention decoupling processing on the image prompt samples and the text prompt samples based on an image prompt adapter.

[0027] In step S100 of some embodiments, image prompt information and text prompt information are obtained.

[0028] It's important to note that image cues are typically presented in a visual form. These cues are an existing image that can contain specific elements, colors, composition, and other visual information, serving as a reference or foundation for generating the target image. For example, a picture of a mountain outline provides visual cues of the mountain range if a complete landscape image is to be generated.

[0029] Textual prompts are typically presented in textual form, and are a description that uses natural language to convey the requirements, themes, styles, and elements of the desired generated image. For example, "A brown deer runs through a forest, surrounded by colorful flowers and tall trees, with sunlight filtering through the leaves, casting dappled shadows."

[0030] Furthermore, textual hints often guide the overall generation process: Textual hints often guide the overall image generation process from a macro perspective, determining the image's theme, general style, and key elements. The model constructs the overall image framework and content distribution based on the textual description.

[0031] Image cues often provide specific visual references: Image cues provide specific visual references for the generated image, particularly in terms of style, element form, and color scheme. If the image cues contain a unique texture or color effect, the model will attempt to incorporate this effect into the generated image. Furthermore, image cues may limit the scope and possibilities of generated images. For example, if a specific building exterior image is given as a cues, the model will adhere to the architectural structure and exterior features of the image to a certain extent when generating a complete scene containing that building.

[0032] Text prompts are easy to modify and adjust. Modifying a text prompt is relatively simple; all you need to do is change the text. You can adjust the description details, add new requirements, or change the subject matter at any time. For example, if you change "sunny beach" to "cloudy beach," the model will generate an image of the corresponding style based on the new text prompt.

[0033] Image hints are often adjusted in conjunction with text: To achieve the desired generation effect, adjustments are often needed in conjunction with text hints. For example, if you have an image hint of a river and want to change the text hint from "calm river" to "turbulent river," the model needs to consider both the text change and the characteristics of the image itself to generate the new image.

[0034] In step S200 of some embodiments, the image prompt information and the text prompt information are input into a trained image generation model for processing to generate a target image.

[0035] Among them, the target image is represented as an image that conforms to the text description in the text prompt information; the image generation model is obtained by training a neural network using image prompt samples, text prompt samples and decoupled cross-attention outputs; the decoupled cross-attention output is obtained by performing attention decoupling processing on the image prompt samples and the text prompt samples based on an image prompt adapter.

[0036] It can be understood that by inputting the image prompt information and text prompt information input by the user into the trained image generation model, the target image required by the image prompt information and the text prompt information can be directly generated. The target image corresponds to the combination of the image prompt information and the text prompt information, and not only conforms to the image described in the text prompt information, but also meets the image description conditions in the image prompt information.

[0037] Furthermore, in some embodiments of the present invention, the decoupled cross-attention output is obtained by performing attention decoupling processing on the image prompt sample and the text prompt sample based on an image prompt adapter, including: Performing computational processing on the text prompt sample to obtain a text cross-attention output corresponding to the text prompt sample; Performing computational processing on the image prompt sample to obtain an image cross-attention output corresponding to the image prompt sample; Based on the image prompt adapter, attention decoupling processing is performed on the text cross-attention output and the image cross-attention output to obtain the decoupled cross-attention output.

[0038] Furthermore, in some embodiments of the present invention, the image generation model includes at least a text encoder and a cross-attention layer.

[0039] The calculating and processing the text prompt sample to obtain a text cross-attention output corresponding to the text prompt sample includes: Using the text encoder to perform text encoding processing on the text prompt sample to obtain text features; The text features are input into the cross attention layer for calculation and processing to obtain the text cross attention output.

[0040] It is understood that first, a text prompt sample to be processed is input into the text encoder. The text prompt sample can be one or more sentences, phrases, etc., which are the objects of subsequent encoding processing.

[0041] Next, the text encoder performs tokenization on the text prompt sample, breaking the text into individual words or subword units. It then maps each word or subword unit to a corresponding vector representation by searching a pretrained vocabulary. For example, the word "apple" may correspond to a specific vector in the vocabulary. In this way, the entire text prompt sample can be represented as a sequence of vectors.

[0042] Finally, the text encoder uses a pre-trained language model (such as Word2Vec or BERT) to extract features and encode the vector sequence. Taking BERT as an example, it uses a Transformer architecture, a multi-layer bidirectional self-attention mechanism, and a feedforward neural network to capture the complex semantic information and contextual relationships in the text. During this process, the vector representation of each word or subword unit is continuously updated and enriched, ultimately obtaining a text feature representation that reflects the text's semantic information.

[0043] In some embodiments of the present invention, the text features obtained by the text encoder are input into the cross-attention layer. The text features are typically a two-dimensional vector matrix, where each row represents the feature vector of a word or subword unit, and each column represents a different position in a sample.

[0044] In the cross-attention layer, attention weights are calculated between different elements in the text feature. Specifically, attention weights are calculated by comparing each element in the text feature pairwise based on their correlation or similarity. For example, the dot product operation can be used to calculate the attention weight between two vectors. That is, the larger the dot product of the two vectors, the stronger the correlation between them, and the greater the attention weight.

[0045] Finally, a weighted sum is performed on all elements in the text feature based on the calculated attention weights. Specifically, the feature vector of each element is multiplied by the corresponding attention weight, and all products are summed to obtain the final cross-text attention output. The cross-text attention output is a vector representation that integrates the relationships between each element in the text prompt sample.

[0046] It should be noted that the attention weights between different elements in a text prompt sample are calculated using the attention mechanism. The attention mechanism considers the relevance of each element to other elements and calculates the degree of attention each element pays to other elements based on certain rules (such as dot products and concatenation). This is called the attention weight.

[0047] In some embodiments of the present invention, a pre-trained text encoder converts text prompts into text embedding vectors, which is done during the pre-training phase, and the weights of the text encoder remain frozen during this process.

[0048] In the original SD model, the text features from the CLIP text encoder are inserted into the UNet model through the input cross attention layer. Given the query features and text features , the text cross attention output is , can be defined by the following equation: in, is the text cross attention output, , , are the query matrix, key matrix, and value matrix of the attention operation, respectively, and , , is the weight matrix of the trainable linear projection layer.

[0049] By using a pre-trained language model for feature extraction and encoding, the present invention automatically learns semantic information and contextual relationships within text. For example, when processing the sentence "I like to eat apples," the model can understand the relationship between "I," "like," "eat," and "apples," as well as information such as "apples" being a type of fruit.

[0050] By calculating the attention weights between different elements in text features, we can explore the semantic associations and dependencies between words or subword units in text prompt samples. For example, in a sentence describing a person's action, we can find the relationship between the person who performs the action, the action itself, and the person who receives the action, which helps to better understand the semantics of the sentence.

[0051] The size of the attention weight reflects the importance of each element in the text. By weighted summing up the cross-attention output of the text, elements containing important information can be more prominently displayed in the output, thus helping the model better focus on key content.

[0052] In some embodiments of the present invention, the image generation model includes at least an image encoder, a linear layer, and a normalization layer; The performing computational processing on the image prompt sample to obtain an image cross attention output corresponding to the image prompt sample includes: Performing image encoding processing on the image prompt sample by using the image encoder to obtain initial image features; Inputting the initial image features into the linear layer for linear transformation processing to obtain corresponding linear features; wherein the linear features are determined by multiplying each initial image feature input into the linear layer by an image weight matrix and adding a bias term; Inputting the linear features into the normalization layer for normalization processing to obtain target image features; The target image features are input into the cross attention layer for calculation and processing to obtain the image cross attention output.

[0053] It is understandable that the image prompt samples to be processed are input into the image encoder. The image prompt samples can be one or more images, which contain rich visual information and are the objects of subsequent encoding processing.

[0054] The image encoder performs feature extraction on the image prompt sample to obtain an initial feature representation of the image. This is typically achieved using methods such as convolutional neural networks (CNNs). For example, a series of convolutional layers are used to perform convolution operations on the image, extracting local features such as edges, texture, and shape. After multiple layers of convolution operations, initial image features are obtained, which reflect the basic content and structure of the image.

[0055] Furthermore, the extracted initial image features are input into the linear layer. The linear layer is a fully connected layer that performs a linear transformation on the input features.

[0056] In the linear layer, each input image feature is multiplied by the image weight matrix and a bias term is added to produce the corresponding linear feature. Specifically, if the initial image feature is represented as a vector x, the image weight matrix is ​​represented as W, and the bias term is represented as b, then the linear feature y = Wx + b. In this way, the initial image features are further processed and transformed.

[0057] Furthermore, the linear features obtained through the linear transformation are input into the normalization layer. The purpose of the normalization layer is to normalize the input features so that they conform to a certain distribution law.

[0058] In the normalization layer of the embodiments of the present invention, linear features are typically processed using methods such as batch normalization, layer normalization, or instance normalization. For example, in batch normalization, the mean and variance of the linear features across the entire batch are calculated. The mean is then subtracted from each feature value and divided by the square root of the variance to ensure that the processed features conform to a standard normal distribution. After normalization, the target image features are obtained, which have better stability and convergence, helping to improve the training effect of the model.

[0059] The normalized target image features are input into the crisscross attention layer, which is used to calculate the correlation and importance between image features.

[0060] In the cross-attention layer of the embodiment of the present invention, attention weights are calculated between different elements in the target image feature. Specifically, the attention weights are calculated by comparing each element in the target image feature pairwise based on the correlation or similarity between them. For example, a dot product operation can be used to calculate the attention weight between two vectors. That is, the larger the dot product of the two vectors, the stronger the correlation between them, and the greater the attention weight.

[0061] Based on the calculated attention weights, a weighted sum is performed on all elements in the target image feature. Specifically, the feature vector of each element is multiplied by the corresponding attention weight, and all products are summed to obtain the final image cross-attention output. This image cross-attention output is a vector representation that summarizes the relationships between the elements in the target image feature.

[0062] The linear transformation processing provided by the embodiments of the present invention can further explore the potential information of image features and enhance the expressiveness of features by linearly transforming the initial image features. Linear transformation can change the feature space to better adapt to different task requirements.

[0063] The normalization processing provided by the embodiment of the present invention can make the input features conform to a certain distribution law, reduce the differences between features, thereby accelerating the training process of the model and making the model converge to the optimal solution more quickly.

[0064] The embodiments provided by the present invention calculate the attention weights between different elements in the target image features, thereby exploring the semantic associations and dependencies between regions in the image cue samples. For example, when processing an image containing a person and a scene, the relationship between the person and the scene, as well as the relationship between different parts of the person, can be found, which helps to better understand the semantics of the image.

[0065] In the embodiments provided by the present invention, the size of the attention weight reflects the importance of each element in the image. By weighted summing the cross-attention output of the image, elements containing important information can be more prominently displayed in the output, thereby helping the model better focus on key content.

[0066] Furthermore, in some embodiments of the present invention, a pre-trained CLIP image encoder model is used to extract image features from image prompt samples. During the training phase, the CLIP image encoder is fixed. In order to effectively decompose the global image embedding, the present invention uses a projection network consisting of a linear layer and a normalization layer to project the image embedding into a length N (in the present invention, ) feature sequence.

[0067] The dimension of the target image features is the same as the dimension of the text features in the pre-trained diffusion model. In the calculation of the attention mechanism, image features and text features of the same dimension can more conveniently perform similarity calculation and weight assignment.

[0068] In this paper, a new cross attention layer is added to each cross attention layer in the original UNet model to insert the target image features, separating the cross attention layers of text features and target image features. , the new image cross attention output The calculation of is as follows: in, is the image cross attention output, , , are the query matrix, key matrix, and value matrix from the target image features, 、 is the corresponding weight matrix.

[0069] The same query is used in image cross-attention and text cross-attention. To speed up the convergence, and Corresponding respectively and Initialization OK.

[0070] Furthermore, in some embodiments of the present invention, the image-based prompt adapter performs attention decoupling processing on the text cross-attention output and the image cross-attention output, and the formula for obtaining the decoupled cross-attention output is as follows: Where, Denoted as disentangled cross-attention output, Represented as text cross attention output, Denoted as image cross attention output, , , , , , since the original UNet model is frozen, only and It is trainable, that is, the attention weights of the target image features need to be trained.

[0071] Furthermore, in some embodiments of the present invention, the text cross-attention output and the image cross-attention output are input into an image prompt adapter.

[0072] Within the image cueing adapter, the text cross-attention output is first analyzed to determine which parts of the attention are strongly related to the image content and the distribution of this attention. For example, by observing the size of the attention weights and the corresponding text feature regions, it can be determined which words or phrases in the text are based on specific objects or scenes in the image. A similar analysis is also performed on the image cross-attention output to identify parts that are closely related to the text content and their attention distribution characteristics.

[0073] Based on the results of the analysis, the attention part of the text cross-attention output that is directly related to the image content is separated from the attention part that is related to the internal logic of the text itself. For example, if the text describes the color and position of an object, where the color information is related to the color representation in the image, and the position information may rely more on the logical order of the text, these two parts of attention need to be decoupled. For the image cross-attention output, the internal attention related to the image features themselves and the attention affected by the text are separated. For example, the internal attention corresponding to the texture features of a region in the image is distinguished from the attention generated by the attention to that region guided by the text description.

[0074] After completing the above decoupling operation, the separated attention components of the text and image are recombined according to certain rules to generate a decoupled cross-attention output. This output can more clearly represent the attention relationship between text and image, as well as within each, providing a foundation for more accurate understanding and generation of multimodal content.

[0075] In some embodiments of the present invention, inputting the image prompt information and the text prompt information into a trained image generation model for processing to generate a target image includes: Inputting the image prompt information and the text prompt information into a trained image generation model for processing to generate an initial image; When the initial image is denoised, the initial image is optimized based on the decoupled cross-attention output to generate the target image.

[0076] First of all, it should be noted that the image generation model includes a U-Net network and an embedded image prompt adapter. The U-Net network includes a text encoder and an image encoder, and the embedded image prompt adapter includes a cross-attention layer, a linear layer, and a normalization layer.

[0077] Therefore, it is usually only necessary to train the embedded image prompt adapter to obtain a trained image generation model. The image generation model is a model based on the U-Net network and the embedded image prompt adapter to generate images based on text.

[0078] As you can understand, first, image hints (e.g., image style, subject outline, color distribution, etc.) and text hints (e.g., descriptions of objects, scenes, actions, etc.) containing information about the desired image are fed into a trained image generation model. These hints provide the model with the basic basis and constraints for generating images.

[0079] For example, if the text prompt information is "a brown deer running in the forest", the image prompt information may be about the background template of the forest, the style features of previous deer-like pictures, etc. The model will combine this information to preliminarily understand the content of the image to be generated.

[0080] Based on the input prompt information, the trained image generation model begins generating images through its internal generation mechanism (such as a deep learning-based neural network structure, including multiple convolutional layers in the generator part, activation functions, and embedded image prompt adapter components). Based on the learned patterns and regularities, the model converts the text and image prompts into specific image pixel information, creating a preliminary image, the initial image. This initial image may already contain some basic elements and features, but may still contain noise or lack clarity and accuracy.

[0081] For example, the initial image may show the general shape of a fawn and the general background of a forest, but the details of the fawn may be blurred, or the color, light and shadow of the forest may not be realistic enough.

[0082] Furthermore, the initial image may contain noise due to various factors during the model generation process. Denoising involves using certain technical methods to reduce noise interference in the image and improve image quality. Common denoising methods include filter-based methods (such as Gaussian filtering and median filtering) and neural network-based denoising methods (such as using a neural network module with denoising capabilities).

[0083] For example, for some of the mottled spots (noise) on the fawn in the initial image, Gaussian filtering can smooth these spots, making the fawn's fur look cleaner and neater while retaining the fawn's basic morphological features.

[0084] Decoupling the cross-attention outputs clarifies the attention relationships between text and images, as well as within each. Using this information to optimize the initial image means the model can adjust different parts of the image more precisely.

[0085] For example, based on the decoupled cross-attention output, if the attention weights corresponding to the text prompt "brown fawn" are unevenly distributed across the fawn in the initial image, the optimization process can strengthen the connection between the brown areas and the corresponding parts of the fawn's body, making the color more accurate and rich. At the same time, if the position or details of the forest background elements in the image prompt (such as trees and grass) are not accurate in the initial image, the decoupled attention information can better adjust these background elements to match the scene described in the text.

[0086] After denoising and optimization based on decoupled cross-attention outputs, the initial image is significantly improved in both quality and alignment with the prompt, ultimately resulting in a target image that meets the requirements. This target image more accurately reflects the text prompt description and offers better visual quality, with more coherent and unified image elements.

[0087] In summary, by processing and optimizing the image generation model, we can gradually obtain high-quality target images from the initial rough images. The whole process not only reflects the advanced nature of image generation technology, but also demonstrates the ability to finely control details in complex tasks.

[0088] Figure 2 This is the second schematic diagram of the image generation method of the embedded image prompt adapter provided by the present invention. First, the text prompt sample is input into the text encoder to obtain text features, and then the text features are input into the cross-attention layer to determine the text cross-attention output.

[0089] The image prompt sample is input into the image encoder for encoding to obtain the initial image features. The initial image features are then input into the linear layer and the normalization layer for processing to obtain the target image features. The target image features are then input into the cross-attention layer to determine the image cross-attention output.

[0090] Based on the image prompt adapter, the text cross-attention output and the image cross-attention output are subjected to attention decoupling processing, that is, the corresponding image cross-attention output and text cross-attention output are input into each layer of the denoising U-Net network respectively, thereby obtaining the decoupled cross-attention output ( Figure 2 The orange and blue arrows in the middle correspond to the two cross-attention outputs in each channel, respectively).

[0091] The denoising U-Net network is trained using the decoupled cross-attention output to obtain a trained image generation model. The trained image generation model can be used to generate a target image based on the input image prompt information and text prompt information.

[0092] In some embodiments of the present invention, further comprising: During training, only the optimization Figure 2 The orange part in , keeps the parameters of the pre-trained diffusion model unchanged. Use the same training objective as the original diffusion model: Where, Represented as text features, Represented as target image features, represents a simplified variant of the variational bound, represents the real image, represents random noise, represents the noise predicted by the model, represents the time step, ò is the noised data at step t of the forward process.

[0093] Randomly drop image conditions during training to enable classifier-free bootstrapping during inference: Where, Represented as text features, Represented as target image features, represents a simplified variant of the variational bound, represents the real image, represents random noise, represents the noise predicted by the model, represents the time step, ò is the noised data of the t-th step in the forward process, w Indicates the weight value.

[0094] Here, if the image condition is discarded, the CLIP image embedding is simply set to zero.

[0095] Since the text criss-cross attention and image criss-cross attention are separate, the weight of the image condition can also be adjusted during the inference phase: in, Denoted as disentangled cross-attention output, , , are the query matrix, key matrix, and value matrix of the attention operation, respectively. , , are the query matrix, key matrix, and value matrix from the target image features, is the weight factor, when , the model becomes the original text-to-image diffusion model.

[0096] The image prompt adapter is a tool specifically designed to process image prompt information in image generation models. It acts as a bridge between the model and the image prompt information, and can convert various forms of image prompts, such as static images, dynamic image sequences, image feature vectors, etc., into a form that the model can understand and use. Figure 2 The orange portion in the middle performs linear and normalization layers on the initial image features output by the image encoder, and decouples the cross-attention outputs of the text and image. The image generation model in this embodiment of the present invention is a model that embeds the aforementioned image prompt adapter.

[0097] The embodiments of the present invention provide an efficient and lightweight image prompting adapter for implementing image prompting for a pre-trained text-to-image diffusion model, which has at least the following technical effects: This paper introduces a decoupled crisscross attention mechanism to achieve a more effective image cue adapter. The proposed image cue adapter remains simple and compact, but outperforms previous adapter methods and is even comparable to fine-tuned models.

[0098] The proposed image cue adapter can achieve comparable or even better performance than a fully fine-tuned image cue model. It can be generalized to other custom models fine-tuned from the same base model and can be used with existing controllable tools for controllable generation. Thanks to the decoupled cross-attention strategy, image cues can also be well combined with text cues to achieve multimodal image generation. The image generation model designed in this paper only requires training the newly added adapter module (shown in orange in the figure), while most of the parameters of the main model remain frozen. This significantly reduces the number of parameters that need to be updated, thereby reducing storage and computational costs.

[0099] The present invention provides an image generation method, device, equipment and storage medium embedded in an image prompt adapter. The method obtains image prompt information and text prompt information; inputs the image prompt information and the text prompt information into a trained image generation model for processing to generate a target image; wherein the target image is characterized as an image that conforms to the text description in the text prompt information; the image generation model is obtained by training a neural network using image prompt samples, text prompt samples and decoupled cross-attention outputs; and the decoupled cross-attention outputs are obtained by performing attention decoupling processing on the image prompt samples and the text prompt samples based on the image prompt adapter. The method is used to solve the technical defects of poor image quality and high computational complexity in the prior art, realize the embedding of an image prompt adapter, and generate a target image with higher image quality through the decoupled cross-attention output determined by the image prompt adapter, while significantly reducing computational complexity.

[0100] The image generating device of the embedded image prompt adapter provided by the present invention is described below. The image generating device of the embedded image prompt adapter described below and the image generating method of the embedded image prompt adapter described above can be referred to each other.

[0101] like Figure 3 FIG. 1 is a schematic diagram of the structure of an image generating device embedded in an image prompt adapter provided by the present invention. The image generating device embedded in an image prompt adapter includes the following modules: An information acquisition module 310 is used to acquire image prompt information and text prompt information; A target image generation module 320 is used to input the image prompt information and the text prompt information into a trained image generation model for processing to generate a target image; wherein, the target image is represented as an image that conforms to the text description in the text prompt information; the image generation model is obtained by training a neural network using image prompt samples, text prompt samples and decoupled cross-attention outputs; the decoupled cross-attention outputs are obtained by performing attention decoupling processing on the image prompt samples and the text prompt samples based on an image prompt adapter.

[0102] Preferably, the image generation device embedded in the image prompt adapter provided by the present invention is specifically used to perform computational processing on the text prompt sample to obtain a text cross-attention output corresponding to the text prompt sample; Performing computational processing on the image prompt sample to obtain an image cross-attention output corresponding to the image prompt sample; Based on the image prompt adapter, attention decoupling processing is performed on the text cross-attention output and the image cross-attention output to obtain the decoupled cross-attention output.

[0103] Preferably, the image generation device embedded in the image prompt adapter provided by the present invention is specifically used in that the image generation model includes at least a text encoder and a cross attention layer; Using the text encoder to perform text encoding processing on the text prompt sample to obtain text features; The text features are input into the cross attention layer for calculation and processing to obtain the text cross attention output.

[0104] Preferably, the image generation device embedded in the image prompt adapter provided by the present invention is specifically used in that the image generation model comprises at least an image encoder, a linear layer, and a normalization layer; Performing image encoding processing on the image prompt sample by using the image encoder to obtain initial image features; Inputting the initial image features into the linear layer for linear transformation processing to obtain corresponding linear features; wherein the linear features are determined by multiplying each initial image feature input into the linear layer by an image weight matrix and adding a bias term; Inputting the linear features into the normalization layer for normalization processing to obtain target image features; The target image features are input into the cross attention layer for calculation and processing to obtain the image cross attention output.

[0105] Preferably, the image generation device embedded with the image prompt adapter provided by the present invention is specifically used to perform attention decoupling processing on the text cross-attention output and the image cross-attention output based on the image prompt adapter, and the formula for obtaining the decoupled cross-attention output is as follows: Where, Denoted as disentangled cross-attention output, Represented as text cross attention output, Represented as image crisscross attention output.

[0106] Preferably, the image generation device embedded in the image prompt adapter provided by the present invention is specifically used to input the image prompt information and the text prompt information into a trained image generation model for processing to generate an initial image; When the initial image is denoised, the initial image is optimized based on the decoupled cross-attention output to generate the target image.

[0107] The present invention provides an image generation method, device, equipment and storage medium embedded in an image prompt adapter. The method obtains image prompt information and text prompt information; inputs the image prompt information and the text prompt information into a trained image generation model for processing to generate a target image; wherein the target image is characterized as an image that conforms to the text description in the text prompt information; the image generation model is obtained by training a neural network using image prompt samples, text prompt samples and decoupled cross-attention outputs; and the decoupled cross-attention outputs are obtained by performing attention decoupling processing on the image prompt samples and the text prompt samples based on the image prompt adapter. The method is used to solve the technical defects of poor image quality and high computational complexity in the prior art, realize the embedding of an image prompt adapter, and generate a target image with higher image quality through the decoupled cross-attention output determined by the image prompt adapter, while significantly reducing computational complexity.

[0108] Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communications bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other via the communications bus 440. The processor 410 may call logic instructions in the memory 430 to execute an image generation method embedded in an image prompt adapter, the method comprising: obtaining image prompt information and text prompt information; inputting the image prompt information and the text prompt information into a trained image generation model for processing to generate a target image; wherein the target image is represented as an image that conforms to the text description in the text prompt information; the image generation model is obtained by training a neural network using image prompt samples, text prompt samples, and decoupled cross-attention outputs; and the decoupled cross-attention outputs are obtained by performing attention decoupling processing on the image prompt samples and the text prompt samples based on the image prompt adapter.

[0109] Furthermore, the logic instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0110] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the image generation method embedded in the image prompt adapter provided by the above methods, the method including: obtaining image prompt information and text prompt information; inputting the image prompt information and the text prompt information into a trained image generation model for processing to generate a target image; wherein, the target image is represented as an image that conforms to the text description in the text prompt information; the image generation model is obtained by training a neural network using image prompt samples, text prompt samples and decoupled cross-attention output; the decoupled cross-attention output is obtained by performing attention decoupling processing on the image prompt samples and the text prompt samples based on the image prompt adapter.

[0111] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the image generation method of the embedded image prompt adapter provided by the above-mentioned methods, the method comprising: obtaining image prompt information and text prompt information; inputting the image prompt information and the text prompt information into a trained image generation model for processing to generate a target image; wherein, the target image is represented as an image that conforms to the text description in the text prompt information; the image generation model is obtained by training a neural network using image prompt samples, text prompt samples and decoupled cross-attention outputs; the decoupled cross-attention outputs are obtained by performing attention decoupling processing on the image prompt samples and the text prompt samples based on the image prompt adapter.

[0112] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0113] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An image generation method embedded in an image prompt adapter, characterized in that: include: Get image prompt information and text prompt information; Inputting the image prompt information and the text prompt information into a trained image generation model for processing to generate a target image; Among them, the target image is represented as an image that conforms to the text description in the text prompt information; the image generation model is obtained by training a neural network using image prompt samples, text prompt samples and decoupled cross-attention outputs; the decoupled cross-attention output is obtained by performing attention decoupling processing on the image prompt samples and the text prompt samples based on an image prompt adapter.

2. The image generation method of the embedded image prompt adapter according to claim 1, characterized in that: The decoupled cross-attention output is obtained by performing attention decoupling processing on the image prompt sample and the text prompt sample based on the image prompt adapter, including: Performing computational processing on the text prompt sample to obtain a text cross-attention output corresponding to the text prompt sample; Performing computational processing on the image prompt sample to obtain an image cross-attention output corresponding to the image prompt sample; Based on the image prompt adapter, attention decoupling processing is performed on the text cross-attention output and the image cross-attention output to obtain the decoupled cross-attention output.

3. The image generation method of the embedded image prompt adapter according to claim 2, characterized in that: The image generation model includes at least a text encoder and a cross attention layer; The calculating and processing the text prompt sample to obtain a text cross-attention output corresponding to the text prompt sample includes: Using the text encoder to perform text encoding processing on the text prompt sample to obtain text features; The text features are input into the cross attention layer for calculation and processing to obtain the text cross attention output.

4. The image generation method of the embedded image prompt adapter according to claim 3, characterized in that: The image generation model includes at least an image encoder, a linear layer, and a normalization layer; The performing computational processing on the image prompt sample to obtain an image cross attention output corresponding to the image prompt sample includes: Performing image encoding processing on the image prompt sample by using the image encoder to obtain initial image features; Inputting the initial image features into the linear layer for linear transformation processing to obtain corresponding linear features; wherein the linear features are determined by multiplying each initial image feature input into the linear layer by an image weight matrix and adding a bias term; Inputting the linear features into the normalization layer for normalization processing to obtain target image features; The target image features are input into the cross attention layer for calculation and processing to obtain the image cross attention output.

5. The image generation method of the embedded image prompt adapter according to claim 4, characterized in that: The image prompt adapter performs attention decoupling processing on the text cross-attention output and the image cross-attention output, and the formula for obtaining the decoupled cross-attention output is as follows: Where, Denoted as disentangled cross-attention output, Represented as text cross attention output, Represented as image crisscross attention output.

6. The image generation method of the embedded image prompt adapter according to any one of claims 1 to 5, characterized in that: The step of inputting the image prompt information and the text prompt information into a trained image generation model for processing to generate a target image includes: Inputting the image prompt information and the text prompt information into a trained image generation model for processing to generate an initial image; When the initial image is denoised, the initial image is optimized based on the decoupled cross-attention output to generate the target image.

7. An image generating device embedded in an image prompting adapter, characterized in that: include: An information acquisition module is used to acquire image prompt information and text prompt information; A target image generation module is used to input the image prompt information and the text prompt information into a trained image generation model for processing to generate a target image; wherein, the target image is represented as an image that conforms to the text description in the text prompt information; the image generation model is obtained by training a neural network using image prompt samples, text prompt samples and decoupled cross-attention outputs; the decoupled cross-attention outputs are obtained by performing attention decoupling processing on the image prompt samples and the text prompt samples based on an image prompt adapter.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the image generation method of the embedded image prompt adapter according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image generation method of the embedded image prompt adapter according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the image generation method of the embedded image prompt adapter according to any one of claims 1 to 6 is implemented.