An interactive and convenient multifunctional image generation method
By receiving image generation control conditions input by the user, adaptive positioning and feature fusion are performed using cross-attention maps and visual control adapters, the problems of inconvenient interaction and poor generation quality in the prior art are solved, and high-quality and high-degree of freedom control of multifunctional image generation are realized.
Patent Information
- Application Number
- CN202510045748.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-13
AI Technical Summary
The existing image generation technology has problems such as inconvenient interaction, poor image quality and single function. Especially in the ControlNet or T2I-Adapter-based method, users need to provide scene-level visual conditions, which is difficult to operate, and the model is difficult to balance the impact of different modes on image generation, ignoring text semantic information, resulting in poor quality of generated images.
By receiving the image input from the user, the control conditions, including text prompts, entity condition diagrams and background diagrams, the cross attention map in the generative model realizes adaptive positioning of the local control area, and performs multi-level feature-level fusion, combining the visual control adapter and attention map optimization function to achieve image generation.
It realizes higher freedom of control selection, improves the quality and accuracy of image generation, solves the one-to-one correspondence between the input of entity conditional graphs and the generated image position, improves user experience, takes into account text semantic information and entity morphology control, and generates images with multi-level control that meets multiple conditions.
Smart Images

Figure CN119444912B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image generation, and more particularly to a multifunctional image generation method with convenient interaction. Background Art
[0002] With the continuous development of deep learning technology, the research of generative artificial intelligence (AIGC) has received close attention. In the field of AIGC, text-to-image synthesis technology has become a research hotspot. Stable Diffusion is an image generation model based on the diffusion process. The emergence of Stable Diffusion allows people to synthesize a high-definition image related to text semantics by inputting text prompts. This technology is currently widely used in art creation, advertising design, etc. In response to the growing user demand, people gradually hope to get the image generation effect they want with more different forms of input and more flexible and convenient methods.
[0003] Existing diffusion model-based research, such as Imegen, Stable Diffusion, and DALL-E, has revolutionized the task of text-to-image generation. These models are good at generating high-quality images, but often lack finer control over the generated content. To address this limitation, researchers have explored various methods to enhance the user's control over generative AI. One solution is to use text-driven control methods, such as adjusting the text prompt or adjusting the effect of cross-attention maps on the final generated image.
[0004] Another solution to controllable image generation is to combine additional input modality information, such as sketches or layout information. Layout Guidance Diffusion uses user-defined markers and bounding boxes to guide the distribution of cross-attention scores in specified areas, thereby indirectly controlling the entity generation location. The Vincent graph model eDiff-I imposes structural constraints on the generated image by calculating the similarity gradient between the target sketch and the intermediate model features. The Vincent graph algorithm ControlNet guides image generation through specific conditions by adding an additional conditional encoding network to the frozen pre-trained generation model. T2I-Adapter introduces a lightweight adapter that combines the internal knowledge of the text-to-image model with external control signals.
[0005] At present, some inventions apply ControlNet to various practical fields to perform more sophisticated and flexible image generation control. The patent application with application number CN202311827908.5 proposes a control method and system for image generation. The method converts the edge line map into a close-to-real image based on ControlNet. This method solves the problems of the lack of flexibility, accuracy and versatility of the traditional GANs model and the difficulty in satisfying user customization. The patent application with application number CN202311646175.5 proposes a posture generation method based on a diffusion model and ControlNet. The method can generate a human image that conforms to a certain posture by inputting a human posture node. The patent application with application number CN202311183764.4 proposes an image processing method and device based on generative artificial intelligence technology. The invention technically uses SAM to segment people or objects in the image, and then redraws the image by image restoration to achieve the effect of modifying the background or content in a specific position.
[0006] However, the above methods still have some problems to be solved, which seriously affect the quality of image generation and user experience:
[0007] (1) Inconvenient interaction. Existing research has not paid enough attention to this issue, especially in methods based on ControlNet or T2I-Adapter, where users need to provide scene-level visual conditions, which is difficult for users to operate. In complex multi-entity scenes, the condition information provided by users not only needs to be highly accurate in position, but also requires one-to-one correspondence between different modal conditions. This method is not only time-consuming, but also difficult to ensure the generation of ideal results.
[0008] (2) Poor quality of generated images. When the image generation model involves multimodal input, due to the limitations of the model network structure or deficiencies in the training process, the model often finds it difficult to achieve a balance between the impact of different modalities on image generation. In practical applications, users usually expect the image modality to supplement the text information, while the main semantic information of the image is still dominated by the text. However, the ControlNet or T2I-Adapter methods often focus too much on the layout information of the image modality and ignore the semantic information of the text, resulting in poor quality of generated images.
[0009] (3) Single function. In most current image generation methods based on visual control, the types of visual conditions are relatively single, and the trained model can usually only process a certain type of visual input, which greatly limits the generation ability and flexibility of the model.
[0010] Therefore, how to solve the problems existing in the above-mentioned existing image generation technologies is the key to improving image generation quality and user experience. Summary of the invention
[0011] In view of the above problems, the present invention provides a multifunctional image generation method with convenient interaction to at least solve some of the technical problems mentioned in the above background technology.
[0012] In order to achieve the above object, the present invention adopts the following technical solution:
[0013] The embodiment of the present invention provides a multifunctional image generation method with convenient interaction, comprising the following steps:
[0014] S1, receiving image generation control conditions input by a user, and preprocessing the image generation control conditions; the image generation control conditions include: text prompts, entity condition graphs and background graphs;
[0015] S2, using random noise that conforms to the standard normal distribution to form an initial noise image; in the early denoising stage, based on the preprocessed text prompt, the initial noise image is globally guided to denoise through the generative model to obtain a noise image;
[0016] S3, using the cross-attention map in the generative model to achieve adaptive positioning of the local control area;
[0017] S4, according to the local control area after positioning, the pre-processed entity condition map and background map are fused at a multi-level feature level to obtain a multi-modal coding feature of a unified space;
[0018] S5. In the later denoising stage, the multimodal coding features are used to obtain visual control features through a visual control adapter; the visual control features and the global intermediate layer features in the generative model are used to jointly guide the generative model to denoise the noisy image to achieve image generation.
[0019] Furthermore, in the step S1, the input of the entity condition diagram supports multiple modes.
[0020] Furthermore, in the step S1, the image generation control condition is preprocessed, specifically including:
[0021] (1) Use the Transformer-based text encoder to perform feature mapping on the text prompt p to obtain the text embedding feature sequence , where k is the total length of the text embedding feature sequence;
[0022] (2) For each entity condition graph Extract the largest polygon that can surround the area where the entity is located as the entity contour map of the corresponding entity condition map ;
[0023] (3) Background image Perform edge detection and binarization to obtain a binary background image .
[0024] Furthermore, in the step S2, based on the preprocessed text prompt, the initial noisy image is globally guided to be denoised by generating a model, specifically:
[0025] Taking the text embedding feature sequence C as the image generation condition, the initial noisy image is globally guided for denoising through the generative model.
[0026] Furthermore, the step S3 specifically includes:
[0027] (1) The cross attention map is expressed as:
[0028]
[0029] in, represents the i-th entity condition graph; express The corresponding text tag; represents the cross attention map, and Only in Time calculation, where The time is when the previous denoising stage is completed; n represents the nth head index in the multi-head attention mechanism; N represents a total of N head indexes; t represents the time step in the diffusion process; Express The latent variable of generates the query vector in cross attention using a linear function; Represents the text tag feature Use a linear function and generate the key vector in the cross attention after transposition operation; d represents the feature embedding dimension;
[0030] (2) Obtain the threshold of the segmentation attention score through the OTSU algorithm; binarize the cross attention map according to the threshold to obtain the entity area mask image of the area where the entity is located in the generated image; the entity area mask image R i It is expressed as:
[0031]
[0032] Among them, R i Represents a solid area mask image; Represents the entity region mask image R i The median coordinate is The pixel value of Represents the cross attention map The median coordinate is Pixel value; OTSU(·) represents the OTSU algorithm;
[0033] (3) Mask image R for each entity region i The localization area of the corresponding entity is approximated to obtain the entity localization area bounding box of the generated image. ,in and They are the upper left horizontal and vertical coordinates of the entity positioning area bounding box respectively; and They are the lower right horizontal and vertical coordinates of the entity positioning area bounding box.
[0034] Furthermore, when the number of entity condition graphs input by the user is multiple, the entity positioning area bounding box determined by the cross attention graph needs to be Make adjustments; specifically include:
[0035] (1) One of the multiple entity condition maps to be located and scaled is used as the target entity condition map, and the location area of the entity in the generated image in the target entity condition map is recorded as the target location area;
[0036] (2) The location area of the entity in the remaining entity condition graph in the generated image is recorded as the remaining location area;
[0037] (3) Calculate the intersection and union ratio of the target positioning area and the remaining positioning areas;
[0038] (4) If the intersection-to-union ratio is greater than a preset value, the target positioning area is reduced in size and moved away from the central intersection of all positioning areas to adjust the target positioning area.
[0039] Furthermore, the step S4 specifically includes:
[0040] (1) Locate the bounding box of the region based on the entity , respectively for the corresponding entity condition graph and solid outlines Perform positioning and scaling processing to obtain the processed local entity condition map and entity location contours ;
[0041] (2) Use an image encoder composed of zero convolutions to encode the local entity condition graph And the binary background image Encode to obtain multimodal encoding features of unified space .
[0042] Furthermore, the step S5 specifically includes:
[0043] (1) obtaining a visual control feature by applying the multimodal encoding feature to a visual control adapter;
[0044] (2) In the later denoising stage, the visual control features are weighted summed with the global intermediate layer features in the generative model conditioned on the text embedding feature sequence C; the weighted sum result is used as the updated intermediate layer features of the generative model;
[0045] (3) Guide the generative model through the updated intermediate layer features of the generative model to denoise the noisy image to achieve image generation.
[0046] Furthermore, in the later denoising stage, an attention map optimization function is obtained based on the cross-attention map and the noisy image is updated; the attention map optimization function is composed of a text tag level function and a pixel level function, which is specifically expressed as:
[0047]
[0048] in, Represents text token level function; represents the pixel-level function; l represents the lth index in the network layer of the generative model; L represents the total number of network layers of the generative model; m represents the total number of entities that need to be locally controlled; i represents the index of the entity that needs to be locally controlled; It represents the entity positioning contour image after the entity contour image is processed by positioning and scaling; express The pixel value with coordinates (x, y); express The pixel value with coordinates (x, y); Represents the local entity condition graph after positioning and scaling processing; Represents the cross attention map corresponding to the i-th entity index of the l-th network layer at the denoising time point t; and are the height and width of the cross attention map of the lth network layer respectively; BCE(.) represents the binary cross entropy operation; represents the attention map optimization function; Represents the hyperparameter that controls the ratio of text tag level function and pixel level function, with a value range of [0,1];
[0049] The gradient value is calculated by the attention graph optimization function, and the noise image is updated according to the gradient value to obtain a new noise image; it is expressed as:
[0050]
[0051] in, represents a noisy image; represents the gradient operation; Represents the updated noise image The learning rate; represents the new noise image updated by the attention map optimization function;
[0052] By performing a finite number of cyclic optimization updates on the noisy image, based on the visual control features, the new noisy image is subjected to post-denoising processing time-step by time-step through the generative model to obtain the generated image.
[0053] Furthermore, in the later denoising stage, the visual control features and the intermediate layer features of the generative model are weighted and calculated by introducing adaptive weights to serve as updated intermediate layer features of the generative model; expressed as:
[0054]
[0055] in, represents the adaptive weight function that changes with denoising time; represents the noise image at time t; represents the multimodal coding feature; a, b and c all represent the control adaptive weight function Hyperparameters of the direction; Represents the updated intermediate layer features of the generative model; Represents the UNet network layer; Represents the basic generative model Weight; represents the drawing encoder weights; represents zero convolution in the visual control adapter, and the corresponding weight is .
[0056] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a multifunctional image generation method with convenient interaction, which has the following beneficial effects:
[0057] The present invention supports text prompts, entity condition diagrams of multiple different modalities, and background images as input. Based on this, according to user needs, image generation can be freely controlled by pure text, single-visual entity local control, multi-visual entity local control, or multi-visual entity background mixed control, which reflects that the present invention has a higher degree of freedom in control selection and is more convenient to interact with users.
[0058] The present invention introduces a local entity condition graph as one of the interactive conditions for the existing image generation methods. The present invention can intelligently control the text-generated graph process according to the entity condition graph provided by the user, comply with the text semantic information in semantic control, and realize the trade-off between the spatial form and semantics specified by the entity condition graph in entity form. Compared with previous methods, this method takes into account the semantic information of the text to a great extent, controls the entity form without affecting the semantics, and does not require artificially specifying the entity position information or providing global scene-based image conditions. Technically, it solves the problem that the entity condition graph input needs to correspond one-to-one with the generated image position, realizes the real entity condition graph-assisted text-generated graph effect, greatly improves the image generation quality and image generation accuracy, and enhances the user experience.
[0059] In the early denoising stage, the present invention uses text prompts as global control conditions, uses a generative model to perform global guided denoising on the initial noisy image, and uses a cross-attention map to achieve adaptive positioning of the local control area, so as to perform multi-level feature fusion of multimodal visual control conditions in this positioning area to obtain multimodal coding features; in the later denoising stage, the multimodal coding features obtain visual control features through a visual control adapter and use them to control the image generation process; the denoising process is optimized by using an attention map optimization function and an adaptive enhancement strategy, and finally a multi-level control generated image that meets multiple conditions is obtained.
[0060] Other features and advantages of the present invention will be set forth in the description which follows, and in part will be apparent from the description, or may be learned by practice of the present invention.
[0061] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0063] Figure 1 A schematic flow chart of a multifunctional image generation method with convenient interaction provided by an embodiment of the present invention.
[0064] Figure 2 A schematic diagram of the plain text control generation mode flow chart provided in an embodiment of the present invention.
[0065] Figure 3A schematic diagram of the flow of a single visual entity local control generation mode provided in an embodiment of the present invention (taking the sketch mode as an example).
[0066] Figure 4 A schematic diagram of the flow chart of the multi-visual entity local control generation mode provided in an embodiment of the present invention (taking the sketch mode as an example).
[0067] Figure 5 A schematic diagram comparing the effects before and after processing the highly overlapping phenomenon of multi-entity positioning areas provided by an embodiment of the present invention.
[0068] Figure 6 A schematic flow chart of a multi-visual entity background hybrid control generation mode provided for an embodiment of the present invention (taking sketch mode as an example).
[0069] Figure 7 The multi-entity local control generation mode provided for the embodiment of the present invention is applied to conditional image inputs of various different modalities (taking segmentation maps, posture maps, depth maps and edge maps as examples).
[0070] Figure 8 A schematic diagram of multiple sample visualization effects of a single visual entity local control generation mode (taking the sketch mode as an example) provided in an embodiment of the present invention.
[0071] Fig. 9 A schematic diagram of multiple sample visualization effects of a multi-visual entity local control generation mode (taking the sketch mode as an example) provided in an embodiment of the present invention.
[0072] Fig.10 A schematic diagram of multiple sample visualization effects of a multi-visual entity background hybrid control generation mode (taking the sketch mode as an example) provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0073] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0074] See also Figure 1 As shown, the embodiment of the present invention discloses a multifunctional image generation method with convenient interaction, including:
[0075] S1, receiving image generation control conditions input by a user, and preprocessing the image generation control conditions; the image generation control conditions include: text prompts, entity condition graphs and background graphs;
[0076] S2, using random noise that conforms to the standard normal distribution to form an initial noise image; in the early denoising stage, based on the preprocessed text prompt, the initial noise image is globally guided to denoise through the generative model to obtain a noise image;
[0077] S3, using the cross-attention map in the generative model to achieve adaptive positioning of the local control area;
[0078] S4, according to the local control area after positioning, the pre-processed entity condition map and background map are fused at a multi-level feature level to obtain a multi-modal coding feature of a unified space;
[0079] S5. In the later denoising stage, the multimodal encoding features are used to obtain visual control features through a visual control adapter. The visual control features and the global intermediate layer features in the generative model are used to jointly guide the generative model to denoise the noisy image and achieve image generation.
[0080] Among them, the entity condition map and the background map are optional inputs; and the entity condition map can be represented in a multimodal form, which specifically includes a hand-drawn sketch, a posture map, a segmentation map, a depth map, etc.;
[0081] In practical applications, the above steps can be reduced to a certain extent according to the image generation control conditions input by the user, and the required control generation mode can be selected to meet the user's control requirements. Based on the input control information, it can be roughly divided into the following four control generation modes: 1. Pure text control generation mode; 2. Single visual entity local control generation mode; 3. Multi-visual entity local control generation mode; 4. Multi-visual entity background mixed control generation mode. Next, the above four image generation control modes are described in detail.
[0082] 1. Plain text control generation mode:
[0083] See also Figure 2 As shown, the image control information input by the user for controlling the semantics of generated images includes only a text prompt; the text prompt is "A dog is on the beach", which is used to describe a dog on the beach; based on the image control information, the pure text control generation mode is used to generate the image:
[0084] (1) Use the Transformer-based text encoder to perform feature mapping on the text prompt p to obtain the text embedding feature sequence , where k is the total length of the text embedding feature sequence;
[0085] (2) Use a standard normal distribution The random noise of the initial noise image ; The text embedding feature sequence C is used as the image generation condition. Within the preset denoising time T, the initial noise image is generated by the generative model. Perform denoising step by step to obtain the generated image .
[0086] For example, the total denoising time is set to T=50, and the initial noisy image is denoised within T=50 denoising time steps. Perform cyclic denoising; after the first denoising time step is completed, a noisy image is obtained , and repeat the denoising cycle 50 times to obtain the final generated image .
[0087] 2. Single visual entity local control generation mode (taking sketch mode as an example):
[0088] See also Figure 3 As shown in the figure, the image control information input by the user for controlling the semantics of the generated image includes a text prompt and a single entity condition image; the text prompt is "A car and an airplane move in front of the Mount Fuji", which is used to describe a car and an airplane moving in front of Mount Fuji; the single entity condition image is a car, and the entity condition image input here is a single-channel binary image with a resolution of 768×768; based on the image control information, the single visual entity local control generation mode is used to generate the image:
[0089] (1) Preprocessing of text prompts and single entity condition graphs:
[0090] ① Use the Transformer-based text encoder to perform feature mapping on the text prompt p to obtain the text embedding feature sequence , where k is the total length of the text embedding feature sequence;
[0091] ② For entity condition diagram Extract the largest polygon that can surround the area where the entity is located as the entity contour map corresponding to the entity condition map ;
[0092] (2) Use a standard normal distribution The random noise of the initial noise image ; Use text encoding features as image generation conditions, set total denoising time T, early denoising time point , in the preset early denoising stage Inside, the initial noise image is generated by the model Perform preliminary denoising at each time step to obtain a noisy image ;
[0093] For example, set the total denoising time T=50, the early denoising time point , then in The initial noise image in the time period Perform denoising step by step to obtain a noisy image ;
[0094] (3) Adaptive positioning of local control areas using the cross-attention map in the generative model; specifically:
[0095] ① Get the cross-attention map corresponding to the entity condition map.
[0096] ② Obtain the threshold of the segmentation attention score through the OTSU algorithm; binarize the cross attention map according to the threshold to obtain the entity area mask image R of the area where the entity is located in the generated image i .
[0097] ③ Mask image R for each entity region i The localization area of the corresponding entity is approximated to obtain the entity localization area bounding box of the generated image. ,in They are the upper left and lower right horizontal and vertical coordinates of the entity positioning area bounding box respectively.
[0098] The generative model realizes the information interaction between text and generated images through cross-attention. The calculated cross-attention value distribution can, to a certain extent, reflect the distribution of image semantic information under text control. Recent studies have shown that the structure of the image has been determined in the early steps of the diffusion process, so the visualization of the attention map between text and image can reflect the location of the entities in the final generated image. Through the adaptive positioning method of the local control area based on the cross-attention map, the relevant area of the entity in the final generated image can be obtained after the early denoising stage, and local control can be performed based on this positioning area in the subsequent process to better follow the text semantic control distribution.
[0099] (4) Perform multi-level feature-level fusion on the entity condition graph according to the local control area, and obtain multi-modal encoding features in a unified space; specifically,
[0100] ① Locate the bounding box of the region based on the obtained entity , respectively for the corresponding entity condition graph and solid outlines Perform positioning and scaling processing to obtain the processed local entity condition map and entity location contours ;
[0101] ② Use an image encoder composed of zero convolutions to encode the local entity condition graph Encode to obtain multimodal encoding features of unified space .
[0102] (5) In the later denoising stage, the multimodal encoding features The visual control features are obtained through the visual control adapter, and together with the global intermediate layer features in the generative model, the generative model is guided to perform accurate image generation, that is, the generative model is guided to perform post-denoising processing on the above-mentioned noisy image to achieve image generation;
[0103] exist Figure 3 In the above single-vision entity local control generation mode, an image of a car and an airplane in front of Mount Fuji is finally generated, where the shape of the "car" is controlled by the entity condition diagram input by the user, and the "airplane" entity is generated by text control.
[0104] 3. Multi-visual entity local control generation mode (taking sketch mode as an example):
[0105] See also Figure 4 As shown in the figure, the image control information input by the user for controlling the semantics of the generated image includes a text prompt and multiple entity condition pictures; the text prompt is "A bear and deer are in the forest", which is used to describe a bear and a deer in the forest; the multiple entity condition pictures include the entity condition picture of the bear and the entity condition picture of the deer, and the entity condition picture input here requires a single-channel binary image with a resolution of 768×768; based on the image control information, the multi-visual entity local control generation mode is used to generate the image:
[0106] (1) Preprocess the text prompt and each entity condition graph:
[0107] ① Use the Transformer-based text encoder to perform feature mapping on the text prompt p to obtain the text embedding feature sequence , where k is the total length of the text embedding feature sequence;
[0108] ②For each entity condition diagram Extract the largest polygon that can surround the area where the entity is located as the entity contour map of the corresponding entity condition map ;
[0109] (2) Use a standard normal distribution The random noise of the initial noise image ; Use text encoding features as image generation conditions, set total denoising time T, early denoising time point , in the preset early denoising stage Inside, the initial noise image is generated by the model Perform preliminary denoising at each time step to obtain a noisy image ;
[0110] For example, set the total denoising time T=50, the early denoising time point , then in The initial noise image in the time period Perform denoising step by step to obtain a noisy image ;
[0111] (3) Adaptive positioning of local control areas using the cross-attention map in the generative model; specifically:
[0112] ① Obtain the cross-attention map corresponding to the entity condition map;
[0113] ② Obtain the threshold of the segmentation attention score through the OTSU algorithm; binarize the cross attention map according to the threshold to obtain the entity area mask image R of the area where the entity is located in the generated image i ;
[0114] ③ Mask image R for each entity region i The localization area of the corresponding entity is approximated to obtain the entity localization area bounding box of the generated image. ,in They are the upper left and lower right horizontal and vertical coordinates of the entity positioning area bounding box respectively;
[0115] In the content of paragraph (3) under this mode, when calculating the cross-attention map, it is necessary to extract the cross-attention maps corresponding to the entities in multiple different entity condition maps, and calculate the potential positioning areas of different entities respectively; when there are multiple entities that need to be controlled, due to the text semantics, random number seeds and model parameters themselves, relying on the cross-attention map for positioning will result in a high degree of overlap in the positioning areas, which will lead to the problem of extremely poor subsequent generation effects. Therefore, in an embodiment of the present invention, in the content of paragraph (3) under this mode, after obtaining the positioning area, further adjustments are made based on the intersection between different entity positioning areas to avoid the influence of overlap between local entity condition maps; the schematic diagram of the comparison of the effects before and after the processing of the high overlap of multiple entity positioning areas can be seen in Figure 5 The specific entity positioning area adjustment process is as follows:
[0116] 1) One of the multiple entity condition maps to be located and scaled is used as the target entity condition map, and the location area of the entity in the generated image in the target entity condition map is recorded as the target location area;
[0117] 2) The location area of the entity in the remaining entity condition graph in the generated image is recorded as the remaining location area;
[0118] 3) Calculate the intersection and union ratio of the target positioning area and the remaining positioning areas;
[0119] 4) If the intersection-to-union ratio is greater than a preset value, the target positioning area is reduced in size and moved away from the central intersection of all positioning areas to adjust the target positioning area.
[0120] By cyclically executing the above process, the intersection-over-union ratio between the overlapping positioning areas meets the preset value, thereby ensuring better image generation quality.
[0121] (4) Locate the bounding box of the region based on the obtained entity , multi-level feature-level fusion is performed on the entity condition graph, and multi-modal encoding features of the unified space are obtained; specifically, the following are included:
[0122] ① According to the local control area, the corresponding entity condition diagram and solid outlines Perform positioning and scaling processing to obtain the processed local entity condition map and entity location contours ;
[0123] ② Use an image encoder composed of zero convolutions to encode the local entity condition graph And the binary background image Encode to obtain multimodal encoding features of unified space .
[0124] (5) In the later denoising stage, the multimodal encoding features The visual control features are obtained through the visual control adapter, and together with the global intermediate layer features in the generative model, the generative model is guided to perform accurate image generation; that is, the generative model is guided to perform post-denoising processing on the above-mentioned noisy image to achieve image generation;
[0125] exist Figure 4 In the above multi-visual entity local control generation mode, an image of a bear and a deer in a forest is finally generated, where the shapes of "bear" and "deer" are controlled by the entity condition graph input by the user, while the "forest" entity is generated by text.
[0126] 4. Multi-visual entity background mixed control generation mode (taking sketch mode as an example):
[0127] See also Figure 6As shown in the figure, the image control information input by the user for controlling the semantics of the generated image includes a text prompt, multiple entity condition images and a background image; the text prompt is "A swan floats on the river next to a boat", which is used to describe a swan floating on the river with a boat next to it; the multiple entity condition images are a swan and a boat, and the entity condition images input here require a single-channel binary image with a resolution of 768×768; based on the image control information, the multi-visual entity background mixed control generation mode is used to generate the image:
[0128] (1) Preprocess the text prompt and each entity condition graph:
[0129] ① Use the Transformer-based text encoder to perform feature mapping on the text prompt p to obtain the text embedding feature sequence , where k is the total length of the text embedding feature sequence;
[0130] ②For each entity condition diagram Extract the largest polygon that can surround the area where the entity is located as the entity contour map of the corresponding entity condition map ;
[0131] ③ Background image Perform edge detection and binarization to obtain a binary background image ;
[0132] (2) Use a standard normal distribution The random noise of the initial noise image ; Use text encoding features as image generation conditions, set total denoising time T, early denoising time point , in the preset early denoising stage Inside, the initial noise image is generated by the model Perform preliminary denoising at each time step to obtain a noisy image ;
[0133] For example, set the total denoising time T=50, the early denoising time point , then in The initial noise image in the time period Perform denoising step by step to obtain a noisy image ;
[0134] (3) Adaptive positioning of local control areas using the cross-attention map in the generative model; specifically:
[0135] ① Get the cross-attention map corresponding to the entity condition map.
[0136] ② Obtain the threshold of the segmentation attention score through the OTSU algorithm; binarize the cross attention map according to the threshold to obtain the entity area mask image R of the area where the entity is located in the generated image i .
[0137] ③ Mask image R for each entity region i The localization area of the corresponding entity is approximated to obtain the entity localization area bounding box of the generated image. ,in They are the upper left and lower right horizontal and vertical coordinates of the entity positioning area bounding box respectively.
[0138] (4) Perform multi-level feature-level fusion on the entity condition graph according to the local control area, and obtain multi-modal encoding features in a unified space; specifically,
[0139] ① According to the local control area, the corresponding entity condition diagram and solid outlines Perform positioning and scaling processing to obtain the processed local entity condition map and entity location contours ;
[0140] ② Use an image encoder composed of zero convolutions to encode the local entity condition graph And the binary background image Encode to obtain multimodal encoding features of unified space .
[0141] Since the conditions used by the present invention to control the final generated image include global text, background image and local entity condition image, directly mixing and controlling multiple levels of conditions will cause distortion of the local edges of the generated image. The global conditions and local conditions can be organically combined through the multi-level feature-level fusion method, and the control of the image by different entities is taken into account, effectively ensuring the consistency of the multi-level conditions and the generated content.
[0142] (5) In the later denoising stage, the multimodal encoding features The visual control features are obtained through the visual control adapter, and together with the global intermediate layer features in the generative model, the generative model is guided to generate accurate images.
[0143] exist Figure 6 In the example, the above multi-visual entity background mixed control generation mode finally generates an image of a goose and a boat with a river as the background.
[0144] In another embodiment, in the above-mentioned pure text control generation mode, single visual entity local control generation mode, multi-visual entity local control generation mode and multi-visual entity background mixed control generation mode, the text encoder of the Transformer used is the CLIP text encoder.
[0145] In another embodiment, in the above-mentioned pure text control generation mode, single visual entity local control generation mode, multi-visual entity local control generation mode and multi-visual entity background mixed control generation mode, the generation model adopted is constructed by a stable diffusion model (Stable Diffusion model); the generation model is initialized using SD V2.1 weights.
[0146] In another embodiment, in the above-mentioned single visual entity local control generation mode, multi-visual entity local control generation mode and multi-visual entity background mixed control generation mode, the cross attention map is represented as:
[0147]
[0148] Among them, i represents the i-th entity index. Since each entity condition graph corresponds to an entity index, represents the i-th entity condition graph; express The corresponding text tag; represents the cross attention map, and Only in The calculation is performed at the moment (i.e., when the previous denoising stage is completed); n represents the nth head index in the multi-head attention mechanism; N represents a total of N head indexes; t represents the time step in the diffusion process; Express The latent variable of generates the query vector in cross attention using a linear function; Represents the text tag feature The key vector in the cross attention is generated using a linear function and a transpose operation; d represents the feature embedding dimension.
[0149] In another embodiment, in the above-mentioned single-vision entity local control generation mode, multi-vision entity local control generation mode and multi-vision entity background mixed control generation mode, the entity area mask image R i It is expressed as:
[0150]
[0151] Among them, R i Represents a solid area mask image; Represents the entity region mask image R i The median coordinate is The pixel value of Represents the cross attention map The median coordinate is ; OTSU(·) represents the OTSU algorithm.
[0152] In another embodiment, in the above-mentioned single-vision entity local control generation mode, multi-vision entity local control generation mode and multi-vision entity background mixed control generation mode, since the cross-attention map in the generation model can reflect the approximate position of the entity in the final generated image to a certain extent, in order to better control the morphology of the entity to be controlled through the entity condition map in the later denoising stage, solve the problem that the entity morphology initially generated by the global control process of the text in the early denoising stage is quite different from the entity morphology in the entity condition map and there is a certain probability of inaccurate positioning, the noise image is optimized by the attention map optimization function, thereby indirectly optimizing the cross-attention map, controlling the entity position in the final image to conform to the text distribution while its morphology is consistent with the entity condition map. Figure 1 To better improve the quality of image generation; specifically:
[0153] In the later denoising stage, the attention map optimization function is obtained based on the cross attention map, and a limited number of loop iterations are performed through the attention map optimization function to optimize the noisy image ; The attention map optimization function consists of a text tag level function and a pixel level function, which is specifically expressed as:
[0154]
[0155] in, Represents text token level function; represents the pixel-level function; l represents the lth index in the network layer of the generative model; L represents the total number of network layers of the generative model; m represents the total number of entities that need to be locally controlled; i represents the index of the entity that needs to be locally controlled; It represents the entity positioning contour image after the entity contour image is processed by positioning and scaling; express The pixel value with coordinates (x, y); express The pixel value with coordinates (x, y); Represents the local entity condition graph after positioning and scaling processing; Represents the cross attention map corresponding to the i-th entity index of the l-th network layer at the denoising time point t; and are the height and width of the cross attention map of the lth network layer respectively; BCE(.) represents the binary cross entropy operation; represents the attention map optimization function; Represents the hyperparameter that controls the ratio of text tag level function and pixel level function, with a value range of [0,1];
[0156] The gradient value is calculated by the attention graph optimization function, and the noise image is updated according to the gradient value to obtain a new noise image; it is expressed as:
[0157]
[0158] in, represents a noisy image; represents the gradient operation; Represents the updated noise image The learning rate, in specific use, The possible value is 20; represents the new noise image updated by the attention map optimization function;
[0159] By performing a finite number of loop optimization updates on the noise image, based on the visual control features, the new noise image is subjected to post-denoising processing time by time step through the generative model to obtain the generated image. After the finite number of loop optimization updates on the noise image is completed, it does not participate in the calculation in the post-denoising stage to reduce computing resources. In specific use, the number of loops is 15.
[0160] In another embodiment, in the above-mentioned single visual entity local control generation mode, multi-visual entity local control generation mode and multi-visual entity background mixed control generation mode, since different users have different understandings of the object morphology, in practical applications, the abstract levels of the entity condition graphs provided by different users for the same object are inconsistent, which requires the model to have good generalization performance. In order to solve this problem, an adaptive control enhancement strategy is used in the later denoising stage to enhance the generalization ability of the model. Specifically, the present invention introduces adaptive weights to weight the visual control features and the intermediate features of the network layer of the generation model, thereby adjusting the visual control strength of the input entity condition graph on the later denoising stage. The specific formula is as follows:
[0161]
[0162] in, represents the adaptive weight function that changes with denoising time; represents the noise image at time t; represents the multimodal coding feature; a, b and c all represent the control adaptive weight function Hyperparameters of the direction; Represents the updated intermediate layer features of the generative model; Represents the UNet network layer; Represents the basic generative model Weight; represents the drawing encoder weights; represents zero convolution in the visual control adapter, and the corresponding weight is .
[0163] Figure 7 A visualization diagram of the effect of applying the multi-entity local control generation mode provided in an embodiment of the present invention to the input generation of entity condition images of various different modalities; wherein the entity condition image modalities, text prompts, and entity vocabulary corresponding to different entity condition images inputted include:
[0164] Taking the segmentation graph as the entity conditional graph modality, “An apple is near a banana on the table.” is used to describe that there is an apple and a banana on the table, where the given segmentation graphs correspond to the entity words “apple” and “banana” respectively;
[0165] Taking the posture graph as the entity condition graph modality, "A woman and a man are walking on the road." is used to describe a woman and a man walking on the road, where the given posture graphs correspond to the entity words "woman" and "man" respectively;
[0166] Using the depth map as the entity conditional graph modality, “A mouse is in the desert near amushroom.” is used to describe a banana next to a mouse in the desert, where the given depth maps correspond to the entity words “mouse” and “mushroom” respectively;
[0167] Taking the edge graph as the entity conditional graph modality, “A chicken with a gray rabbit.” is used to describe a hen and a gray rabbit, where the given edge graphs correspond to the entity words “chicken” and “rabbit” respectively;
[0168] The interactive and convenient multifunctional image generation method proposed by the present invention supports local entity condition image input of multiple different modalities and supports global text and background image input. The model has high generalization ability and flexibility of use. Users can use images of various modalities as input, which greatly facilitates the use of users and solves the problem of inconvenient interaction in the existing image generation field. In terms of control mode, you can freely choose a pure text control generation mode, a single visual entity local control generation mode, a multi-visual entity local control generation mode, and a multi-visual entity background mixed control generation mode. Its control mode is richer than that of previous image generation methods, solving the problem of single function in the image generation field. The present invention further optimizes the later denoising process through adaptive positioning of local control areas based on cross-attention maps, attention map optimization functions, and adaptive enhancement strategies, and realizes multi-modal and multi-level image control, thereby ensuring high-quality output of images and solving the problem of poor quality of existing image generation.
[0169] Figure 8 A schematic diagram of multiple sample visualization effects of a single visual entity local control generation mode provided by an embodiment of the present invention (taking the sketch mode as an example); the entity vocabulary corresponding to the text prompt and the sketch includes:
[0170] “The cup is sitting on the table” is used to describe that there is a cup on the table, where the given sketch corresponds to the entity word “cup”;
[0171] “The dog is barking loudly in the park” is used to describe that the dog is barking loudly in the park, where the given sketch corresponds to the entity word “dog”;
[0172] “A bear is standing by the lake” is used to describe a bear standing by the lake, where the given sketch corresponds to the entity word “bear”;
[0173] “A crab encounters a shell on the beach” is used to describe a crab encountering a shell on the beach, where the given sketch corresponds to the entity word “crab”;
[0174] “A woman is skiing down a snowy slope” is used to describe a woman skiing on a snowy slope, where the given sketch corresponds to the entity word “woman”;
[0175] “An oddlyshaped black toilet lid in a bathroow” is used to describe an oddly shaped black toilet lid in a bathroom, where the given sketch corresponds to the entity word “toilet”;
[0176] “A white pitcher is holding flowers in a window sill” is used to describe a bouquet of flowers placed in a white pot on the windowsill, where the given sketch corresponds to the entity word “pitcher”;
[0177] “There is a tower clock that is in the middle of the city” is used to describe that there is a clock tower in the center of the city, where the given sketch corresponds to the entity word “tower”;
[0178] “Creamy cheesecake dessert with whip cream and caramel” is used to describe a cream cheesecake dessert with whipped cream and caramel, where the given sketch corresponds to the entity word “cheesecake”;
[0179] “A brown horse is walking outside in the grass” is used to describe a brown horse walking outside on the grass, where the given sketch corresponds to the entity word “horse”;
[0180] “An elephant is surrounded by stones” is used to describe an elephant surrounded by stones, where the given sketch corresponds to the entity word “elephant”;
[0181] “A motorcycle with its brake extended standing outside” is used to describe a motorcycle parked outside, where the given sketch corresponds to the entity word “motorcycle”;
[0182] Fig. 9 A schematic diagram of multiple sample visualization effects of the multi-visual entity local control generation mode provided by an embodiment of the present invention (taking the sketch mode as an example); the entity vocabulary corresponding to the text prompt and the sketch includes:
[0183] “The pear rests beside the rabbit” is used to describe that the pear is next to the rabbit, where the given sketches correspond to the entity words “pear” and “rabbit” respectively;
[0184] “The cat naps beside the ripe pineapple on the ketchen counter” is used to describe the cat napping beside the ripe pineapple on the kitchen counter, where the given sketches correspond to the entity words “cat” and “pineapple” respectively;
[0185] “The tiger is prowling near the truck” is used to describe a tiger prowling near a truck, where the given sketches correspond to the entity words “tiger” and “truck” respectively;
[0186] “The squirrel investigates the pizza left on the picnic blanket” is used to describe the squirrel investigating the pizza left on the picnic blanket, where the given sketches correspond to the entity words “squirrel” and “pizza” respectively;
[0187] “The airplane soars above the castle below” is used to describe the airplane flying above the castle, where the given sketches correspond to the entity words “airplane” and “castle” respectively;
[0188] “The chicken is pecking at the ground, while the dog is sleeping” is used to describe that the chicken is pecking at the ground while the dog is sleeping, where the given sketches correspond to the entity words “chicken” and “dog” respectively;
[0189] “A woman is in gear skiing down a snowy slope with a dog” is used to describe a woman skiing down a snowy slope with a dog, where the given sketches correspond to the entity words “woman” and “dog” respectively;
[0190] “Asleeping dog is lying on a round bed next to a rocking chair” is used to describe a dog sleeping on a round bed next to a rocking chair, where the given sketches correspond to the entity words “dog” and “chair” respectively;
[0191] “The shoe is beside the teapot and scissors” is used to describe that the shoe is beside the teapot and scissors, where the given sketches correspond to the entity words “shoe”, “teapot” and “scissors” respectively;
[0192] “The rabbit rests near the penguin and pineapple” is used to describe the rabbit resting near the penguin and pineapple, where the given sketches correspond to the entity words “rabbit”, “penguin” and “pineapple” respectively;
[0193] “In the rural field, a motorcycle speeds past a cow and follows adog” is used to describe that in a rural field, a motorcycle follows a dog and passes by a cow. The given sketches correspond to the entity words “motorcy”, “cow” and “dog” respectively.
[0194] “A car and a giraffe and an airplane are on the road” is used to describe that there is a car, a giraffe and an airplane on the road, where the given sketches correspond to the entity words “car”, “giraffe” and “airplane” respectively;
[0195] Fig.10 A schematic diagram of multiple sample visualization effects of a multi-visual entity background hybrid control generation mode provided in an embodiment of the present invention (taking the sketch mode as an example); the entity vocabulary corresponding to the text prompt and the sketch includes:
[0196] “A zebra is next to the Leaning Tower of Pisa” is used to describe a zebra next to the Leaning Tower of Pisa, where the given sketch corresponds to the entity word “zebra”;
[0197] “A cow walks by the pyrawid” is used to describe a cow walking by the pyramids, where the given sketch corresponds to the entity word “cow”;
[0198] “There is a parrot on the shopping street” is used to describe that there is a parrot on the shopping street, where the given sketch corresponds to the entity word “parrot”;
[0199] “The dog is in the desert” is used to describe a dog in the desert, where the given sketch corresponds to the entity word “dog”;
[0200] “A red chair is on the grass” is used to describe that there is a red chair on the grass, where the given sketch corresponds to the entity word “chair”;
[0201] “A car is parking in front of an ancient Japanese buliding” is used to describe a car parked in front of an ancient Japanese building, where the given sketch corresponds to the entity word “car”;
[0202] “An airplane is flying pass the skyscraper” is used to describe an airplane flying over a skyscraper, where the given sketch corresponds to the entity word “airplane”;
[0203] “A cat is lying on the beach” is used to describe a cat lying on the beach, where the given sketch corresponds to the entity word “cat”;
[0204] “The lion and bird are standing on Mars” is used to describe a lion and a bird standing on Mars, where the given sketches correspond to the entity words “lion” and “bird” respectively;
[0205] “The lion and a deer are in the forest” is used to describe that there is a lion and a deer in the forest, where the given sketches correspond to the entity words “lion” and “deer” respectively;
[0206] “A butterfly and a camel are in front of the Eiffel Tower” is used to describe a butterfly and a crow in front of the Eiffel Tower, where the given sketches correspond to the entity words “butterfly” and “camel” respectively;
[0207] “A car and a horse are in front of the Big Ben” is used to describe that there is a car and a horse in front of Big Ben, where the given sketches correspond to the entity words “car” and “horse” respectively;
[0208] Next, the multifunctional visual control text image method based on the diffusion model provided by the present invention (taking the sketch mode as an example) is compared with the existing multimodal image generation method in terms of objective indicators. The comparison methods include: a stable diffusion model (SD) using only text; T2I-Adapter, ControlNet and UniControl using text and sketch modes; GLIGEN and InstanceDiffusion using text and entity bounding boxes. The comparison indicators include: CLIP (used to evaluate the semantic consistency between the generated image and the input text), FID (used to evaluate the image quality of the generated image), DINO (used to evaluate the semantic consistency between the generated image and the original image) and ACC (used to evaluate the accuracy of the generated entity). Among them, w / GT indicates whether the original image entity location in the dataset is used as the input prior, LoD indicates the use of sketch information consistent with the original image entity position, and BBoxes indicates the use of bounding box information consistent with the original image entity position. This method uses sketch information consistent with the original image entity position and sketch information located by the cross attention map for comparison. The comparison results are shown in Table 1 below:
[0209] Table 1
[0210]
[0211] From the above Table 1, we can see that: i) Compared with SD, the FID values of other methods are lower, indicating that adding sketch or entity bounding box information enhances the spatial alignment between the generated image and the real image. ii) Regardless of whether the real position or the predicted position is used, this method can achieve a lower FID value, which reflects its robustness and superiority. iii) SD obtains the highest CLIP score, which may be due to the conflict between multiple conditions or the limitation of sketch extraction technology. This method solves this problem by optimizing the attention map and adaptive control strategy during the reasoning process. iv) This method obtains the highest DINO score under different settings, indicating that the generated image is closest to the real image in visual features, further proving its practicality. v) InstanceDiffusion performs well in entity semantics and obtains the highest ACC, but its visual information is limited to the entity bounding box, resulting in the loss of local detail information. In contrast, this method better balances local details and global semantic information.
[0212] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multifunctional image generation method with convenient interaction, characterized in that: The steps include: S1, receiving image generation control conditions input by a user, and preprocessing the image generation control conditions; the image generation control conditions include: text prompts, entity condition graphs and background graphs; S2, using random noise that conforms to the standard normal distribution to form an initial noise image; in the early denoising stage, based on the preprocessed text prompt, the initial noise image is globally guided to denoise through the generative model to obtain a noise image; S3, using the cross-attention map in the generative model to achieve adaptive positioning of the local control area; S4, according to the local control area after positioning, the pre-processed entity condition map and background map are fused at a multi-level feature level to obtain a multi-modal coding feature of a unified space; S5. In the later denoising stage, the multimodal coding features are passed through the visual control adapter to obtain visual control features; the visual control features and the global intermediate layer features in the generation model jointly guide the generation model to perform denoising on the noisy image to achieve image generation; The step S3 specifically includes: (1) The cross attention graph is expressed as: Among them, s i represents the i-th entity condition graph; c i Indicates i The corresponding text tag; represents the cross attention map, and It is only calculated at time t = τ, where t = τ is when the previous denoising stage is completed; n represents the nth head index in the multi-head attention mechanism; N represents a total of N head indices; t represents the time step in the diffusion process; Q n (s i ) indicates that i The latent variable K generates the query vector in the cross attention using a linear function; n (c i ) represents the text tag feature c i Use a linear function and generate the key vector in the cross attention after transposition operation; d represents the feature embedding dimension; (2) Obtain the threshold of the segmentation attention score through the OTSU algorithm; binarize the cross attention map according to the threshold to obtain the entity area mask image of the area where the entity is located in the generated image; the entity area mask image R i It is expressed as: Among them, R i Represents the entity area mask image; R i (x r ,y r ) represents the entity region mask image R i The median coordinate is (x r ,y r )’s pixel value; Represents the cross attention map The median coordinate is (x r ,y r ) pixel value; OTSU(·) represents the OTSU algorithm; (3) For each entity region mask image R i The localization area is approximated to obtain the entity localization area bounding box of the corresponding entity in the generated image. in and are the upper left horizontal and vertical coordinates of the entity positioning area bounding box; and They are the lower right horizontal and vertical coordinates of the entity positioning area bounding box.
2. The interactive and convenient multifunctional image generation method according to claim 1, characterized in that: In the step S1, the input of the entity condition graph supports multiple modes.
3. The interactive and convenient multifunctional image generation method according to claim 1, characterized in that: In the step S1, the image generation control condition is preprocessed, specifically including: (1) Use the Transformer-based text encoder to perform feature mapping on the text prompt p and obtain the text embedding feature sequence C = {c1, c2, …, c k }, where k is the total length of the text embedding feature sequence; (2) For each entity condition graph s i Extract the largest polygon that can surround the area where the entity is located as the entity contour map of the corresponding entity condition map (3) Background image s bg Perform edge detection and binarization to obtain the binary background image s' bg .
4. The interactive and convenient multifunctional image generation method according to claim 3, characterized in that: In step S2, based on the preprocessed text prompt, the initial noisy image is globally guided to be denoised by generating a model, specifically: Taking the text embedding feature sequence C as the image generation condition, the initial noisy image is globally guided for denoising through the generative model.
5. The interactive and convenient multifunctional image generation method according to claim 1, characterized in that: When the number of entity condition graphs input by the user is multiple, the entity positioning area bounding box B determined by the cross attention graph needs to be i Make adjustments; specifically include: (1) One of the multiple entity condition maps to be located and scaled is used as a target entity condition map, and the location area of the entity in the generated image in the target entity condition map is recorded as the target location area; (2) The location area of the entity in the generated image in the remaining entity condition graph is recorded as the remaining location area; (3) Calculate the intersection and union ratio of the target positioning area and the remaining positioning areas; (4) If the intersection-to-union ratio is greater than a preset value, the target positioning area is reduced in size and moved in a direction away from the central intersection of all positioning areas to adjust the target positioning area.
6. The interactive and convenient multifunctional image generation method according to claim 1, characterized in that: The step S4 specifically includes: (1) Based on the entity positioning area bounding box B i , respectively for the corresponding entity condition graph s i and solid outlines Perform positioning and scaling processing to obtain the processed local entity condition graph s′ i and entity location contours (2) Use an image encoder composed of zero convolutions to respectively decode the local entity condition graph s′ i and the binary background image s′ bg Encode to obtain the multimodal encoding feature z of the unified space S .
7. The interactive and convenient multifunctional image generation method according to claim 3, characterized in that: The step S5 specifically includes: (1) obtaining a visual control feature by applying the multimodal coding feature to a visual control adapter; (2) In the later denoising stage, the visual control features are weighted summed with the global intermediate layer features in the generative model conditioned on the text embedding feature sequence C; the weighted sum result is used as the updated intermediate layer features of the generative model; (3) Guide the generative model through the updated intermediate layer features of the generative model to denoise the noisy image to achieve image generation.
8. The interactive and convenient multifunctional image generation method according to claim 5, characterized in that: In the later denoising stage, an attention map optimization function is obtained based on the cross-attention map and the noisy image is updated; the attention map optimization function consists of a text tag level function and a pixel level function, which is specifically expressed as: in, Represents text token level function; represents the pixel-level function; l represents the lth index in the network layer of the generative model; L represents the total number of network layers of the generative model; m represents the total number of entities that need to be locally controlled; i represents the index of the entity that needs to be locally controlled; It represents the entity positioning contour image after the entity contour image is processed by positioning and scaling; express The pixel value with coordinates (x, y); express The pixel value with coordinates (x, y); s′ i Represents the local entity condition graph after positioning and scaling processing; represents the cross attention map corresponding to the i-th entity index of the l-th network layer at the denoising time point t; h l and w l are the height and width of the cross attention map of the lth network layer, respectively; BCE(·) represents the binary cross entropy operation; represents the attention map optimization function; λ represents the hyperparameter that controls the ratio of text tag level function to pixel level function, and its value range is [0,1]; The gradient value is calculated by the attention graph optimization function, and the noise image is updated according to the gradient value to obtain a new noise image; it is expressed as: Among them, z t represents a noisy image; represents the gradient operation; α represents the update of the noise image z t The learning rate z′ t represents the new noise image updated by the attention map optimization function; By performing a finite number of cyclic optimization updates on the noisy image, based on the visual control features, the new noisy image is subjected to post-denoising processing time-step by time-step through the generative model to obtain the generated image.
9. The interactive and convenient multifunctional image generation method according to claim 3, characterized in that: In the later denoising stage, the visual control features and the intermediate layer features of the generative model are weighted and calculated by introducing adaptive weights to serve as the updated intermediate layer features of the generative model; expressed as: Among them, γ(t) represents the adaptive weight function that changes with the denoising time; z t represents the noise image at time t; z s represents the multimodal encoding feature; a, b and c are all hyperparameters that control the direction of the adaptive weight function γ(t); y S Represents the updated intermediate layer features of the generative model; represents the UNet network layer; θ Φ represents the weight of the basic generative model Φ; Θ S represents the drawing encoder weights; represents the zero convolution in the visual control adapter, with the corresponding weight θ z2 .
Citation Information
Patent Citations
Image processing method and device based on generative artificial intelligence technology
CN117252753A
Attitude generation method based on diffusion model and ControlNet
CN117576260A
Image generation control method and system
CN117765113A
Method and system for generating image by text based on lightweight dynamic refinement
CN116912367A
Multi-instance controllable image generation method based on cross attention redistribution
CN118628611A