An image generation system and method oriented towards multi-object spatial conditions
By constructing a multi-object spatial conditional dataset and introducing object-level structural control and control relaxation submodules, the semantic consistency and creative controllability issues of object-level image generation in existing technologies are solved, achieving precise control and high-quality generation of multiple objects, and improving the model's versatility and generation effect.
Patent Information
- Application Number
- CN202511157744.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-08-19
AI Technical Summary
Existing object-oriented spatial conditional image generation methods face challenges in terms of semantic consistency and creative controllability, making it difficult to achieve precise control and high-quality generation of multiple objects.
By constructing a multi-object spatial conditional dataset, introducing object-level structural control submodules and control relaxation submodules, fusing object-level semantic features and ControlNet original features, and adopting a staged training strategy, the object-level control accuracy and model versatility of image generation are improved.
It significantly improves the object-level control precision and quality of image generation, supports multiple types of spatial input conditions, enhances the compatibility of the model and the flexibility of image generation, and improves the image generation quality and control precision.
Smart Images

Figure CN120655767B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and deep learning technology, and in particular to an image generation system and method for multi-object spatial conditions. Background Technology
[0002] In the field of image generation, text-guided diffusion models can generate highly matched and high-quality diverse images based on text descriptions, attracting widespread attention from academia and industry. To further enhance the model's pixel-level control over generated images, existing technologies have successfully incorporated spatial conditions such as depth information and edge features into the diffusion model framework, and these techniques have been successfully applied to fields such as film special effects production, product design, and digital art creation.
[0003] In text-condition-guided diffusion image generation models, introducing spatial input conditions enables more precise control over the generated results. ControlNet cleverly integrates spatial conditions from multimodal guided images (such as Canny edges, depth maps, or human poses) into a trainable denoising U-Net variant, and combines the advantages of zero-padding convolutions with the base model to generate images that are both visually appealing and contextually coherent.
[0004] Several studies have been conducted to optimize ControlNet, including multi-condition fusion, network optimization, control relaxation, and frequency domain control. For example, Reference 1 (Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion Models[J], Computer Vision and Pattern Recognition, 2023) discloses a Uni-ControlNet model that innovatively adopts a dual-adaptor architecture. This allows for the flexible and composable use of different local controls (e.g., edge maps, depth maps, segmentation masks) and global controls (e.g., CLIP image embedding) within a single model. Furthermore, it only requires fine-tuning two additional adapters on a frozen pre-trained text-to-image diffusion model, thus eliminating the significant cost of training from scratch and overcoming the limitation of traditional ControlNet requiring separate control models for each spatial condition. This design not only enables flexible switching between different conditions but also supports multi-condition guided collaborative optimization, significantly improving the model's applicability and computational efficiency. Cocktail proposes a control normalization technique to maintain semantic alignment between multiple spatial conditions and designs a cross-attention modulation strategy to avoid the misgeneration of objects outside the attention region. SmartControl employs an adaptive control scale predictor to reconcile semantic conflicts between visual conditions and textual cues.
[0005] Unlike the conditional image generation methods mentioned above that strictly adhere to visual input, Multi-Object Generation (MIG) techniques focus on precise control over the position and attributes of multiple objects. Current MIG methods are mainly divided into two categories: training-free methods and tuning-based methods. The former avoids additional training on pre-trained models, instead using feature manipulation or gradient-based guidance to tune the activation regions of instances in the cross-attention layer. However, this type of method is prone to instability in the latent space, thus reducing the quality of the generated images.
[0006] Recent works have focused on achieving precise semantic control in image generation under complex visual conditions, including spatial conditions with multiple instances. For example, Reference 2 (3DIS: Depth-Driven Decoupled Instance Synthesis for Text-to-Image Generation, Computer Vision and Pattern Recognition, 2024) designed a depth-instance decoupled framework, 3DIS, aiming to achieve precise semantic control of instances by generating coarse-grained instance depth maps. However, its rendering process still relies on ControlNet and suffers from semantic obfuscation. Semantic segmentation-based methods can utilize masks to achieve independent denoising control for different objects, and this method can be combined with multi-instance generation techniques to improve generation accuracy, but it is difficult to handle application scenarios with conflicting text-spatial conditions. In addition, some methods based on weakened control conditions use gradient guidance and feature operations to adjust spatial condition weights, but their ability to finely adjust at the object level is limited. Spatial feature masking is also an alternative method, but it leads to a decrease in image quality. Despite the significant progress made in existing work, many key issues remain to be addressed in image generation methods for object-level spatial conditions.
[0007] In summary, current research on object-oriented spatial condition image generation methods mainly faces two key scientific problems: (1) Semantic consistency problem: It is necessary to establish an accurate mapping relationship between the semantic information of the object and the corresponding foreground structural features; (2) Creative controllability challenge: Although non-traditional element combinations can produce innovative results, the existing visual structural constraint mechanism limits the user's ability to achieve the expected effect through text modification, and the algorithm needs to have the dynamic adaptability of local structural features. Summary of the Invention
[0008] The purpose of this invention is to provide an image generation system and method for multi-object spatial conditions. This system and method fuse user-input spatial condition maps, multi-object layout information, and text descriptions to generate high-quality images that highly match these conditions. This improves the object generation capability under complex visual input, significantly enhances the object-level control precision of the generated images, and can be adapted to different ControlNet backbone networks. It also supports various types of spatial input conditions such as depth maps and edge detection maps.
[0009] An embodiment provides an image generation system for multi-object spatial conditions, comprising the following modules:
[0010] The dataset construction module is used to extract object-level bounding box information from the original image dataset, perform image editing within a given bounding box region, extract aligned and unaligned spatial conditions from the original image and the image editing respectively, and extract the image text description of each image from the original image dataset to construct a multi-object spatial condition dataset.
[0011] The feature extraction module is used to extract object-level text features from image text descriptions, extract object-level location features from object-level bounding box information, and concatenate object-level text features and object-level location features to obtain object-level semantic features.
[0012] The multi-object spatial condition image generation model uses the ControlNet backbone network as the basic framework. It introduces an object-level structural control submodule to fuse object-level semantic features and ControlNet original features to obtain object-level structural features. At the same time, it introduces an object-level control relaxation submodule to establish object-level control features. It measures the similarity between object-level semantic features and ControlNet original features, fuses ControlNet original features, object-level structural features and object-level control features, and adds the fused features to the original U-Net to generate multi-object spatial condition images.
[0013] The model training and inference generation module is used to construct a loss function to train the multi-object spatial condition image generation model in stages, taking the multi-object spatial condition dataset as input. The spatial conditions, multi-object layout information and text description input by the user are input into the trained multi-object spatial condition image generation model to generate images that conform to the multi-object spatial conditions.
[0014] In one embodiment, the constructed multi-object spatial condition dataset includes: the original image, image text description, object-level label list, object-level bounding box information, and aligned and unaligned spatial conditions.
[0015] In one embodiment, the feature extraction module extracts object-level text features from the image text description by: first, using a word segmenter to vectorize the image text description, and then using a text CLIP encoder. Extracting object-level text features using MLP layers;
[0016] The extraction of object-level location features from object-level bounding box information includes: converting the object-level bounding box information into an object-level mask, and using an image CLIP encoder. The MLP layer extracts object-level location features.
[0017] In one embodiment, in a multi-object spatial conditional image generation model, the object-level structure control submodule fuses object-level semantic features and ControlNet original features to obtain object-level structural features, including:
[0018] The object-level structure control submodule inputs object-level semantic features and ControlNet raw features into the cross-attention layer to obtain object-level structure features, calculated as follows:
[0019] ,
[0020] ,
[0021] In the formula, This represents the index of each object in the image. Represents object-level structural features. This indicates that the query features in the ControlNet backbone network are the original ControlNet features. Representing object-level semantic features, express The transformed object-level structure control submodule key features Representing feature dimension, Represents the transpose matrix. This represents the feature map of the ControlNet backbone network. express The transformed object-level structural control submodule value characteristics Represents the query matrix. This represents the weight matrix corresponding to the key features in the object-level structural control submodule. This represents the weight matrix corresponding to the value features in the object-level structural control submodule.
[0022] Among them, object-level structural features are used to describe the correlation between object-level text descriptions and corresponding local spatial conditions, enabling precise control over image-related locations and text.
[0023] In one embodiment, in a multi-object spatial conditional image generation model, the object-level control relaxation submodule establishes object-level control features and measures the similarity between object-level semantic features and ControlNet original features, including:
[0024] The object-level control relaxation submodule controls by querying features Dot product operation is performed between the key features to calculate the attention similarity between the object-level semantic features and the original ControlNet features in the pixel dimension, and average pooling is performed on the features along the channel dimension to obtain the object-level control features:
[0025] ,
[0026] In the formula, Indicates object-level control characteristics. This indicates the average pooling operation. This indicates that the query features in the ControlNet backbone network are the original ControlNet features. Representing feature dimension, Represents the transpose matrix. express The transformed object-level control relaxation submodule key features are obtained through... The weight matrix corresponding to the key features in the object-level control relaxation submodule is learned;
[0027] Among them, object-level control features are used to measure the correlation between object-level semantic features and ControlNet original features, and to alleviate the conflict between object-level text description and corresponding local spatial conditions.
[0028] In one embodiment, the image generation model under multi-object spatial conditions fuses ControlNet raw features, object-level structural features, and object-level control features, including:
[0029] In object-level structural features A masked object-level structural feature Summing is performed, and then the results are fused through a self-attention layer to obtain the global structural features.
[0030] ,
[0031] in, Represents global structural features. Indicates the self-attention layer. Represent each bounding box Process and obtain object-level masks;
[0032] In object-level control features A masked object-level control feature Summation is performed, and then global control features are obtained through convolution and the sigmoid activation function:
[0033] ,
[0034] in, Indicates global control characteristics. This represents the summation operation. Indicates the convolution operation;
[0035] The global structural features and ControlNet original features are compared according to the global control features. We perform weighted fusion to obtain the fused features:
[0036] ,
[0037] in, Indicates the characteristics after fusion. Indicates the self-attention layer. Represents the original features of ControlNet. This indicates the fusion weight.
[0038] In one embodiment, the model training and inference generation module, which uses a multi-object spatial condition dataset as input and constructs a loss function to train the multi-object spatial condition image generation model in stages, includes:
[0039] Spatial conditions in a multi-object spatial condition dataset Global text hints in image text descriptions Bounding box and object-level text Using the input as the model input, a loss function is constructed to train the multi-object spatial condition image generation model in stages. Noise is added to the input based on the multi-object spatial condition dataset, enabling the model to reconstruct an image that matches the input. The constructed loss function is expressed as follows:
[0040] ,
[0041] in, This represents the constructed loss function. Represents the features of the original input image. Represents a timestamp. Represents the noise latent vector. Represents real noise. This represents the predicted noise.
[0042] Furthermore, in the model training and inference generation module, the phased training includes:
[0043] In the first stage, an image generation model with multiple object spatial conditions is trained using unaligned spatial conditions, enabling the model to learn object-level semantic features.
[0044] In the second stage, a multi-object spatial condition image generation model is trained using a combination of 50% unaligned spatial conditions and 50% aligned spatial conditions, with bounding box features initialized to zero to preserve the original structure.
[0045] Furthermore, during phased training, the parameters of the ControlNet backbone network and U-Net are frozen, and only the object-level structural control submodule and the object-level control relaxation submodule are trained. Each phase is optimized for 20 epochs (batch size of 32, gradient accumulation of 8).
[0046] On the other hand, the present invention also provides an image generation method for multi-object spatial conditions, comprising the following steps:
[0047] Extract object-level bounding box information from the original image dataset, perform image editing within the given bounding box region, and extract aligned and unaligned spatial conditions from the original image and the image editing respectively. At the same time, extract the image text description of each image from the original image dataset to construct a multi-object spatial condition dataset.
[0048] Extract object-level text features from image text descriptions, extract object-level location features from object-level bounding box information, and concatenate object-level text features and object-level location features to obtain object-level semantic features.
[0049] Using the ControlNet backbone network as the basic framework, an object-level structural control submodule is introduced to fuse object-level semantic features and ControlNet original features to obtain object-level structural features. At the same time, an object-level control relaxation submodule is introduced to establish object-level control features. The similarity between object-level semantic features and ControlNet original features is measured. The ControlNet original features, object-level structural features and object-level control features are fused and the fused features are added to the original U-Net to generate images with multi-object spatial conditions.
[0050] Using a multi-object spatial condition dataset as input, a loss function is constructed to train the multi-object spatial condition image generation model in stages. The spatial conditions, multi-object layout information, and text description input by the user are fed into the trained multi-object spatial condition image generation model to generate multi-object spatial condition images.
[0051] Compared with the prior art, the beneficial effects of the present invention include at least the following:
[0052] (1) This invention introduces object-level semantic features on the basis of the original ControlNet backbone network, which greatly improves the flexibility of user input text description, enhances the compatibility of existing models, and supports various types of spatial input conditions such as depth maps, edge detection maps, sketches, doodles, and line drawings, and has excellent model versatility.
[0053] (2) The present invention trains the model by constructing a multimodal dataset and evaluates the model through quantitative and qualitative experiments. The results show that the image generation results are superior in multiple indicators, including image-text consistency, image quality and structure control preservation, etc., which confirms that the present invention has significant advantages in object-level image generation tasks and is superior to existing methods in terms of image generation quality and control accuracy.
[0054] (3) The present invention can combine visual language models such as GPT-4o to improve the end-to-end image generation experience for users. Attached Figure Description
[0055] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0056] Figure 1 This is a schematic diagram of the structure of the image generation system for multi-object spatial conditions provided by the present invention.
[0057] Figure 2 This is a schematic diagram of the structure of the multi-object spatial condition image generation model provided in the embodiment.
[0058] Figure 3 The flowchart illustrates the operation of the dataset construction module provided in this embodiment.
[0059] Figure 4 This is a schematic diagram of the structure of each sub-module in the multi-object spatial condition image generation model provided in the embodiment.
[0060] Figure 5 This is a flowchart illustrating the image generation method for multi-object spatial conditions provided by the present invention.
[0061] Figure 6 This is a comparison chart of the model generation results of this invention and existing technologies.
[0062] Figure 7 The generation results on the SDXL model are provided for the example.
[0063] Figure 8 The generation result of combining the model of the present invention with a visual language model is provided for the embodiment.
[0064] Figure 9Visual analysis of object-level structure control provided for the embodiments. Detailed Implementation
[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and given in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0066] This invention provides an image generation system oriented towards multi-object spatial conditions, such as... Figure 1 As shown, the system includes the following modules: dataset construction module, feature extraction module, multi-object spatial conditional image generation model, and model training and inference generation module. The first three modules complete the model development, and the fourth module realizes the model application.
[0067] In the embodiments, such as Figure 2 The diagram shown illustrates the specific structure of an image generation system that incorporates multi-object spatial conditions, serving as an example for the development and application of image generation models oriented towards multi-object spatial conditions.
[0068] 1. The dataset construction module is used to extract object-level bounding box information from the original image dataset, perform image editing within a given bounding box region, and extract alignment and non-alignment spatial conditions from the original image and the image editing respectively. At the same time, it extracts the image text description of each image from the original image dataset to construct a multi-object spatial condition dataset.
[0069] In this embodiment, the original text-image data was first collected. Specifically, nearly 3 million text-image datasets were crawled from the open-source LAION-Aesthetics dataset, and 1 million high-quality original text-image datasets were selected based on the aesthetic scores in the text-image datasets.
[0070] Next, a multi-object spatial condition dataset is constructed. Since this invention aims to solve the task of generating high-quality images under multi-object spatial conditions, the constructed model must be able to achieve text-image consistency under both aligned and unaligned spatial input conditions.
[0071] Alignment spatial conditions can be obtained directly from the original image using traditional preprocessing techniques (such as MIDIS or Canny edge detection), while the latter's non-alignment conditions require object-level deletion or modification operations on the original image before preprocessing. Therefore, the construction of a multi-object spatial condition dataset involves three stages, such as... Figure 3 As shown:
[0072] i. Object-level label information generation: The image labeling model RAM is used to extract the object-level label list of each image from the existing publicly available high-quality image dataset LAION-Aesthetics, and the semantic segmentation model (Grounded-SAM) is combined to obtain object-level bounding box information.
[0073] ii. Image Editing: Use the image editing model (SDXL-Inpainting) to remove or modify objects within a given bounding box area.
[0074] Compared to existing methods, this model exhibits better stability in image editing.
[0075] iii. Alignment and non-alignment spatial condition extraction: Spatial condition processing techniques (such as Canny edge detection and depth maps) are used to extract fine and coarse-grained spatial condition maps from the original and edited images, respectively.
[0076] Finally, a multi-object spatial conditional dataset of nearly 400,000 images was constructed by summarizing the original images, image text descriptions, object-level label lists, object-level bounding boxes, and aligned and unaligned spatial structure graphs.
[0077] 2. Feature extraction module, which is used to extract object-level text features from image text description, extract object-level position features from object-level bounding box information, and concatenate object-level text features and object-level position features to obtain object-level semantic features.
[0078] In this embodiment, text and image encoders are used to obtain object-level semantic features. Specifically, given a text description of each object, a word segmenter is used to vectorize the image text description, and then a text CLIP encoder is used. And a trainable MLP layer to extract object-level text features .
[0079] Given the bounding box information of each object, it is transformed into an object-level mask using traditional image processing, and then a CLIP encoder is used. And a trainable MLP layer to extract object-level location features .
[0080] Finally, object-level text features and object-level location features Concatenation is performed to obtain object-level semantic features.
[0081] 3. A multi-object spatial condition image generation model is used as the basic framework of the ControlNet backbone network. By introducing an object-level structure control submodule, object-level semantic features and ControlNet original features are fused to obtain object-level structural features. At the same time, an object-level control relaxation submodule is introduced to establish object-level control features. The similarity between object-level semantic features and ControlNet original features is measured. The ControlNet original features, object-level structural features and object-level control features are fused and added to the original U-Net to generate images that meet the multi-object spatial conditions.
[0082] In this embodiment, the present invention injects two basic design modules into the ControlNet backbone network: an object-level structure control module and an object-level control relaxation module.
[0083] The object-level structure control submodule employs a multi-branch cross-attention fusion mechanism to effectively fuse object-level semantic features with spatial structure features. By incorporating additional cross-attention features into the ControlNet backbone, the object-level structure control submodule independently computes the associated features of all objects and combines these features based on their bounding boxes. The resulting hybrid features are then weighted together with the original features to enhance object-based spatial understanding. The object-level control relaxation submodule aims to address the inconsistency between textual descriptions and structural conditions. It consists of several object-level feature scaling units, each utilizing object-level text... and object-level bounding boxes This allows for precise control of local differences. Since this invention relies on the position of user-provided objects, it can directly predict this information by combining a visual language model. The following sections will elaborate on the object-level structure control submodule and the object-level control relaxation submodule.
[0084] Object-level structure control submodule: The cross-attention layer is the only layer in the Stable Diffusion generative model that connects text and image features. By manipulating the features in this layer, precise control over image-related locations and text can be achieved. Therefore, an object-level structure control submodule is introduced, such as... Figure 4 As shown.
[0085] This submodule accepts object-level bounding boxes. Object-level text As conditional input. Passed through the corresponding text CLIP encoder. and image CLIP encoder After processing by the MLP layer, the conditional features are concatenated and represented as object-level semantic features. :
[0086]
[0087] in, This represents the index of each object in the image. Represent each bounding box Process to obtain object-level masks, This indicates a splicing operation.
[0088] The object-level semantic features and the original ControlNet features are input into a learnable cross-attention layer to obtain the object-level structural features, calculated as follows:
[0089]
[0090]
[0091] In the formula, This represents the index of each object in the image. Represents object-level structural features. This indicates that the query features in the ControlNet backbone network are the original ControlNet features. Representing object-level semantic features, express The transformed object-level structure control submodule key features Representing feature dimension, Represents the transpose matrix. This represents the feature map of the ControlNet backbone network. express The transformed object-level structural control submodule value characteristics Represents the query matrix. This represents the weight matrix corresponding to the key features in the object-level structural control submodule. This represents the weight matrix corresponding to the value features in the object-level structural control submodule.
[0092] Among them, object-level structural features are used to describe the correlation between object-level text descriptions and corresponding local spatial conditions, enabling precise control over image-related locations and text. Since multiple objects of the same class can exist under the same spatial input condition, the object-level structural control submodule requires location information to process object-level structural features. Unlike traditional methods that rely on Fourier transform and MLP encoding of location features, this invention uses a frozen image encoder from the bounding box... Extract location features from them.
[0093] Object-level control relaxation submodule: Conflicts between object descriptions and spatial input conditions are common in creative image generation. To address this issue, this invention introduces an object-level control relaxation submodule, such as... Figure 4 As shown.
[0094] This submodule predicts the object-level control scale, and then uses this scale to measure the similarity between the original backbone features and the object-level structural features. We do this by querying features... Dot product operation is performed between the key features to calculate the attention similarity between the object-level semantic features and the original ControlNet features in the pixel dimension, and average pooling is performed on the features along the channel dimension to obtain the object-level control features:
[0095]
[0096] In the formula, Indicates object-level control characteristics. This indicates the average pooling operation. This indicates that the query features in the ControlNet backbone network are the original ControlNet features. Representing feature dimension, Represents the transpose matrix. express The transformed object-level control relaxation submodule key features are obtained through... The weight matrix corresponding to the key features in the object-level control relaxation submodule is learned;
[0097] Among them, object-level control features are used to measure the correlation between object-level semantic features and ControlNet original features, and to alleviate the conflict between object-level text description and corresponding local spatial conditions.
[0098] Object-level feature fusion submodule: Through the above two modules, object-level structural features are obtained. and object-level control features To effectively aggregate structural features across different objects, ControlNet raw features, object-level structural features, and object-level control features are fused together, such as... Figure 4 (As shown).
[0099] First, for each object-level structural feature Masked object-level structural features Summing is performed, and then the results are fused through a self-attention layer to obtain the global structural features.
[0100]
[0101] in, Represents global structural features. Indicates the self-attention layer. Represent each bounding box Process to obtain object-level masks.
[0102] Next, regarding the object-level control features... A masked object-level control feature Summation is performed, and then global control features are obtained through convolution and the sigmoid activation function:
[0103]
[0104] in, Indicates global control characteristics. This represents the summation operation. This indicates a convolution operation.
[0105] Finally, the global structural features and the original ControlNet features are compared according to the global control features. We perform weighted fusion to obtain the fused features:
[0106]
[0107] in, Indicates the characteristics after fusion. Indicates the self-attention layer. Represents the original features of ControlNet. The fusion weights are represented by hyperparameters. The settings allow users to manually adjust aggregation features to achieve the desired effect.
[0108] 4. The model training and inference generation module is used to construct a loss function to train the multi-object spatial condition image generation model in stages, taking the multi-object spatial condition dataset as input. The spatial conditions, multi-object layout information and text description input by the user are input into the trained multi-object spatial condition image generation model to generate multi-object spatial condition images.
[0109] In this embodiment, a multi-object spatial condition dataset is used to train the method proposed in this invention. Specifically, the spatial conditions in the multi-object spatial condition dataset are used... Global text hints in image text descriptions Bounding box and object-level text When constructing a loss function for a multi-object spatial conditional image generation model and training samples in stages, this invention uses the following training objective to predict noise:
[0110]
[0111] in, This represents the constructed loss function. Represents the features of the original input image. Represents a timestamp. Represents the noise latent vector. Represents real noise. This represents the predicted noise.
[0112] By gradually towards Add noise and obtain the noise latent vector Subsequently, this invention utilizes this learning objective for training, enabling the multi-object spatial conditional image generation model to reconstruct images that conform to the input. It is worth noting that both the ControlNet backbone network and the U-Net network will be frozen.
[0113] To improve the consistency between images and text when ControlNet handles aligned and unaligned spatial conditions, this invention designs a two-stage training strategy to optimize the multi-object spatial condition image generation model constructed in this invention.
[0114] In the first stage, a multi-object spatially conditional image generation model is trained using unaligned spatial conditions. However, when aligned spatial conditions are input, the original structural information may be lost.
[0115] Therefore, the second stage further refines the multi-object spatial conditional image generation model by using an aligned spatial condition to address this issue. Specifically, the multi-object spatial conditional image generation model is trained using a mix of 50% unaligned and 50% aligned spatial conditions. When handling the aligned spatial condition, bounding box features are initialized to zero to mitigate their potential impact on the original structural information. Both stages undergo 20 epochs of optimization (batch size 32, gradient accumulation 8) to ensure reliable performance improvements. The trained multi-object spatial conditional image generation model is then trained on ControlNet (StableDiffusion-v1.4) and ControlNet (StableDiffusionXL), respectively.
[0116] On the other hand, the present invention also provides an image generation method for multi-object spatial conditions, comprising the following steps:
[0117] Extract object-level bounding box information from the original image dataset, perform image editing within the given bounding box region, and extract aligned and unaligned spatial conditions from the original image and the image editing respectively. At the same time, extract the image text description of each image from the original image dataset to construct a multi-object spatial condition dataset.
[0118] Extract object-level text features from image text descriptions, extract object-level location features from object-level bounding box information, and concatenate object-level text features and object-level location features to obtain object-level semantic features.
[0119] Using the ControlNet backbone network as the basic framework, an object-level structural control submodule is introduced to fuse object-level semantic features and ControlNet original features to obtain object-level structural features. At the same time, an object-level control relaxation submodule is introduced to establish object-level control features. The similarity between object-level semantic features and ControlNet original features is measured. The ControlNet original features, object-level structural features and object-level control features are fused and the fused features are added to the original U-Net to generate images with multi-object spatial conditions.
[0120] Using a multi-object spatial condition dataset as input, a loss function is constructed to train the multi-object spatial condition image generation model in stages. The spatial conditions, multi-object layout information, and text description input by the user are fed into the trained multi-object spatial condition image generation model to generate multi-object spatial condition images.
[0121] To illustrate the feasibility and excellent generation effect of the image generation system and method for multi-object spatial conditions provided by the present invention, a comparison was made with the prior art.
[0122] First, a visual comparison of the results is presented: In this embodiment, the method proposed in this invention is compared with the latest layout-to-image generation methods and the ControlNet family of methods. L2I methods include MultiDiffusion (MD), BoxDiff, 3DIS, GLIGEN, MIGC, and InstanceDiffusion (InsDiff). ControlNet family of methods includes ControlNet (CNet), ControlNetMask (CNetMask), and SmartControl (SmaCtrl). For aligned spatial conditions, the embodiment combines existing L2I methods with CNet. For unaligned spatial conditions, the embodiment integrates L2I methods with CNetMask and SmaCtrl. All methods are implemented using official code and default configurations.
[0123] Visual comparison results of the generated images are as follows Figure 6As shown, compared to CNet, SmaCtrl encounters more severe semantic confusion when handling multi-object generation tasks. Despite global textual cues containing semantic information, the generated images are rendered incorrectly. The L2I method renders objects using misaligned visual input, which degrades image quality. Furthermore, due to the lack of spatial information to guide rendering, the rendering of specific objects does not match the realistic image. In contrast, our method performs better in all aspects. Notably, our method generates high-quality images while better preserving the pose information of objects in the spatial input. Our method's generation results on SDXL also demonstrate its superiority. Figure 7 ).
[0124] Furthermore, since this invention relies on the position of user-provided objects, it can be combined with a visual language model to directly predict this information. When the user inputs a global text prompt... and spatial conditions In this context, visual language models can leverage their understanding of spatial relationships to automatically identify the location of each object, regardless of whether those objects actually exist. By incorporating visual language models, users can more easily perform object-level image conditional generation, thereby enhancing the practicality and flexibility of the method (e.g., Figure 8 (As shown).
[0125] like Figure 9 As shown, in the depth map, green boxes represent local structures that match a given object description, while red boxes represent mismatched parts. The method provided by this invention achieves accurate text structure association while maintaining structural controllability (first row). Furthermore, it can adaptively control local structural information to modify local appearance and pose, thereby improving text-image consistency (second row).
[0126] To further demonstrate the superiority of the image generation method for multi-object spatial conditions provided by this invention, quantitative results were compared.
[0127] In this embodiment, a set of metrics is used to measure the performance of object-level image conditional generation:
[0128] i. Spatial fidelity: To evaluate whether the position of the generated object is consistent with the position of the object in the original spatial input conditions, Ground-DINO is used to predict the bounding box information of each object in the generated result, and the average accuracy (mAP) of all classes is calculated based on the GroundTruth bounding boxes of all classes.
[0129] ii. Structural Similarity: To preserve the structural information of the original space, the global mean square error (G-MSE) between the generated depth image and the original input depth image was calculated. If the spatial condition type is a Canny edge map or a Hed edge map, F1 scoring was used for structural similarity evaluation.
[0130] iii. Text Figure 1 Consistency: Text-image alignment quality is measured using CLIP scores (G-CLIP) and Local CLIP scores (L-CLIP).
[0131] iv. Image quality: FID was used to evaluate image quality.
[0132] Based on the aforementioned evaluation metrics, 3000 samples containing non-aligned spatial conditions were randomly selected from LAION-Aesthetics and COCO-2014 to form the LAION-3K and COCO-3K-R evaluation datasets. After data processing in the first module, each sample contains the original image, precise visual conditions, coarse visual conditions, bounding boxes, object category information, and global text descriptions.
[0133] Table 1 Quantitative comparison under misalignment depth conditions
[0134]
[0135] Table 2 Comparison results under Canny and Hed space conditions
[0136]
[0137] As shown in Table 1, the method provided in this invention achieves state-of-the-art performance on CLIP, MSE, and FID datasets. Because SmaCtrl relies on generative prior knowledge to relax structural features, it is prone to semantic confusion in complex scenes, leading to poor performance. Despite injecting MIGC or GLIGEN into U-Net, SmaCtrl still considers producing high-quality results challenging. This invention achieves results comparable to MIGC in terms of spatial fidelity and structural similarity. Although this invention may sacrifice spatial fidelity for higher image quality and structural similarity, it still achieves better results than GLIGEN. As shown in Table 2, compared with existing methods, the multi-object spatial condition image generation model provided in this invention demonstrates superior performance under different spatial conditions.
[0138] Furthermore, the terms "upper," "lower," "inner," "outer," "front," and "rear" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Unless otherwise specifically stated, the relative steps, numerical expressions, and values of components and steps described in these embodiments do not limit the scope of the invention. Of course, the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of the invention. All equivalent changes or modifications made to the structures, features, and principles described in the claims of this invention should be included within the scope of the claims of this invention.
[0139] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An image generation system oriented towards multi-object spatial conditions, characterized in that, Includes the following modules: The dataset construction module is used to extract object-level bounding box information from the original image dataset, perform image editing within a given bounding box region, extract aligned and unaligned spatial conditions from the original image and the image editing respectively, and extract the image text description of each image from the original image dataset to construct a multi-object spatial condition dataset. The feature extraction module is used to extract object-level text features from image text descriptions, extract object-level location features from object-level bounding box information, and concatenate object-level text features and object-level location features to obtain object-level semantic features. The multi-object spatial condition image generation model uses the ControlNet backbone network as the basic framework. It introduces an object-level structural control submodule to fuse object-level semantic features and ControlNet original features to obtain object-level structural features. At the same time, it introduces an object-level control relaxation submodule to establish object-level control features. It measures the similarity between object-level semantic features and ControlNet original features, fuses ControlNet original features, object-level structural features and object-level control features, and adds the fused features to the original U-Net to generate multi-object spatial condition images. The model training and inference generation module is used to construct a loss function to train the multi-object spatial condition image generation model in stages, taking the multi-object spatial condition dataset as input. The spatial conditions, multi-object layout information and text description input by the user are input into the trained multi-object spatial condition image generation model to generate images that conform to the multi-object spatial conditions.
2. The image generation system for multi-object spatial conditions according to claim 1, characterized in that, The constructed multi-object spatial condition dataset includes: original images, image text descriptions, object-level label lists, object-level bounding box information, and aligned and unaligned spatial conditions.
3. The image generation system for multi-object spatial conditions according to claim 2, characterized in that, In the feature extraction module, the extraction of object-level text features from the image text description includes: first, using a word segmenter to vectorize the image text description, and then using a text CLIP encoder. Extracting object-level text features using MLP layers; The extraction of object-level location features from object-level bounding box information includes: converting the object-level bounding box information into an object-level mask, and using an image CLIP encoder. The MLP layer extracts object-level location features.
4. The image generation system for multi-object spatial conditions according to claim 3, characterized in that, In the multi-object spatial condition image generation model, the object-level structure control submodule fuses object-level semantic features and ControlNet original features to obtain object-level structural features, including: The object-level structure control submodule inputs object-level semantic features and ControlNet raw features into the cross-attention layer to obtain object-level structure features, calculated as follows: , , In the formula, This represents the index of each object in the image. Represents object-level structural features. This indicates that the query features in the ControlNet backbone network are the original ControlNet features. Representing object-level semantic features, express The transformed object-level structure control submodule key features Representing feature dimension, Represents the transpose matrix. This represents the feature map of the ControlNet backbone network. express The transformed object-level structural control submodule value characteristics Represents the query matrix. This represents the weight matrix corresponding to the key features in the object-level structural control submodule. This represents the weight matrix corresponding to the value features in the object-level structural control submodule. Among them, object-level structural features are used to describe the correlation between object-level text descriptions and corresponding local spatial conditions, enabling precise control over image-related locations and text.
5. The image generation system for multi-object spatial conditions according to claim 4, characterized in that, In multi-object spatial conditional image generation models, the object-level control relaxation submodule establishes object-level control features and measures the similarity between object-level semantic features and ControlNet original features, including: The object-level control relaxation submodule controls by querying features Dot product operation is performed between the key features to calculate the attention similarity between the object-level semantic features and the original ControlNet features in the pixel dimension, and average pooling is performed on the features along the channel dimension to obtain the object-level control features: , In the formula, Indicates object-level control characteristics. This indicates the average pooling operation. This indicates that the query features in the ControlNet backbone network are the original ControlNet features. Representing feature dimension, Represents the transpose matrix. express The transformed object-level control relaxation submodule key features are obtained through... The weight matrix corresponding to the key features in the object-level control relaxation submodule is learned; Among them, object-level control features are used to measure the correlation between object-level semantic features and ControlNet original features, and to alleviate the conflict between object-level text description and corresponding local spatial conditions.
6. The image generation system for multi-object spatial conditions according to claim 5, characterized in that, In image generation models with multi-object spatial conditions, ControlNet raw features, object-level structural features, and object-level control features are fused, including: In object-level structural features A masked object-level structural feature Summing is performed, and then the results are fused through a self-attention layer to obtain the global structural features. , in, Represents global structural features. Indicates the self-attention layer. Represent each bounding box Process and obtain object-level masks; In object-level control features A masked object-level control feature Summation is performed, and then global control features are obtained through convolution and the sigmoid activation function: , in, Indicates global control characteristics. This represents the summation operation. Indicates the convolution operation; The global structural features and ControlNet original features are compared according to the global control features. We perform weighted fusion to obtain the fused features: , in, Indicates the characteristics after fusion. Indicates the self-attention layer. Represents the original features of ControlNet. This indicates the fusion weight.
7. The image generation system for multi-object spatial conditions according to claim 6, characterized in that, In the model training and inference generation module, the step of using a multi-object spatial condition dataset as input and constructing a loss function to train the multi-object spatial condition image generation model in stages includes: Spatial conditions in a multi-object spatial condition dataset Global text hints in image text descriptions Bounding box and object-level text Using the input as the model input, a loss function is constructed to train the multi-object spatial condition image generation model in stages. Noise is added to the input based on the multi-object spatial condition dataset, enabling the model to reconstruct an image that matches the input. The constructed loss function is expressed as follows: , in, This represents the constructed loss function. Represents the features of the original input image. Represents a timestamp. Represents the noise latent vector. Represents real noise. This represents the predicted noise.
8. The image generation system for multi-object spatial conditions according to claim 7, characterized in that, In the model training and inference generation module, the phased training includes: In the first stage, an image generation model with multiple object spatial conditions is trained using unaligned spatial conditions, enabling the model to learn object-level semantic features. In the second stage, a multi-object spatial conditional image generation model is trained using a combination of unaligned and aligned spatial conditions, with bounding box features initialized to zero to preserve the original structure.
9. An image generation method oriented towards multi-object spatial conditions, characterized in that, The image generation method uses the image generation system for multi-object spatial conditions as described in any one of claims 1 to 8, and includes the following steps: Extract object-level bounding box information from the original image dataset, perform image editing within the given bounding box region, and extract aligned and unaligned spatial conditions from the original image and the image editing respectively. At the same time, extract the image text description of each image from the original image dataset to construct a multi-object spatial condition dataset. Extract object-level text features from image text descriptions, extract object-level location features from object-level bounding box information, and concatenate object-level text features and object-level location features to obtain object-level semantic features. Using the ControlNet backbone network as the basic framework, an object-level structural control submodule is introduced to fuse object-level semantic features and ControlNet original features to obtain object-level structural features. At the same time, an object-level control relaxation submodule is introduced to establish object-level control features. The similarity between object-level semantic features and ControlNet original features is measured. The ControlNet original features, object-level structural features and object-level control features are fused and the fused features are added to the original U-Net to generate images with multi-object spatial conditions. Using a multi-object spatial condition dataset as input, a loss function is constructed to train the multi-object spatial condition image generation model in stages. The spatial conditions, multi-object layout information and text description input by the user are fed into the trained multi-object spatial condition image generation model to generate images that conform to the multi-object spatial conditions.
Citation Information
Patent Citations
Novel image generation method jointly driven by text and semantic segmentation map
CN117557683A
Method for controllably generating spherical image
CN118037606A