A method for controllable generation of multi-instance object images based on decoupled contour representation
Through the method based on decoupled contour representation, the problem of inaccurate object position and shape control in multi-object image generation is solved, and high-precision multi-instance object image generation is realized, which is suitable for industrial design and virtual scene construction and other fields.
Patent Information
- Application Number
- CN202510570635.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-06
AI Technical Summary
Existing image generation technology is difficult to accurately control the details of each object in multi-object scenes, resulting in unreasonable location and inaccurate shape of the object in the generated image, affecting the usability and authenticity of the image, especially in applications such as autonomous driving and virtual reality.
Using a method based on decoupled contour representation, the precise generation of multi-instance object images is achieved through data acquisition and organization, data augmentation, contour representation processing, contour fusion module calculation, global information injection, contour position coding and object space perception calculation, combined with layout diffusion model and guidance technology.
It realizes accurate representation and decoupling of the outline information of different objects in the same coordinate space, avoids object confusion and dislocation, and the generated images are rich in details and meet expectations. It is suitable for areas with high image accuracy requirements such as industrial design and virtual scene construction.
Smart Images

Figure CN120088363B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image generation, and particularly to a method for controllable generation of multi-instance object images based on decoupled contour representation. Background Art
[0002] In the field of computer vision, image generation has always been an important research direction that has received much attention, and its development process has witnessed the replacement and evolution of many technologies. In the early days, the generative adversarial network (GANs) series occupied an important position in image generation. Through the adversarial training mechanism of the generator and discriminator, GANs can generate images with a certain degree of authenticity. For example, certain achievements have been made in some simple image scenarios such as the generation of specific-style textures and the reconstruction of low-resolution images. However, GANs have problems such as unstable training and mode collapse, which limit their application in complex image generation tasks. The variational autoencoder (VAE) series is also a way of image generation. VAE attempts to learn the latent distribution of data and generate images through the encoding and decoding processes. VAE shows certain advantages in some application scenarios of data compression and generation, but there is still a gap between the details and quality of the generated images and the actual requirements.
[0003] In recent years, image generation technologies based on diffusion models have emerged and become a research hotspot. The basic principle of the diffusion model is to gradually add noise to the image until the image becomes pure Gaussian noise, and then reverse this process to convert the random noise back into an image. This mechanism enables the diffusion model to learn the complex distribution of images and demonstrates powerful capabilities in image generation tasks. Some advanced text-to-image models can generate images of complex scenarios such as "a unicorn running in the forest" and "a bustling street scene in a future city", meeting people's needs for customized image content to a certain extent.
[0004] However, in the multi-object generation scenario, existing technologies face many difficulties. Due to the ambiguity and limitations of text conditions, text-to-image models are difficult to precisely control the details of each object. For example, for a description like "there are fruits and a vase on the table", problems such as unreasonable fruit placement and inaccurate vase shape may occur. Although the "layout-image" related research uses layout diagrams for spatial modeling, as high-level semantic information, layout diagrams lack the ability to finely control the shape of objects, easily resulting in deviations in the shape of the generated objects. In practical applications, in fields that require a large number of multi-object images such as the simulation image generation of autonomous driving scenarios and the construction of virtual reality scenarios, these problems seriously affect the usability and authenticity of the images, and there is an urgent need for new technologies to break through these bottlenecks and achieve more precise and efficient generation of multi-instance object images. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a multi-instance object image controllable generation method based on decoupled contour representation with simple algorithm, high precision and efficiency.
[0006] The technical solution of the present invention to solve the above technical problems is: a multi-instance object image controllable generation method based on decoupled contour representation, comprising the following steps:
[0007] S1, data acquisition and arrangement: acquire image data from the dataset, and label the contour sequences and category information of each object in the image;
[0008] S2, data augmentation: perform multi-dimensional data augmentation on the acquired images;
[0009] S3, contour representation processing: preprocess and encode the contour sequences and category information to obtain a comprehensive embedding;
[0010] S4, pre-training of the contour encoder;
[0011] S5, contour fusion module calculation: input the comprehensive embedding obtained through contour representation processing into the contour fusion module composed of multiple layers of self-attention modules, and perform self-attention calculation according to the adjusted weights in each layer of the multiple layers of self-attention modules to promote information exchange and feature fusion between different contour objects;
[0012] S6, global information injection module fuses global and image features;
[0013] S7, contour position encoding: regard the image patch as a special square region with contour coordinates, so as to realize contour position encoding;
[0014] S8, object space perception calculation: in the object space perception module, adopt the cross-attention mechanism and combine self-attention calculation to enhance the interaction between different image patches;
[0015] S9, model construction: build the PolyDiffusion model with layout diffusion as the backbone network, and combine the guidance technology to control the generation effect;
[0016] S10, image generation: input the preprocessed contour sequences and category information into the trained PolyDiffusion model, and output the multi-instance object image.
[0017] In the above multi-instance object image controllable generation method based on decoupled contour representation, in the step S2, the process of data augmentation is:
[0018] At the image level, perform horizontal and vertical flipping operations; horizontal flipping is to mirror the image along the vertical central axis to generate a new image sample; vertical flipping is to perform a mirror flipping operation along the horizontal central axis.
[0019] For the contour representation, data augmentation is carried out. First, according to the model input requirements, the contour point sequence is fixed to a specific length. Then, using a random number generation algorithm, random offsets are introduced for some contour points. For each contour point to be adjusted , where x and y are the abscissa and ordinate of the contour point respectively, and random offsets are generated within the set range , and the new coordinates become , 、 are the offsets of the abscissa and ordinate respectively, causing the contour to deform.
[0020] For the above-mentioned controllable generation method of multi-instance object images based on decoupled contour representation, the specific process of step S3 is as follows:
[0021] S31: Load the contour sequence of the image. If the length of the contour sequence is less than the specified number N, interpolation algorithm is used to complete it;
[0022] S32: Preprocess the contour sequence. By calculating the distance between each point and the upper left corner, the point with the closest distance is adjusted to the starting point of the sequence to ensure the consistency of contour encoding;
[0023] S33: In the contour encoding link, a contour encoder module based on relative position is adopted. Each contour point is regarded as a token, and a center point is added to represent the global feature, forming a token sequence with a length of l + 1;
[0024] S34: Through the projection layer map the contour points to the specified dimension dim, represents a real number, is the feature dimension, and relative position embeddings are added for positions except the starting head token , , is the l-th element in;
[0025] S35: Send the token sequence into the contour encoder composed of multiple layers of self-attention and cross-attention. The self-attention layer focuses on the correlation of data at different coordinate points in the sequence, and the cross-attention layer focuses on the interaction between category data and the contour sequence. After multiple layers of interaction, the starting token represents the feature encoding of the object contour information;
[0026] S36: For the class encoder, use a simple linear layer Implementation; After performing linear layer operations on the class feature c and the contour feature and adding them in the channel dimension, the comprehensive embedding L is obtained, i.e., L = B + C, where B is the contour vector after passing through the contour encoder and C is the class vector. , , where CEM is the contour encoder module and P is the set of contour polygon sequences.
[0027] In the above multi-instance object image controllable generation method based on decoupled contour representation, in step S4, the pre-training tasks include contour sequence reconstruction, predicting the target frame, and object classification.
[0028] In the contour sequence reconstruction task, using the random masking technique, masks are constructed at a set proportion of coordinate points, and when calculating the attention, these masked tokens are ignored, thereby realizing self-supervised learning and enabling the model to deeply learn the shape and structure information of the contour.
[0029] The predicting target frame task requires the model to predict the corresponding target frame image based on the given contour information, thereby learning the context association information between the contour and the target frame.
[0030] The object classification task prompts the model to combine the contour with high-level semantic information, and through training with labeled data, enables the model to accurately identify the object categories corresponding to different contours.
[0031] In the above multi-instance object image controllable generation method based on decoupled contour representation, the specific process of step S5 is as follows:
[0032] S51: Take the comprehensive embedding obtained after contour representation processing as the input of the contour fusion module. represents the nth element in L. First, calculate the in each token for the of other tokens , and the calculation formula is:
[0033]
[0034] where , are calculation functions based on the attention mechanism, is the feature dimension;
[0035] S52: To better capture the relationships between objects, create an n×n relationship graph S based on the distance between the centers of two objects; when calculating the distance between the centers of objects, assume that the center coordinates of object are , and the center coordinates of object are , then the distance between the two center points , according to Construct the relationship graph S, in the relationship graph S represents and prior correlation;
[0036] S53: Adjust the original attention weights based on the relationship graph S, and the adjusted attention weights are , where is an adjustable parameter;
[0037] S54: After the adjustment is completed, re-normalize the weights to ensure that the sum of all weights is 1;
[0038] S55: Feed the adjusted comprehensive embedding into the contour fusion module composed of multiple self-attention modules. In the multiple self-attention modules, self-attention calculations are performed at each layer according to the adjusted weights to promote information exchange and feature fusion between different contour objects.
[0039] The above multi-instance object image controllable generation method based on decoupled contour representation, the specific process of step S6 is as follows:
[0040] S61: After completing the calculation of the contour fusion module, record the first token output by the contour fusion module as and extract it. Regard as the global information of the entire layout, while other tokens are regarded as individual object representations embedded with local information of related objects;
[0041] S62: Directly combine representing the global information of the contour map with image features at different resolutions; Assume that the representations of image features at different resolutions are , represent low, medium, and high resolution image features respectively. Project into the same dimensional space as the image features through the projection matrix W, and then perform an addition operation, that is , , , represent the low, medium, and high resolution image features after projection and addition operations respectively;
[0042] S63: Linearly combine the time step with the global information of the contour; Add the time step information to the image feature fusion result at each resolution to obtain the final global token, that is ; , , respectively represent the low, medium, and high-resolution image features after linear combination;
[0043] S64: Using the adaptive layer normalization technique, embed them into the corresponding resolution features respectively; Adaptive layer normalization dynamically adjusts the normalization parameters according to the statistical information of the input features, so that the global information is integrated into the image generation process, promoting the integration of local and global context details of the contour map in the image generation process.
[0044] For the above multi-instance object image controllable generation method based on decoupled contour representation, the specific process of step S7 is as follows:
[0045] Let represent the feature map of the height h, width w, and channels of the entire image. For each image patch in the image, define as the image patch located at the u-th row and v-th column of the image. Obtain the boundary point set , as the l-th sampling point. During the sampling process, determine the sampling interval according to the size and shape of the image patch to ensure that the sampling points are evenly distributed on the contour, so as to integrate the contour position and shape information of the image patch and the object into a unified space; The coordinate calculation method of the sampling point is:
[0046]
[0047] where is the original coordinate on the contour of the image patch.
[0048] For the above multi-instance object image controllable generation method based on decoupled contour representation, the specific process of step S8 is as follows:
[0049] S81: Perform self-attention calculation on the image patch, and calculate the attention scores of the query vector Q, key vector K, and value vector V of the image features respectively; When calculating , through the splicing function perform a splicing operation on the Q and K values of the image features and the image patch contour embedding in the channel dimension, that is , ; Directly take the V value of the image feature , where respectively represent the Q value, K value, and V value of the image feature, which are obtained by performing convolution and normalization operations on the image feature, that is , , Convolution and normalization operations are represented respectively, and I represents an image;
[0050] S82: Calculate the self-attention output , and the formula is:
[0051]
[0052] Through self-attention calculation, information interaction is carried out between different image patches;
[0053] S83: Calculate the cross-attention between the image patch and the contour embedding; first calculate the Q value of the cross-embedding , the K value of the cross-embedding , and the V value of the cross-embedding , , , , where is the image patch after self-attention calculation, obtained by performing convolution and normalization on , that is , are the K value and V value of the global contour embedding respectively, obtained by performing convolution operations on the contour embedding and C, that is , ; the contour embedding represents 's first token;
[0054] S84: Calculate the cross-attention map CA, and the formula is , and obtain the final output ;
[0055] S85: After the calculation is completed, collect cross-attention maps CA of different sizes, process these cross-attention maps, first calculate the average cross-attention map, that is, average the cross-attention maps of different sizes according to the set weights, and then perform normalization and aggregation operations on the averaged cross-attention map;
[0056] S86: Use the sigmoid function to convert the processed cross-attention map into an approximate binary mask, and calculate the DiceLoss of the loss function between this approximate binary mask and the true mask , and the form of the loss function is , where is the alignment loss, is the DiceLoss loss function, represents the cross-attention map after logical processing, , represents the logical function, represents the cross-attention map after upsampling operation, , represents the cross-attention map with a feature scale of , is the weight coefficient, is the upsampling function, and b is the number of feature scales;
[0057] S87: Integrate into the final loss function and optimize the final loss function to ensure the consistency of the generated image with the input conditions at the pixel level.
[0058] In the above multi-instance object image controllable generation method based on decoupled contour representation, in step S9, the PolyDiffusion model is built with layout diffusion as the backbone network. Layout diffusion constructs a standard U-net network structure based on the residual network Resnet and transformer, providing basic feature extraction and generation capabilities for the entire PolyDiffusion model;
[0059] According to the basic principle of the diffusion model and the defined objective function, the model is trained; during the training process, the AdamW optimizer is used, and the classifier-free guidance technique is adopted. By constructing an empty-set contour containing empty-set contour points , during training, the contour condition is randomly discarded with a set probability and replaced with an empty-set contour ; at the same time, in the fine-tuning stage, noise is added to the image generated by the pre-trained PolyDiffusion model through the diffusion forward operation to obtain the image ; when the added noise is small enough, predicting is equivalent to sampling the perturbed image once and only performing the last step of denoising. According to this principle, the reward loss , , where is the loss function, represents the image finally generated by the model, represents the real image, and is combined with the diffusion training loss , the alignment loss to form the final loss function :
[0060]
[0061]
[0062] where and They are all hyperparameters used to adjust the contribution degree of different loss terms to the overall loss. is the current time step, is a hyperparameter used to control the less than or equal to when calculation method; is the current time step the real noise added at represents the noise prediction neural network, represents the input at the current time step ; by continuously optimizing the final loss function, the parameters of the PolyDiffusion model are adjusted.
[0063] In the above multi-instance object image controllable generation method based on decoupled contour representation, in step S10, the preprocessed contour sequence and category information are input into the trained PolyDiffusion model. The PolyDiffusion model first performs contour representation processing on the input contour sequence, converts it into an effective feature encoding, and combines the category information to obtain a comprehensive embedding. The comprehensive embedding is successively processed by a contour fusion module, a global information injection module, a contour position encoding module, and an object space perception module. In the contour fusion module, information between different contour objects is interacted and fused. The global information injection module integrates global information into the image generation process. The contour position encoding module unifies the spatial representation of image patches and contours. The object space perception module further improves the consistency between the generated image and the input conditions at the pixel level through a cross-attention mechanism. After the collaborative processing of the contour fusion module, the global information injection module, the contour position encoding module, and the object space perception module, the PolyDiffusion model finally outputs a multi-instance object image.
[0064] The beneficial effects of the present invention are as follows: Through the encoding method based on the coordinates of contour boundary points, the present invention realizes the accurate representation and decoupling of different object contour information in the same coordinate space, enabling the model to clearly distinguish each object and avoiding common problems such as object confusion and misalignment in traditional methods. When generating a scene image containing multiple objects with different shapes, the contour of each object can be accurately presented, and its position and size can be precisely controlled. The generated image has rich details and meets expectations, and has great application potential in fields with high requirements for image accuracy such as industrial design and virtual scene construction. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 is the flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] The present invention will be further described below with reference to the drawings and embodiments.
[0067] As shown Figure 1 in the figure, a controllable generation method for multi-instance object images based on decoupled contour representation includes the following steps:
[0068] S1, Data acquisition and arrangement: Image data is collected from the COCO-stuff and VOC-2012 datasets. For the COCO-stuff dataset, images containing 1-9 non-human objects or containing 3-9 objects with only one person are selected, obtaining 28,292 training images and 1,682 validation images, among which 18,794 do not contain people. At the same time, the contour sequences and category information of each object in the images are labeled.
[0069] S2, Data augmentation: To improve the generalization ability of the model and its adaptability to contour changes, multi-dimensional data augmentation is performed on the collected images.
[0070] The process of data augmentation is as follows:
[0071] At the image level, horizontal and vertical flipping operations are performed; horizontal flipping is achieved by mirroring the image along the vertical central axis to generate new image samples; vertical flipping is the mirroring operation along the horizontal central axis; this not only expands the number of image samples but also enables the model to learn the feature representations of the image in different directions.
[0072] For the contour representation, data augmentation is carried out; first, according to the model input requirements, the contour point sequence is fixed to a specific length; then, using the random number generation algorithm, random offsets are introduced for some contour points. For each contour point to be adjusted , where x and y are the abscissa and ordinate of the contour point respectively, random offsets are generated within the set range , and the new coordinates become , , are the offsets of the abscissa and ordinate respectively, causing the contour to deform; in this way, the diversity of contour data is increased, enabling the model to learn the contour features under different deformation conditions and enhancing the processing ability for contour changes.
[0073] S3, Contour representation processing: The contour sequence and category information are preprocessed and encoded to obtain a comprehensive embedding.
[0074] The specific process of step S3 is as follows:
[0075] S31: Load the contour sequence of the image. If the length of the contour sequence is less than the specified number N, use an interpolation algorithm to complete it. Common interpolation methods such as linear interpolation are used to insert reasonable coordinate values at the positions of missing points according to the distribution law of existing contour points to ensure that the lengths of all contour sequences are the same.
[0076] S32: Preprocess the contour sequence. By calculating the distance of each point from the upper left corner, adjust the point with the closest distance to the starting point of the sequence to ensure the consistency of contour encoding.
[0077] S33: In the contour encoding step, use a contour encoder module based on relative positions. Treat each contour point as a token and add a center point to represent the global feature, forming a token sequence of length l + 1.
[0078] S34: Through the projection layer map the contour points to the specified dimension dim, representing real numbers, is the feature dimension, and add relative position embeddings for positions other than the starting head token , , is the l-th element in
[0079] S35: Feed the token sequence into a contour encoder composed of multiple layers of self-attention and cross-attention. The self-attention layer focuses on the correlation of data at different coordinate points in the sequence, and the cross-attention layer focuses on the interaction between category data and the contour sequence. After multiple interactions, the starting token represents the feature encoding of the object contour information.
[0080] S36: For the class encoder, use a simple linear layer to implement; after performing a linear layer operation on the class feature c and the contour feature and adding them in the channel dimension, obtain the comprehensive embedding L, that is, L = B + C, where B is the contour vector after passing through the contour encoder and C is the category vector. , , CEM is the contour encoder module, and P is the set of contour polygon sequences.
[0081] S4, Pre-train the contour encoder. The pre-training tasks include contour sequence reconstruction, predicting the target frame, and object classification.
[0082] In the contour sequence reconstruction task, use the random masking technique to construct masks at a set proportion of coordinate points. When calculating the attention, ignore these masked tokens to achieve self-supervised learning and enable the model to deeply learn the shape and structure information of the contour.
[0083] The task of predicting the target frame requires the model to predict the corresponding target frame image based on the given contour information, so as to learn the contextual association information between the contour and the target frame;
[0084] The object classification task then prompts the model to combine the contour with high-level semantic information. Through training with labeled data, the model can accurately identify the object categories corresponding to different contours, laying a solid foundation for subsequent image generation tasks.
[0085] S5. The contour fusion module calculates: Input the comprehensive embedding obtained after contour representation processing into the contour fusion module composed of multiple self-attention modules. In each layer of the multiple self-attention modules, self-attention calculation is performed according to the adjusted weights to promote information exchange and feature fusion between different contour objects.
[0086] The specific process of the step S5 is as follows:
[0087] S51: Take the comprehensive embedding obtained after contour representation processing as the input of the contour fusion module, denote the nth element in L. First, calculate the in each token for the original attention weights for , and the calculation formula is:
[0088]
[0089] where , are calculation functions based on the attention mechanism, is the feature dimension;
[0090] S52: To better capture the relationships between objects, create an n×n relationship graph S based on the distance between the centers of two objects. When calculating the distance between the centers of objects, assume that the center coordinates of object are , and the center coordinates of object are , then the distance between the two centers . Based on construct the relationship graph S, where represents and the prior correlation;
[0091] S53: Adjust the original attention weights based on the relationship graph S. The adjusted attention weight is , where is an adjustable parameter used to control the influence degree of prior knowledge on weight adjustment;
[0092] S54: After the adjustment is completed, re-normalize the weights to ensure that the sum of all weights is 1, making the weight distribution more reasonable;
[0093] S55: Feed the adjusted comprehensive embedding into the contour fusion module composed of multiple self-attention modules. In the multiple self-attention modules, self-attention calculation is performed at each layer according to the adjusted weights. In this way, promote the information exchange and feature fusion between different contour objects, enabling the model to better understand the relationships between objects in a multi-object scene and providing richer context information for subsequent image generation.
[0094] S6, The global information injection module fuses global and image features.
[0095] The specific process of the step S6 is as follows:
[0096] S61: After completing the calculation of the contour fusion module, record the first token output by the contour fusion module as and extract it. Regard as the global information of the entire layout, while other tokens are regarded as individual object representations embedded with local information of relevant objects;
[0097] S62: Directly combine representing the global information of the contour map with image features at different resolutions; Assume that the representations of the image features at different resolutions are , representing low, medium, and high resolution image features respectively. Project into the same dimensional space as the image features through the projection matrix W, and then perform an addition operation, that is , , , represent the low, medium, and high resolution image features after projection and addition operations respectively;
[0098] S63: Linearly combine the time step with the global information of the contour; Add the time step information to the fused result of the image features at each resolution to obtain the final global token, that is ; , , represent the low, medium, and high resolution image features after linear combination respectively;
[0099] S64: Use the adaptive layer normalization technique to Embed them into the corresponding resolution features respectively; Adaptive layer normalization dynamically adjusts the normalization parameters according to the statistical information of the input features, integrates the global information into the image generation process, promotes the integration of local and global context details of the contour map in the image generation process, and makes the generated image more reasonable and natural in terms of overall layout and details.
[0100] S7, contour position encoding: Regard the image patch as a special square area with contour coordinates to achieve contour position encoding.
[0101] The specific process of step S7 is as follows:
[0102] Let represent the height h, width w and channel feature map of the whole image. For each image patch in the image, define as the image patch located in the u-th row and v-th column of the image. Obtain the boundary point set , as the l-th sampling point. During the sampling process, determine the sampling interval according to the size and shape of the image patch to ensure that the sampling points are evenly distributed on the contour, so as to integrate the contour position and shape information of the image patch and the object into a unified space; for example, for a larger image patch, the number of sampling points can be appropriately increased; for a smaller image patch, the number of sampling points can be correspondingly reduced. The coordinate calculation method of the sampling point is:
[0103]
[0104] where is the original coordinate on the contour of the image patch.
[0105] The contour coordinates of the image patch share the same spatial representation with other contour tokens, replacing the traditional position encoding method. This unified spatial representation enables the model to more intuitively understand the spatial relationship between the image and the contour information when processing the image and the contour, so as to more accurately locate and shape the object when generating the image, improving the accuracy and quality of image generation.
[0106] S8, object space perception calculation: In the object space perception module, adopt the cross-attention mechanism and combine self-attention calculation to enhance the interaction between different image patches.
[0107] The specific process of step S8 is as follows:
[0108] S81: Perform self-attention calculation on the image patch, and calculate the attention scores of the query vector Q, key vector K, and value vector V of the image features respectively ; When calculating , through the splicing function Perform splicing operations on the Q and K values of the image features in the channel dimension and the image patch contour embedding That is , ; Directly take the V value of the image feature , that is , where respectively represent the Q value, K value, and V value of the image feature, which are obtained by performing convolution and normalization operations on the image feature, that is , , respectively represent the convolution and normalization operations, and I represents the image;
[0109] S82: Calculate the self-attention output , and the formula is:
[0110]
[0111] Through self-attention calculation, information interaction is carried out between different image patches, enabling the model to better capture the relationships between image patches;
[0112] S83: Calculate the cross-attention between the image patch and the contour embedding; first calculate the Q value of the cross-embedding , the K value of the cross-embedding , and the V value of the cross-embedding , , , , where is the image patch after self-attention calculation, which is obtained by performing convolution and normalization on , that is , are respectively the K value and V value of the global contour embedding, which are obtained by performing convolution operations on the contour embedding and C, that is , ; The contour embedding represents 's first token;
[0113] S84: Calculate the cross-attention map CA, and the formula is , and obtain the final output ;
[0114] S85: After the calculation is completed, collect the cross-attention maps CA of different sizes (8×8, 16×16, 32×32), process these cross-attention maps, first calculate the average cross-attention map, that is, average the cross-attention maps of different sizes according to the set weights, and then perform normalization and aggregation operations on the averaged cross-attention map;
[0115] S86: Use the sigmoid function to convert the processed cross-attention map into an approximate binary mask, and calculate the DiceLoss of the loss function between the approximate binary mask and the ground truth mask The form of the loss function is , where is the alignment loss, is the DiceLoss function, represents the cross-attention map after logical processing, , represents the logical function, represents the cross-attention map after the upsampling operation, , represents the cross-attention map with a feature scale of , is the weight coefficient, is the upsampling function, and b is the number of feature scales;
[0116] S87: Integrate into the final loss function and optimize the final loss function to ensure the consistency of the generated image with the input conditions at the pixel level.
[0117] S9, Model construction: Build the PolyDiffusion model with Layout Diffusion as the backbone network and combine the guidance technology to control the generation effect.
[0118] In step S9, build the PolyDiffusion model with Layout Diffusion as the backbone network. Layout Diffusion constructs a standard U-net network structure based on the residual network Resnet and the transformer, providing basic feature extraction and generation capabilities for the entire PolyDiffusion model;
[0119] Train the model according to the basic principle of the diffusion model and the defined objective function; during the training process, use the AdamW optimizer, and set the learning rate to , , are the estimated exponential decay rates at the first and second moments respectively; adopt the classifier-free guidance technology, and by constructing an empty set contour containing the contour points of the empty set , randomly discard the contour conditions with a set probability during training and replace them with the empty set contour , this method can learn both conditional and unconditional generation modes during the training process, reduce the training cost, and improve the generalization ability of the model. At the same time, according to the designed reward strategy, in the fine-tuning stage, add noise to the image generated by the pre-trained PolyDiffusion model through the diffusion forward operation Obtain an image ; When the added noise is small enough, predicting the original image is equivalent to performing one-step sampling on the perturbed image and only performing the last step of denoising. According to this principle, calculate the reward loss , , where is the loss function, represents the image finally generated by the model, represents the real image, and combine with the diffusion training loss , the alignment loss to form the final loss function :
[0120]
[0121]
[0122] where and are both hyperparameters used to adjust the contribution degree of different loss terms to the overall loss, is the current time step, is a hyperparameter used to control the calculation method when is less than or equal to ; ; is the real noise added at the current time step , represents the noise prediction neural network, represents the input at the current time step ; By continuously optimizing the final loss function, adjust the parameters of the PolyDiffusion model.
[0123] S10, Image generation: Input the preprocessed contour sequence and class information into the trained PolyDiffusion model to output a multi-instance object image.
[0124] In the step S10, the preprocessed contour sequence and category information are input into the trained PolyDiffusion model. The PolyDiffusion model first performs contour representation processing on the input contour sequence, converts it into an effective feature encoding, and combines the category information to obtain a comprehensive embedding. The comprehensive embedding is sequentially processed by a contour fusion module, a global information injection module, a contour position encoding module, and an object space perception module. In the contour fusion module, information between different contour objects is interacted and fused. The global information injection module integrates global information into the image generation process. The contour position encoding module unifies the spatial representation of image patches and contours. The object space perception module further improves the consistency between the generated image and the input conditions at the pixel level through a cross-attention mechanism. After the collaborative processing of the contour fusion module, the global information injection module, the contour position encoding module, and the object space perception module, the PolyDiffusion model finally outputs a multi-instance object image.
[0125] Table 1 is a performance comparison table of the method of the present invention and other methods (Grid2Im, PLGAN, InstanceDiffusion, LayoutDiffusion) when the COCOStuff dataset is selected. The COCOStuff dataset is constructed based on the COCO2017 dataset. It covers a total of 172 categories, including 80 "thing" categories, 91 "stuff" categories, and 1 "unlabeled" category.
[0126] Table 2 is a performance comparison table of the method of the present invention and other methods (Grid2Im, PLGAN, InstanceDiffusion, LayoutDiffusion) when the VOC-2012 dataset is selected. The VOC-2012 dataset is a key dataset in the Visual Object Classes Challenge. This dataset covers 20 object categories and 1 background category, with a total of 11,530 images, of which the training set and the validation set contain 10,582 images, and the test set has 948 images.
[0127]
[0128]
[0129] In Tables 1 and 2, FID is the Frechet inception distance, which is an index to measure the similarity between the distribution of generated images and the distribution of real images; IS is an evaluation index for the quality of generated images; CAS is the classification accuracy score; AP mask is the average precision, which is an evaluation index for the image segmentation task and is used to measure the matching degree between the generated mask and the real mask; ARmask is the average recall rate, which is used in conjunction with AP mask to measure the coverage ability of the generated mask for the true mask. As can be seen from Table 1 and Table 2, the method of the present invention has achieved good results in most metrics, especially in AR mask and AP mask metrics, achieving the best results.
[0130] The multi-instance object image controllable generation method based on decoupled contour representation proposed by the present invention can effectively solve many problems faced by traditional image generation technologies in dealing with complex multi-object scenarios, bringing significant technical improvements and innovative application values to the field of image generation.
Claims
1. A method for controllably generating multi-instance object images based on decoupled contour representation, characterized in that It includes the following steps: S1, Data collection and collation: Collect image data from the dataset and label the contour sequences and category information of each object in the images; S2, Data augmentation: Perform multi-dimensional data augmentation on the collected images; S3, Contour representation processing: Preprocess and encode the contour sequences and category information to obtain a comprehensive embedding; The specific process of step S3 is as follows: S31: Load the contour sequence of the image. If the length of the contour sequence is less than the specified number N, use an interpolation algorithm to complete it; S32: Preprocess the contour sequence. By calculating the distance between each point and the upper left corner, adjust the point with the closest distance to the starting point of the sequence to ensure the consistency of contour encoding; S33: In the contour encoding link, use a contour encoder module based on relative positions. Treat each contour point as a token and add a center point to represent global features, forming a token sequence with a length of l + 1; S34: Through the projection layer Map the contour points to the specified dimension dim, where represents a real number, d is the feature dimension, and relative position embeddings R = {r1, r2, …, r l} are added for positions other than the starting head token, and r l is the l-th element in R; S35: Send the token sequence into a contour encoder composed of multiple layers of self-attention and cross-attention. The self-attention layer focuses on the correlation of data at different coordinate points in the sequence, and the cross-attention layer focuses on the interaction between category data and the contour sequence. After multiple layers of interaction, the starting token represents the feature encoding of the object contour information; S36: For the class encoder, use a simple linear layer to implement; after performing linear layer operations on the class feature c and the contour feature, add them in the channel dimension to obtain the comprehensive embedding L, that is, L = B + C, where B is the contour vector passed through the contour encoder, C is the class vector, C = cW c , B = CEM((P × W B ) + R, C), CEM is the contour encoder module, and P is the set of contour polygon sequence; S4, Pre-training of the contour encoder; S5, Calculation of the contour fusion module: Input the comprehensive embedding obtained through contour representation processing into a contour fusion module composed of multiple layers of self-attention modules. In each layer of the multiple layers of self-attention modules, perform self-attention calculations according to the adjusted weights to promote information exchange and feature fusion between different contour objects; S6, Global information injection module fuses global and image features; S7, Contour position encoding: Regard the image patch as a special square area with contour coordinates to achieve contour position encoding; S8, Object space perception calculation: In the object space perception module, use a cross-attention mechanism and combine self-attention calculations to enhance the interaction between different image patches; S9, Model construction: Build the PolyDiffusion model with layout diffusion as the backbone network and combine guiding techniques to control the generation effect; S10, Image generation: Input the preprocessed contour sequences and category information into the trained PolyDiffusion model to output multi-instance object images.
2. The method for controllably generating multi-instance object images based on decoupled contour representation according to claim 1, wherein In step S2, the process of data augmentation is as follows: At the image level, perform horizontal and vertical flipping operations; Horizontal flipping is to generate a new image sample by mirroring the image along the vertical central axis; Vertical flipping is to perform a mirroring operation along the horizontal central axis; For contour representation, data augmentation is carried out. First, according to the model input requirements, the contour point sequence is fixed to a specific length. Then, using a random number generation algorithm, random offsets are introduced for some contour points. For each contour point (x, y) that needs to be adjusted, where x and y are the abscissa and ordinate of the contour point respectively, random offsets (Δx, Δy) are generated within a set range, and the new coordinates become (Δx + x, Δy + y), where Δx and Δy are the offsets of the abscissa and ordinate respectively, causing the contour to deform.
3. The multi-instance object image controllable generation method based on decoupled contour representation according to claim 1, wherein In step S4, the pre-training tasks include contour sequence reconstruction, predicting the target frame, and object classification. In the contour sequence reconstruction task, using the random masking technique, masks are constructed at a set proportion of coordinate points. When calculating the attention, these masked tokens are ignored, thereby achieving self-supervised learning and enabling the model to deeply learn the shape and structure information of the contour. The predicting the target frame task requires the model to predict the corresponding target frame image based on the given contour information, thereby learning the context association information between the contour and the target frame. The object classification task prompts the model to combine the contour with high-level semantic information. Through training with labeled data, the model can accurately identify the object categories corresponding to different contours.
4. The multi-instance object image controllable generation method based on decoupled contour representation according to claim 1, wherein The specific process of step S5 is as follows: S51: Use the comprehensive embedding L = (O1, O2,..., O n ) obtained through contour representation processing as the input of the contour fusion module. O n represents the nth element in L. First, calculate the O in each token j for the O in other tokens i of the original attention weight a ji , and the calculation formula is: Among them, Q_proj and K_proj are computational functions based on the attention mechanism, and d is the feature dimension. S52: To better capture the relationships between objects, an n×n relationship graph S is created based on the distance between the center points of two objects. When calculating the distance from the center point of the object to be calculated, assume the object O i The center point coordinates of are For object O j The center point coordinates of are Then the distance between the two center points According to d ij Construct the relationship graph S. In the relationship graph S, S ij represents the prior correlation between O i and O j ; S53: Adjust the original attention weights based on the relationship graph S, and the adjusted attention weights are a ji = a ji × (1 + λ × S ij ), where λ is an adjustable parameter; S54: After the adjustment, the weights are re-normalized to ensure that the sum of all weights is 1. S55: The adjusted comprehensive embedding is sent into the contour fusion module composed of multiple self-attention modules. In the multiple self-attention modules, each layer performs self-attention calculation according to the adjusted weights, promoting information exchange and feature fusion between different contour objects.
5. The method for controllably generating a multi-instance object image based on decoupled contour representation according to claim 4, wherein The specific process of step S6 is as follows: S61: After completing the calculation of the contour fusion module, the first token output by the contour fusion module is denoted as O1 and extracted. O1 is regarded as the global information of the entire layout, while the other tokens are regarded as individual object representations embedded with local information of related objects. S62: directly combine O′1 representing the global information of the contour map with image features of different resolutions; assume the representations of the image features at different resolutions are I low 、I mid 、I high ,I low 、I mid 、I high represent the low, medium, and high-resolution image features respectively. Project O′1 onto the same dimensional space as the image features through the projection matrix W, and then perform an addition operation, that is, I′ low =I low +O′1W, I′ mid =I mid +O′1W, I′ high =I high +O′1W, I′ low 、I′ mid 、I′ high represent the low, medium, and high-resolution image features after the projection and addition operations respectively; S63: Linearly combine the time step with the global information of the contour; add the time step information t to the image feature fusion result at each resolution to obtain the final global token, i.e., I″ low = I′ low + t, I″ mid = I′ mid + t, I″ high = I′ high + t; I″ low 、I″ mid 、I″ high respectively represent the low, medium, and high-resolution image features after linear combination; S64: Using the adaptive layer normalization technique, embed I″ low , I″ mid , I″ high into the corresponding resolution features respectively; Adaptive layer normalization dynamically adjusts the normalization parameters according to the statistical information of the input features, enabling the global information to be integrated into the image generation process and promoting the integration of local and global context details of the contour map in the image generation process.
6. The multi-instance object image controllable generation method based on decoupled contour representation according to claim 5, wherein, The specific process of step S7 is as follows: Let \(I\) denote the height \(h\), width \(w\), and channel feature map of the entire image. For each image patch in the image, define \(I(u, v)\) as the image patch located at the \(u\)-th row and \(v\)-th column of the image, and obtain the set of boundary points \(S\) by uniformly sampling from the image patch contour uv =\(\{m_1, m_2, \ldots, m\) l \}\), where \(m\) l is the \(l\)-th sampling point. During the sampling process, determine the sampling interval according to the size and shape of the image patch to ensure that the sampling points are evenly distributed on the contour, thereby integrating the contour position and shape information of the image patch and the object into a unified space; the coordinate calculation method of the sampling points is as follows: where (x l , y l ) is the original coordinate of m l on the contour of the image block.
7. The method for controllable generation of multi-instance object images based on decoupled contour representation according to claim 5, wherein The specific process of step S8 is as follows: S81: Perform self-attention calculation on the image block, and calculate the attention scores S of the query vector Q, key vector K, and value vector V of the image feature respectively. Q , S K , S V ; Calculate S Q ,S K When the splicing function Φ is used to embed the Q and K values of the image features and the image block contour into P in the channel dimension I Perform splicing operation, that is, S Q =Φ(Q I ,P I ), S K =Φ(K I ,P I );S V Directly take the V value V of the image feature I , that is, S V =V I , where Q I ,K I ,V I They represent the Q value, K value, and V value of the image features, respectively, which are obtained by convolution and normalization of the image features, that is, Q I ,K I ,V I =Conv(Norm(I)), Conv and Norm represent convolution and normalization operations respectively, and I represents the image; S82: Calculate the self-attention output O′, and the formula is: Through self-attention calculation, information interaction occurs between different image patches. S83: Calculate the cross-attention of the image patch and the contour embedding; first calculate the Q value C of the cross-embedding Q , the K value C of the cross-embedding K , the V value C of the cross-embedding V , C Q = Φ(Q U , P I ), C K = Φ(K L , P I ), C V = V L , where Q U is the image patch after self-attention calculation, obtained by convolving and normalizing O′, i.e., Q U = Conv(Norm(O′)), K L , V L are the K value and V value of the global contour embedding respectively, obtained by convolving the contour embedding L′ and C, i.e., K L = Conv(L′), V L = Conv(Norm(C)); the contour embedding L′ represents the first token of L; S84: Calculate the cross-attention map CA, with the formula and obtain the final output O = CA · C V ; S85: After the calculation is completed, collect the cross-attention maps CA of different sizes, process these cross-attention maps, first calculate the average cross-attention map, that is, average the cross-attention maps of different sizes according to the set weights, and then perform normalization and aggregation operations on the averaged cross-attention map. S86: Convert the processed cross-attention map into an approximate binary mask using the sigmoid function, and calculate the DiceLoss of the approximate binary mask and the true mask mask T The loss function DiceLoss has the form where Loss2 is the alignment loss, Dice is the DiceLoss function, represents the cross-attention map after logical processing, LogisticFun represents the logistic function, represents the cross-attention map after the upsampling operation, CA m′ represents the cross-attention map with a feature scale of m′, α m′ is the weight coefficient, Up is the upsampling function, and b is the number of feature scales; S87: Integrate Loss2 into the final loss function and optimize the final loss function to ensure the consistency of the generated image with the input conditions at the pixel level.
8. The method for controllably generating multi-instance object images based on decoupled contour representation according to claim 7, wherein In step S9, the PolyDiffusion model is built with layout diffusion as the backbone network. Layout diffusion constructs a standard U-net network structure based on the residual network Resnet and the transformer, providing the basic feature extraction and generation capabilities for the entire PolyDiffusion model. The model is trained according to the basic principles of the diffusion model and the defined objective function; during the training process, the AdamW optimizer is used, and the classifier-free guidance technique is adopted. By constructing an empty set contour c containing the contour points of the empty set, φ during training, the contour condition is randomly discarded with a set probability and replaced with the empty set contour c φ ; meanwhile, in the fine-tuning stage, noise ∈ is added to the image x0 generated by the pre-trained PolyDiffusion model through a diffusion forward operation to obtain the image When the added noise ∈ is small enough, predict is equivalent to performing one-step sampling on the perturbed image and only performing the last step of denoising. According to this principle, the reward loss Loss3 is calculated, where MSE is the MSE loss function, represents the image finally generated by the model, Img represents the real image, and Loss3 is combined with the diffusion training loss Loss1 and the alignment loss Loss2 to form the final loss function Loss: Loss1 = ||∈ t -∈ θ (x t ,t)|| 2 where both λ1 and λ2 are hyperparameters used to adjust the contribution degree of different loss terms to the overall loss, t is the current time step, t thre is a hyperparameter used to control the calculation method of Loss when t is less than or equal to t thre ; ∈ t is the real noise added at the current time step t, ∈ θ represents the noise prediction neural network, x t represents the input at the current time step t; by continuously optimizing the final loss function, the parameters of the PolyDiffusion model are adjusted.
9. The method for controllably generating a multi-instance object image based on a decoupled contour representation according to claim 8, wherein In step S10, the preprocessed contour sequence and class information are input into the trained PolyDiffusion model. The PolyDiffusion model first processes the input contour sequence for contour representation, converting it into an effective feature encoding, and combines the class information to obtain a comprehensive embedding. The comprehensive embedding is successively processed by a contour fusion module, a global information injection module, a contour position encoding module, and an object space perception module. In the contour fusion module, information between different contour objects is interacted and fused. The global information injection module integrates global information into the image generation process. The contour position encoding module unifies the spatial representation of image patches and contours. The object space perception module further improves the consistency of the generated image with the input conditions at the pixel level through the cross-attention mechanism. After the collaborative processing of the contour fusion module, the global information injection module, the contour position encoding module, and the object space perception module, the PolyDiffusion model finally outputs a multi-instance object image.
Citation Information
Patent Citations
Remote sensing image building extraction and contour optimization method based on deep learning
CN113516135A
Marine multi-target ship instance segmentation method
CN117197452A