Multi-instance object image controllable generation method based on decoupling contour representation

Through the multi-instance object image controllable generation method based on the decoupled contour representation, the PolyDiffusion model is constructed, which solves the problem of difficulty in object detail control in the multi-instance object generation scenario, and realizes precise control and high-quality generation of multi-instance object images.

CN120088363AActive Publication Date: 2025-06-03HUNAN UNIV OF SCI & TECH

Patent Information

Application Number
CN202510570635.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-03
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

Existing image generation techniques are difficult to accurately control the details of each object in multi-object generation scenarios, resulting in deviations in the shape and layout of the generated images, affecting the usability and authenticity of the images.

Method used

The multi-instance object image controllable generation method based on decoupled contour representation is adopted. Through steps such as data acquisition and organization, data enhancement, contour representation processing, contour fusion module calculation, global information injection module fusion, contour position coding and object space perception calculation, the PolyDiffusion model is constructed to generate multi-instance object images.

Benefits of technology

Accurate control of images of multiple instance objects is achieved, ensuring that the outline of each object is accurately presented, and the position and size can be accurately controlled. The generated image is rich in details and meets expectations, and is suitable for areas with high requirements for image accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088363A_ABST
    Figure CN120088363A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-instance object image controllable generation method based on decoupling contour representation, and belongs to the field of image generation, and the method comprises the following steps: data collection and arrangement; data enhancement; contour representation processing; pre-training a contour encoder; calculating by a contour fusion module; the global information injection module fuses global and image features; contour position coding; object space perception calculation; constructing a model; and generating an image. According to the invention, through a coding mode based on contour boundary point coordinates, accurate representation and decoupling of contour information of different objects in the same coordinate space are realized, so that the model can clearly distinguish each object, and common problems of object confusion, dislocation and the like in a traditional method are avoided. When a scene image containing a plurality of objects in different shapes is generated, the contour of each object can be accurately presented, the position and the size of each object can be accurately controlled, and the generated image is rich in details and accords with expectation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image generation, and particularly to a method for controllably generating multi-instance object images based on decoupled contour representation. Background Art

[0002] In the field of computer vision, image generation has always been an important research direction that has received much attention, and its development process has witnessed the replacement and evolution of many technologies. In the early days, the generative adversarial network (GANs) series occupied an important position in image generation. Through the adversarial training mechanism of the generator and the discriminator, GANs can generate images with a certain degree of authenticity. For example, certain achievements have been made in some simple image scenarios such as the generation of specific-style textures and the reconstruction of low-resolution images. However, GANs have problems such as unstable training and mode collapse, which limit their application in complex image generation tasks. The variational autoencoder (VAE) series is also a way of image generation. VAE attempts to learn the latent distribution of data and generate images through the encoding and decoding processes. VAE shows certain advantages in some application scenarios of data compression and generation, but there is still a gap between the details and quality of the generated images and the actual requirements.

[0003] In recent years, image generation technologies based on diffusion models have emerged and become a research hotspot. The basic principle of the diffusion model is to gradually add noise to the image until the image becomes pure Gaussian noise, and then reverse this process to convert the random noise back into an image. This mechanism enables the diffusion model to learn the complex distribution of images and demonstrates powerful capabilities in image generation tasks. Some advanced text-to-image models can generate images of complex scenes such as "a unicorn running in the forest" and "a bustling street scene in a future city", which to a certain extent meet the needs of people for customized image content.

[0004] However, in the multi-object generation scenario, existing technologies face many dilemmas. Due to the ambiguity and limitations of text conditions, text-to-image models are difficult to precisely control the details of each object. For example, for a description like "there are fruits and a vase on the table", problems such as unreasonable placement of fruits and inaccurate vase shapes may occur. Although the "layout-image" related research uses layout diagrams for spatial modeling, as high-level semantic information, layout diagrams lack the ability to finely control the shape of objects, which easily leads to deviations in the shapes of the generated objects. In practical applications, in fields that require a large number of multi-object images such as the simulation image generation of autonomous driving scenarios and the construction of virtual reality scenarios, these problems seriously affect the usability and authenticity of the images, and there is an urgent need for new technologies to break through these bottlenecks and achieve more precise and efficient generation of multi-instance object images. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a multi-instance object image controllable generation method based on decoupled contour representation, which has a simple algorithm, is accurate and efficient.

[0006] The technical solution of the present invention to solve the above technical problems is: a multi-instance object image controllable generation method based on decoupled contour representation, including the following steps: S1, data collection and arrangement: Collect image data from the dataset, and label the contour sequence and category information of each object in the image; S2, data augmentation: Perform multi-dimensional data augmentation on the collected images; S3, contour representation processing: Preprocess and encode the contour sequence and category information to obtain a comprehensive embedding; S4, pre-training of the contour encoder; S5, contour fusion module calculation: Input the comprehensive embedding obtained through contour representation processing into the contour fusion module composed of multiple layers of self-attention modules. In each layer of the multiple layers of self-attention modules, self-attention calculation is performed according to the adjusted weights to promote information exchange and feature fusion between different contour objects; S6, global information injection module fuses global and image features; S7, contour position encoding: Regard the image patch as a special square area with contour coordinates, so as to realize contour position encoding; S8, object space perception calculation: In the object space perception module, adopt the cross-attention mechanism and combine self-attention calculation to enhance the interaction between different image patches; S9, model construction: Build the PolyDiffusion model with layout diffusion as the backbone network and combine the guidance technology to control the generation effect; S10, image generation: Input the preprocessed contour sequence and category information into the trained PolyDiffusion model, and output a multi-instance object image.

[0007] In the above multi-instance object image controllable generation method based on decoupled contour representation, in the step S2, the process of data augmentation is as follows: At the image level, perform horizontal and vertical flipping operations; Horizontal flipping is to mirror the image along the vertical central axis to generate a new image sample; Vertical flipping is to perform a mirror flipping operation along the horizontal central axis; For the contour representation, perform data augmentation; First, according to the model input requirements, fix the contour point sequence to a specific length; Then, use the random number generation algorithm to introduce random offsets for some contour points. For each contour point that needs to be adjusted , where x and y are the abscissa and ordinate of the contour point respectively, and random offsets are generated within the set range , the new coordinates become , 、 are the offsets of the abscissa and ordinate respectively, which cause the contour to deform.

[0008] In the above multi-instance object image controllable generation method based on decoupled contour representation, the specific process of step S3 is as follows: S31: Load the contour sequence of the image. If the length of the contour sequence is less than the specified number N, use the interpolation algorithm to complete it; S32: Preprocess the contour sequence. By calculating the distance between each point and the upper left corner, adjust the point with the closest distance to the starting point of the sequence to ensure the consistency of contour encoding; S33: In the contour encoding link, use a contour encoder module based on relative position. Treat each contour point as a token, and add a center point to represent the global feature to form a token sequence of length l + 1; S34: Through the projection layer map the contour points to the specified dimension dim, representing real numbers, is the feature dimension, and add relative position embeddings for positions except the starting head token , , is the l-th element in; S35: Send the token sequence into a contour encoder composed of multiple layers of self-attention and cross-attention. The self-attention layer focuses on the correlation of data at different coordinate points in the sequence, and the cross-attention layer focuses on the interaction between category data and the contour sequence. After multiple layers of interaction, the starting token represents the feature encoding of the object contour information; S36: For the class encoder, use a simple linear layer to implement; after performing a linear layer operation on the class feature c and the contour feature, add them in the channel dimension to obtain the comprehensive embedding L, that is, L = B + C, where B is the contour vector after passing through the contour encoder and C is the category vector, , , CEM is the contour encoder module, and P is the set of contour polygon sequences.

[0009] In the above multi-instance object image controllable generation method based on decoupled contour representation, in step S4, the pre-training tasks include contour sequence reconstruction, predicting the target frame, and object classification; In the contour sequence reconstruction task, use the random masking technique to construct masks on a set proportion of coordinate points. When calculating attention, ignore these masked tokens to achieve self-supervised learning and enable the model to deeply learn the shape and structure information of the contour; The prediction target frame task requires the model to predict the corresponding target frame image based on the given contour information, so as to learn the context association information between the contour and the target frame; The object classification task prompts the model to combine the contour with high-level semantic information. Through training with labeled data, the model can accurately identify the object categories corresponding to different contours.

[0010] For the above multi-instance object image controllable generation method based on decoupled contour representation, the specific process of step S5 is as follows: S51: Use the comprehensive embedding obtained after contour representation processing as the input of the contour fusion module. Denote the nth element in L. First, calculate the in each token for the original attention weights for other tokens. The calculation formula is:

[0011] where , are computational functions based on the attention mechanism, is the feature dimension; S52: To better capture the relationships between objects, create an n×n relationship graph S based on the distance between the center points of two objects; when calculating the distance between the center points of objects, assume that the center point coordinates of object are , and the center point coordinates of object are , then the distance between the two center points . Based on , construct the relationship graph S. In the relationship graph S, represents and 's prior correlation; S53: Adjust the original attention weights based on the relationship graph S. The adjusted attention weight is , where is an adjustable parameter; S54: After the adjustment, perform re-normalization on the weights to ensure that the sum of all weights is 1; S55: Send the adjusted comprehensive embedding into the contour fusion module composed of multiple self-attention modules. In the multiple self-attention modules, each layer performs self-attention calculations according to the adjusted weights to promote information exchange and feature fusion between different contour objects.

[0012] The above-mentioned controllable generation method for multi-instance object images based on decoupled contour representation, the specific process of step S6 is as follows: S61: After completing the calculation of the contour fusion module, record the first token output by the contour fusion module as and extract it. Regard as the global information of the entire layout, while other tokens are regarded as individual object representations embedded with local information of relevant objects; S62: Directly combine representing the global information of the contour map with image features of different resolutions; Assume the representations of the image features at different resolutions are , representing the low, medium, and high-resolution image features respectively. Project onto the same dimensional space as the image features through the projection matrix W, and then perform an addition operation, that is , , , represent the low, medium, and high-resolution image features after projection and addition operations respectively; S63: Linearly combine the time step with the global information of the contour; Add the time step information to the image feature fusion results at each resolution to obtain the final global token, that is ; , , represent the low, medium, and high-resolution image features after linear combination respectively; S64: Use the adaptive layer normalization technique to embed into the corresponding resolution features respectively; Adaptive layer normalization dynamically adjusts the normalization parameters according to the statistical information of the input features, so that the global information is integrated into the image generation process, and promotes the integration of local and global context details of the contour map in the image generation process.

[0013] The above-mentioned controllable generation method for multi-instance object images based on decoupled contour representation, the specific process of step S7 is as follows: Let represent the height h, width w, and channel feature map of the entire image. For each image patch in the image, define as the image patch located in the u-th row and v-th column of the image. Uniformly sample the boundary point set from the image patch contour, For the l-th sampling point, during the sampling process, the sampling interval is determined according to the size and shape of the image patch to ensure that the sampling points are evenly distributed on the contour, so as to integrate the position and shape information of the image patch and the object contour into a unified space; the coordinate calculation method of the sampling point is as follows:

[0014] where is the original coordinate on the contour of the image patch.

[0015] For the above-mentioned multi-instance object image controllable generation method based on decoupled contour representation, the specific process of step S8 is as follows: S81: Perform self-attention calculation on the image patch, and calculate the attention scores of the query vector Q, key vector K, and value vector V of the image features ; When calculating , through the splicing function perform a splicing operation on the Q and K values of the image features and the image patch contour embedding in the channel dimension, that is , ; Directly take the V value of the image feature , that is , where respectively represent the Q value, K value, and V value of the image feature, which are obtained by performing convolution and normalization operations on the image feature, that is , , respectively represent convolution and normalization operations, and I represents the image; S82: Calculate the self-attention output , and the formula is:

[0016] Through self-attention calculation, information interaction is carried out between different image patches; S83: Calculate the cross-attention between the image patch and the contour embedding; first calculate the Q value of the cross-embedding, the K value of the cross-embedding, and the V value of the cross-embedding, , , , where is the image patch after self-attention calculation, which is obtained by performing convolution and normalization on , that is , are respectively the K value and V value of the global contour embedding, which are obtained by performing convolution operations on the contour embedding and C, that is , ; contour embedding represents the first token of; S84: Calculate the cross-attention map CA, with the formula and obtain the final output ; S85: After the calculation is completed, collect the cross-attention maps CA of different sizes, process these cross-attention maps, first calculate the average cross-attention map, that is, average the cross-attention maps of different sizes according to the set weights, and then perform normalization and aggregation operations on the averaged cross-attention map; S86: Use the sigmoid function to convert the processed cross-attention map into an approximate binary mask, and calculate the loss function DiceLoss of this approximate binary mask and the true mask The loss function is in the form of where is the alignment loss, is the DiceLoss function, represents the cross-attention map after logical processing, , represents the logical function, represents the cross-attention map after the upsampling operation, , represents the feature scale as the cross-attention map of, is the weight coefficient, is the upsampling function, b is the number of feature scales; S87: Integrate into the final loss function and optimize the final loss function to ensure the consistency of the generated image with the input conditions at the pixel level.

[0017] In the above multi-instance object image controllable generation method based on decoupled contour representation, in step S9, a PolyDiffusion model is built with layout diffusion as the backbone network. Layout diffusion constructs a standard U-net network structure based on the residual network Resnet and the transformer, providing basic feature extraction and generation capabilities for the entire PolyDiffusion model; Train the model according to the basic principle of the diffusion model and the defined objective function; during the training process, use the AdamW optimizer, adopt the classifier-free guidance technique, and construct an empty set contour containing empty set contour points , and randomly discard the contour conditions with a set probability during training and replace them with the empty set contour ; Meanwhile, in the fine-tuning stage, noise is added to the images generated by the pre-trained PolyDiffusion model through diffusion forward operations to obtain the images ; When the added noise is small enough, the prediction is equivalent to taking one step of sampling on the perturbed image and only performing the last step of denoising. Based on this principle, the reward loss is calculated, , , where is the loss function, represents the image finally generated by the model, represents the real image. By combining with the diffusion training loss and the alignment loss , the final loss function is formed:

[0018]

[0019] where and are both hyperparameters used to adjust the contribution degree of different loss terms to the overall loss. is the current time step, is a hyperparameter used to control the calculation method when is less than or equal to ; ; is the real noise added at the current time step , represents the noise prediction neural network, represents the input at the current time step ; By continuously optimizing the final loss function, the parameters of the PolyDiffusion model are adjusted.

[0020] The above-mentioned multi-instance object image controllable generation method based on decoupled contour representation. In step S10, the preprocessed contour sequence and category information are input into the trained PolyDiffusion model. The PolyDiffusion model first performs contour representation processing on the input contour sequence, converts it into an effective feature encoding, and combines the category information to obtain a comprehensive embedding. The comprehensive embedding is sequentially processed through a contour fusion module, a global information injection module, a contour position encoding module, and an object space perception module. In the contour fusion module, information between different contour objects is interacted and fused. The global information injection module integrates global information into the image generation process. The contour position encoding module unifies the spatial representation of image patches and contours. The object space perception module further improves the consistency between the generated image and the input conditions at the pixel level through a cross-attention mechanism. After the collaborative processing of the contour fusion module, the global information injection module, the contour position encoding module, and the object space perception module, the PolyDiffusion model finally outputs a multi-instance object image.

[0021] The beneficial effects of the present invention are as follows: Through the encoding method based on the coordinates of contour boundary points, the present invention realizes the accurate representation and decoupling of the contour information of different objects in the same coordinate space, enabling the model to clearly distinguish each object and avoiding common problems such as object confusion and misalignment in traditional methods. When generating a scene image containing multiple objects with different shapes, the contour of each object can be accurately presented, and its position and size can also be precisely controlled. The generated image has rich details and meets expectations, and has great application potential in fields with high requirements for image accuracy such as industrial design and virtual scene construction. Brief Description of the Drawings

[0022] Figure 1 It is a flowchart of the present invention. Detailed Embodiments

[0023] The present invention will be further described below with reference to the drawings and embodiments.

[0024] As Figure 1 shown, a multi-instance object image controllable generation method based on decoupled contour representation includes the following steps: S1, Data collection and arrangement: Image data is collected from the COCO-stuff and VOC-2012 datasets. For the COCO-stuff dataset, images containing 1-9 non-human objects or containing 3-9 objects and only one person are selected, obtaining 28,292 training images and 1,682 validation images, of which 18,794 do not contain people. At the same time, the contour sequence and category information of each object in the image are labeled.

[0025] S2, Data Augmentation: To improve the generalization ability of the model and its adaptability to contour changes, multi-dimensional data augmentation is performed on the collected images.

[0026] The process of data augmentation is as follows: At the image level, horizontal and vertical flipping operations are carried out. Horizontal flipping is achieved by mirroring the image along the vertical central axis to generate new image samples; vertical flipping is performed by mirroring along the horizontal central axis. This not only expands the number of image samples but also enables the model to learn the feature representations of images in different directions.

[0027] For contour representation, data augmentation is carried out. First, according to the model input requirements, the contour point sequence is fixed to a specific length. Then, using a random number generation algorithm, random offsets are introduced for some contour points. For each contour point to be adjusted , where x and y are the abscissa and ordinate of the contour point respectively, random offsets are generated within a set range , and the new coordinates become , , are the offsets of the abscissa and ordinate respectively, causing the contour to deform. In this way, the diversity of contour data is increased, enabling the model to learn contour features under different deformation conditions and enhancing the processing ability for contour changes.

[0028] S3, Contour Representation Processing: Preprocess and encode the contour sequence and class information to obtain a comprehensive embedding.

[0029] The specific process of step S3 is as follows: S31: Load the contour sequence of the image. If the length of the contour sequence is less than the specified number N, an interpolation algorithm is used to complete it. Common interpolation methods such as linear interpolation insert reasonable coordinate values at the missing point positions according to the distribution law of the existing contour points to ensure that all contour sequences have the same length; S32: Preprocess the contour sequence. By calculating the distance of each point from the upper left corner, the point with the closest distance is adjusted to the starting point of the sequence to ensure the consistency of contour encoding; S33: In the contour encoding step, a contour encoder module based on relative position is adopted. Each contour point is regarded as a token, and a center point is added to represent the global feature, forming a token sequence of length l + 1; S34: Through a projection layer map the contour points to the specified dimension dim, where represents a real number, is the feature dimension, and relative position embeddings are added for positions other than the starting head token , , is the l-th element in; S35: Feed the token sequence into a contour encoder composed of multi-layer self-attention and cross-attention. The self-attention layer focuses on the correlation of data at different coordinate points in the sequence, and the cross-attention layer focuses on the interaction between category data and the contour sequence. After multiple layers of interaction, the starting token represents the feature encoding of the object contour information; S36: For the class encoder, use a simple linear layer to implement; After performing a linear layer operation on the class feature c and the contour feature and adding them in the channel dimension, the comprehensive embedding L is obtained, that is, L = B + C, where B is the contour vector after passing through the contour encoder and C is the category vector. , , CEM is the contour encoder module, and P is the set of contour polygon sequences.

[0030] S4, Contour encoder pre-training. The pre-training tasks include contour sequence reconstruction, predicting the target frame, and object classification; In the contour sequence reconstruction task, use the random masking technique to construct masks at a set proportion of coordinate points. When calculating the attention, ignore these masked tokens, thereby implementing self-supervised learning and enabling the model to deeply learn the shape and structure information of the contour; The predicting target frame task requires the model to predict the corresponding target frame image based on the given contour information, thereby learning the context association information between the contour and the target frame; The object classification task prompts the model to combine the contour with high-level semantic information. Through training with labeled data, the model can accurately identify the object categories corresponding to different contours, laying a solid foundation for subsequent image generation tasks.

[0031] S5, Contour fusion module calculation: Feed the comprehensive embedding obtained after contour representation processing into a contour fusion module composed of multi-layer self-attention modules. In each layer of the multi-layer self-attention module, self-attention calculation is performed according to the adjusted weights to promote information exchange and feature fusion between different contour objects.

[0032] The specific process of the said step S5 is as follows: S51: Feed the comprehensive embedding obtained after contour representation processing as the input of the contour fusion module, denotes the n-th element in L. First, calculate the in each token for the original attention weight for , and the calculation formula is:

[0033] Among them 、 is a computational function based on the attention mechanism, where is the feature dimension; S52: To better capture the relationships between objects, create an n×n relationship graph S based on the distance between the center points of two objects; when calculating the distance between the center points of objects, assume that the center point coordinates of object are and the center point coordinates of object are , then the distance between the two center points . Based on , construct the relationship graph S, where in the relationship graph S represents and 's prior correlation; S53: Adjust the original attention weights based on the relationship graph S, and the adjusted attention weights are , where is an adjustable parameter used to control the influence degree of prior knowledge on weight adjustment; S54: After the adjustment, renormalize the weights to ensure that the sum of all weights is 1, making the weight distribution more reasonable; S55: Feed the adjusted comprehensive embedding into the contour fusion module composed of multiple self-attention modules. In the multiple self-attention modules, each layer performs self-attention calculation according to the adjusted weights. In this way, promote information exchange and feature fusion between different contour objects, enabling the model to better understand the relationships between objects in a multi-object scene and providing richer context information for subsequent image generation.

[0034] S6, the global information injection module fuses global and image features.

[0035] The specific process of the said step S6 is as follows: S61: After completing the calculation of the contour fusion module, record the first token output by the contour fusion module as and extract it. Regard as the global information of the entire layout, while other tokens are regarded as individual object representations embedded with local information of relevant objects; S62: Directly combine representing the global information of the contour map with image features at different resolutions; assume that the representations of the image features at different resolutions are , representing low, medium, and high resolution image features respectively. Through the projection matrix W, combine ​Project it onto the same dimensional space as the image features, and then perform an addition operation, that is , , , respectively represent the low, medium, and high-resolution image features after projection and addition operations; S63: Linearly combine the time step with the global information of the contour; add the time step information to the image feature fusion result at each resolution , to obtain the final global token, that is ; , , respectively represent the low, medium, and high-resolution image features after linear combination; S64: Use the adaptive layer normalization technique to embed into the corresponding resolution features respectively; Adaptive layer normalization dynamically adjusts the normalization parameters according to the statistical information of the input features, so that the global information is integrated into the image generation process, promoting the integration of local and global context details of the contour map in the image generation process, making the generated image more reasonable and natural in terms of overall layout and details.

[0036] S7, Contour position encoding: Treat the image patch as a special square region with contour coordinates to achieve contour position encoding.

[0037] The specific process of the step S7 is as follows: Let represent the height h, width w, and channel feature map of the entire image. For each image patch in the image, define as the image patch located in the u-th row and v-th column of the image, and obtain the boundary point set by uniformly sampling from the image patch contour, is the l-th sampling point. During the sampling process, determine the sampling interval according to the size and shape of the image patch to ensure that the sampling points are evenly distributed on the contour, so as to integrate the contour position and shape information of the image patch and the object into a unified space; for example, for a larger image patch, the number of sampling points can be appropriately increased; for a smaller image patch, the number of sampling points can be correspondingly reduced. The coordinate calculation method of the sampling points is:

[0038] where is the original coordinate on the image patch contour.

[0039] The contour coordinates of the image patches share the same spatial representation with other contour tokens, replacing the traditional position encoding method. This unified spatial representation enables the model to more intuitively understand the spatial relationship between images and contour information when processing them, thus more accurately positioning and shaping objects during image generation, improving the accuracy and quality of image generation.

[0040] S8, Object Space Perception Computation: In the object space perception module, a cross-attention mechanism is adopted and combined with self-attention calculation to enhance the interaction between different image patches.

[0041] The specific process of step S8 is as follows: S81: Perform self-attention calculation on the image patches, and calculate the attention scores of the query vector Q, key vector K, and value vector V of the image features respectively ; Calculate When calculating perform a concatenation operation on the Q and K values of the image features and the image patch contour embedding in the channel dimension through the concatenation function , that is ; Directly take the V value of the image feature , that is , where respectively represent the Q value, K value, and V value of the image feature, which are obtained by performing convolution and normalization operations on the image feature, that is , , respectively represent convolution and normalization operations, and I represents the image; S82: Calculate the self-attention output , and the formula is:

[0042] Through self-attention calculation, information interaction occurs between different image patches, enabling the model to better capture the relationship between image patches; S83: Calculate the cross-attention between the image patch and the contour embedding; first calculate the Q value of the cross-embedding, the K value of the cross-embedding, and the V value of the cross-embedding, , , , where is the image patch after self-attention calculation, which is obtained by performing convolution and normalization on , that is , are respectively the K value and V value of the global contour embedding, which are obtained by performing convolution operations on the contour embedding and C, that is , ; contour embedding represents the first token of; S84: Calculate the cross-attention map CA, with the formula and obtain the final output ; S85: After the calculation is completed, collect the cross-attention maps CA of different sizes (8×8, 16×16, 32×32), process these cross-attention maps, first calculate the average cross-attention map, that is, average the cross-attention maps of different sizes according to the set weights, and then perform normalization and aggregation operations on the averaged cross-attention map; S86: Use the sigmoid function to convert the processed cross-attention map into an approximate binary mask, and calculate the loss function DiceLoss of this approximate binary mask and the true mask The loss function has the form where is the alignment loss, is the DiceLoss function, represents the cross-attention map after logical processing, , represents the logical function, represents the cross-attention map after the upsampling operation, , represents the cross-attention map with a feature scale of is the weight coefficient, is the upsampling function, b is the number of feature scales; S87: Integrate into the final loss function and optimize the final loss function to ensure the consistency of the generated image with the input conditions at the pixel level.

[0043] S9, Model construction: Build the PolyDiffusion model with layout diffusion as the backbone network and combine the guidance technology to control the generation effect.

[0044] In the step S9, build the PolyDiffusion model with layout diffusion as the backbone network. The layout diffusion constructs a standard U-net network structure based on the residual network Resnet and the transformer, providing the basic feature extraction and generation capabilities for the entire PolyDiffusion model; Train the model according to the basic principle of the diffusion model and the defined objective function; during the training process, use the AdamW optimizer and set the learning rate to ,​ , are the exponentially decaying rates estimated at the first and second moments respectively; By adopting the classifier-free guidance technique, an empty set contour is constructed by including the contour points of the empty set , and during training, the contour conditions are randomly discarded with a set probability and replaced with an empty set contour . In this way, both conditional and unconditional generation modes can be learned simultaneously during training, reducing the training cost and improving the generalization ability of the model. At the same time, according to the designed reward strategy, during the fine-tuning stage, noise is added to the image generated by the pre-trained PolyDiffusion model through the diffusion forward operation to obtain the image ; When the added noise is small enough, predicting the original image is equivalent to performing one-step sampling on the perturbed image and only performing the last step of denoising. Based on this principle, the reward loss is calculated, , where is the loss function, represents the image finally generated by the model, represents the real image. Combine with the diffusion training loss , the alignment loss to form the final loss function :

[0045]

[0046] where and are both hyperparameters used to adjust the contribution degree of different loss terms to the overall loss. is the current time step, is a hyperparameter used to control the calculation method when is less than or equal to ; is the current time step of the added real noise, represents the noise prediction neural network, represents the input at the current time step . By continuously optimizing the final loss function, the parameters of the PolyDiffusion model are adjusted.

[0047] S10, Image generation: Input the preprocessed contour sequence and class information into the trained PolyDiffusion model to output a multi-instance object image.

[0048] In step S10, the preprocessed contour sequence and class information are input into the trained PolyDiffusion model. The PolyDiffusion model first performs contour representation processing on the input contour sequence, converts it into an effective feature encoding, and combines the class information to obtain a comprehensive embedding. The comprehensive embedding is successively processed by a contour fusion module, a global information injection module, a contour position encoding module, and an object space perception module. In the contour fusion module, information between different contour objects is interacted and fused. The global information injection module integrates global information into the image generation process. The contour position encoding module unifies the spatial representation of image patches and contours. The object space perception module further improves the consistency between the generated image and the input conditions at the pixel level through the cross-attention mechanism. After the collaborative processing of the contour fusion module, the global information injection module, the contour position encoding module, and the object space perception module, the PolyDiffusion model finally outputs a multi-instance object image.

[0049] Table 1 is a performance comparison table of the method of the present invention and other methods (Grid2Im, PLGAN, InstanceDiffusion, LayoutDiffusion) when using the COCOStuff dataset. The COCOStuff dataset is constructed based on the COCO2017 dataset. It covers a total of 172 categories, including 80 "thing" categories, 91 "stuff" categories, and 1 "unlabeled" category.

[0050] Table 2 is a performance comparison table of the method of the present invention and other methods (Grid2Im, PLGAN, InstanceDiffusion, LayoutDiffusion) when using the VOC-2012 dataset. The VOC-2012 dataset is a key dataset in the Visual Object Classes Challenge. This dataset covers 20 object categories and 1 background category, with a total of 11,530 images, of which the training set and the validation set contain 10,582 images, and the test set has 948 images.

[0051]

[0052]

[0053] In Table 1 and Table 2, FID is the Frechet Inception Distance, which is an index to measure the similarity between the generated image distribution and the real image distribution; IS is the generated image quality evaluation index; CAS is the classification accuracy score; AP maskis the average precision, which is an evaluation metric for image segmentation tasks and is used to measure the matching degree between the generated mask and the ground truth mask; AR mask is the average recall rate, which is used in conjunction with AP mask to measure the coverage ability of the generated mask for the ground truth mask. As can be seen from Table 1 and Table 2: The method of the present invention has achieved good results in most metrics, especially in AR mask and AP mask metrics, achieving the best results.

[0054] The multi-instance object image controllable generation method based on decoupled contour representation proposed by the present invention can effectively solve many problems faced by traditional image generation technologies in dealing with complex multi-object scenes, bringing significant technological improvements and innovative application values to the field of image generation.

Claims

1. A controllable generation method of multi-instance object images based on decoupled contour representation, characterized in that: The following steps are involved: S1, data collection and organization: collect image data from the dataset and annotate the contour sequence and category information of each object in the image; S2, data enhancement: multi-dimensional data enhancement is performed on the collected images; S3, contour representation processing: preprocessing and encoding the contour sequence and category information to obtain a comprehensive embedding; S4, contour encoder pre-training; S5, contour fusion module calculation: the comprehensive embedding obtained after contour representation processing is input into the contour fusion module composed of multi-layer self-attention modules. In the multi-layer self-attention module, each layer performs self-attention calculation according to the adjusted weights to promote information exchange and feature fusion between different contour objects; S6, global information injection module fuses global and image features; S7, contour position encoding: The image block is regarded as a special square area with contour coordinates to achieve contour position encoding; S8, object space perception calculation: In the object space perception module, the cross attention mechanism is adopted and combined with self-attention calculation to enhance the interaction between different image blocks; S9, model construction: build the PolyDiffusion model with layout diffusion as the backbone network, and combine it with guidance technology to control the generation effect; S10, image generation: input the preprocessed contour sequence and category information into the trained PolyDiffusion model to output a multi-instance object image.

2. The method for controllable generation of multi-instance object images based on decoupled contour representation according to claim 1, characterized in that: In step S2, the process of data enhancement is as follows: At the image level, horizontal and vertical flipping operations are performed; horizontal flipping is to generate a new image sample by mirroring the image along the vertical center axis; vertical flipping is to mirror flip along the horizontal center axis; Data enhancement is performed for contour representation. First, the contour point sequence is fixed to a specific length according to the model input requirements. Then, a random number generation algorithm is used to introduce random offsets for some contour points. For each contour point that needs to be adjusted, , x and y are the horizontal and vertical coordinates of the contour point respectively, and the offset is randomly generated within the set range , the new coordinates become , , They are the offset of the horizontal and vertical coordinates, respectively, which cause the contour to deform.

3. The controllable generation method of multi-instance object images based on decoupled contour representation according to claim 1 is characterized in that: The specific process of step S3 is: S31: Load the contour sequence of the image. If the length of the contour sequence is less than the specified number N, an interpolation algorithm is used to fill it. S32: preprocessing the contour sequence, by calculating the distance between each point and the upper left corner, adjusting the point with the closest distance as the starting point of the sequence, to ensure the consistency of the contour encoding; S33: In the contour encoding stage, a relative position-based contour encoder module is used to treat each contour point as a token, and a center point is added to represent the global feature to form a token sequence of length l+1; S34: Through the projection layer Map the contour points to the specified dimension dim, represents a real number, For feature dimensions, add relative position embedding for positions other than the starting token , , for The lth element in ; S35: Send the token sequence to a contour encoder composed of multiple layers of self-attention and cross-attention. The self-attention layer focuses on the correlation of data at different coordinate points in the sequence, and the cross-attention layer focuses on the interaction between the category data and the contour sequence. After multiple layers of interaction, the starting token represents the feature encoding of the object contour information. S36: For the class encoder, a simple linear layer is used Implementation: After the class feature c and the contour feature are operated through the linear layer, they are added in the channel dimension to obtain the comprehensive embedding L, that is, L=B+C, where B is the contour vector after the contour encoder, and C is the category vector. , , CEM is the contour encoder module, and P is the set of contour polygon sequences.

4. The method for controllable generation of multi-instance object images based on decoupled contour representation according to claim 1, characterized in that: In step S4, the pre-training tasks include contour sequence reconstruction, target frame prediction, and object classification; In the contour sequence reconstruction task, the random masking technique is used to construct a mask on the coordinate points of a set ratio. When calculating the attention, these masked tokens are ignored, so as to achieve self-supervised learning and enable the model to deeply learn the shape and structure information of the contour; The target frame prediction task requires the model to predict the corresponding target frame image based on the given contour information, so as to learn the contextual association information between the contour and the target frame; The object classification task prompts the model to combine contours with high-level semantic information, and through training with labeled data, the model can accurately identify the object categories corresponding to different contours.

5. The method for controllable generation of multi-instance object images based on decoupled contour representation according to claim 3, characterized in that: The specific process of step S5 is as follows: S51: The comprehensive embedding obtained by contour representation processing As the input of the contour fusion module, Represents the nth element in L. First, calculate the For other tokens The original attention weight , the calculation formula is: ; in , It is a calculation function based on the attention mechanism. is the feature dimension; S52: In order to better capture the relationship between objects, create an n×n relationship graph S according to the distance between the center points of two objects; When calculating the distance of the center point of an object, it is assumed that the object The center point coordinates are , object The center point coordinates are , then the distance between the two center points ,according to Construct a relationship graph S, in which express and A priori correlation of S53: Adjust the original attention weight based on the relationship graph S. The adjusted attention weight is ,in is an adjustable parameter; S54: After the adjustment is completed, the weights are renormalized to ensure that the sum of all weights is 1; S55: The adjusted comprehensive embedding is sent to a contour fusion module composed of a multi-layer self-attention module. In the multi-layer self-attention module, each layer performs self-attention calculation according to the adjusted weights to promote information exchange and feature fusion between different contour objects.

6. The method for controllable generation of multi-instance object images based on decoupled contour representation according to claim 5, characterized in that: The specific process of step S6 is as follows: S61: After completing the contour fusion module calculation, the first token output by the contour fusion module is recorded as And extract it, is regarded as the global information of the entire layout, while other tokens are regarded as individual object representations that embed local information of related objects; S62: Directly represent the global information of the contour map Combined with image features of different resolutions; assuming that the image features at different resolutions are represented as , Represent low, medium and high resolution image features respectively, and are transformed into Project it to the same dimensional space as the image features, and then perform the addition operation, that is, , , , Respectively represent the low, medium and high resolution image features after projection and addition operations; S63: Linearly combine the time step with the global information of the contour; add the time step information to the image feature fusion result at each resolution , get the final global token, that is ; , , Respectively represent the low, medium and high resolution image features after linear combination; S64: Using adaptive layer normalization technology, They are embedded into the corresponding resolution features respectively; Adaptive layer normalization dynamically adjusts normalization parameters according to the statistical information of input features, so that global information is integrated into the image generation process, promoting the integration of local and global context details of contour maps in the image generation process.

7. The method for controllable generation of multi-instance object images based on decoupled contour representation according to claim 6, characterized in that: The specific process of step S7 is as follows: set up Represents the feature map of the height h, width w and channels of the entire image. For each image block in the image, define For the image block located at the uth row and vth column of the image, the boundary point set is obtained by uniformly sampling from the image block contour , is the lth sampling point. During the sampling process, the sampling interval is determined according to the size and shape of the image block to ensure that the sampling points are evenly distributed on the contour, so as to integrate the contour position and shape information of the image block and the object into a unified space; the coordinate calculation method of the sampling point is: ; in for The original coordinates on the image patch outline.

8. The method for controllable generation of multi-instance object images based on decoupled contour representation according to claim 6, characterized in that: The specific process of step S8 is as follows: S81: Perform self-attention calculation on the image block, and calculate the attention scores of the query vector Q, key vector K, and value vector V of the image feature respectively. ;calculate When, through the splicing function Embedding the Q, K values ​​of image features and image patch contours in the channel dimension Perform splicing operation, that is , ; Directly take the V value of the image feature ,Right now ,in They represent the Q value, K value, and V value of the image features, respectively, and are obtained by performing convolution and normalization operations on the image features, that is, , , Respectively represent convolution and normalization operations, and I represents the image; S82: Calculate self-attention output , the formula is: ; Through self-attention calculation, information interaction is carried out between different image blocks; S83: Calculate the cross attention of the image patch and contour embedding; first calculate the Q value of the cross embedding , K value of cross embedding , cross-embedded V value , , , ,in, is the image block after self-attention calculation, through After convolution and normalization, we get , They are the K value and V value of global contour embedding respectively. And C are convolved to obtain, that is , ; Contour embedding represent The first token of S84: Calculate the cross attention map CA, the formula is , and get the final output ; S85: After the calculation is completed, collect cross-attention maps CA of different sizes, process these cross-attention maps, first calculate the average cross-attention map, that is, average the cross-attention maps of different sizes according to the set weights, and then normalize and aggregate the averaged cross-attention maps; S86: Use the sigmoid function to convert the processed cross attention map into an approximate binary mask, and calculate the difference between the approximate binary mask and the true mask. The loss function DiceLoss, the loss function form is ,in is the alignment loss, is the DiceLoss loss function, represents the cross attention map after logical processing, , represents a logical function, represents the cross attention map after upsampling operation, , The characteristic scale is The cross attention map, is the weight coefficient, is the upsampling function, b is the number of feature scales; S87: Integrate it into the final loss function and optimize the final loss function to ensure that the generated image is consistent with the input conditions at the pixel level.

9. The method for controllable generation of multi-instance object images based on decoupled contour representation according to claim 8, characterized in that: In the step S9, a PolyDiffusion model is built with layout diffusion as the backbone network. Layout diffusion builds a standard U-net network structure based on the residual network Resnet and transformer to provide basic feature extraction and generation capabilities for the entire PolyDiffusion model. The model is trained according to the basic principle of the diffusion model and the defined objective function. During the training process, the AdamW optimizer is used, and the classifier-free guidance technology is adopted to construct an empty set contour containing empty set contour points. , randomly discard contour conditions with a set probability during training and replace them with empty set contours ; At the same time, in the fine-tuning stage, the images generated by the pre-trained PolyDiffusion model Perform a diffusion forward operation to add noise Get the image ; When the noise added Enough hours to predict It is equivalent to sampling the perturbed image in one step and performing only the last step of denoising. According to this principle, the reward loss is calculated , ,in for Loss function, represents the image finally generated by the model, Represents the real image, With diffusion training loss , alignment loss Combined to form the final loss function : ; ; in and They are all hyperparameters used to adjust the contribution of different loss items to the overall loss. is the current time step, is a hyperparameter used to control Less than or equal to hour Calculation method of is the current time step The real noise added, represents the noise prediction neural network, Indicates the current time step The parameters of the PolyDiffusion model are adjusted by continuously optimizing the final loss function.

10. The method for controllable generation of multi-instance object images based on decoupled contour representation according to claim 9, characterized in that: In the step S10, the preprocessed contour sequence and category information are input into the trained PolyDiffusion model. The PolyDiffusion model first performs contour representation processing on the input contour sequence, converts it into an effective feature code, and obtains a comprehensive embedding in combination with the category information; the comprehensive embedding is processed in sequence by a contour fusion module, a global information injection module, a contour position encoding module, and an object space perception module. In the contour fusion module, information between different contour objects interacts and fuses. The global information injection module incorporates global information into the image generation process. The contour position encoding module unifies the spatial representation of image blocks and contours. The object space perception module further improves the consistency of the generated image at the pixel level with the input conditions through a cross-attention mechanism. After the coordinated processing of the contour fusion module, the global information injection module, the contour position encoding module, and the object space perception module, the PolyDiffusion model finally outputs a multi-instance object image.

Citation Information

Patent Citations

  • Motion blur restoration method based on contour enhancement strategy

    CN111815536A

  • Remote sensing image building extraction and contour optimization method based on deep learning

    CN113516135A

  • Marine multi-target ship instance segmentation method

    CN117197452A

  • Target instance segmentation model establishment method based on automatic prompt learning and application thereof

    CN118279320A

  • Arbitrary-shape scene text detection method based on contour feature enhancement

    CN119296094A

Cited By

  • Generalized contour data processing method, object cognition method and device

    CN120823231A

  • A generalized contour data processing method, object recognition method and device

    CN120823231B