Joint synthesis and placement of objects in a scene

By using a combination of multiple generator models and discriminator models in machine learning, the problem of difficult context consistency when objects are inserted in the scene in the prior art is solved, and the realistic synthesis and diversity of objects are achieved.

CN110880203BActive Publication Date: 2025-06-20NVIDIA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201910826859.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-11-27
Filing Date
2019-09-03
Publication Date
2025-06-20
Estimated Expiration
2040-02-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively insert objects into the scene in scenes such as image synthesis, augmented reality, and virtual reality in machine learning to maintain context consistency.

Method used

By using a combination of multiple generator models and discriminator models, the bounding boxes and shapes of the generated objects are represented based on the semantics of the image, and the diversity of the generator models is enhanced by joint training and supervision paths.

Benefits of technology

Realistic synthesis and reasonable placement of objects in scenes such as image synthesis and augmented reality, improving the diversity and context consistency of object positions and shapes in the scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110880203B_ABST
    Figure CN110880203B_ABST
Patent Text Reader

Abstract

The present invention discloses the joint synthesis and placement of objects in a scene. Specifically, an embodiment of a method includes: applying a first generator model to a semantic representation of an image to generate an affine transformation, where the affine transformation represents a bounding box associated with at least one region within the image. The method further includes: applying a second generator model to the affine transformation and the semantic representation to generate the shape of an object. The method further includes inserting the object into the image based on the bounding box and the shape.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 726,872, filed on September 4, 2018, entitled "CONTEXT-AWARE SYNTHESIS AND PLACEMENT OF OBJECT INSTANCES", the subject matter of which is incorporated herein by reference. BACKGROUND OF THE INVENTION

[0003] Objects can be inserted into scenarios of real-world applications, which include but are not limited to image synthesis in machine learning, augmented reality, virtual reality, and / or domain randomization. For example, a machine learning model may insert pedestrians and / or cars into an image containing a road for subsequent use in training an autonomous driving system and / or generating a video game or virtual reality environment. Inserting objects into a scenario in a realistic and / or contextually meaningful way presents many technical challenges. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] Accordingly, the above-described features of the various embodiments can be understood in more detail, and a more specific description of the inventive concept briefly outlined above can be obtained with reference to the various embodiments, some of which are illustrated in the drawings. However, it should be noted that the drawings only illustrate typical embodiments of the inventive concept and should not be considered to limit its scope in any way. There are other equally effective embodiments.

[0005] Figure 1 is a block diagram showing a system configured to implement one or more aspects of the various embodiments.

[0006] Figure 2 is according to the various embodiments Figure 1 a more detailed description of the training engine and the execution engine.

[0007] Figure 3 is a flowchart showing the method steps for performing joint synthesis and placement of objects in a scenario according to the various embodiments.

[0008] Figure 4 is a flowchart showing the method steps for training a machine learning model that performs joint synthesis and placement of objects in a scenario according to the various embodiments.

[0009] Figure 5 is a block diagram of a computer system configured to implement one or more aspects of the various embodiments.

[0010] Figure 6is according to each embodiment of Figure 5 is a block diagram of a parallel processing unit (PPU) included in a parallel processing subsystem of

[0011] Figure 7 is according to each embodiment of Figure 6 is a block diagram of a general processing cluster (GPC) included in a parallel processing unit (PPU) of Detailed implementation

[0012] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of each embodiment. However, it will be apparent to one of ordinary skill in the art that the inventive concept may be practiced without one or more of these specific details.

[0013] System review

[0014] Figure 1 Describes a computing device 100 configured to implement one or more aspects of each embodiment. In one embodiment, the computing device 100 can be a desktop computer, laptop computer, smartphone, personal digital assistant (PDA), tablet computer, or any other type of computing device configured to receive input, process data, and selectively display images, and is adapted to practice one or more embodiments. The computing device 100 is configured to run a training engine 122 and an execution engine 124 residing in memory 116. It should be noted that the computing devices described herein are illustrative, and any other technically feasible configuration falls within the scope of the present disclosure.

[0015] In one embodiment, the computing device 100 includes, but is not limited to: an interconnect (bus) 112 connecting one or more processing units 102, an input / output (I / O) device interface 104 coupled to one or more input / output (I / O) devices 108, a memory 116, a storage 114, and a network interface 106. One or more processing units 102 can be any suitable processor implemented as: a central processing unit (CPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), an artificial intelligence (AI) accelerator, any other type of processing unit, or a combination of different processing units (such as a CPU configured to operate in conjunction with a GPU). Generally speaking, one or more processing units 102 can be any technically feasible hardware unit capable of processing data and / or executing software applications. Additionally, in the context of the present disclosure, the computing elements shown in the computing device 100 can correspond to a physical computing system (e.g., a system in a data center), or can be a virtual computing instance executing in a computing cloud.

[0016] In one embodiment, the I / O device 108 includes devices capable of providing input, such as a keyboard, a mouse, a touch-sensitive screen, etc., and devices capable of providing output, such as a display device. In addition, the I / O device 108 may include devices capable of receiving input and providing output, such as a touch screen, a Universal Serial Bus (USB) port, and so on. The I / O device 108 is configurable to receive various types of input from an end user (e.g., a designer) of the computing device 100 and provide various types of output to the end user of the computing device 100, such as a displayed digital image or digital video or text. In some embodiments, one or more of the I / O devices 108 are configured to couple the computing device 100 to the network 110.

[0017] In one embodiment, the network 110 is any technically feasible type of communication network that allows for the exchange of data between the computing device 100 and an external entity or device, such as a web server or other networked computing device. For example, the network 110 may include a Wide Area Network (WAN), a Local Area Network (LAN), a wireless (WiFi) network, and / or the Internet, etc.

[0018] In one embodiment, the memory 114 includes non-volatile memory for applications and data, which may include fixed or removable disk drives, flash devices, and CD-ROM, DVD-ROM, Blu-ray, HD-DVD, or other magnetic, optical, or solid-state storage devices. The training engine 122 and the execution engine 124 may be stored in the memory 114 and loaded into the memory 116 when executed.

[0019] In one embodiment, the memory 116 includes Random Access Memory (RAM) modules, flash memory cells, or any other type of storage cells or combinations thereof. One or more processing units 102, the I / O device interface 104, and the network interface 106 are configured to read data from and write data to the memory 116. The memory 116 includes various software programs executable by one or more processors 102 and application data associated with the software programs, including the training engine 122 and the execution engine 124.

[0020] In one embodiment, the training engine 122 generates a machine learning model for inserting objects into a scene. The scene may include a semantic representation of an image, such as a segmentation map that associates individual pixels in the image with semantic labels. For example, a segmentation map for an outdoor scene may include pixel regions assigned to labels such as "road," "sky," "building," "bridge," "tree," "ground," "car," and "pedestrian." In turn, the machine learning model created by the training engine 122 can be used to identify reasonable locations for objects in the scene, as well as reasonable sizes and shapes for objects at that location. In various embodiments, the machine learning model can learn "where" objects can be inserted in the scene and "what" objects look like so that the objects remain contextually consistent with the scene.

[0021] In one embodiment, the execution engine 124 executes the machine learning model to perform joint synthesis of objects and placement of objects into the scene. The joint synthesis of objects and placement of objects into the scene may involve joint learning of the position and scale of each object in a given scene, and the shape of each object given the corresponding position and scale. Continuing with the above example, the execution engine 124 may apply the first generator model generated by the training engine 122 to the semantic representation of the outdoor scene to identify reasonable locations where cars, pedestrians, and / or other types of objects can be inserted into the scene. The execution engine 124 may then apply the second generator model created by the training engine 122 to the semantic representation and positions identified by the first generator model to generate realistic shapes of objects at the identified locations. The training engine 122 and the execution engine 124 are described below with respect to Figure 2 Described in more detail in .

[0022] Joint compositing and placement of objects in a scene

[0023] Figure 2 According to various embodiments Figure 1 206 and 210. In the illustrated embodiment, the training engine 122 creates a number of generator models, such as generator models 202-204, and a number of discriminator models, such as an affine discriminator 206, a layout discriminator 208-210, and a shape discriminator 212, to perform joint synthesis and placement of objects based on the semantic representation 200 of the image 258. The discriminator models are collectively referred to herein as "discriminator models" 206-212. In various embodiments, the execution engine 124 applies the generator models 202-204 to additional images 258 from the image repository 264 to insert objects 260 into the image 258 in a realistic and / or semantically reasonable manner.

[0024] The generator models 202-204 and / or the corresponding discriminator models 206-212 can be in any technically feasible form of machine learning model. For example, the generator models 202-204, the affine discriminator 206, the layout discriminators 208-210, and / or the shape discriminator 212 can include recurrent neural networks (RNNs), convolutional neural networks (CNNs), deep neural networks (DNNs), deep convolutional networks (DCNs), deep belief networks (DBNs), restricted Boltzmann machines (RBMs), long short-term memory (LSTM) units, gated recurrent units (GRUs), generative adversarial networks (GANs), self-organizing maps (SOMs), and / or other types of artificial neural networks or components of artificial neural networks. In another embodiment, the generator models 202-204, the affine discriminator 206, the layout discriminators 208-210, and / or the shape discriminator 212 can include functions that perform clustering, principal component analysis (PCA), latent semantic analysis (LSA), Word2vec, and / or other unsupervised learning techniques. In a third example, the generator models 202-204, the affine discriminator 206, the layout discriminators 208-210, and / or the shape discriminator 212 can include regression models, support vector machines, decision trees, random forests, gradient-boosted trees, naive Bayes classifiers, Bayesian networks, hierarchical models, and / or ensemble models.

[0025] As described above, the semantic representation 200 of the image 200 can include the pixels 218 in the image 200 and the labels 220 that associate the pixels 218 with different categories. For example, the semantic representation of an outdoor scene can include a "segmentation map" of the regions containing the pixels 218, and these pixel regions are mapped to labels 220 such as "sky", "ground", "tree", "water", "road", "sidewalk", "building", "structure", "car", and / or "pedestrian". In various embodiments, each region of the pixels 218 is mapped to one or more of the labels 220.

[0026] In one embodiment, the training engine 122 inputs the semantic representation 200 of the scene into the generator model 202. The generator model 202 outputs an affine transformation 230, which represents the bounding box of the object 260 that can be inserted into the scene. For example, the training engine 122 can input the segmentation map of an outdoor scene into the generator model 202, and the generator model 202 can define the bounding boxes of cars, pedestrians, and / or other objects to be inserted into the outdoor scene as an affine transformation matrix applied to the unit bounding box in the scene.

[0027] In various embodiments, the affine transformation matrix can include translations, scalings, rotations, and / or other types of affine transformations that are applied to a unit bounding box to generate a bounding box at certain positions and scale in the scene. In these embodiments, given a two-dimensional (2D) semantic representation of a scene with a unit bounding box of 1 pixel x 1 pixel, the bounding box can be calculated using the following equation:

[0028]

[0029] In the above equation, x and y represent the coordinates of each point in the unit bounding box, and x' and y' represent the corresponding points in the bounding box recognized by the generator model 202 in the scene. a, b, c, d, t x and t y represent the parameters of the affine transformation applied to x and y to generate x' and y'.

[0030] In one embodiment, the generator model 202 includes a variational autoencoder (VAE) 224 and / or a spatial transformer network (STN) 228. In this embodiment, the encoder portion of the VAE 224 is applied to the semantic representation of the scene and a random input 214 to generate a vector in the latent space. The vector is then input into the STN 228 to generate one or more affine transformations 230 that represent the bounding boxes of the objects in the scene. Thus, each affine transformation can specify the position and scale of the corresponding object in the scene.

[0031] For example, the random input 214 can include a random vector having a standard normal distribution that is combined (e.g., associated) with a given semantic representation of the image to generate an input to the VAE 224. The training engine 122 can apply the encoder portion of the VAE 224 to the combination of the random input 214 and the semantic representation to generate a vector in the latent space that also has a standard normal distribution. The training engine 122 can then use the STN 228 to transform the vector into an affine transformation that represents the bounding box of the object to be inserted into the scene.

[0032] In one embodiment, the training engine 122 inputs the semantic representation 200 of the scene and the corresponding affine transformation 230 generated by the generator model 202 into the generator model 204. The generator model 204 outputs the shape 232 of the object 260 within the bounding box represented by the affine transformation 230. In one embodiment, the generator model 204 includes another VAE 226. In this embodiment, the encoder portion of the VAE 226 is applied to the semantic representation of the scene, which includes one or more affine transformations 230 and a random input 214 output by the generator model 202, to generate a vector in the latent space. Then, the vector is input into the decoder portion of the VAE 226 to generate one or more shapes 232 that fit within the bounding box represented by the affine transformation 230.

[0033] For example, the random input 216 can include a random vector having a standard normal distribution, which is combined (e.g., associated) with the semantic representation of the image to generate an input to the VAE 226. The semantic representation can be updated to include the region of the pixels 218 and the corresponding label 220 for the bounding box represented by the affine transformation 230. The training engine 122 can apply the encoder portion of the VAE 226 to the input to generate a vector in the latent space, which also has a standard normal distribution. Then, the training engine 122 can apply the decoder portion of the VAE 226 to the vector to generate a binary mask that includes the shape of the object within the bounding box represented by the affine transformation 230.

[0034] In various embodiments, the affine transformation 230 output by the STN 228 can be used as a differentiable link between the generator models 202-204. Thus, the training engine 122 can perform joint training and / or update of the generator models 202-204 using the differentiable link. In this way, the generator models 202-204 operate as an end-to-end machine learning model that learns the joint distribution of the positions and shapes of different types of objects 260 using the semantic representation 200 of the scene.

[0035] More specifically, in one embodiment, the training engine 122 combines the outputs of the generator models 202-204 with the corresponding predictions 234-240 from the affine discriminator 206, layout discriminators 208-210, and shape discriminator 212 to train and / or update the generator models 202-204. In this embodiment, the training engine 122 inputs the ground truth of the affine transformation 230 output by the generator model 202 and / or the object positions in the semantic representation 200 into the affine discriminator 206 and layout discriminator 208. The affine discriminator 206 can output a prediction 234 that classifies the parameters of the affine transformation 230 as true or false, and the layout discriminator 208 can output a prediction 236 that classifies the placement of the corresponding bounding box in the semantic representation 200 as true or false.

[0036] In one embodiment, the training engine 122 also inputs the shape 232 output by the generator model 204 and the ground truth 222 of the object shape 232 in the semantic representation 200 into the layout discriminator 210 and shape discriminator 212. The layout discriminator can output a prediction 238 that classifies the placement of the shape 232 in the semantic representation 200 as true or false, and the shape discriminator 212 can output a prediction 240 that classifies the generated shape 232 as true or false. Then the training engine 122 calculates losses 242-248 based on the predictions 234-240 from the discriminator models and uses the gradients associated with the losses 242-248 to update the parameters of the generator models 202-204.

[0037] In various embodiments, the training engine 122 can combine the generator models 202-204, affine discriminator 206, layout discriminators 208-210, and shape discriminator 212 into a GAN, where each generator model and the corresponding discriminator model are trained against each other. For example, the generator models 202-204, affine discriminator 206, layout discriminators 208-210, and shape discriminator 212 can be included in a convolutional GAN, conditional GAN, recurrent GAN, Wasserstein GAN, and / or other types of GANs. In turn, the generator models 202-204 can generate more realistic affine transformations 230 and shapes 232, while the discriminator models can learn to better distinguish between true and false object positions and shapes in the semantic representation 200 of the scene.

[0038] In one embodiment, to increase the diversity of the affine transformations 230 and shapes 232 generated by the generator models 202-204, the training engine 122 can update the generator models 202-204 via both the supervised paths 250-252 and the unsupervised paths 254-256. The supervised path 250 can include the ground truth 222 of the bounding boxes of the objects that can be inserted into the corresponding semantic representations 200 as an additional input to the generator model 202. Similarly, the supervised path 252 can include the ground truth 222 of the shapes 232 of the objects that can be inserted into the corresponding semantic representations as an additional input to the generator model 204.

[0039] Thus, the supervised paths 250-252 can allow the generator models 202-204 to learn additional positions and / or shapes of the objects in the scene other than those generated via the unsupervised paths 254-256, which lack the ground truth 222 of the positions and shapes of the objects. For example, training the generator models 202-204 via the unsupervised paths 254-256 may cause the generator models 202-204 to effectively ignore the random inputs 214-216 during the generation of the corresponding affine transformations 230 and / or shapes 232. By adding the supervised paths 250-252 to the pipeline for training the generator models 202-204, the training engine 122 can train the VAEs 224-226 and / or the STNs 228 to reconstruct the corresponding ground truth 222 input to the generator models 202-204, thereby allowing different values of the random inputs 214-216 to generate different bounding boxes and shapes 232 of the objects in a given scene.

[0040] In one embodiment, the training of the generator model 202 can be performed using a minimax game among the generator model 202, the affine discriminator 208, and the layout discriminator 208, and has the following loss function:

[0041]

[0042] In the above function, L1 represents the loss associated with the GAN that includes the generator model 202, the affine discriminator 206, and the layout discriminator 208; G l represents the generator model 202; and D l represents the discriminators associated with the output of the generator model 202 (i.e., the affine discriminator 206 and the layout discriminator 208). The loss includes three components: the unsupervised adversarial loss represented by L l adv the reconstruction loss represented by L l recon and the loss represented by L l supThe supervised adversarial loss shown. The unsupervised adversarial loss (e.g., loss 244) is determined based on the generator model 202 and the layout discriminator 208, and is represented by The reconstruction loss is determined based on the generator model 202. The supervised adversarial loss (e.g., loss 242) is determined based on the generator model 202 and the affine discriminator 206, and is represented by D affine representation.

[0043] In one embodiment, the training engine 122 uses the unsupervised adversarial loss to update the generator model 202 via the unsupervised path 254. For example, the following equation can be used to calculate the unsupervised adversarial loss:

[0044]

[0045] In the above equation, z l represents the random input 214, x represents the semantic representation input into the generator model 202, A(b) represents the affine transformation matrix A, which is applied to the unit bounding box b to generate the true bounding box of the object. represents the prediction A of the generator model 202.

[0046] In one embodiment, the reconstruction loss is also used to update the generator model 202 via the unsupervised path 254. For example, the reconstruction loss can be calculated using the following equation:

[0047]

[0048] In the above equation, x' and z′ l represent the reconstructions of x and z l respectively generated from the latent vectors produced by the VAE 224. Therefore, the reconstruction loss can be used to ensure that the random input 214 and the semantic representation input into the generator model 202 are encoded in the latent vectors.

[0049] In one embodiment, the supervised adversarial loss is used to update the generator model 202 via the supervised path 250. For example, the following equation can be used to calculate the supervised adversarial loss:

[0050]

[0051] In the above equation, A is the affine transformation that generates a realistic bounding box given the ground truth, is the affine transformation of the prediction generated via the supervised path 250, z A represents the vector encoded from the parameters of the ground truth bounding box of the object. E A represents the encoder that encodes the parameters of the input affine transformation A, K Lrepresents the Kullback-Leibler divergence, L sup,adv represents the adversarial loss, which focuses on predicting the true Conversely, this equation can be used to update the generator model 202 so that the generator model 202 maps z A to A for each ground truth.

[0052] In one embodiment, the training of the generator model 204 can be performed using a minimax game among the generator model 204, the layout discriminator 210, and the shape discriminator 212, and has the following loss function:

[0053]

[0054] L s represents the loss associated with the GAN that includes the generator model 204, the layout discriminator 210, and the shape discriminator 212; G s represents the generator model 204; and D s represents the discriminators associated with the output of the generator model 204 (i.e., the layout discriminator 210 and the shape discriminator 212). Similar to the loss function for updating the generator model 202, the above loss function includes three parts: the unsupervised adversarial loss represented by L s adv the reconstruction loss represented by L s recon and the supervised adversarial loss represented by L s sup The unsupervised adversarial loss (e.g., loss 246) is determined based on the generator model 204 and the layout discriminator 210, denoted by . The reconstruction loss is determined based on the generator model 202. The supervised adversarial loss (e.g., loss 248) is determined based on the generator model 204 and the shape discriminator 212, denoted by D shape .

[0055] In one embodiment, the roles of the components in the loss function for updating the generator model 204 are similar to the roles of the corresponding components in the loss function for updating the generator model 204. That is, the unsupervised adversarial loss and the reconstruction loss are used to update the generator model 204 via the unsupervised path 256, and the supervised adversarial loss is used to update the generator model 204 via the supervised path 252. On the other hand, the supervised adversarial loss can be used to train the generator model 204 to reconstruct the true shape of the object rather than the true bounding box and / or position of the object. Additionally, one or more losses 246-248 associated with the generator model 204, the layout discriminator 210, and the shape discriminator 212 can be backpropagated through the VAE 226 of the generator model 204 and the STN 228 of the generator model 202, so the losses 246-248 associated with the generated shape 232 are used to adjust the parameters of the generator models 202-204.

[0056] In one embodiment, after the training of the generator models 202-204 is completed, the execution engine 124 applies the generator models 202-204 to additional images 258 in the image repository 264 to insert the object 260 into the images 258. For example, the execution engine 124 can execute the unsupervised path 254 that includes the generator model 202, the affine discriminator 206, and the layout discriminator 208 to generate an affine transformation 230 representing the bounding box of the object 260 based on the random input 214 and the semantic representation 200 of the image 258. Then, the execution engine 124 can execute the unsupervised path 256 that includes the generator model 204, the layout discriminator 210, and the shape discriminator 212 to generate a shape 232 suitable for the bounding box based on the random input 216, the semantic representation 200, and the affine transformation 230. Finally, the execution engine 124 can apply the affine transformation 230 to the corresponding shape 232 to insert the object 260 into the image 258 at the predicted position.

[0057] Figure 3 is a flowchart of method steps for performing joint synthesis and placement of objects in a scene. Although the method steps are described in conjunction with Figure 1 and Figure 2 systems, those skilled in the art should understand that any system configured to execute the method steps in any order falls within the scope of the present disclosure.

[0058] As shown, the execution engine 124 applies a first generator model to the semantic representation of the image to generate an affine transformation (302) representing a bounding box associated with at least one region in the image. For example, the first generator model may include a VAE and an STN. The input to the VAE may include the semantic representation and a random input (e.g., a random vector). In turn, the encoder in the VAE can generate a latent vector from the input, and the STN can transform the latent vector into an affine transformation that specifies the position and scale of the object to be inserted into the image.

[0059] Next, the execution engine 124 applies a second generator model to the affine transformation and the semantic representation to generate the shape of the object (304). For example, the second generator model may also include a VAE. The input to the VAE may include the semantic representation, the affine transformation, and a random input (e.g., a random vector). The VAE in the second generator model can generate a shape representing the object based on the input and suitable for the position and scale indicated by the affine transformation.

[0060] Then, the execution engine 124 inserts the object into the image based on the bounding box and the shape (306). For example, the execution engine 124 may apply the affine transformation to the shape to obtain the pixel region in the image that contains the object. Then, the execution engine 124 can insert the object into the image by updating the semantic representation to include the mapping from the pixel region to the label of the object.

[0061] Figure 4 is a flowchart of method steps for training a machine learning model that performs joint synthesis and placement of objects in a scene. Although the method steps are described in conjunction with Figure 1 and Figure 2 systems, those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of the present disclosure.

[0062] As shown, the training engine 122 calculates an error associated with two generator models based on the output of the discriminator models from the generator models (402). For example, the discriminator models whose output represents the affine transformation of the bounding box of the object may include a layout discriminator model and an affine discriminator model, where the layout discriminator model classifies the position of the bounding box in the image as true or false, and the affine discriminator model classifies the affine transformation as true or false. In another example, the discriminator models of another generator model (which outputs the shape of the object in the bounding box) may include a layout discriminator model and a shape discriminator model, where the layout discriminator model classifies the position of the shape in the image as true or false, and the shape discriminator model classifies the shape as true or false. The error may include a loss calculated based on the predictions of the discriminator models and / or the outputs of the corresponding generator models.

[0063] Second, the training engine 122 executes an unsupervised path to update the parameters of each generator model based on a first error (404). The training engine 122 also executes a supervised path that includes the ground truth of each generator model to update the parameters of the generator model based on a second error (406).

[0064] For example, the first error may include an unsupervised adversarial loss calculated by a first discriminator model of the generator model and / or a reconstruction loss associated with the random input of the generator model, and the second error may include a supervised adversarial loss calculated by a second discriminator model of the generator model. Thus, the unsupervised path can be used to improve the performance of the generator model in generating true bounding boxes and / or shapes of objects in the corresponding images, and the supervised path can be used with the ground truth to increase the diversity of the bounding boxes and / or shapes generated by the generator model.

[0065] Example Hardware Architecture

[0066] Figure 5 is a block diagram of a computer system 500 configured to implement one or more aspects of the various embodiments. In some embodiments, the computer system 500 is a server machine running in a data center or a cloud computing environment that provides scalable computing resources as a service over a network.

[0067] In various embodiments, the computer system 500 includes, but is not limited to, a central processing unit (CPU) 502 and a system memory 504 coupled to a parallel processing subsystem 512 via a memory bridge 505 and a communication path 513. The memory bridge 505 is further coupled to an I / O (input / output) bridge 507 via a communication path 506, and the I / O bridge 507 is in turn coupled to a switch 516.

[0068] In one embodiment, the I / O bridge 507 is configured to receive user input information from an optional input device 508, such as a keyboard or a mouse, and forward the input information to the CPU 502 for processing via the communication path 506 and the memory bridge 505. In some embodiments, the computer system 500 may be a server machine in a cloud computing environment. In such embodiments, the computer system 500 may not have an input device 508. Instead, the computer system 500 may receive equivalent input information by receiving commands in the form of messages sent over the network and received via a network adapter 518. In one embodiment, the switch 516 is configured to provide connections between the I / O bridge 507 and other components of the computer system 500, such as the network adapter 518 and various plug-in cards 520 and 521.

[0069] In one embodiment, the I / O bridge 507 is coupled to a system disk 514, which may be configured to store content, applications, and data for use by the CPU 502 and the parallel processing subsystem 512. In one embodiment, the system disk 514 provides non-volatile storage for applications and data and may include a fixed or removable hard disk drive, a flash memory device, and a CD-ROM (Compact Disc Read-Only Memory), a DVD-ROM (Digital Versatile Disc-ROM), a Blu-ray disc, an HD-DVD (High Definition DVD), or other magnetic, optical, or solid-state storage devices. In various embodiments, other components such as a universal serial bus or other port connection, a compact disc drive, a digital versatile disc drive, a film recording device, etc. may also be connected to the I / O bridge 507.

[0070] In various embodiments, the memory bridge 505 may be a northbridge chip and the I / O bridge 507 may be a southbridge chip. Additionally, the communication paths 506 and 513 and other communication paths within the computer system 500 may be implemented using any technically suitable protocol, including but not limited to AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol known in the art.

[0071] In some embodiments, the parallel processing subsystem 512 includes a graphics subsystem that delivers pixels to an optional display device 510, which may be any conventional cathode ray tube, liquid crystal display, light emitting diode display, or the like. In such embodiments, the parallel processing subsystem 512 includes circuitry optimized for graphics and video processing, including, for example, video output circuitry. As described in more detail below in connection with Figure 6 and Figure 7 such circuitry may be included across one or more parallel processing units (PPUs, also referred to as parallel processors) included in the parallel processing subsystem 512. In other embodiments, the parallel processing subsystem 512 includes circuitry optimized for general and / or computing processing. Again, such circuitry may be included across one or more PPUs included in the parallel processing subsystem 512, which are configured to perform such general and / or computing operations. In other embodiments, one or more PPUs included in the parallel processing subsystem 512 may be configured to perform graphics processing, general processing, and computing processing operations. The system memory 504 includes at least one device driver configured to manage the processing operations of one or more PPUs in the parallel processing subsystem 512.

[0072] In various embodiments, the parallel processing subsystem 512 may be associated with Figure 5The parallel processing subsystem 512 may be integrated with one or more other elements of the CPU 502 to form a single system. For example, the parallel processing subsystem 512 may be integrated with the CPU 502 and other connected circuits on a single chip to form a system on a chip (SoC).

[0073] In one embodiment, CPU 502 is the main processor of computer system 500, controlling and coordinating the operation of other system components. In one embodiment, CPU 502 issues commands that control the operation of the PPUs. In some embodiments, communication path 513 is a PCI Express link in which a dedicated lane is allocated to each PPU as known in the art. Other communication paths may also be used. The PPUs advantageously implement a highly parallel processing architecture. The PPUs may have any number of local parallel processing memories (PP memories).

[0074] It should be understood that the system shown herein is illustrative, and variations and modifications are possible. The connection topology (including the number and arrangement of bridges, the number of CPUs 502, and the number of parallel processing subsystems 512) can be modified as needed. For example, in some embodiments, the system memory 504 can be directly connected to the CPU 502, rather than being connected through the memory bridge 505, and other devices will communicate with the system memory 504 via the memory bridge 505 and the CPU 502. In other embodiments, the parallel processing subsystem 512 can be connected to the I / O bridge 507 or directly connected to the CPU 502, rather than being connected to the memory bridge 505. In other embodiments, the I / O bridge 507 and the memory bridge 505 can be integrated into a single chip, rather than existing as one or more discrete devices. Finally, in some embodiments, Figure 5 One or more of the components shown may not be present. For example, the switch 516 may be removed, and the network adapter 518 and the add-in cards 520 and 521 may be connected directly to the I / O bridge 507.

[0075] Figure 6 According to various embodiments, Figure 5 A block diagram of a parallel processing unit (PPU) 602 included in the parallel processing subsystem 512 of FIG. As described above, although Figure 6 One PPU 602 is depicted, but parallel processing subsystem 512 may include any number of PPUs 602. As shown, PPU 602 is coupled to a local parallel processing (PP) memory 604. PPU 602 and PP memory 604 may be implemented using one or more integrated circuit devices, such as a programmable processor, an application specific integrated circuit (ASIC), or a memory device, or in any other technically feasible manner.

[0076] In some embodiments, the PPU 602 includes a Graphics Processing Unit (GPU) that may be configured to implement a graphics rendering pipeline to perform various operations related to generating pixel data based on graphics data provided by the CPU 502 and / or the system memory 504. When processing graphics data, the PP memory 604 may be used as a graphics memory that stores one or more conventional frame buffers and, if needed, one or more other render targets. Additionally, the PP memory 604 can be used to store and update pixel data and transfer the final pixel data or display frame to an optional display device 510 for display. In some embodiments, the PPU 602 may also be configured for general-purpose processing and computing operations. In certain embodiments, the computer system 500 may be a server machine in a cloud computing environment. In these embodiments, the computer system 500 may not have a display device 510. Instead, the computer system 500 can generate equivalent output information by sending commands in the form of messages over a network via a network adapter 518.

[0077] In some embodiments, the CPU 502 is the main processor of the computer system 500 that controls and coordinates the operation of other system components. In one embodiment, the CPU 502 issues commands that control the operation of the PPU 602. In some embodiments, the CPU 502 writes a command stream for the PPU 602 to a data structure ( Figure 5 or Figure 6 not explicitly shown in the figure) that may be located in the system memory 504, the PP memory 604, or another storage location accessible to both the CPU 502 and the PPU 602. A pointer to the data structure is written to a command queue (also referred to herein as a push buffer) to initiate the processing of the command stream in the data structure. In one embodiment, the PPU 602 reads the command stream from the command queue and then executes the commands asynchronously with respect to the operation of the CPU 502. In embodiments where multiple push buffers are generated, an application may specify an execution priority for each push buffer via a device driver to control the scheduling of different push buffers.

[0078] In one embodiment, the PPU 602 includes an I / O (Input / Output) unit 605 that communicates with the rest of the computer system 500 via the communication path 513 and the memory bridge 505. In one embodiment, the I / O unit 605 generates data packets (or other signals) for transmission on the communication path 513 and also receives all incoming data packets (or other signals) from the communication path 513, directing the incoming data packets to the appropriate components of the PPU 602. For example, commands related to processing tasks can be directed to the host interface 606, while commands related to memory operations (e.g., reading from or writing to the PP memory 604) can be directed to the crossbar unit 610. In one embodiment, the host interface 606 reads each command queue and sends the command stream stored in the command queue to the front end 612.

[0079] As described above in connection with Figure 5 what has been said, the connection of the PPU 602 to the rest of the computer system 500 can be different. In some embodiments, the parallel processing subsystem 512 (which includes at least one PPU 602) is implemented as an insert card that can be inserted into an expansion slot of the computer system 500. In other embodiments, the PPU 602 can be integrated on a single chip with a bus bridge, such as the memory bridge 505 or the I / O bridge 507. Similarly, in other embodiments, some or all of the elements of the PPU 602 can be included with the CPU 502 in a single integrated circuit or chip system (SoC).

[0080] In one embodiment, the front end 612 sends processing tasks received from the host interface 606 to a work distribution unit (not shown) within the task / work unit 607. In one embodiment, the work distribution unit receives pointers to the processing tasks, which are encoded as task metadata (TMD) and stored in the memory. The pointers to the TMD are included in a command stream, which is stored as a command queue and received by the front end unit 612 from the host interface 606. Processing tasks that can be encoded as TMD include an index associated with the data to be processed and status parameters and commands that define how the data is to be processed. For example, the status parameters and commands can define a program to be executed on the data. Also for example, the TMD can specify the number and configuration of a set of cooperative thread arrays (CTA). Generally, each TMD corresponds to one task. The task / work unit 607 receives tasks from the front end 612 and ensures that the GPC 608 is configured to an active state before the processing tasks specified by each TMD are launched. A priority can also be specified for each TMD for scheduling the execution of the processing tasks. The processing tasks can also be received from the processing cluster array 630. Optionally, the TMD can include a parameter that controls whether the TMD is added to the head or tail of the processing task list (or added to the list of pointers to the processing tasks), thus providing another layer of control over the execution priority.

[0081] In one embodiment, the PPU 602 implements a highly parallel processing architecture based on the processing cluster array 630, which includes a set of C general processing clusters (GPC) 608, where C≥1. Each GPC 608 is capable of simultaneously executing a large number (e.g., hundreds or thousands) of threads, where each thread is an instance of a program. In various applications, different GPCs 608 can be allocated to process different types of programs or perform different types of computations. The allocation of the GPC 608 can vary according to the workload generated by each type of program or computation.

[0082] In one embodiment, the memory interface 614 includes a set of D partitioning units 615, where D≥1. Each partitioning unit 615 is coupled to one or more dynamic random access memories (DRAM) 620 residing in the PP memory 604. In some embodiments, the number of partitioning units 615 is equal to the number of DRAMs 620, and each partitioning unit 615 is coupled to a different DRAM 620. In other embodiments, the number of partitioning units 615 can be different from the number of DRAMs 620. Those of ordinary skill in the art will understand that the DRAM 620 can be replaced with any other technically suitable storage device. In operation, various rendering targets (such as texture maps and frame buffers) can be stored on the DRAM 620, allowing the partitioning units 615 to write portions of each rendering target in parallel, thus effectively using the available bandwidth of the PP memory 604.

[0083] In one embodiment, a given GPC 608 may process data for any DRAM 620 to be written into the PP memory 604. In one embodiment, the crossbar unit 610 is configured to route the output of each GPC 608 to the input of any partition unit 615 or to any other GPC 608 for further processing. The GPC 608 communicates with the memory interface 614 via the crossbar unit 610 to read from or write to the respective DRAMs 620. In some embodiments, the crossbar unit 610 is connected to the I / O unit 605 and is also connected to the PP memory 604 via the memory interface 614, enabling the processing cores in different GPCs 608 to communicate with the system memory 504 or other memories local to the non-PPU 602. In Figure 6 an embodiment, the crossbar unit 610 is directly connected to the I / O unit 605. In various embodiments, the crossbar unit 610 may use virtual channels to separate the traffic flow between the GPC 608 and the partition unit 615.

[0084] In one embodiment, the GPC 608 may be programmed to perform processing tasks related to various applications, including but not limited to linear and non-linear data transformations, filtering of video and / or audio data, modeling operations (e.g., applying physical laws to determine the position, velocity, and other properties of an object), image rendering operations (e.g., tessellation shaders, vertex shaders, geometry shaders, and / or pixel / fragment shader programs), general computing operations, etc. In operation, the PPU 602 is configured to transfer data from the system memory 504 and / or the PP memory 604 to one or more on-chip memory units, process the data, and write the resulting data back to the system memory 504 and / or the PP memory 604. Then, other system components (including the CPU 502, another PPU 602 in the parallel processing subsystem 512, or another parallel processing subsystem 512 in the computer system 500) may access the resulting data.

[0085] In one embodiment, any number of PPUs 602 may be included in the parallel processing subsystem 512. For example, multiple PPUs 602 may be provided on a single plug-in card, or multiple plug-in cards may be connected to the communication path 513, or one or more PPUs 602 may be integrated into a bridge chip. The PPUs 602 in a multi-PPU system may be the same or different from each other. For example, different PPUs 602 may have different numbers of processing cores and / or different numbers of PP memories 604. In implementations where there are multiple PPUs 602, these PPUs may operate in parallel to process data at a higher throughput than is possible with a single PPU 602. Systems including one or more PPUs 602 may be implemented in a variety of configurations and form factors, including but not limited to desktops, laptops, handheld personal computers or other handheld devices, servers, workstations, game consoles, embedded systems, and the like.

[0086] Figure 7 According to various embodiments, Figure 6 6. A block diagram of a general processing cluster (GPC) included in a parallel processing unit (PPU) 602 of FIG. As shown, GPC 608 includes, but is not limited to, a pipeline manager 705, one or more texture units 715, a pre-raster operation unit 725, a work distribution crossbar 730, and an L1.5 cache 735.

[0087] In one embodiment, GPC 608 may be configured to execute a large number of threads in parallel to perform graphics processing, general processing and / or computing operations. As used herein, "thread" refers to an instance of a specific program executed on a specific input data set. In some embodiments, a single instruction, multiple data (SIMD) instruction issuance technology is used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In other embodiments, a public instruction unit configured to issue instructions to a group of processing engines in GPC 608 is used, and a single instruction, multiple threads (SIMT) technology is used to support the parallel execution of a large number of threads that are usually synchronized. Different from the SIMD execution mechanism that all processing engines usually execute the same instruction, SIMT execution allows different threads to more easily follow different execution paths through a given program. It will be appreciated by those of ordinary skill in the art that the SIMD processing mechanism represents a functional subset of the SIMT processing mechanism.

[0088] In one embodiment, the operation of GPC 608 is controlled via pipeline manager 705, which distributes processing tasks received from a work distribution unit (not shown) in task / work unit 607 to one or more streaming multiprocessors (SMs) 710. Pipeline manager 705 may also be configured to control work distribution crossbar 730 by specifying the destination of processed data output by SM 710.

[0089] In various embodiments, GPC 608 includes a set of M SMs 710, where M ≥ 1. Additionally, each SM 710 includes a set of functional execution units (not shown), such as execution units and load-store units. The processing operations specific to any functional execution unit can be pipelined, which enables new instructions to be issued for execution before previous instructions have completed execution. Any combination of the functional execution units in a given SM 710 can be provided. In various embodiments, the functional execution units can be configured to support a variety of different operations, including integer and floating-point arithmetic (e.g., addition and multiplication), comparison operations, boolean operations (AND, OR, XOR), shift operations, and calculation of various algebraic functions (e.g., planar interpolation and trigonometric, exponential, and logarithmic functions, etc.). Advantageously, the same functional execution units can be configured to perform different operations.

[0090] In various embodiments, each SM 710 includes a plurality of processing cores. In one embodiment, SM 710 includes a large number (e.g., 128, etc.) of different processing cores. Each core can include a fully pipelined, single-precision, double-precision, and / or mixed-precision processing unit that includes a floating-point arithmetic logic unit and an integer arithmetic logic unit. In one embodiment, the floating-point arithmetic logic unit implements the IEEE 754-2008 standard for floating-point arithmetic. In one embodiment, the core includes 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.

[0091] In one embodiment, the tensor cores are configured to perform matrix operations, and in one embodiment, one or more tensor cores are included in the core. In particular, the tensor cores are configured to perform deep learning matrix arithmetic, such as convolution operations for neural network training and inference. In one embodiment, each tensor core operates on 4×4 matrix operations and performs matrix multiplication and accumulation operations D = A×B + C, where A, B, C, and D are 4×4 matrices.

[0092] In one embodiment, the matrix multiplication inputs A and B are 16-bit floating-point matrices, and the accumulation matrices C and D can be 16-bit floating-point matrices or 32-bit floating-point matrices. The tensor core performs operations on 16-bit floating-point input data using 32-bit floating-point accumulation. The 16-bit floating-point multiplication requires 64 operations to obtain a full-precision product, which is then added to other intermediate products using 32-bit floating-point addition to obtain a 4×4×4 matrix multiplication. In practice, the tensor core is used to perform larger two-dimensional or higher-dimensional matrix operations constructed from these smaller elements. APIs such as the CUDA 9 C++ API expose dedicated matrix load, matrix multiplication and accumulation, and matrix store operations to effectively use the tensor core from a CUDA-C++ program. At the CUDA level, the warp-level interface assumes a matrix of size 16×16, which spans all 32 threads of a warp.

[0093] Neural networks rely heavily on matrix math operations, and complex multi-layer networks require a large amount of floating-point performance and bandwidth to improve efficiency and speed. In various embodiments, thousands of processing cores optimized for matrix math operations are employed and provide performance up to dozens to hundreds of TFLOPS. The SM 710 provides a computing platform that can provide the performance required for artificial intelligence and machine learning applications based on deep neural networks.

[0094] In various embodiments, each SM 710 may also include multiple special function units (SFUs) that perform special functions (e.g., attribute evaluation, reciprocal square root, etc.). In one embodiment, the SFU may include a tree traversal unit configured to traverse a hierarchical tree data structure. In one embodiment, the SFU may include a texture unit configured to perform texture mapping filtering operations. The texture unit is configured to load a texture map (e.g., a two-dimensional texture pixel array) from memory and sample the texture map to produce a sampled texture value for use in a shading program executed by the SM. In various embodiments, each SM 710 also includes multiple load / store units (LSUs) that implement load and store operations between the shared memory / L1 cache and the register file inside the SM 710.

[0095] In one embodiment, each SM 710 is configured to process one or more thread groups. As used herein, a "thread group" or "warp" refers to a set of threads that execute the same program on different input data simultaneously, where one thread in the group is assigned to a different execution unit in the SM 710. The number of threads included in a thread group can be less than the number of execution units in the SM 710. In this case, when processing this thread group, some execution units may be idle during a cycle. A thread group can also include more threads than the number of execution units in the SM 710. In this case, processing may occur in consecutive clock cycles. Since each SM 710 can support up to G thread groups simultaneously, up to G*M thread groups can be executing in the GPC 608 at any given time.

[0096] In addition, in one embodiment, multiple related thread groups can be simultaneously active (in different execution phases) in the SM 710. This set of thread groups is referred to herein as a "cooperative thread array" ("CTA") or "thread array". The size of a particular CTA is equal to m*k, where k is the number of threads that execute simultaneously in a thread group, which is typically an integer multiple of the number of execution units in the SM 710, and m is the number of thread groups that are simultaneously active in the SM 710. In some embodiments, a single SM 710 can support multiple CTAs simultaneously, where the granularity of these CTAs is the granularity at which work is assigned to the SM 710.

[0097] In one embodiment, each SM 710 includes a level 1 (L1) cache, or uses space in a corresponding L1 cache external to the SM 710 to support load and store operations and the like performed by the execution units. Each SM 710 can also access a level 2 (L2) cache (not shown) shared among all GPCs 608 in the PPU 602. The L2 cache can be used to transfer data between threads. Finally, the SM 710 can also access off-chip "global" memory, which can include the PP memory 604 and / or the system memory 504. It should be understood that any memory external to the PPU 602 can be used as global memory. In addition, as Figure 7 shown, a level 1.5 (L1.5) cache 735 can be included in the GPC 608 and is configured to receive and hold data requested by the SM 710 from memory via the memory interface 614. Such data can include, but is not limited to, instructions, unified data, and constant data. In embodiments where there are multiple SM 710s in the GPC 608, the SM 710s can beneficially share the general instructions and data cached in the L1.5 cache 735.

[0098] In one embodiment, each GPC 608 may have an associated memory management unit (MMU) 720 configured to map virtual addresses to physical addresses. In various embodiments, the MMU 720 may reside within the GPC 608 or the memory interface 614. The MMU 720 includes a set of page table entries (PTEs) for mapping virtual addresses to the physical addresses of tiles or memory pages and an optional cache line index. The MMU 720 may include a translation lookaside buffer (TLB) or a cache that may reside within the SM 710, one or more L1 caches, or the GPC 608.

[0099] In one embodiment, in graphics and compute applications, the GPC 608 may be configured such that each SM 710 is coupled to a texture unit 715 to perform texture mapping operations, such as determining texture sampling locations, reading texture data, and filtering texture data.

[0100] In one embodiment, each SM 710 sends the processed tasks to a work distribution crossbar 730 to provide the processed tasks to another GPC 608 for further processing or store the processed tasks in an L2 cache (not shown), the parallel processing memory 604, or the system memory 504 via the crossbar unit 610. Additionally, a pre-raster operation (preROP) unit 725 is configured to receive data from the SM 710, direct the data to one or more raster operation (ROP) units within the partition unit 615, perform color mixing optimizations, organize pixel color data, and perform address translations.

[0101] It should be understood that the architectures described herein are illustrative and may be changed and modified. Additionally, any number of processing units, such as SM 710s, texture units 715, or preROP units 725, may be included within the GPC 608. Further, as described in connection with Figure 6 the PPU 602 may include any number of GPC 608s that are configured to be functionally similar to one another such that the execution behavior does not depend on which GPC 608 receives a particular processing task. Additionally, each GPC 608 operates independently of the other GPC 608s within the PPU 602 to perform the tasks of one or more applications.

[0102] In summary, the disclosed technology utilizes multiple generator models to synthesize and place objects in a scene based on a semantic representation of the scene in an image. A differentiable affine transformation representing a bounding box of an object can be passed between the generator models. An error resulting from predictions of a discriminator model associated with the generator models can be used together with the affine transformation to jointly update parameters of the generator models. A supervised path containing ground truth instances of the generator models can be additionally used for training to increase the diversity of the generator model outputs.

[0103] One technical advantage of the disclosed technology is that the affine transformation provides a differentiable link between two generator models. Thus, the differentiable link can be used to perform joint training and / or updating of the generator models so that the generator models operate as an end-to-end machine learning model that learns the joint distribution of the positions and shapes of different types of objects in the semantic representation of an image. Another technical advantage of the disclosed technology includes increasing the diversity of the generator model outputs, which is achieved by training the generator models using both a supervised path and an unsupervised path. Accordingly, the disclosed technology provides a technical improvement in machine learning models, computer systems, applications, and / or the training, execution, and performance of inserting object context into an image and / or scene.

[0104] 1. In some embodiments, a method includes: applying a first generator model to a semantic representation of an image to generate an affine transformation, wherein the affine transformation represents a bounding box associated with at least one region within the image; applying a second generator model to the affine transformation and the semantic representation to generate a shape of an object; and inserting the object into the image based on the bounding box and the shape.

[0105] 2. The method according to clause 1, further comprising calculating one or more errors associated with the first generator model and the second generator model based on an output from a discriminator model associated with at least one of the first generator model and the second generator model; and updating parameters of at least one of the first generator model and the second generator model based on the one or more errors.

[0106] 3. The method according to clauses 1-2, wherein updating the parameters includes: performing an unsupervised path to update the parameters of the first generator model and the second generator model based on a first error among the one or more errors; and performing a supervised path to update the parameters of the first generator model and the second generator model based on a second error among the one or more errors, the supervised path including ground truth of the first generator model and the second generator model.

[0107] 4. The method according to clauses 1-3, wherein the first error includes an unsupervised adversarial loss calculated from a first discriminator model of at least one of the first generator model and the second generator model.

[0108] 5. The method according to clauses 1-4, wherein the second error includes a supervised adversarial loss calculated from a second discriminator model of at least one of the first generator model and the second generator model.

[0109] 6. The method according to clauses 1-5, wherein the first error includes a reconstruction loss associated with a random input of at least one of the first generator model and the second generator model.

[0110] 7. The method according to clauses 1-6, wherein the first discriminator model associated with the first generator model includes a layout discriminator model or an affine discriminator model, the layout discriminator model classifying the position of the bounding box as true or false; the affine discriminator model classifying the affine transformation as true or false.

[0111] 8. The method according to clauses 1-7, wherein the first discriminator model associated with the second generator model includes a layout discriminator model or a shape discriminator model, the layout discriminator model classifying the position of the shape as true or false, the shape discriminator model classifying the shape as true or false.

[0112] 9. The method according to clauses 1-8, wherein inserting the object into the image based on the bounding box and the shape includes: applying the affine transformation to the shape.

[0113] 10. The method according to clauses 1-9, wherein each of the first generator model and the second generator model includes at least one of a variational autoencoder (VAE) and a spatial transformation network.

[0114] 11. In some embodiments, a non-transitory computer-readable medium stores instructions that, when executed by a processor, cause the processor to at least: apply a first generator model to a semantic representation of an image to generate an affine transformation, wherein the affine transformation represents a bounding box associated with at least one region within the image; apply a second generator model to the affine transformation and the semantic representation to generate a shape of an object; and insert the object into the image based on the bounding box and the shape.

[0115] 12. The non-transitory computer-readable medium according to clause 11 further includes program instructions to cause the processor to: calculate one or more errors associated with the first generator model and the second generator model based on an output from a discriminator model associated with at least one of the first generator model and the second generator model; and update parameters of at least one of the first generator model and the second generator model based on the one or more errors.

[0116] 13. For the non-transitory computer-readable medium according to any one of clauses 11-12, updating the parameters includes: performing an unsupervised path to update the parameters of the first generator model and the second generator model based on a first error among the one or more errors; and performing a supervised path to update the parameters of the first generator model and the second generator model based on a second error among the one or more errors, the supervised path including ground truths of the first generator model and the second generator model.

[0117] 14. For the non-transitory computer-readable medium according to any one of clauses 11-13, wherein the first error and the second error include at least one of an unsupervised adversarial loss, a supervised adversarial loss, and a reconstruction loss, the unsupervised adversarial loss is calculated by a first discriminator model, the supervised adversarial loss is calculated by a second discriminator model, and the reconstruction loss is associated with a random input of at least one of the first generator model and the second generator model.

[0118] 15. For the non-transitory computer-readable medium according to any one of clauses 11-14, wherein the discriminator model includes: a layout discriminator model that classifies a position of the bounding box as true or false; an affine discriminator model that classifies the affine transformation as true or false; a layout discriminator model that classifies a position of the shape as true or false; and a shape discriminator model that classifies the shape as true or false.

[0119] 16. In some embodiments, a system includes: a memory storing one or more instructions; and a processor that executes the instructions to at least: apply a first generator model to a semantic representation of an image to generate an affine transformation, wherein the affine transformation represents a bounding box associated with at least one region within the image; apply a second generator model to the affine transformation and the semantic representation to generate a shape of an object; and insert the object into the image based on the bounding box and the shape.

[0120] 17. The system according to clause 16, wherein the processor executes the instructions to: calculate one or more errors associated with the first generator model and the second generator model based on the output from a discriminator model associated with at least one of the first generator model and the second generator model; update parameters of at least one of the first generator model and the second generator model based on the one or more errors.

[0121] 18. The system according to any one of clauses 16 - 17, wherein updating the parameters includes: executing an unsupervised path to update the parameters of the first generator model and the second generator model based on a first error among the one or more errors; and executing a supervised path to update the parameters of the first generator model and the second generator model based on a second error among the one or more errors, the supervised path including ground truths of the first generator model and the second generator model.

[0122] 19. The system according to any one of clauses 16 - 18, wherein the first error and the second error include at least one of an unsupervised adversarial loss, a supervised adversarial loss, and a reconstruction loss, the unsupervised adversarial loss is calculated by a first discriminator model, the supervised adversarial loss is calculated by a second discriminator model, and the reconstruction loss is associated with a random input of at least one of the first generator model and the second generator model.

[0123] 20. The system according to any one of clauses 16 - 19, wherein inserting the object into the image based on the bounding box and the shape includes: applying the affine transformation to the shape.

[0124] Any combination of claim elements recited in any claim and / or any element described in any way in this application falls within the scope contemplated by this disclosure and protection.

[0125] The description of the various embodiments has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the embodiments.

[0126] Aspects of the present embodiment may be embodied as a system, a method, or a computer program product. Accordingly, various aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware (commonly referred to herein as a "module", "system"). In addition, any hardware and / or software technologies, processes, functions, components, engines, modules, or systems described in the present disclosure may be implemented as a circuit or a set of circuits. Further, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon.

[0127] Any combination of one or more computer-readable media may be used. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. For example, the computer-readable storage medium includes, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0128] Aspects of the present disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, enable the functions / acts specified in one or more blocks of the flowchart and / or block diagram to be implemented. Such a processor may be, but is not limited to, a general purpose processor, a special purpose processor, an application specific processor, or a field programmable gate array.

[0129] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of code, which includes one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative embodiments, the functions noted in the blocks may occur in the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block in the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or combinations of dedicated hardware and computer instructions.

[0130] Although the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the present disclosure may be devised without departing from its basic scope, the scope thereof being determined by the claims.

Claims

1. A method, comprising: Apply a first generator model to an input including a semantic representation of an image to generate an affine transformation, where the affine transformation represents a bounding box associated with at least one region within the image; Apply a second generator model to an input including (a) the affine transformation and (b) the semantic representation to generate a shape of an object; And Insert the object into the image based on the bounding box and the shape.

2. The method according to claim 1, further comprising: Calculate one or more errors associated with the first generator model and the second generator model based on an output from a discriminator model associated with at least one of the first generator model and the second generator model; Update parameters of at least one of the first generator model or the second generator model based on the one or more errors.

3. The method according to claim 2, wherein updating the parameter comprises: Execute an unsupervised path to update parameters of the first generator model and the second generator model based on a first error among the one or more errors; And Execute a supervised path to update the parameters of the first generator model and the second generator model based on a second error among the one or more errors, the supervised path including ground truths of the first generator model and the second generator model.

4. The method according to claim 3, wherein the first error comprises an unsupervised adversarial loss calculated from a first discriminator model of at least one of the first generator model or the second generator model.

5. The method according to claim 4, wherein the second error comprises a supervised adversarial loss calculated from a second discriminator model of at least one of the first generator model or the second generator model.

6. The method according to claim 3, wherein the first error comprises a reconstruction loss associated with a random input of at least one of the first generator model or the second generator model.

7. The method according to claim 2, wherein the first discriminator model associated with the first generator model comprises a layout discriminator model or an affine discriminator model, the layout discriminator model classifying the position of the bounding box as true or false, and the affine discriminator model classifying the affine transformation as true or false.

8. The method according to claim 2, wherein the first discriminator model associated with the second generator model comprises a layout discriminator model or a shape discriminator model, the layout discriminator model classifying the position of the shape as true or false, and the shape discriminator model classifying the shape as true or false.

9. The method according to claim 1, wherein inserting the object into the image based on the bounding box and the shape comprises: Apply the affine transformation to the shape.

10. The method according to claim 1, wherein each of the first generator model and the second generator model comprises at least one of a variational autoencoder (VAE) or a spatial transformation network.

11. The method according to claim 1, wherein the affine transformation is generated using a neural network.

12. A non - transitory computer - readable medium stores instructions that, when executed by a processor, cause the processor to at least: Apply a first generator model to an input including a semantic representation of an image to generate an affine transformation, where the affine transformation represents a bounding box associated with at least one region within the image; Apply a second generator model to an input including (a) the affine transformation and (b) the semantic representation to generate a shape of an object; And Insert the object into the image based on the bounding box and the shape.

13. The non - transitory computer - readable medium according to claim 12, further comprising program instructions to cause the processor to: Calculate one or more errors associated with the first generator model or the second generator model based on an output from a discriminator model associated with at least one of the first generator model and the second generator model; Update parameters of at least one of the first generator model or the second generator model based on the one or more errors.

14. The non - transitory computer - readable medium according to claim 13, wherein updating the parameters includes: Execute an unsupervised path to update parameters of the first generator model and the second generator model based on a first error among the one or more errors; And Execute a supervised path to update the parameters of the first generator model and the second generator model based on a second error among the one or more errors, the supervised path including ground truths of the first generator model and the second generator model.

15. The non - transitory computer - readable medium according to claim 14, wherein, The first error and the second error include at least one of an unsupervised adversarial loss, a supervised adversarial loss, or a reconstruction loss, the unsupervised adversarial loss is calculated by a first discriminator model, the supervised adversarial loss is calculated by a second discriminator model, and the reconstruction loss is associated with a random input of at least one of the first generator model or the second generator model.

16. The non - transitory computer - readable medium according to claim 13, wherein the discriminator model includes: A layout discriminator model that classifies a position of the bounding box as true or false; an affine discriminator model that classifies the affine transformation as true or false; a layout discriminator model that classifies a position of the shape as true or false; and a shape discriminator model that classifies the shape as true or false.

17. A system, comprising: A memory that stores one or more instructions; And A processor that executes the instructions to at least: Apply a first generator model to an input including a semantic representation of an image to generate an affine transformation, where the affine transformation represents a bounding box associated with at least one region within the image; Apply a second generator model to an input including (a) the affine transformation and (b) the semantic representation to generate a shape of an object; And Insert the object into the image based on the bounding box and the shape.

18. The system according to claim 17, wherein the processor further executes the instructions to: Calculate one or more errors associated with the first generator model or the second generator model based on an output from a discriminator model associated with at least one of the first generator model and the second generator model; Update parameters of at least one of the first generator model or the second generator model based on the one or more errors.

19. The system according to claim 18, wherein updating the parameter comprises: Based on a first error among the one or more errors, perform an unsupervised path to update parameters of the first generator model and the second generator model; and Based on a second error among the one or more errors, perform a supervised path to update the parameters of the first generator model and the second generator model, the supervised path including ground truths of the first generator model and the second generator model.

20. The system according to claim 19, wherein the first error and the second error comprise at least one of an unsupervised adversarial loss, a supervised adversarial loss, or a reconstruction loss, the unsupervised adversarial loss is calculated by a first discriminator model, the supervised adversarial loss is calculated by a second discriminator model, and the reconstruction loss is associated with a random input of at least one of the first generator model or the second generator model.

21. The system according to claim 17, wherein inserting the object into the image based on the bounding box and the shape comprises: Apply the affine transformation to the shape.

Citation Information

Patent Citations

  • Method for manipulating image through sliding attribute based on attenuation network

    CN107330954A

  • Image display apparatus, image display method, and computer program product

    US20150116349A1

  • Joint object and object part detection using web supervision

    US20170330059A1