Generating composite image and / or training machine learning model based on composite image

By identifying and rendering the size of the 3D object model in the synthetic image, generating the synthetic image and automatically assigning the label, the problem of the domain gap between the synthetic image and the real image in the prior art is solved, and the performance of the machine learning model is improved.

CN119991898APending Publication Date: 2025-05-13GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411897077.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-11-16
Filing Date
2019-11-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art has significant domain gap problems when generating synthetic images and training machine learning models based on synthetic images, resulting in poor performance of models when predicting real visual data.

Method used

By identifying the size of the foreground 3D object model in the foreground layer of the synthetic image, and rendering the background 3D object model in the corresponding position in the background layer, combining the foreground and background layers to generate the composite image, and automatically assigning ground truth tags, we provide training examples for training machine learning models.

Benefits of technology

It effectively reduces the domain gap, improves the performance of machine learning models on real visual data, and reduces the computing and network resource consumption of generating training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991898A_ABST
    Figure CN119991898A_ABST
Patent Text Reader

Abstract

Particular techniques for generating composite images and / or for training machine learning model (s) based on the generated composite images. For example, a machine learning model is trained based on training instances each including a generated composite image and ground truth tag (s) for the generated composite image. After the training of the machine learning model is completed, the trained machine learning model may be deployed on one or more robots and / or one or more computing devices.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application with the application date of November 15, 2019, application number 201980040171.8, and invention name "Generating synthetic images and / or training (one or more) machine learning models based on synthetic images". Technical Field

[0002] The present application relates to a method, system, computer program product, transient or non-transitory computer-readable storage medium implemented by one or more processors for generating synthetic images and / or training (one or more) machine learning models based on synthetic images. Background Art

[0003] Detecting and / or classifying objects in challenging environments is a necessary skill for many machine vision and / or robotics tasks. For example, in order for a robot to manipulate (e.g., grasp, push, and / or pull) an object, the robot must be able to at least detect the object in the visual data (e.g., determine a 2D and / or 3D bounding box corresponding to the object). As another example, a robot can use object detection and classification to identify a certain type(s) of object(s) and avoid collisions with that certain type(s) of object(s).

[0004] Various machine learning models have been proposed for object detection and / or classification. For example, deep convolutional architectures such as Faster R-CNN, SSD, R-FCN, Yolo9000, and RetinaNet have been proposed for object detection. The training of such models (which may include millions of parameters) requires a large amount of labeled training data to achieve state-of-the-art results.

[0005] Training data including real images and corresponding human-assigned labels (e.g., labeled (one or more) bounding boxes) have been used to train such models. However, generating such training data can take up a lot of computing and / or network resources. For example, when generating (one or more) labels assigned by humans for real images, the real images must be transmitted to the client device used by the corresponding human reviewer. The real images are rendered on the client device, and then the human reviewer must use the client device to review the images and provide (one or more) user interface inputs to assign (one or more) labels. Then, the (one or more) labels assigned by humans are transmitted to the server, where they are paired with the real images and used to train the corresponding model. When considering labeling hundreds of thousands (or even millions) of real images, the transmission to and from the client device consumes a lot of network resources, and the rendering of the images and the processing of (one or more) user interface inputs consume a lot of client device resources. Moreover, the labels assigned by humans can include errors (e.g., incorrectly placed bounding boxes) and human labeling can be a time-consuming process. In addition, setting up various real scenes and capturing real images can also be resource-intensive.

[0006] Synthetic training data including synthetic images and automatically assigned labels have also been used to train such models. Synthetic training data can overcome some of the deficiencies of training data including real images and human-assigned labels. However, training machine learning models primarily or solely on synthetic training data with synthetic images generated according to various prior art techniques still results in a significant domain gap. This can be due to, for example, differences between synthetic images and real images. The domain gap can result in poor performance of machine learning models trained with synthetic training data when the machine learning models are used to make predictions based on real visual data. Summary of the invention

[0007] Embodiments disclosed herein are directed to specific techniques for generating synthetic images and / or for training (one or more) machine learning models (e.g., neural network models) based on the generated synthetic images (e.g., training based on training instances, each training instance comprising a generated synthetic image and (one or more) ground truth labels for the generated synthetic image).

[0008] In some embodiments, a method implemented by one or more processors is provided, the method comprising identifying a size at which a foreground three-dimensional (3D) object model is rendered in a foreground layer of a composite image. For each of a plurality of randomly selected background 3D object models, the method further comprises: rendering the background 3D object model at a corresponding background position in the background layer of the composite image, with a corresponding rotation and at a corresponding size determined based on the size at which the foreground 3D object model is rendered. The method further comprises rendering the foreground 3D object model at a foreground position in the foreground layer. The rendering of the foreground 3D object model is at the size and at a given rotation of the foreground 3D object model. The method further comprises generating a composite image based on fusing the background layer and the foreground layer, and assigning a ground truth label for rendering the foreground 3D object model to the composite image. The method further comprises providing a training instance, the training instance comprising a composite image paired with a ground truth label, for training at least one machine learning model based on the training instance.

[0009] These and other implementations of the technology disclosed herein may include one or more of the following features.

[0010] In some embodiments, the method further comprises determining a range of scaling values ​​based on the size of the rendered foreground 3D object model. In those embodiments, for each of the selected background 3D object models, rendering the selected background 3D object model at a corresponding size comprises: selecting a corresponding scaling value from the range of scaling values; scaling the selected background 3D object model based on the corresponding scaling value to generate a corresponding scaled background 3D object model; and rendering the scaled background 3D object model at a corresponding background position in the background layer. In some versions of those embodiments, determining the range of scaling values ​​comprises determining a lower scaling value of the scaling value and determining an upper scaling value of the scaling value. In some of these versions, determining the lower scaling value is determined based on: if the lower scaling value is used to scale any one of the background 3D object models before rendering, then the corresponding size will be at a lower percentage limit of the foreground size. In those versions, determining the upper scaling value is determined based on: if the upper scaling value is used to scale any one of the background 3D object models before rendering, then the corresponding size will be at an upper percentage limit of the foreground size. The foreground size may be based on the size of the foreground 3D object model to be rendered in the foreground layer and / or the size(s) of the additional foreground 3D object model(s) to be rendered. For example, the foreground size may be the same as the size at which the foreground 3D object model is rendered, or may be a function of that size and at least one additional size of at least one additional foreground 3D object model also rendered in the foreground layer. The lower percentage limit may be, for example, between 70% and 99% and / or the upper percentage limit may be, for example, between 100% and 175%. Optionally, for each of the selected background 3D object models, selecting the corresponding scaling value comprises randomly selecting the corresponding scaling value from among all scaling values ​​within the range of scaling values.

[0011] In some embodiments, for each of a plurality of selected background 3D object models, rendering the selected background 3D object model at a corresponding background position includes selecting the background position based on other background 3D objects that have not yet been rendered at the background position. In some of those embodiments, rendering the selected background 3D object model is iteratively performed each time for the addition of the selected background 3D object model. The iterative rendering of the selected background 3D object model may be performed until it is determined that one or more coverage conditions are met. The (one or more) coverage conditions may include, for example, that all positions of the background layer have content rendered thereon, or may include that there is no exposed area greater than a threshold size (e.g., n consecutive pixel sizes).

[0012] In some embodiments, the method further comprises: selecting an additional background 3D object model; identifying a random position within a bounded area that bounds the rendering of the foreground 3D object model; and rendering the additional background 3D object model in the random position and an occlusion layer of the composite image. Rendering the additional background 3D object model may optionally include scaling the additional background 3D object model prior to rendering so as to occlude only a portion of the rendering of the foreground 3D object model. In those embodiments, generating the composite image is based on fusing the background layer, the foreground layer, and the occlusion layer. In some versions of those embodiments, the degree of occlusion may be based on the size of the foreground object to be rendered.

[0013] In some embodiments, the foreground 3D object model is selected from a corpus of foreground 3D object models, the background 3D object model is randomly selected from a corpus of background 3D object models, and the corpus of foreground objects is disjoint from the corpus of background objects.

[0014] In some embodiments, the method further includes generating an additional synthetic image, the additional synthetic image including a foreground 3D object model rendered at a smaller size than the size at which the foreground 3D object is rendered in the synthetic image. The additional synthetic image also includes an alternative background 3D object model rendered at a corresponding alternative size, the alternative size being determined based on the smaller size at which the foreground 3D object model is rendered in the additional synthetic image. In those embodiments, the method further includes: assigning an additional ground truth label to the additional synthetic image for rendering the foreground 3D object model in the additional synthetic image; and providing an additional training instance, the additional training instance including the additional synthetic image paired with the additional ground truth label for further training of the at least one machine learning model. Further training of the at least one machine learning model based on the additional training instance can be based on the foreground object rendered at a smaller size, after training of the at least one machine learning model based on the training instance. Optionally, the additional synthetic image can include an occluding object that occludes the rendering of the 3D object model to a greater extent than any occlusion of the rendering of the 3D object model in the synthetic image. This greater degree of occlusion can be based on the foreground object rendered at a smaller size in the additional synthetic image. Optionally, the method may further include training the machine learning model based on the training instance, and training the machine learning model based on additional training instances after training the machine learning model based on the training instance.

[0015] In some embodiments, the ground truth label includes a bounding shape of the foreground object, a six-dimensional (6D) pose of the foreground object, and / or a classification of the foreground object. For example, the ground truth label may include a bounding shape, and the bounding shape may be a two-dimensional bounding box.

[0016] In some implementations, rendering the foreground 3D object model at the foreground position in the foreground layer includes randomly selecting the foreground position from a plurality of foreground positions that do not yet have a foreground 3D object model rendered.

[0017] In some embodiments, a method implemented by one or more processors is provided, the method comprising: selecting a foreground three-dimensional (3D) object model; and generating a plurality of first-scale rotations of the foreground 3D object model at a first scale of the foreground 3D object model. The method further comprises, for each of the plurality of first-scale rotations of the foreground 3D object model, rendering the foreground 3D object model at a corresponding one of the first-scale rotations and at a first scale in a corresponding randomly selected position in a corresponding first-scale foreground layer. The method further comprises generating a first-scale synthetic image. Generating each of the corresponding first-scale synthetic images comprises: fusing a corresponding one of the corresponding first-scale foreground layers with a corresponding one of a plurality of non-intersecting first-scale background layers, the plurality of non-intersecting first-scale background layers each comprising a corresponding rendering of a corresponding randomly selected background 3D object model. The method further comprises generating first-scale training instances, each first-scale training instance comprising a corresponding one of the first-scale synthetic images, and a corresponding ground truth label for rendering the foreground 3D object model in a corresponding one of the first-scale synthetic images. The method further comprises generating a plurality of second-scale rotations of the foreground 3D object model at a second scale of the foreground 3D object model that is a scale smaller than the first scale. The method also includes, for each of a plurality of second scale rotations of the foreground 3D object model: rendering the foreground 3D object model in a corresponding randomly selected position in a corresponding second scale foreground layer at a corresponding one of the second scale rotations and at a second scale. The method also includes generating a second scale synthetic image. Generating each of the corresponding second scale synthetic images includes fusing a corresponding one of the corresponding second scale foreground layers with a corresponding one of a plurality of non-intersecting second scale background layers, each of the plurality of non-intersecting second scale background layers including a corresponding rendering of a corresponding randomly selected background 3D object model. The method also includes generating second scale training instances, each second scale training instance including a corresponding one of the second scale synthetic images, and a corresponding ground truth label for rendering the foreground 3D object model in the corresponding one of the second scale synthetic images. The method also includes training a machine learning model based on the first scale training instance before training the machine learning model based on the second scale training instance.

[0018] These and other implementations of the technology disclosed herein may include one or more of the following features.

[0019] In some embodiments, in the first scale background layer, the corresponding rendering sizes of the randomly selected background 3D object models are smaller than the corresponding rendering sizes of the randomly selected background 3D object models in the second scale background layer.

[0020] In some embodiments, in a first-scale background layer, corresponding renderings of randomly selected background 3D object models are all within a threshold percentage range of the first scale; and in a second-scale background layer, corresponding renderings of randomly selected background 3D object models are all within a threshold percentage range of the second scale.

[0021] In some embodiments, a method implemented by one or more processors is provided, the method comprising: training a machine learning model using first-scale training instances, each first-scale training instance comprising a corresponding first-scale composite image and at least one corresponding label. The corresponding first-scale composite images each comprise one or more corresponding first-scale foreground objects, each of which is within a first size range. The method also comprises, after training the machine learning model using the first-scale training instances, and based on having trained the machine learning model using the first-scale training instances: further training the machine learning model using second-scale training instances. Each second-scale training instance comprises a corresponding second-scale composite image and at least one corresponding label. The corresponding second-scale composite images each comprise one or more corresponding second-scale foreground objects, each of which is within a second size range. The sizes of the second size range are all smaller than the sizes of the first size range. Optionally, the corresponding first-scale composite images of the first-scale training instances do not have any foreground objects within the second size range.

[0022] Other implementations of the methods and techniques disclosed herein can each optionally include one or more of the following features.

[0023] In some embodiments, the corresponding first-scale composite image includes a corresponding first-scale foreground object's corresponding first-scale occlusion extent that is less than (on average, or on an individual basis) a corresponding second-scale foreground object's corresponding second-scale occlusion extent.

[0024] Other embodiments may include one or more non-transitory computer-readable storage media storing instructions executable by a processor (e.g., a central processing unit (CPU) or a graphics processing unit (GPU)) to perform a method such as one or more of the methods described herein. Another embodiment may include a system of one or more computers including one or more processors operable to execute stored instructions to perform a method such as one or more (e.g., all) aspects of one or more of the methods described herein.

[0025] In some embodiments, a method implemented by one or more processors is provided, the method comprising: for each of a plurality of selected background 3D object models: rendering the background 3D object model at a corresponding random background position in a background layer and with a random hue, for each of one or more selected foreground 3D object models: rendering the foreground 3D object model at a corresponding random foreground position in the foreground layer; for each of one or more selected occlusion 3D object models: rendering the occlusion 3D object model at a corresponding random occlusion position in the occlusion layer; wherein, during the rendering of the background object model, one or more foreground 3D object models and the occlusion 3D object model, at least one light source is randomly applied with a random light color; generating a composite image based on a fusion of the rendering of the background layer, the foreground layer and the occlusion layer; assigning ground truth labels for the rendering of one or more foreground 3D object models to the composite image; and providing a training instance comprising the composite image paired with the ground truth labels for training at least one machine learning model based on the training instance.

[0026] In some embodiments, a system is provided, comprising: a memory storing instructions; and one or more processors executing the instructions so that the one or more processors: for each of a plurality of selected background 3D object models: render the background 3D object model at a corresponding random background position in a background layer and with a random hue, for each of one or more selected foreground 3D object models: render the foreground 3D object model at a corresponding random foreground position in the foreground layer; for each of one or more selected occlusion 3D object models: render the occlusion 3D object model at a corresponding random occlusion position in the occlusion layer; wherein, during the rendering of the background object model, one or more foreground 3D object models and the occlusion 3D object model, at least one light source is randomly applied with a random light color; a synthetic image is generated based on the rendering of the fused background layer, the foreground layer and the occlusion layer; a ground truth label for the rendering of one or more foreground 3D object models is assigned to the synthetic image; and a training instance comprising the synthetic image paired with the ground truth label is provided for training at least one machine learning model based on the training instance.

[0027] In some embodiments, a method implemented by one or more processors is provided, the method comprising: identifying a size at which a foreground three-dimensional (3D) object model is rendered in a foreground layer of a composite image; for each of a plurality of randomly selected background 3D object models: rendering the background 3D object model at a corresponding background position in the background layer of the composite image, with a corresponding rotation and at a corresponding size determined based on the size at which the foreground 3D object model is rendered; rendering the foreground 3D object model at a foreground position in the foreground layer, the foreground 3D object model being rendered with the size and at a given rotation of the foreground 3D object model; selecting additional background 3D object models; and selecting the background 3D object models in the foreground layer. Identifying random locations within a determined bounded area of ​​rendering; rendering an additional background 3D object model in the identified random location within the determined bounded area and in an occlusion layer of a synthetic image, wherein rendering the additional background 3D object model includes scaling the additional background 3D object model prior to rendering so as to occlude only a portion of the rendering of the foreground 3D object model; generating a synthetic image based on fusing the background layer with the foreground layer and the occlusion layer; assigning a ground truth label for the rendering of the foreground 3D object model to the synthetic image; and providing the synthetic image paired with the ground truth label as a training instance for training at least one machine learning model for object detection and / or classification tasks based on the training instance.

[0028] In some embodiments, a method implemented by one or more processors is provided, the method comprising: identifying a size at which a foreground three-dimensional (3D) object model is rendered in a foreground layer of a synthetic image; for each of a plurality of randomly selected background 3D object models: rendering the background 3D object model at a corresponding background position in the background layer of the synthetic image, with a corresponding rotation and at a corresponding size determined based on the size at which the foreground 3D object model is rendered, wherein rendering the selected background 3D object model at the corresponding background position comprises selecting the background position based on the fact that no other background 3D object has been rendered at the background position, and wherein rendering is performed iteratively, each time for an additional one of the background 3D object models, until it is determined that all positions of the background layer have content rendered thereon; rendering the foreground 3D object model at the foreground position in the foreground layer, the rendering of the foreground 3D object model being at the size and at a given rotation of the foreground 3D object model; generating a synthetic image based on fusing the background layer and the foreground layer; assigning a ground truth label for the rendering of the foreground 3D object model to the synthetic image; and providing a training instance comprising the synthetic image paired with the ground truth label for training at least one machine learning model based on the training instance.

[0029] In some embodiments, a method implemented by one or more processors is provided, the method comprising: selecting a foreground three-dimensional (3D) object model; generating multiple first-scale rotations of the foreground 3D object model with the foreground 3D object model at a first scale; for each of the multiple first-scale rotations of the foreground 3D object model: rendering the foreground 3D object model at a corresponding one of the first-scale rotations and at the first scale in a corresponding randomly selected position in a corresponding first-scale foreground layer; generating a first-scale synthetic image, wherein generating each of the corresponding first-scale synthetic images comprises: fusing a corresponding one of the corresponding first-scale foreground layers with a corresponding one of a plurality of non-intersecting first-scale background layers, the plurality of non-intersecting first-scale background layers each comprising a corresponding rendering of a corresponding randomly selected background 3D object model; generating first-scale training instances, the first-scale training instances each comprising a corresponding one of the first-scale synthetic images, and a corresponding ground truth label for the rendering of the foreground 3D object model in the corresponding one of the first-scale synthetic images; generating multiple first-scale synthetic images of the foreground 3D object model with the foreground 3D object model at a second scale that is smaller than the first scale. Two-scale rotation; for each of multiple second-scale rotations of the foreground 3D object model: render the foreground 3D object model in a corresponding randomly selected position in the corresponding second-scale foreground layer with a corresponding one of the second-scale rotations and at a second scale; generate a second-scale synthetic image, generating each of the corresponding second-scale synthetic images including: fusing a corresponding one of the corresponding second-scale foreground layer with a corresponding one of a plurality of non-intersecting second-scale background layers, the plurality of non-intersecting second-scale background layers each including a corresponding rendering of a corresponding randomly selected background 3D object model; generate second-scale training instances, the second-scale training instances each including a corresponding one of the second-scale synthetic images, and a corresponding ground truth label for the rendering of the foreground 3D object model in a corresponding one of the second-scale synthetic images, wherein, in the first-scale background layer, the corresponding rendering size of the corresponding randomly selected background 3D object model is smaller than the corresponding rendering size of the corresponding randomly selected background 3D object model in the second-scale background layer; and train the machine learning model based on the first-scale training instances before training the machine learning model based on the second-scale training instances.

[0030] It should be appreciated that all combinations of the foregoing concepts and additional concepts described in more detail herein are contemplated as being part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 An example environment is illustrated in accordance with various implementations disclosed herein.

[0032] Figure 2 is a flow chart illustrating an example method of generating a background layer according to various embodiments disclosed herein.

[0033] Figure 3 is a flowchart illustrating an example method of generating a foreground layer, generating a composite image based on fusing the corresponding foreground layer, background layer, and optional occlusion layer, and generating a training example including the composite image according to various embodiments disclosed herein.

[0034] Figure 4 is a flowchart illustrating an example method for training a machine learning model according to a curriculum according to various embodiments disclosed herein.

[0035] Figure 5A , 5B , 5C and 5D illustrate example composite images according to various embodiments disclosed herein.

[0036] Figure 6 An example architecture of a robot is schematically depicted.

[0037] Figure 7 An example architecture of a computer system is schematically depicted. DETAILED DESCRIPTION

[0038] Some embodiments disclosed herein generate purely synthetic training instances, each training instance comprising a synthetic image and corresponding (one or more) synthetic ground truth labels. The (one or more) synthetic ground truth labels may include, for example, (one or more) 2D bounding boxes of (one or more) foreground objects in the synthetic image, (one or more) classifications of (one or more) foreground objects, and / or (one or more) other labels. In some of those embodiments, each synthetic image is generated by fusing / blending three image layers: (1) a purely synthetic background layer; (2) a purely synthetic foreground layer; (3) an optional purely synthetic occlusion layer. Embodiments for generating each of these three layers are now addressed in turn, starting with a description of generating the background layer.

[0039] The proposed techniques for generating background layers can seek to: maximize background clutter; minimize the risk of having the same background layer present in multiple synthetic images; create background layers with structures similar in scale to the object(s) in the corresponding foreground layer; and / or present foreground and background layers from the same domain. Experiments show that these principles (alone and / or in combination) can create synthetic images that train machine learning models to learn the geometry and visual appearance of objects when used in training examples. Moreover, (one or more) of such principles mitigate the opportunity for training models to learn to simply distinguish synthetic foreground objects from background objects from foreground objects and background objects with different characteristics (e.g., different object sizes and / or noise distributions).

[0040] The background layer can be generated from a corpus / dataset of textured background 3D object models. A large number (e.g., 10,000 or more, 15,000 or more) of background 3D object models can be included in the corpus. Moreover, the corpus of background 3D object models does not intersect with the corpus of foreground 3D object models. In other words, no background 3D object model can be included in the foreground 3D object model. All background 3D object models can optionally be initially demeaned and scaled so that they fit into a unit sphere.

[0041] The background layer can be generated by successively selecting areas in the background where other background 3D object models have not been rendered ("exposed areas") and rendering random background 3D object models onto each selected area. Each background 3D object model can be rendered with random rotation, and the process is repeated until the entire background is covered by the synthetic background object (i.e., until there are no exposed areas). The risk of having the same background layer in multiple background images can be mitigated by randomly selecting background 3D object models, rendering each selected background 3D object model with random rotation and translation, and / or by identifying exposed areas. As used herein, it should be noted that random includes true random and pseudo-random.

[0042] In various embodiments, the size of the projected background objects in the background layer may be determined relative to the size of the foreground object(s) rendered in the foreground layer that is to be fused with the background layer when the composite image is subsequently generated. In other words, the background objects in the background layer may be similar in scale to the foreground object(s) in the corresponding foreground layer. This may enable a machine learning model trained on such composite images to learn the geometry and visual appearance of objects while mitigating the chances of the trained model instead learning to distinguish composite foreground objects from background objects simply based on size differences between the background and foreground objects.

[0043] In some embodiments, a randomized isotropic scale S can be generated when generating projected background objects having a similar size as (one or more) foreground objects (e.g., 90% to 150% of the size, or other size range). The randomized isotropic scale can be applied to the selected background 3D object models before rendering. As mentioned above, the background 3D object models of the corpus can initially all have similar scales (e.g., they can first be de-averaged and scaled to fit into a unit sphere). The randomized isotropic scale applied to the selected background 3D object models can be used to create background objects so that their projected size on the image plane is similar to the foreground size, where the foreground size can be based on the (one or more) size of the (one or more) foreground objects (e.g., the average size of the (one or more) foreground objects). For example, a scale range S = [s min ,s max ], which scale range represents the scaling values ​​that can be applied to the background 3D object models to make them appear within [0.9, 1.5] (or other percentage range) of the foreground size. The foreground size can be calculated by the average projection size of the foreground object size from all projections of all foreground objects rendered in the current image (or by any other statistical average). When generating each background layer, a random subset can be generated This ensures that the background layer is created not only with an even distribution of objects of all sizes, but also with predominantly large or small objects. bg Randomly draw the isotropic scaling value s applied to each background 3D object model in bg , so that the background object sizes in the image are evenly distributed. In other words, when selecting a scaling value for a given background 3D object model to be rendered, the scaling value can be randomly selected from a range of scaling values ​​with an even distribution. Thus, some background layers will include evenly distributed object sizes, other background layers will have mostly (or only) large (relative to the scaling range) objects, and other background layers will have mostly (or only) small (relative to the scaling range) objects.

[0044] In some embodiments, for each background layer, the texture of each rendered object can be converted into a hue, saturation, value (HSV) space, the hue value of the object is randomly changed, and after changing the hue value, the HSV space can be converted back to a red, green, blue (RGB) space. This can diversify the background layers and ensure that the background colors are well distributed. Any other (one or more) foreground and / or background color transforms can be applied additionally and / or alternatively. Thus, by applying (one or more) color transforms, the risk of having the same background layer in multiple composite images is further mitigated.

[0045] Now turn to generating each foreground layer, each foreground layer may include (one or more) renderings of (one or more) foreground 3D object models. As described below, the rotation of the foreground 3D object model in the rendering (a combination of in-plane rotation and out-of-plane rotation) can be selected from the rotation set generated for the foreground 3D object model, and the size of the rendering can be selected based on its consistency with the rendering used to generate the corresponding background layer. In other words, for the foreground 3D object model, it will be rendered in multiple different foreground layers (once per layer), and each rendering will include the corresponding object with a disparate rotation and / or at a different size (relative to other renderings). The rotation set generated for the foreground 3D object model can be determined based on the desired pose space to be covered for the object. In other words, the rotation set can cover the pose space of the object together, for which it is expected that once trained, the machine learning model can be used to predict (one or more) values ​​of the object. The rotation and size of the corresponding object can be optionally determined based on ensuring that each foreground object is rendered with multiple disparate rotations and / or different sizes across multiple synthetic images, and used for foreground generation. For example, the same n completely different rotations can be generated at each of a plurality of completely different scales, and each rotation, scale pairing of each foreground object can be rendered in at least one (and optionally only one) foreground layer. In these and other ways, each foreground object will appear in multiple composite images at completely different rotations and different sizes. Moreover, the training of the machine learning model can be initially based on composite images with larger sized foreground objects, then based on composite images with smaller sized foreground objects, and then based on composite images with even smaller sized foreground objects. As described herein, training in this manner can result in improved performance (e.g., accuracy and / or recall) of the trained machine learning model.

[0046] For rendering purposes, a certain degree of cropping of foreground objects at image boundaries may be allowed (e.g., up to 50% cropping or other cropping threshold). Additionally, a certain degree of overlap between pairs of rendered foreground objects may be allowed for rendering purposes (e.g., up to 30% overlap or other threshold overlap). For each object, it may be placed at a random location, and if the initial attempted random location(s) fail (e.g., due to excessive cropping at image boundaries and / or too much overlap), then additional attempts at placement may be made. For example, up to n=100 (or other threshold) random placement attempts may be performed. If the foreground object being processed cannot be placed within the foreground layer within a threshold number (and / or duration) of attempts due to violation of cropping constraints, overlap constraints, and / or other constraints(s) - then processing of the current foreground layer may be stopped, and the foreground object being processed may be placed instead in the next foreground layer to be processed. In other words, in some embodiments, multiple foreground 3D object models may be rendered in the foreground layer, and rendering of new objects may continue until it is determined that no more foreground objects can be placed through random attempts (e.g., after 100 attempts or other threshold) without violating (one or more) constraints.

[0047] As mentioned above, for each foreground 3D object model, a large rotation set can be generated. The rotation set can, for example, uniformly cover the rotation space in which the corresponding object is expected to be detected. As an example of generating a large rotation set for a foreground object, an icosahedron can be recursively divided, which is the largest convex regular polyhedron of the foreground object 3D model. This can produce evenly distributed vertices on a sphere and each vertex represents a distinct view defined by two out-of-plane rotations. In addition to these two out-of-plane rotations, in-plane rotations can also be sampled equally. In addition, the distance at which the foreground object is rendered can be sampled inversely proportional to the projected size of the foreground object to ensure that the pixel coverage of the projected object changes approximately linearly between continuous scale levels.

[0048] In contrast to background layer generation, the rendering of foreground objects can optionally occur based on a curriculum strategy. In other words, this means that there can be a determined schedule of what step size each foreground object and rotation should be rendered (or at least what step size the corresponding synthetic image should be provided for training). For example, rendering can start with the scale closest to the camera and then gradually move to the farthest scale. Therefore, each object initially looks the largest in the initial synthetic image and is therefore easier to learn for the machine learning model being trained. As learning proceeds, objects become smaller in subsequent initial synthetic images and are more difficult to learn for the machine learning model being trained. For each scale of the foreground object, all considered out-of-plane rotations can be iterated through, and for each out-of-plane rotation, all considered in-plane rotations can be iterated through, thereby creating multiple rotations for the foreground object at the corresponding scale. Once a rotation for a scale has been generated for all foreground objects, all foreground objects can be iterated during the generation of the foreground layer, and each of them is rendered with a given rotation at a random position using a uniform distribution. As described herein, when generating a composite image, a foreground layer having rendered objects at corresponding sizes (one or more) may be fused with a background layer having background object sizes based on the sizes (one or more). After all foreground objects have been processed at all rotations for a given size / scale level, the process may be repeated for the next (smaller) scale level.

[0049] Turning now to occlusion layer generation, an occlusion layer can be generated in which random objects (e.g., from a corpus of background 3D objects) can partially occlude (one or more) foreground objects by including them in corresponding positions in the corresponding foreground layer. In some embodiments, this is accomplished by determining a bounding box (or other bounding shape) for each rendered foreground object in the foreground layer and by rendering randomly selected occluding objects within this bounding box but at uniformly random positions in the occlusion layer. The occluding objects can be randomly scaled so that their projections cover a certain percentage of the corresponding foreground object (e.g., within a range of 10% to 30% coverage of the foreground object). The rotation and / or color of the occluding objects can optionally be randomized (e.g., in the same manner as background objects). In some embodiments, whether (one or more) occlusions of (one or more) foreground objects are generated for a composite image, the number of (one or more) foreground objects that are occluded for the composite image, and / or the degree of coverage of the occlusions can depend on the size (one or more) of the (one or more) foreground objects. For example, for synthetic images with large foreground objects, less occlusion can be exploited than for synthetic images with relatively small foreground objects. This can be used as part of the curriculum strategy described in this paper to enable machine learning models to initially learn based on synthetic images with less occlusion and then learn on "harder" synthetic images with more occlusion.

[0050] By having corresponding background, foreground, and occlusion layers, all three layers can be fused to generate a combined pure synthetic image. For example, an occlusion layer can be rendered on top of a foreground layer, and the result can be rendered on top of a background layer. In some embodiments, random light sources are added during rendering, optionally with random perturbations in the light color. Additionally or alternatively, white noise can be added and / or the synthetic image can be blurred with a kernel (e.g., a Gaussian kernel) in which both the kernel size and the standard deviation are randomly selected. Thus, the background, foreground, and occlusion portions share the same image characteristics, in contrast to other methods that mix real images and synthetic renderings. This can make it impossible for a machine learning model to be trained to distinguish between foreground and background only on attributes specific to its domain. In other words, this can force the machine learning model to effectively learn to detect foreground objects and / or one or more characteristics of foreground objects based on their geometry and visual appearance.

[0051] By utilizing the techniques described above and / or elsewhere herein, synthetic images are generated, each of which can be paired with ground truth automatically generated (one or more) labels to generate corresponding synthetic training instances. Training a machine learning model using such synthetic training instances (and optionally using only synthetic training instances, or 90% or more of synthetic training instances) can result in a trained model that outperforms a corresponding model that was instead trained based on the same number of real images and human-provided labels. Various techniques disclosed above and / or elsewhere herein can improve the performance of machine learning models trained based on synthetic training instances, such as, for example, the curriculum techniques described herein, the relative proportions of background objects to foreground objects, the use of synthetic background objects, and / or the use of random colors and blurs.

[0052] Thus, various embodiments disclosed herein create purely synthetic training data for training machine learning models, such as object detection machine learning models. Some of those embodiments take advantage of large datasets of 3D background models and render them densely in a background layer using global randomization. This produces a background layer with locally realistic background clutter with realistic shapes and textures, on top of which foreground objects of interest can be rendered. Optionally, during training, a curriculum strategy can be followed that ensures that all foreground models are presented equally to the network under all possible rotations and with increasing complexity. Optionally, randomized lighting, blur and / or noise are added during the generation of synthetic images. Various embodiments disclosed herein do not require complex scene compositions, or real background images to provide the necessary background clutter, as in difficult photo-realistic image generation.

[0053] Now go to Figure 1 , illustrates an example environment in which embodiments disclosed herein may be implemented. The example environment includes a synthetic training example system 110. The training example system 110 may be implemented by one or more computing devices, such as a cluster of one or more servers. The training example system 110 includes a background engine 112, a foreground engine 114, an occlusion engine 116, a fusion engine 118, and a labeling engine 120.

[0054] The background engine 112 generates a background layer for synthesizing an image. When generating the background layer, the background engine 112 may utilize background 3D object models from the background 3D object model database 152. The background 3D object model database 152 may include a large number (e.g., 10,000 or more) of background 3D object models that may be selected and utilized by the background engine 112 when generating the background layer. All background 3D object models may optionally be initially de-meaned and scaled so that they fit within a unit sphere.

[0055] In some implementations, the background engine 112 may execute Figure 2 One or more (e.g., all) blocks of method 200 (described below) of . In various embodiments, the background engine 112 generates a background layer by successively selecting areas in the background where no other background 3D object models have been rendered ("bare areas") and rendering random background 3D object models to each selected area using random rotations. The background engine 112 can repeat this process until the entire background is covered with synthetic background objects. In various embodiments, the background engine 112 determines the size of the projected background objects used in generating the background layer based on the size of (one or more) foreground objects rendered in the foreground layer to be merged with the background layer when a synthetic image is subsequently generated. In some embodiments, for each background layer, the background engine 112 converts the texture of each rendered object to HSV space, randomly changes the hue value in the HSV space, and then converts back to RGB space.

[0056] The foreground engine 114 generates a foreground layer for synthesizing an image. In generating the foreground layer, the foreground engine 114 may utilize foreground 3D object models from the foreground 3D object model database 154. The foreground 3D object models of the foreground 3D object model database 154 may optionally be disjoint from those of the background 3D object model database 152. All foreground 3D object models may optionally be initially de-meaned and scaled so that they fit within a unit sphere.

[0057] When generating the foreground layer, the foreground engine 114 may include (one or more) renderings of the foreground 3D object model(s). The foreground engine 114 may select the rotation of the foreground 3D object model in the rendering from the set of rotations generated for the foreground 3D object model, and the size of the rendering may be selected based on its compliance with the rotation used to generate the corresponding background layer. The foreground engine 114 may optionally determine the rotation and size of the corresponding object according to the course strategy to ensure that each foreground object is rendered with multiple disparate rotations and / or different sizes across multiple composite images.

[0058] When rendering foreground 3D object models in the foreground layer, the foreground engine 114 may allow foreground objects to be cropped to a certain extent at the image boundary, and / or may allow for a certain degree of overlap (the same or additional degree) between pairs of rendered foreground objects. For each object, the foreground engine 114 may place it at a random position, and if the (one or more) initial attempted random positions fail (e.g., due to violating cropping and / or overlap constraints), then additional attempts at placement are made. If the foreground object being processed cannot be placed within the foreground layer within a threshold number of attempts (and / or duration), then the foreground engine 114 may consider the processing of the current foreground layer to be completed, and may render the currently processed foreground object in the next foreground layer to be processed. In some embodiments, the foreground engine 114 may perform Figure 3 One or more (e.g., all) blocks 302, 304, 306, 308, 310, and / or 312 of method 300 (described below).

[0059] The occlusion engine 116 generates an occlusion layer for a composite image. In some embodiments, the occlusion engine 116 may generate the occlusion layer by determining a corresponding bounding box (or other bounding shape) of one or more rendered foreground objects in the corresponding foreground layer and rendering a corresponding randomly selected occlusion object within the corresponding bounding box but at a uniform random position in the occlusion layer. The occlusion object may be, for example, a background 3D object model selected from the background 3D object model database 152. The occlusion engine 116 may scale the object so that its projected coverage is less than an upper percentage limit (e.g., 30%) and / or greater than a lower percentage limit (e.g., 5%) of the corresponding foreground object. The occlusion engine 116 may determine a random rotation of the occlusion object and / or may randomly color the occlusion object (e.g., using the HSV adjustment technique described herein).

[0060] The fusion engine 118 generates a composite image by fusing the corresponding background layers, foreground layers, and occlusion layers. For example, the fusion engine 118 may render the occlusion layer on top of the foreground layer and then render the result on top of the background layer. In some embodiments, the fusion engine 118 adds random light sources during rendering, optionally adding random perturbations in the light colors. Additionally or alternatively, the fusion engine 118 adds white noise and / or blurs the composite image (e.g., with a Gaussian kernel in which the kernel size and standard deviation are randomly selected). In some embodiments, the fusion engine 118 may perform Figure 3 314 of method 300 (described below).

[0061] The labeling engine 120 generates label(s) for each composite image generated by the fusion engine 118. For example, the labeling engine 120 may generate label(s) for the composite image, such as labels including corresponding 2D bounding boxes (or other bounding shapes) for each rendered foreground object and / or a classification of each rendered foreground object. The labeling engine 120 may determine the labels from, for example, the foreground engine 114 when the foreground engine 114 determines the 3D objects and their rotations and positions when generating the foreground layer. The labeling engine 120 provides each pair of composite images and corresponding label(s) as a training example to be stored in the training example database 156. In some embodiments, the labeling engine 120 may perform Figure 3 316 of method 300 (described below).

[0062] Figure 1 Also included is a training engine 130 that trains a machine learning model 165 based on the training instances of the training instance database 156. The machine learning model 165 can be configured to process the image (e.g., can have an input layer that conforms to the dimensions of the synthetic image or a scale thereof) to generate one or more predictions based on the image (e.g., predictions corresponding to the labels generated by the labeling engine 120). As some non-limiting examples, the machine learning model 165 can be a Faster R-CNN, SSD, R-FCN, Yolo9000, or RetinaNet Faster model. When training the machine learning model, the training engine 130 can process the synthetic images of the training instances to generate predictions, compare those predictions with the labels of the corresponding training instances to determine errors, and update the weights of the machine learning model 165 based on the errors. The training engine 130 can optionally utilize batch processing techniques during training. As described herein, in various embodiments, the training engine 130 can utilize a curriculum strategy during training, wherein during training, training instances including synthetic images with larger foreground objects are first utilized, followed by training instances including synthetic images with relatively smaller foreground objects (relative to the larger foreground objects), and optionally followed by one or more instances of additional training instances including synthetic images with relatively smaller foreground objects (relative to the foreground objects of the immediately preceding instances). In some embodiments, the training engine 130 can perform Figure 4 One or more (e.g., all) blocks of method 400 (described below).

[0063] After training of the machine learning model 165 is complete, the trained machine learning model can be deployed on one or more robots 142 and / or one or more computing devices 144. The robot(s) 142 can share one or more aspects with the robot 620 described below. The computing device(s) can share one or more aspects with the computing device 710 described below. The training of the machine learning model 165 can be determined to be complete in response to satisfying one or more training criteria. The training criteria can include, for example, performing a threshold number of training epochs, training based on a threshold number of (e.g., all available) training instances, determining that performance criteria and / or training criteria for the machine learning model are satisfied.

[0064] Figure 2 2 is a flowchart illustrating an example method 200 for generating a background layer according to various embodiments disclosed herein. For convenience, the operations of the flowchart are described with reference to a system that performs the operations. This system may include one or more processors, such as implementing the background engine 112 ( Figure 1 ) of one or more processors. Although the operations of method 200 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.

[0065] At block 202, the system identifies the size(s) at which to render the foreground 3D object(s). For example, at a given iteration, the system may identify the size(s) or scale(s) at which to render the foreground 3D object(s) in a foreground layer that will be fused with the background layer when generating a composite image (see Figure 3 ). In some embodiments, the determination at block 202 may be made in each iteration for each background layer generated. For example, the determination may be based on a corresponding particular foreground layer that has been generated, is being generated in parallel, or will be generated shortly after the background is generated. In some other embodiments, the determination at block 202 may be made once for a batch of background layers to be generated and subsequently fused with a corresponding foreground layer in a batch of foreground layers that all have the same or similar size(s).

[0066] Box 202 may optionally include sub-box 202A, where the system determines a scale range based on the size (or sizes) of the rendered (one or more) foreground 3D objects. The scale range may define an upper scale value, a lower scale value, and a scale value between the upper scale value and the lower scale value. The scale values ​​may each be an isotropic scale value that may be applied to the background 3D object model to uniformly "shrink" or "expand" (depending on the specific value) the background 3D object model. In some embodiments, at sub-box 202A, the system determines the scale range based on the scale value that determines the scale range (if applied to the 3D object model) will make the model within a corresponding percentage range of the foreground size. The foreground size is based on the size (or sizes) of the rendered (one or more) foreground 3D objects. For example, the foreground size may be the average of the sizes of multiple foreground 3D objects. The percentage range may be, for example, from 70% to 175%, from 90% to 150%, from 92% to 125%, or other percentage ranges.

[0067] At block 204, the system randomly selects a background 3D object. For example, the system may randomly select a background 3D object from a corpus of background 3D objects, such as a corpus that includes more than 1,000, more than 5,000, or more than 10,000 disparate background 3D objects. Optionally, the corpus of background 3D objects includes (e.g., is limited to) objects that are specific to the environment for which the synthetic image is being generated. For example, if the synthetic image is being generated for a home environment, typical household objects may be included in the background 3D objects in the corpus.

[0068] At block 206, the system renders the background 3D object at the corresponding position in the background layer with a random rotation (e.g., random in-plane and / or out-of-plane rotation) and with a size based on the size(s) at which the foreground 3D object(s) were rendered. It should be noted that when rendering the background 3D object, the rendering of the background object may overlap and / or intersect with other already rendered background object(s). In some embodiments, the system selects the corresponding position based on the area where it is currently exposed (i.e., currently lacking any rendered objects).

[0069] In some embodiments, the system renders the background by isotropically scaling the background 3D object to a size based on the size(s) of block 202 prior to rendering by an isotropic scaling value based on the size. For example, at optional block 206A, the system may scale the background object using a scaling value randomly selected from a scaling range optionally determined at block 202A. In some embodiments, anisotropic scaling may additionally or alternatively be used for scaling of the object(s) (background or foreground). In some embodiments, at block 202, during and / or after rendering the background object, the texture color of the background object is randomly perturbed. As a non-limiting example, the texture of the background object may be converted to HSV space, the hue value of the object randomly changed, and after changing the hue value, the HSV space may be converted back to red, green, blue (RGB) space. Note that while blocks 204 and 206 are described with respect to a single background 3D object for simplicity of description, in some implementations and / or iterations of blocks 204 and 206, multiple background 3D objects may be selected in block 204 and rendered (in corresponding positions) at block 206. Selecting and rendering multiple background 3D objects may improve rendering / data generation throughput.

[0070] At block 208, the system determines whether there are any bare areas remaining in the background layer. In some embodiments, when determining whether a region is a bare area, the system determines whether the region has at least a threshold size. For example, the system can determine that a region is a bare area only if there are at least a threshold number of consecutive bare pixels (in one or more directions) in that region. For example, if the number of bare pixels in that region is greater than a threshold number of pixels, then the region can be determined to be bare, otherwise the region is considered not bare. In some embodiments, the threshold number can be zero, which means that there will be no bare pixels. If the determination at block 208 is yes (i.e., the system determines that there is a bare area), then the system proceeds to block 210 and selects the bare area as the next position, and then returns to block 204, where it randomly selects additional background 3D objects. The system then proceeds to block 206 and renders the additional background 3D objects in the background layer at the next position in block 210 and at a random rotation and at a size based on the size (or sizes). Through multiple iterations, a background layer with no bare areas / full of background clutter can be generated.

[0071] If at the iteration of block 208, the system determines that no exposed areas remain, then the system proceeds to block 212 and saves the background layer including the rendered object from the multiple iterations of block 206. Figure 3As described in the method 300 of , the saved background layer 212 is then merged with the foreground layer, and optionally with the occlusion layer, when generating a composite image.

[0072] At block 214, the system determines whether to generate additional background layers. If so, the system returns to block 202 (however, in a batch processing technique, block 202 may be skipped when the size remains the same for multiple background layer(s)) and performs multiple iterations of blocks 204, 206, 208, and 210 as additional background layers are generated. If the decision at block 214 is no, the system may proceed to block 216 and stop background layer generation.

[0073] Figure 3 1 is a flowchart illustrating an example method 300 according to various embodiments disclosed herein, the method 300 comprising: generating a foreground layer; generating a composite image based on fusing the corresponding foreground layer, background layer, and optional occlusion layer; and generating a training instance comprising the composite image. For convenience, the operations of the flowchart are described with reference to a system performing the operations. This system may include one or more processors, such as implementing the foreground engine 114, the occlusion engine 116, the fusion engine 118, and / or the labeling engine 120 ( Figure 1 ) of one or more processors. Although the operations of method 300 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.

[0074] First, it should be noted that although the Figure 3 Method 300 and Figure 4 400, but in various embodiments, they can be performed in coordination and / or in parallel. For example, the foreground layer generation at iterative blocks 308, 310, and 312 of method 300 can be coordinated with the corresponding background layer generation at iterative blocks 202, 204, 206, 208, and 210 of method 200.

[0075] At box 302, the system selects (one or more) sizes at which to render (one or more) foreground 3D objects. As described herein, in some embodiments, (one or more) larger size (e.g., occupying more pixels) foreground object renderings can be used to generate an initial set of synthetic images, which are used in an initial training instance to initially train the machine learning model. In addition, (one or more) smaller size foreground object renderings can be used to generate a next set of synthetic images, which are used in a next set of training instances to train the machine learning model. (One or more) additional sets of (one or more) synthetic images can be generated, each set including (one or more) smaller size foreground object renderings (relative to the size of the previous set) and used sequentially to additionally train the machine learning model. Training a machine learning model according to this curriculum strategy can result in improved performance of the trained machine learning model and / or achieving a given performance level with a smaller number of training instances. The following is about Figure 4 Method 400 describes its implementation.

[0076] At block 304, the system generates multiple rotations for the selected foreground 3D object, and the selected foreground 3D object is scaled based on the size selected at block 302. Each rotation may include different pairs of in-plane rotations (with respect to the image position in the plane) and out-of-plane rotations. In some (one or more) embodiments, n different combinations of out-of-plane rotations are generated, and for each out-of-plane rotation, m different in-plane rotations are generated. For example, an icosahedron can be recursively divided to generate n different uniformly distributed combinations of out-of-plane rotations, and m different in-plane rotations can be generated for each combination of out-of-plane rotations. In-plane rotations can also be uniformly distributed, with a given discreteness (e.g., 1 degree or other discreteness) between in-plane rotations. The scaling of the size can be based on the isotropic scaling of the foreground 3D object model, and the rotation (one or more) can be generated by manipulating the rotation of the foreground 3D object model. Thus, in various embodiments, at the end of the iteration of block 304, multiple rotations of the foreground 3D object model are generated, each rotation scaling the foreground 3D object model based on the size selected at block 302.

[0077] At block 306, the system determines whether there are any additional foreground 3D object models to process using block 304. If so, the system selects one of the unprocessed additional foreground 3D object models and returns to block 304 to generate a plurality of rotations for that unprocessed model, each rotation scaling the model based on the size selected at block 302.

[0078] After all foreground 3D object models have been processed in multiple iterations of blocks 304 and 306, the system selects a foreground 3D object and rotation (from among the multiple foreground 3D objects generated in multiple iterations of block 304) at block 308. After the foreground 3D object and rotation are selected, that particular foreground 3D object and rotation combination may optionally be marked as "done," thereby preventing it from being selected in subsequent iterations of block 308.

[0079] At box 310, the system renders the selected 3D object in the foreground layer and at a random position with the selected rotation (and size). In some embodiments, the system can ensure that rendering at the random position will not violate one or more constraints before rendering the object at the random position. Constraints may include cropping constraints and / or overlapping constraints mentioned herein. If the system determines that (one or more) constraints are violated, the system may choose to attach a random position. This may continue until a threshold number of attempts have been made and / or until a threshold duration has passed. If a threshold number of attempts have been made and / or a threshold duration has passed, the decision of box 312 (below) may be "no", and the currently selected foreground 3D object and rotation may be used as the initially selected foreground 3D object and rotation for the next iteration of box 308 when generating the next foreground layer.

[0080] At block 312, the system determines whether to render the additional foreground object in the foreground layer. In some embodiments, this determination may be "yes" as long as a threshold number of attempts have not been made and / or a threshold duration has not elapsed in the immediately preceding iteration of block 310. If the decision at block 312 is yes, the system returns to block 308 and selects an additional foreground 3D object and an additional rotation, and then proceeds to block 310 to attempt to render the additional foreground 3D object in the foreground layer with the additional rotation. Optionally, the constraint may prevent rendering the same foreground 3D object more than once (i.e., with different rotations) in the same foreground layer.

[0081] If the decision at block 312 is no, the system proceeds to block 314 and generates a composite image based on fusing the foreground layer with a corresponding background layer. The corresponding background layer may be a composite image generated using method 200 ( Figure 2 ) and may correspond to the foreground layer based at least in part on the size of the background objects in the background layer being similar to the size(s) of the foreground objects in the foreground layer. For example, the background objects in the background layer may be scaled based on the size / scaling of the foreground 3D objects used when generating the foreground layer.

[0082] Box 314 may optionally include sub-box 314A, where the system further generates a composite image based on fusing the occlusion layer with the background layer and the foreground layer. The occlusion layer may be generated based on rendering (one or more) additional background 3D objects within (one or more) bounded areas of the (one or more) foreground 3D objects rendered in the foreground layer. For example, the occlusion layer may be generated by randomly selecting (one or more) background 3D object models and rendering each background 3D object model in a corresponding random position within a corresponding bounding box that defines one of the foreground objects in the foreground layer. The occlusion objects may be randomly scaled so that their projections cover a certain percentage of the corresponding foreground objects.

[0083] In some embodiments, at block 314, the system renders the occlusion layer on top of the foreground layer and renders the result on top of the background layer when generating the composite image. In some of those embodiments, random light sources are added, optionally with random perturbations in light color. Additionally or alternatively, white noise may be added and / or the composite image may be blurred with a Gaussian kernel, such as a Gaussian kernel in which both the kernel size and standard deviation are randomly selected.

[0084] At block 316, the system generates a training instance that includes a synthetic image and (one or more) labels for rendering (one or more) foreground 3D object models in the synthetic image. The (one or more) labels may, for example, include for each rendered foreground object: a corresponding 2D bounding box (or other bounding shape), a corresponding six-dimensional (6D) pose, a corresponding classification, a corresponding semantic label map, and / or any other relevant label data. Since the foreground 3D objects and their rotations are known when generating the foreground layer, the labels can be easily determined.

[0085] At block 318, the system determines whether there are additional unprocessed foreground 3D object, rotation pairs. In other words, whether there are any foreground 3D object, rotation pairs that have not yet been rendered in the foreground layer (and thus included in the composite image). If so, the system returns to block 308 and then performs iterations of blocks 308, 310, 312 when generating additional foreground layers based on the unprocessed foreground 3D objects and rotations, and generates additional composite images based on the additional foreground layers at block 314. In these and other ways, through multiple iterations, a composite image is generated that, for the size of block 302, collectively includes rendering all foreground 3D object models at all generated rotations.

[0086] If the decision at the iteration of block 318 is no, the system proceeds to block 320 and determines whether an additional size should be used for the foreground object. If so, the system returns to block 302 and selects another size (e.g., a smaller size). The blocks of method 300 may then be repeated to generate another batch of composite images having (one or more) foreground objects rendered based on the additional size.

[0087] If the decision at the iteration of block 320 is no, then the system may stop the generation of synthetic images and synthetic training instances.

[0088] Figure 4 4 is a flowchart illustrating an example method 400 for training a machine learning model based on a curriculum according to various embodiments disclosed herein. For convenience, the operations of the flowchart are described with reference to a system that performs the operations. This system may include one or more processors, such as implementing the training engine 130 ( Figure 1 ) of one or more processors. Although the operations of method 400 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.

[0089] At box 402, the system selects training instances having synthetic images that have (one or more) foreground objects of larger size (one or more). For example, the system may select training instances generated according to method 300 that all include synthetic images having the following foreground objects: the foreground objects have a first size or proportion or are within a threshold percentage (e.g., 10%) of the first size or proportion. This can include tens of thousands or hundreds of thousands of training instances. As described herein, in various embodiments, the background objects of the synthetic images of the training instances can have similar sizes as the foreground objects. Also as described herein, in various embodiments, the synthetic images of the training instances can collectively include all foreground objects of interest, and can collectively include each foreground object at a large number of evenly distributed rotations.

[0090] At box 404, the system trains the machine learning model based on the selected training instances. For example, at box 404, the system may iteratively train the machine learning model using a batch of training instances at each iteration. When training the machine learning model, predictions may be made based on processing synthetic images of the training instances, those predictions may be compared with the labels of the training instances to determine errors, and the weights of the machine learning model may be iteratively updated based on the errors. For example, the weights of the machine learning model may be updated using back propagation, and optionally using an error-based loss function. It should be noted that for some (one or more) machine learning models, only certain weights may be updated, while other weights may be fixed / static throughout the training period. For example, some machine learning models may include a pre-trained image feature extractor (which may optionally be fixed during training), and an additional layer that further processes the extracted features and includes nodes with trainable weights. It should also be noted that various machine learning models may be trained, such as object detection and / or classification models. For example, the machine learning model may be a Faster R-CNN model or other object detection and classification model.

[0091] At block 406, the system selects additional training instances having synthetic images having foreground objects of smaller size(s) after block 404 is completed. In an initial iteration of block 406, the smaller size(s) are smaller relative to the larger size(s) of block 402. In subsequent iterations of block 406, the smaller size(s) are smaller relative to the most recent iteration of block 406. For example, the system may select training instances generated according to method 300 that all include synthetic images having foreground objects of a second size or proportion or within a threshold percentage (e.g., 10%) of the second size or proportion. This may include tens or hundreds of thousands of training instances. As described herein, in various embodiments, the background objects of the synthetic images of the training instances may have a size similar to the foreground objects. As also described herein, in various embodiments, the synthetic images of the training instances may collectively include all foreground objects of interest, and may collectively include each foreground object at a large number of evenly distributed rotations.

[0092] At block 408, the system further trains the machine learning model based on the selected additional training examples. For example, at block 404, the system may iteratively further train the machine learning model using a batch of additional training examples at each iteration.

[0093] At box 410, and after box 408 is completed, the system determines whether there are additional training instances of foreground objects whose composite images have (one or more) even smaller sizes (relative to the sizes used in the most recent iteration of box 408). If so, the system returns to box 406. If not, the system proceeds to box 412, where the system can deploy the machine learning model. Optionally, at box 412, the system deploys the machine learning model only after determining that one or more training criteria are met (e.g., in addition to no additional training instances remaining). The machine learning model can be deployed to (one or more) computing devices and / or (one or more) robotic devices (e.g., transmitted to a computing device and / or robotic device or otherwise caused to be stored locally thereon). For example, the machine learning model can be deployed on a robot for the robot to perform various robotic tasks.

[0094] Now go to Figure 5A , 5B , 5C and 5D, illustrate some example composite images 500A-500D. For simplicity, each of the composite images 500A-500D only includes a corresponding detailed view of a corresponding portion 500A1-500D1 of the composite image 500A-500D. Each of the composite images 500A-500D also includes a corresponding background object descriptor 552A-D and a corresponding foreground object descriptor 554A-D of one of the foreground objects. It is noted that the descriptors 552A-D and 554A-D will not actually be included with the composite images 500A-D, but are provided herein for illustrative purposes only.

[0095] First go to Figure 5A , as indicated by background object descriptors 552A, the composite image 500A includes a first set of background objects that are all within a first size range. The composite image 500A may also include one or more foreground objects, one of which (foreground object 501) is described by foreground object descriptors 554A. As indicated by foreground object descriptors 554A, foreground object 501 has a first rotation and a first size corresponding to the first size range of background objects in the composite image 500A. (One or more) other foreground objects (not shown) may also be provided in the composite image 500A, and may be (one or more) different objects. The (one or more) different objects may have a rotation different from the rotation of object 501, but the (one or more) objects will also have the same or similar size / proportion as object 501.

[0096] Portion 500A1 illustrates a first size and a first rotation of object 501. Additionally, portion 500A1 illustrates some of the background objects 511, 512, and 513. As can be determined by looking at background objects 511, 512, and 513, they are of similar size / proportion relative to each other and relative to object 501. For simplicity, representations 521, 522, 523 of other background objects are also shown as different shades in portion 500A1. In other words, representations 521, 522, 523 are merely representative of what would actually be rendered in detail and would have similar size / proportion as objects 511, 512, and 513, but only for the sake of simplicity. Figure 5A Background objects 511, 512, and 513, as well as those represented by representations 521, 522, and 523, collectively cover the background in portion 500A1 (i.e., there are no bare spots) and represent a subset of the background objects of composite image 500A.

[0097] Next go to Figure 5B , as indicated by background object descriptors 552B, the composite image 500B includes a second set of background objects that are all within a first size range. The second set differs from the first set of composite images 500A in that different background objects are included and / or the different background objects are rendered at different rotations. The composite image 500B may also include one or more foreground objects, one of which (foreground object 501) is described by foreground object descriptors 554B. As indicated by foreground object descriptors 554B, foreground object 501 has a second rotation and a first size in the composite image 500B. The second rotation of foreground object 501 in the composite image 500B is different from the first rotation of foreground object 501 in the composite image 500A (i.e., a different in-plane rotation). The first size of foreground object 501 in the composite image 500B is the same as the first size of foreground object 501 in the composite image 500A. (One or more) other foreground objects (not shown) may also be provided in the composite image 500B and may be (one or more) different objects, one or more of which may be different from the different foreground objects in the composite image 500A. The different object(s) may have a different rotation than the rotation of object 501 in composite image 500B, but the different object(s) will also have the same or similar size / proportion as object 501 .

[0098] Portion 500B1 illustrates a first size and a second rotation of object 501. Additionally, portion 500B1 illustrates some of the background objects 514, 515, and 516. As can be determined by looking at background objects 514, 515, and 516, they are of similar size / proportion relative to each other and relative to object 501. For simplicity, representations 524, 524, 526, and 527 of other background objects are also shown as different shades in portion 500B1. In other words, representations 524, 524, 526, and 527 are merely representative of what would actually be rendered in detail and would have similar size / proportion as objects 514, 515, and 516, but only for the sake of simplicity. Figure 5B 500B1 and 500B2. Other background objects are represented as different shades for the sake of simplicity. Background objects 514, 515, and 516, and those represented by representations 524, 524, 526, and 527, collectively cover the background of portion 500B1 (i.e., have no bare spots) and represent a subset of the background objects of composite image 500A. Occlusion object 517 is an object that may be rendered in an occlusion layer as described herein, and partially occludes a portion of object 501.

[0099] Composite images 500A and 500B thus illustrate how object 501 may be provided in multiple composite images at the same size but at different rotations and at different locations and between different background object(s) in the multiple composite images and / or with different occlusion(s) (or no occlusion). Note that object 501 will be included in multiple additional composite images at the same size in accordance with the techniques described herein. In those additional composite images, object 501 will be at different rotations (including those with alternating out-of-plane rotations) and may be at different locations and / or between different background clutter and / or occluded in different ways and / or with different objects.

[0100] Next go to Figure 5C and 5D , composite images 500C and 500D also include object 501, but as described below, include objects of smaller size. Figure 5C In FIG. 500C, as indicated by background object descriptor 552C, composite image 500C includes a third set of background objects that are all within a second size range. The second size range is different from Figure 5A and 5BThe third set of background objects may be a first size range of background objects of the composite image 500A and may include smaller size values. The second size range corresponds to smaller sizes of foreground object 501 and (one or more) other foreground objects in composite image 500C. In addition to size, the third set of background objects may differ from the first set of composite images 500A and the second set of composite images 500B in that different background objects are included and / or the different background objects are rendered with different rotations.

[0101] Composite image 500C may also include one or more foreground objects, one of which (foreground object 501) is described by foreground object descriptor 554C. As indicated by foreground object descriptor 554C, foreground object 501 has a first rotation and a second size in composite image 500C. The first rotation of foreground object 501 in composite image 500C is the same as the first rotation of foreground object 501 in composite image 500A, but foreground object 501 is in a different position in the two composite images 500A and 500C. Moreover, the second size of foreground object 501 in composite image 500C is smaller than its size in composite images 500A and 500B. Other foreground object(s) (not shown) may also be provided in composite image 500C, and may be different object(s). That different object(s) may be in a different rotation than the rotation of object 501 in composite image 500C, but that object(s) will also have the same or similar size / proportion as object 501 in composite image 500C.

[0102] Portion 500C1 illustrates the second size and first rotation of object 501. Additionally, portion 500C1 illustrates some of the background objects 511, 518, and 519. As can be determined by looking at background objects 511, 518, and 519, they are of similar size / proportion relative to each other and relative to object 501. For simplicity, representations 528 and 529 of other background objects are also shown as different shades in portion 500C1. In other words, representations 528 and 529 are merely representative of what would actually be rendered in detail and would be of similar size / proportion as objects 511, 518, and 519, but only for the sake of simplicity. Figure 5C Background objects 511, 518, and 519, as well as those represented by representations 528 and 529, collectively cover the background of portion 500C1 and represent a subset of the background objects of composite image 500C.

[0103] exist Figure 5D500D, as indicated by background object descriptors 552D, the composite image 500D includes a third set of background objects that are all within a second size range. The second size range corresponds to smaller sizes of foreground object 501 and (one or more) other foreground objects in the composite image 500D. The third set of background objects may differ from the first, second, and third sets of composite images 500A, 500B, and 500C in that different background objects are included and / or the different background objects are rendered at different rotations and / or different locations.

[0104] Composite image 500D may also include one or more foreground objects, one of which (foreground object 501) is described by foreground object descriptor 554D. As indicated by foreground object descriptor 554D, foreground object 501 has a third rotation and a second size in composite image 500D. The third rotation of foreground object 501 in composite image 500C is different from the first and second rotations (in-plane) of composite images 500A, 500B, and 500C. The second size of foreground object 501 in composite image 500D is the same as the second size in composite image 500C. Other foreground object(s) (not shown) may also be provided in composite image 500D and may be different object(s). The different object(s) may be in a different rotation than the rotation of object 501 in composite image 500D, but the object(s) will have the same or similar size / proportion as object 501 in composite image 500D.

[0105] Portion 500D1 shows a second size and a first rotation of object 501. Additionally, portion 500D1 illustrates some of the background objects 512, 516, and 519. As can be determined by looking at background objects 512, 516, and 519, they are of similar size / proportion relative to each other and relative to object 501. For simplicity, representations 530, 531, and 532 of other background objects are also shown as different shades in portion 500D1. In other words, representations 530, 531, and 532 are merely representative of what would actually be rendered in detail and would be of similar size / proportion as objects 512, 516, and 519, but only for the sake of simplicity. Figure 5D Background objects 512, 516, and 519, as well as those represented by representations 530, 531, and 532, collectively cover the background of portion 500D1 and represent a subset of the background objects of composite image 500D.

[0106] Composite images 500C and 500D thus illustrate how object 501 may be provided in multiple composite images at the same size (different from the size of composite images 500A and 500B) but at different rotations and at different locations and between different background object(s) and / or with different occlusion(s) (or no occlusion) in the multiple composite images. Note that object 501 will be included in multiple additional composite images at the same second size in accordance with the techniques described herein. In those additional composite images, object 501 will be at different rotations (including those with alternating out-of-plane rotations) and may be at different locations and / or between different background clutter and / or in different ways and / or occluded by different objects. As described herein, training of a machine learning model may be performed based on composite images 500A, 500B, and a number of additional composite images having foreground objects of similar size to the foreground objects of composite images 500A and 500B. After training on such synthetic images, the machine learning model may then be further trained based on synthetic images 500C, 500D, and a large number of additional synthetic images having foreground objects of similar size to the foreground objects of synthetic images 500C and 500D.

[0107] Figure 6 An example architecture of a robot 600 is schematically depicted. The robot 600 includes a robot control system 602, one or more operating components 604a-n, and one or more sensors 608a-m. The sensors 608a-m may include, for example, vision sensors (e.g., camera(s), 3D scanners), light sensors, pressure sensors, position sensors, pressure wave sensors (e.g., microphones), proximity sensors, accelerometers, gyroscopes, thermometers, barometers, etc. Although the sensors 608a-m are depicted as being integrated with the robot 600, this is not meant to be limiting. In some embodiments, the sensors 608a-m may be located external to the robot 600, such as as a stand-alone unit.

[0108] The operating components 604a-n may include, for example, one or more end effectors (e.g., gripping end effectors) and / or one or more servomotors or other actuators to achieve movement of one or more components of the robot. For example, the robot 600 may have multiple degrees of freedom, and each actuator may control actuation of the robot 600 within one or more degrees of freedom in response to control commands (e.g., torque and / or other commands generated based on a control strategy) provided by the robot control system 602. As used herein, the term actuator also encompasses a mechanical or electrical device (e.g., a motor) that produces motion, in addition to any driver (s) that may be associated with the actuator and that translates received control commands into one or more signals for driving the actuator. Thus, providing a control command to an actuator may include providing a control command to a driver that translates the control command into an appropriate signal for driving the electrical or mechanical device to produce the desired motion.

[0109] The robot control system 602 may be implemented in one or more processors (such as a CPU, GPU, and / or other controller(s)) of the robot 600. In some embodiments, the robot 600 may include a "brainbox" that may include all or some aspects of the control system 602. For example, the brainbox may provide real-time bursts of data to the operating components 604a-n, where each real-time burst includes a set of one or more control commands that specify, among other things, one or more parameters of movement of each of one or more of the operating components 604a-n. In various embodiments, the control commands may be selectively generated by at least the control system 602 based at least in part on object detection, object classification, and / or other determination(s) made using a machine learning model stored locally on the robot 620 and trained according to embodiments described herein.

[0110] Although in Figure 6 600, but in some embodiments, all or some aspects of the control system 602 may be implemented in a component separate from but in communication with the robot 600. For example, all or some aspects of the control system 602 may be implemented on one or more computing devices (such as computing device 710) that communicate with the robot 600 in wired and / or wireless communication.

[0111] Figure 77 is a block diagram of an example computing device 710 that may optionally be used to perform one or more aspects of the techniques described herein. The computing device 710 typically includes at least one processor 714 that communicates with a plurality of peripheral devices via a bus subsystem 712. These peripheral devices may include a storage subsystem 724, including, for example, a memory subsystem 725 and a file storage subsystem 726, a user interface output device 720, a user interface input device 722, and a network interface subsystem 716. The input and output devices allow a user to interact with the computing device 710. The network interface subsystem 716 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.

[0112] The user interface input device 722 may include a keyboard, a pointing device (such as a mouse, trackball, touch pad or graphic tablet), a scanner, a touch screen incorporated into a display, an audio input device (such as a voice recognition system, a microphone), and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways to input information into the computing device 710 or a communication network.

[0113] User interface output device 720 may include a display subsystem, a printer, a fax machine, or a non-visual display (such as an audio output device). The display subsystem may include a cathode ray tube (CRT), a flat panel device (such as a liquid crystal display (LCD)), a projection device, or some other mechanism for creating a visual image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from computing device 710 to a user or another machine or computing device.

[0114] Storage subsystem 724 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, storage subsystem 724 may include logic to perform selected aspects of one or more methods described herein.

[0115] These software modules are generally executed by the processor 714 alone or in combination with other processors. The memory 725 used in the storage subsystem 724 may include multiple memories, including a main random access memory (RAM) 730 for storing instructions and data during program execution and a read-only memory (ROM) 732 for storing fixed instructions. The file storage subsystem 726 may provide persistent storage for programs and data files, and may include a hard drive, a floppy disk drive, and associated removable media, a CD-ROM drive, an optical drive, or a removable media box. The modules that implement the functions of certain embodiments may be stored in the storage subsystem 724 by the file storage subsystem 726, or stored in other machines accessible to (one or more) processors 714.

[0116] The bus subsystem 712 provides a mechanism for the various components and subsystems of the computing device 710 to communicate with each other as intended. Although the bus subsystem 1312 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple busses.

[0117] The computing device 710 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, for the purpose of illustrating some embodiments, the computing device 710 is described in detail below. Figure 7 The description of computing device 710 depicted in FIG. 7 is intended only as a specific example. Many other configurations of computing device 710 may have more Figure 7 The computing device depicted may have more or fewer components.

Claims

1. A method implemented by one or more processors, the method comprising: For each of the multiple selected background 3D object models: Rendering the background 3D object model at a corresponding random background position in the background layer and with a random color tone, For each of the one or more selected foreground 3D object models: Rendering a foreground 3D object model at a corresponding random foreground position in the foreground layer; For each of the one or more selected occluding 3D object models: Rendering the occluded 3D object model at the corresponding random occluded position in the occluded layer; wherein during rendering of the background object model, the one or more foreground 3D object models, and the occluded 3D object model, at least one light source is randomly applied with a random light color; Generate a composite image based on rendering that fuses the background layer, the foreground layer, and the occlusion layer; assigning to the composite image ground truth labels for renderings of the one or more foreground 3D object models; and Providing training instances comprising synthetic images paired with ground truth labels for training at least one machine learning model based on the training instances.

2. The method according to claim 1, wherein: During rendering of the one or more foreground 3D object models and / or during rendering of the one or more background 3D object models, white noise is added. The method of claim 1 , further comprising blurring the composite image.

4. The method according to claim 3, wherein: Blurring the composite image includes blurring the composite image using a randomly selected kernel.

5. The method according to claim 1, wherein: Rendering background 3D object models at corresponding random background positions in the background layer includes selecting the background position based on that other background 3D objects have not been rendered at the background position, and wherein rendering is performed iteratively, each time for additional ones of the background 3D object models, until it is determined that all positions of the background layer have content rendered thereon.

6. The method according to claim 1, wherein: Rendering the background 3D object model at the corresponding random background position in the background layer includes rendering the background 3D object model with the random background rotation.

7. The method according to claim 1, wherein: Rendering the foreground 3D object model at the corresponding random foreground position in the foreground layer includes rendering the foreground 3D object model with a random foreground rotation.

8. The method according to claim 1, wherein: Rendering the occlusion 3D object model at the corresponding random foreground position in the occlusion layer includes rendering the foreground 3D object model with the random occlusion rotation.

9. The method according to claim 1, wherein: The ground truth labels include a defined shape of the foreground object, a six-dimensional (6D) pose of the foreground object, and / or a classification of the foreground object.

10. The method according to claim 9, wherein: The ground truth label comprises a bounding shape, and wherein the bounding shape is a two-dimensional bounding box.

11. A system comprising: Memory, which stores instructions; as well as One or more processors that execute instructions to cause the one or more processors to: For each of the multiple selected background 3D object models: Rendering the background 3D object model at a corresponding random background position in the background layer and with a random color tone, For each of the one or more selected foreground 3D object models: Rendering a foreground 3D object model at a corresponding random foreground position in the foreground layer; For each of the one or more selected occluding 3D object models: Rendering the occluded 3D object model at the corresponding random occluded position in the occluded layer; wherein during rendering of the background object model, the one or more foreground 3D object models, and the occluded 3D object model, at least one light source is randomly applied with a random light color; Generate a composite image based on rendering that fuses the background layer, the foreground layer, and the occlusion layer; assigning to the composite image ground truth labels for renderings of the one or more foreground 3D object models; and Providing training instances comprising synthetic images paired with ground truth labels for training at least one machine learning model based on the training instances.

12. The system according to claim 11, wherein: During rendering of the one or more foreground 3D object models, white noise is added.

13. The system according to claim 11, wherein: When executing the instructions, the one or more processors also: Blur composite image.

14. The system according to claim 11, wherein: When executing the instructions, the one or more processors also: Blur the composite image using a randomly chosen kernel.

15. The system according to claim 11, wherein: When rendering background 3D object models at corresponding random background positions in the background layer, one or more processors select the background position based on that other background 3D objects have not been rendered at the background position, and wherein rendering is performed iteratively, each time for additional ones of the background 3D object models, until it is determined that all positions of the background layer have content rendered thereon.

16. The system of claim 11, wherein: The one or more processors render the background 3D object model with a random background rotation when rendering the background 3D object model at a corresponding random background position in the background layer.

17. The system of claim 11, wherein: The one or more processors render the foreground 3D object model with a random foreground rotation when rendering the foreground 3D object model at a corresponding random foreground position in the foreground layer.

18. The system of claim 11, wherein: The one or more processors render the foreground 3D object models with the random occlusion rotations while rendering the occlusion 3D object models at corresponding random foreground positions in the occlusion layer.

19. The system of claim 11, wherein: The ground truth labels include a defined shape of the foreground object, a six-dimensional (6D) pose of the foreground object, and / or a classification of the foreground object.

20. The system of claim 11, wherein: The ground truth label comprises a bounding shape, and wherein the bounding shape is a two-dimensional bounding box.

21. A method implemented by one or more processors, the method comprising: identifying a size of a foreground three-dimensional (3D) object model rendered in a foreground layer of a composite image; For each of multiple randomly selected background 3D object models: rendering the background 3D object model at a corresponding background position in the background layer of the composite image with a corresponding rotation and at a corresponding size determined based on the size of the rendered foreground 3D object model; Rendering a foreground 3D object model at a foreground position in the foreground layer, the rendering of the foreground 3D object model being at the size and at a given rotation of the foreground 3D object model; Select additional background 3D object model; identifying a random location within a determined bounded area that defines a rendering of a foreground 3D object model; rendering the additional background 3D object model in the identified random position within the determined bounded area and in the occlusion layer of the composite image, rendering the additional background 3D object model including scaling the additional background 3D object model prior to rendering so as to occlude only a portion of the rendering of the foreground 3D object model; Generate a composite image based on fusing the background layer, the foreground layer and the occlusion layer; Assigning a ground truth label for a rendering of a foreground 3D object model to the synthetic image; as well as The synthetic images paired with the ground truth labels are provided as training examples for training at least one machine learning model for object detection and / or classification tasks based on the training examples.

22. A method implemented by one or more processors, the method comprising: identifying a size of a foreground three-dimensional (3D) object model rendered in a foreground layer of a composite image; For each of multiple randomly selected background 3D object models: rendering the background 3D object model at a corresponding background position in the background layer of the composite image with a corresponding rotation and at a corresponding size determined based on the size of the rendered foreground 3D object model, Rendering the selected background 3D object model at the corresponding background position includes selecting the background position based on that no other background 3D object has been rendered at the background position, and wherein rendering is performed iteratively, each time for an additional one of the background 3D object models, rendering is performed iteratively until it is determined that all positions of the background layer have content rendered thereon; Rendering a foreground 3D object model at a foreground position in the foreground layer, the rendering of the foreground 3D object model being at the size and at a given rotation of the foreground 3D object model; Generate a composite image based on fusing the background layer and the foreground layer; Assigning a ground truth label for a rendering of a foreground 3D object model to the synthetic image; and Providing training instances comprising synthetic images paired with ground truth labels for training at least one machine learning model based on the training instances.

23. A method implemented by one or more processors, the method comprising: Select a foreground three-dimensional (3D) object model; generating a plurality of first-scale rotations of the foreground 3D object model using the foreground 3D object model at the first scale; For each of the plurality of first scale rotations of the foreground 3D object model: rotating a corresponding one of the foreground 3D object models at a first scale and rendering the foreground 3D object model in a corresponding randomly selected position in a corresponding first scale foreground layer at the first scale; Generating a first-ratio composite image, generating each of the corresponding first-ratio composite images comprises: fusing a corresponding one of the corresponding first-scale foreground layers with a corresponding one of a plurality of non-intersecting first-scale background layers, the plurality of non-intersecting first-scale background layers each including a corresponding rendering of a corresponding randomly selected background 3D object model; generating first-scale training instances, each of the first-scale training instances comprising a corresponding one of the first-scale synthetic images and a corresponding ground truth label for a rendering of a foreground 3D object model in the corresponding one of the first-scale synthetic images; generating a plurality of second scale rotations of the foreground 3D object model for the foreground 3D object model at a second scale that is smaller than the first scale; For each of the plurality of second scale rotations of the foreground 3D object model: rotating a corresponding one of the foreground 3D object models at a second scale and rendering the foreground 3D object model in a corresponding randomly selected position in a corresponding second scale foreground layer at the second scale; Generating a second-ratio composite image, generating each of the corresponding second-ratio composite images comprises: fusing a corresponding one of the corresponding second-scale foreground layers with a corresponding one of a plurality of non-intersecting second-scale background layers, the plurality of non-intersecting second-scale background layers each comprising a corresponding rendering of a corresponding randomly selected background 3D object model; generating second-scale training instances, each of the second-scale training instances comprising a corresponding one of the second-scale synthetic images and a corresponding ground truth label for a rendering of a foreground 3D object model in the corresponding one of the second-scale synthetic images, wherein in the first scale background layer, the corresponding rendering size of the randomly selected background 3D object model is smaller than the corresponding rendering size of the randomly selected background 3D object model in the second scale background layer; and The machine learning model is trained based on the first scale training instances before the machine learning model is trained based on the second scale training instances.

24. A system comprising: Memory, which stores instructions; as well as One or more processors, the one or more processors executing instructions to cause the one or more processors to perform the method of claim 21, 22 or 23.