Unsupervised Learning of Scene Structure for Synthetic Data Generation

By using a probabilistic scene grammar to generate virtual scenes based on predefined rules, the method addresses the challenges of creating realistic and diverse synthetic datasets, achieving improved realism and performance in machine learning applications.

JP7690296B2Active Publication Date: 2025-06-10NVIDIA CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021019433
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-10
Filing Date
2021-02-10
Publication Date
2025-06-10
Estimated Expiration
2041-02-10

AI Technical Summary

Technical Problem

Existing techniques for generating synthetic datasets are cumbersome and time-consuming, requiring manual adjustment of procedural models and referencing real images to create realistic and diverse virtual environments.

Method used

The method involves generating a virtual scene based on a set of rules that define object placement and appearance, using a probabilistic scene grammar to sample scene structures and learn scene parameters, thereby reducing the gap between simulation and reality.

Benefits of technology

This approach enables the generation of synthetic images and datasets that closely match real-world distributions, improving the realism and diversity of virtual environments and enhancing the performance of machine learning models trained on these datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007690296000013
    Figure 0007690296000013
  • Figure 0007690296000014
    Figure 0007690296000014
  • Figure 0007690296000015
    Figure 0007690296000015
Patent Text Reader

Abstract

To make it possible to use a rule set or scene grammar to generate a scene graph that represents the structure and visual parameters of objects in a scene.SOLUTION: A renderer can take this scene graph as input and, with a library of content for assets identified in the scene graph, can generate a synthetic image of a scene that has the desired scene structure without the need for manual placement of the objects in the scene. Images or environments synthesized in this way can be used to, for example, generate training data for real world navigational applications, and to generate virtual worlds for games or virtual reality experiences.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This patent application claims the priority of U.S. Provisional Patent Application No. 62 / 986,614, filed on March 6, 2020, with the title "Bridging the Sim-to-Real Gap: Unsupervised Learning of Scene Structure for Synthetic Data Generation", and incorporates the same in its entirety herein for all purposes.

Background Art

[0002] Applications such as games, animations, simulations, etc. are increasingly relying on more detailed and realistic virtual environments. Often, procedural models are used to synthesize scenes in these environments and to create labeled synthetic datasets for machine learning. To generate realistic and diverse scenes, several parameters for managing the procedural model have to be carefully adjusted by experts. These parameters control both the structure of the generated scene (e.g., how many cars are in the scene) and the parameters for arranging objects in a valid configuration. The complexity and amount of knowledge required to manually determine and adjust these parameters, as well as to configure other aspects of these scenes, can limit widespread adoption and may also limit the realism or scope of the generated environments.

[0003] Various embodiments according to the present disclosure will be described with reference to the drawings.

Brief Description of the Drawings

[0004]

Figure 1A

Figure 1B

Figure 2A

Figure 2B

Figure 2C

Figure 3A

Figure 3B

Figure 3C

Figure 4

Figure 5

Figure 6

Figure 7A

Figure 7B

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15A

Figure 15B

[0005] Techniques according to various embodiments can provide for the generation of synthetic images and datasets. Specifically, various embodiments can generate a virtual scene or environment based at least in part on a set of rules that define the placement and appearance of objects or “assets” within that environment. These datasets can be used not only to generate virtual environments, but also to generate large training datasets related to a target real-world dataset. Synthetic datasets provide an attractive opportunity to train machine learning models for tasks such as perception and planning in autonomous and semi-autonomous driving, indoor scene perception, creation of generated content, and robot control. Synthetic datasets can provide ground truth data for tasks such as segmentation, depth, or material information via a graphics engine, where labels are expensive or even unavailable. As shown in image 100 of FIG. 1A and image 150 of FIG. 1B, this image can include ground truth data for objects rendered in these images, such as bounding boxes or labels for automobiles 102, 152, and person 104 rendered in these images. Adding a new type of label to such a synthetic dataset can be performed by making a call to a renderer, rather than undertaking new tools and the time-consuming annotation work of adopting, training, and monitoring annotators.

[0006] There are various obstacles to creating synthetic datasets using conventional techniques. Content such as 3D computer-aided design (3D CAD) models that make up a scene can be obtained from sources such as online asset stores, but often, artists have to describe complex procedural models that synthesize a scene by arranging these assets in a realistic layout. This often requires referring to a large number of real images to carefully adjust the procedural model, which can be a very time-consuming task. In the case of scenarios such as street scenes, it may be necessary to adjust from scratch a procedural model created for another city in order to create a synthetic scene related to a certain city. Techniques according to various embodiments can attempt to provide automated techniques for handling these and other such tasks.

[0007] In one approach, scene parameters in a scene generated by synthesis can be optimized by leveraging the visual similarity between the generated (e.g., rendered) synthetic data and real data. The scene structure and parameters can be represented by a scene graph, and data is generated by sampling a random scene structure (and parameters) from a given probabilistic grammar of the scene and then modifying the scene parameters using a learned model. Since such techniques only learn scene parameters, the gap between simulation and reality remains within the scene structure. For example, in Manhattan, it is understandable that the density of cars, people, and buildings is higher than that of an ancient Italian village. Other work on generative models of structured data such as graphs and grammar strings requires a large amount of ground-truth data for training to generate realistic samples. However, the scene structure is very cumbersome for annotation and thus not available in most real datasets.

[0008] The methods according to various embodiments can utilize a procedural generation model of synthetic scenes learned without a teacher from real images. In at least one embodiment, one or more scene graphs can be generated for each object by sampling rule expansions from a given probabilistic scene grammar and learning to generate scene parameters. Learning without a teacher for such tasks can be difficult, at least in part due to the discrete nature of the scene structures to be generated and the existence of non-differentiable renderers in the generation process. For this purpose, the generated (e.g., rendered) scenes can be compared to real scenes using feature space divergence that can be determined for individual scenes. Such methods can enable confidence assignment for training using reinforcement learning. Experiments with two synthetic datasets and a real dataset have shown that the method according to at least one embodiment significantly reduces the distribution gap between the scene structures in the generated data and the scene structures in the target data, and improves the human prior distribution regarding scene structures by learning to closely match the distribution of the target structures. In the real dataset, starting from a minimal human prior distribution, the structural distribution of real target scenes can be almost accurately restored, which is notable considering that this model can be trained without labels. Object detectors trained with this generated data have been shown to be superior, for example, to detectors trained with data generated using the human prior distribution, demonstrating an improvement in the measure of distribution similarity of rendered images generated using real data.

[0009] The techniques according to various embodiments can generate new data similar to the target distribution, instead of performing inference for each scene like the prior distribution technique. One technique is to optimize a differentiable simulator using a variational upper bound of an objective function such as a GAN, or to learn to optimize the simulator parameters of a control task by directly comparing the real trajectory and the simulated trajectory. The technique according to at least one embodiment can learn to generate a discrete scene structure constrained by grammar while optimizing the objective function of distribution matching (using reinforcement learning), instead of using adversarial training. Such techniques can be used to generate large and complex scenes, as opposed to images of a single object or face.

[0010] In at least one embodiment, a generative model consisting of graphs and trees can generate a graph with a richer structure that is more flexible than a grammar-based model, but may not be able to generate a syntactically correct graph when accompanied by a defined syntax such as a program and a scene graph. Grammar-based methods have been used for various tasks such as program translation, generation of conditional programs, induction of grammars, and generative modeling of structures with syntax such as molecules. However, these methods assume that the ground-truth graph structure can be accessed for learning. The techniques according to various embodiments can train the model in a teacherless manner without annotation of the ground-truth scene graph.

[0011] In at least one embodiment, a set of rules can be generated or obtained for a virtual scene or environment that is to be generated. For example, an artist may generate or provide a set of rules to be used in a scene, or select from a library of sets of rules for various scenes. This may include, for example, browsing through scene type options via a graphical interface and selecting a scene type having an associated set of rules that may correspond to scene types such as European cities, American countryside, dungeons, etc. For a given virtual or “synthetic” scene, the artist may create or obtain the content of various objects (e.g., “assets”) within the scene, and this content may include models, images, textures, and other features that can be used to render the objects. The rules provided can indicate how these objects should relate to each other within a given scene.

[0012] For example, FIG. 2A shows an exemplary set of rules 200 that can be utilized by various embodiments. This set of rules can include, in some embodiments, any number of rules, or up to a maximum number. Each rule can define a relationship between at least two types of objects that are to be represented in the synthetic scene. This set of rules is applied to a location that has roads and sidewalks. As shown in the set of rules 200, according to the first rule, a road can have lanes. According to an additional rule, a lane can be one lane or multiple lanes, and each lane can be associated with a sidewalk and one or more vehicles. As illustrated, the rules can also define whether an object type can have one or more instances of that type of object associated with a given object type. Another pair of rules indicates that one or more people may be present on the sidewalks within this scene.

[0013] Using such rules, one or more scene structures representing the scenes that would be generated can be generated. FIGS. 2B and 2C show two exemplary scene structures 230, 260. In each of these structures, the road is shown as the main parent node in a hierarchical tree structure. The rules from the rule set 200 and the relationships defined therein determine potential parent-child relationships that can be used to generate different tree structures that conform to those relationships. In at least one embodiment, the generation model can generate these scene structures from the rule set using an appropriate sampling or selection process. In the structure 230 of FIG. 2B, there is a two-lane road with a sidewalk in each lane, and there is a car in one of the lanes. Also, there are trees and people near the sidewalk on the side of the lane with the car. In the structure 260 of FIG. 2C, there is a one-lane road with three cars and a sidewalk, and there are two people and trees near the sidewalk. As can be seen, these structures represent two different scenes generated from the same rule set. Using such techniques, a virtual environment can be constructed, supplemented, extended, generated, or synthesized that includes variations of object structures that conform to all of the selected rule sets. Using such techniques, without the user manually selecting or placing these objects, a single rule set can be used to generate an environment of the desired size with variations that can follow a random or variation policy. Such methods can be used in an unsupervised manner to generate synthetic scenes from real images. In at least one embodiment, such techniques can learn a generation model of the scene structure and render samples therefrom (using additional scene parameters) to create synthetic images and labels. In at least one embodiment, using such techniques, synthetic training data with appropriate labels can be generated using unlabeled real data. The rule set and the unlabeled real data can be provided as input to a generation model that can generate a set of diverse scene structures.

[0014] The approach according to at least one embodiment can learn such a generative model for a synthetic scene. Specifically, in the case of a dataset of real images X R , the problem is to create synthetic data D(θ)=(X(θ),Y(θ)) of images X(θ) and labels Y(θ) that represent X R , where θ represents the parameters of the generative model. In at least one embodiment, it is stipulated that the synthetic data D is an output that creates an abstract scene representation, and by rendering that scene representation using a graphics engine, the progress of the graphics engine and rendering can be utilized. Rendering can ensure that it is not necessary to model the low-level pixel information in X(θ) (and its corresponding annotation Y(θ)). To ensure the semantic validity of the sampled scenes, it may be necessary to impose at least some constraints on their structure. Scene grammars use a set of rules to greatly reduce the space of sampleable scenes and make the learning a more structured and tractable problem. For example, a scene grammar can explicitly enforce that a car can only exist on a road, which does not need to be implicitly learned. The approaches according to various embodiments can partially utilize this by using a probabilistic scene grammar. A scene-graph structure can be sampled from a prior distribution imposed on a probabilistic context-free grammar (PCFG), which is referred to herein as a structure prior distribution. For all nodes in the scene graph, parameters are sampled from a parameter prior distribution and learned to predict new parameters for each node while keeping the structure intact. Thus, the resulting scenes are derived from a (context-free) structure prior distribution and a learned parameter distribution, which may introduce a simulation-reality gap in the scene structure.

[0015] Techniques according to various embodiments can at least mitigate this gap by learning a teacherless context-dependent structure distribution of a synthetic scene from an image. In at least one embodiment, one or more scene graphs can be used as an abstract scene representation, which can be rendered onto a corresponding image using labels. FIGS. 3A-3C show components that can be used at different stages of such a process. FIG. 3A shows a set 300 of logits generated from regular samples, where a given sample is used to determine the next logit. FIG. 3B shows a corresponding mask 330 that can be utilized when generating a scene graph. In the scene graph generation process, the shape of the logits and the mask is T max ×K. In FIG. 3, unpatterned (e.g., solid white) regions represent higher values, and regions filled with a pattern represent lower values. Such a process can autoregressively sample rules at each time step, predict logits for the next rule conditioned on that sample, and capture context-dependence. As shown in FIG. 3C, sampling can be used to generate a scene structure 362 and determine the parameters of the nodes of that scene structure. These parameters can include information such as, for example, location, height, and pose. These and other parameters 366 can be sampled and applied to each node within the scene structure to generate a complete scene graph. Thus, such a process can utilize rules sampled from a grammar and convert them into a graph structure. In this example, only objects that can be rendered are retained from a complete grammar string. The parameters of all nodes can be sampled from a prior distribution or, optionally, learned. The generated scene graph can be rendered as shown. Such a generative model can sequentially sample expansion rules from a given probabilistic scene grammar to generate a rendered scene graph. This model can be trained using reinforcement learning without a teacher using a feature matching-based distribution divergence specially designed to be adaptable to such settings.

[0016] In at least some embodiments, the scene graph can be advantageous because in fields such as computer graphics and vision, it can describe a scene in a concise hierarchical way where each node describes an object in the scene along with its parameters. The parameters can be related to the appearance, such as a 3D asset or pose. The parent-child relationship can define the parameters of the child nodes relative to the parent, enabling direct editing and manipulation of the scene. Additionally, cameras, lighting, weather, and other effects can be encoded in the scene graph. Generating the corresponding pixels and annotations can correspond to placing the objects in the scene within the graphics engine and rendering using the defined parameters.

[0017] In at least one embodiment, a set of rules can be defined as a vector having a length equal to the number of rules. Then, a network can be used to determine which of these rules to deploy, and these rules can be sequentially deployed at different time steps. For each scene that is to be generated, a categorical distribution can be generated over all the relevant rules in the set. Then, the generation network can sample from this categorical distribution to select the rules to use for this scene, where the categorical distribution can be masked so as to be forced to have a zero probability that a particular rule is not selected. The network can also infer which rule or option to deploy for each object. In at least one embodiment, this generation model can be a recurrent neural network (RNN). A latent vector defining the scene can be input into this RNN, and the rules of the scene can be sequentially generated and deployed based on the determined probabilities. The RNN proceeds down a tree or stack until all the rules are processed (or until the maximum number of rules is reached).

[0018] In one or more embodiments, each row of FIG. 3A can correspond to a sampling rule. As shown in FIG. 3, then, the mask 330 can be used to indicate to the model which rule should be deployed at a given time step. This process can be repeatedly executed to generate a valid scene description. Further, geometric constraints on the scene are also provided so that relationships between objects in the scene graph cannot exist outside of the specified relationships, such as a car cannot be located outside of a lane or on a sidewalk. The parameters for the nodes define various visual attributes so that a rural road in New Zealand looks different from a road in a large city in Thailand. In at least some embodiments, there may be ranges set for these various parameters for certain types of objects, such as sidewalks are provided only at a specific width and roads have only a limited number of lanes at most.

[0019] Additional learning can also be performed using these data structures. For example, this data can be used downstream to train a model to detect cars, for example, in captured image data. This structure can be retained with the generated images, for example, so that the model can more quickly identify cars based on where in the scene structure they occur.

[0020] In at least one embodiment, a context-free grammar G can be defined as a list of symbols (e.g., terminal and non-terminal symbols) and expansion rules. Non-terminal symbols have at least one expansion rule to a new set of symbols. Sampling from the grammar can involve expanding the start symbol (or initial symbol or parent symbol) until only non-terminal symbols remain. The total number of expansion rules K can be defined in the grammar G. A scene grammar can be defined and strings can be sampled from the represented grammar using one or more scene graphs. For each scene graph, a structure T can be sampled from the grammar G, and subsequently, corresponding parameters α can be sampled for all nodes in the graph. In at least one embodiment, a convolutional network is used with this scene graph to sample one set of parameters for all single nodes in the graph.

[0021] In various approaches, generative models with grammatically constrained graphs can be utilized. In at least one embodiment, a recurrent neural network can be used to map a latent vector z to an unnormalized probability over all possible grammar rules in an autoregressive manner. In such embodiments, this can continue for a maximum of T max steps. In at least one embodiment, at each time step, one rule r t can be sampled and this rule can be used to predict the logits for the next rule f t+1 . This enables the model to easily capture context dependencies, in contrast to the context-free nature of conventional approaches to scene graphs. For a list of rules sampled up to a maximum of T max , as shown in FIG. 3, the corresponding scene graph can be generated by treating each rule expansion as a node expansion in the graph.

[0022] To ensure the validity of these sampled rules at each time step t, a last-in-first-out (LIFO) stack of unexpanded non-terminal nodes can be maintained. Nodes are popped from the stack, expanded according to the sampled rule expansion, and then the resulting new non-terminal nodes are pushed onto the stack. When a non-terminal symbol is popped, a mask m of size K t can be generated, which is 1 if the rule from that non-terminal symbol is valid and 0 otherwise. Let the logit for the next expansion be f t , then the probability r of the rule t,k is [Number] obtained by

[0023] Sampling from this masked multinomial distribution can ensure that only valid rules are sampled as r t . Let the logit and the sampled rule be (f t , r t ) ∀t ∈ 1...T max , then the probability of the corresponding scene structure T given z is [Number] obtained by

[0024] Putting it all together, an image can be generated by sampling a scene structure T ~ q θ (·|z) from the model, then subsequently sampling the parameters α ~ q(·|T) of all nodes in the scene, and rendering the image v’ = R(T, α) ~ q I . For a given v’ ~ q I , in the case of parameters α and structure T, q I (v’|z) = q(α|T)q θ (T|z) can be assumed as follows.

[0025] For such generative models, various training methods can be utilized. In at least one embodiment, this training can be performed using variational inference or by optimizing a measure of distribution similarity. Variational inference enables the use of a reconstruction-based objective function by introducing an approximate learned posterior distribution. Training such a model using variational inference can be difficult, at least in part, due to the complexity arising from discrete sampling and the presence of a renderer in the generation process. Furthermore, the recognition network here would correspond to performing inverse graphics, which is a very difficult problem in itself. In at least one embodiment, a measure of the distribution similarity between the generated data and the target data can be optimized. Adversarial training of the generative model can be utilized together with reinforcement learning (RL) by, for example, carefully restricting the capabilities of the critic. In at least one embodiment, reinforcement learning can be used to train a discrete generative model of a scene graph. Samples can be calculated for all samples, thereby significantly improving the entire training process.

[0026] The generative model can be trained to match the distribution of the features of real data in the latent space of a certain feature extractor φ. The real feature distribution can be defined as for a certain v~p I with respect to p f s.t F~p f ⇔F = φ(v). Similarly, the generated feature distribution can be defined as for a certain v~q I with respect to q f s.t F~q f ⇔F = φ(v). In at least one embodiment, distribution matching can be achieved by approximately calculating p f , q f from samples and minimizing the KL divergence from p f to q f . In at least one embodiment, the training objective function is

Number

[0027] Using the definition of the above characteristic distribution, the equivalent objective function is

Number

[0028] The true underlying characteristic distributions q f and p f may be difficult to handle for calculation. In at least one embodiment, approximations calculated using kernel density estimation (KDE)

Number

Number

Number

[0029] The generative model according to at least one embodiment can make discrete (e.g., non-differentiable) choices at each step, and as a result, it can be advantageous to optimize the objective function using reinforcement learning techniques. Specifically, this can include using a REINFORCE score function estimator along with a moving average baseline length, whereby the gradient is

Number

Number

Number

[0030] Note that the above gradient needs to calculate the marginal probability q I (v’) of the generated image v’ rather than the conditional q I (v’|z). Calculating the marginal probability of the generated image involves an intractable marginalization over the latent variable z. To avoid this, it is possible to simply marginalize using a fixed finite number of latent vectors from the set Z that are uniformly sampled. This is

Number

[0031] The maximum length T that can be sampled from the grammar maxSince the scene graph is finite, such an approach can also provide sufficient modeling capabilities. Empirically, the probability in regular sampling can compensate for the lost probability in the latent space, so using a single latent vector may be sufficient.

[0032] In at least one embodiment, pre-training can be an important step. On the scene structure, a manually defined prior distribution can be defined. For example, a simple prior distribution can be placing one car on one road in a driving scene. The model can be pre-trained by sampling strings (e.g., scene graphs) from the grammar prior distribution at least partially and training the model to maximize the log-likelihood of these scene graphs. Feature extraction can also be an important step in distribution matching because effective training requires capturing structural scene information such as the number of objects and their contextual spatial relationships.

[0033] During the training of the model, sampling can result in incomplete strings generated in at most T max steps. Therefore, the scene graph T can be repeatedly sampled until its length reaches at most T max . To ensure that this does not require too many trials, the rejection rate r reject (F) of the sampled feature F can be recorded as the average failed sampling trials when sampling a single scene graph used to generate F. A threshold f can be set for r reject (F), which can represent the maximum allowable rejection amount and the weight λ, and this can be added to the original loss obtained by

Equation

[0034] Such an approach can provide unsupervised learning of a generative model of synthetic scene structure by optimizing visual similarity to real data. Inferring scene structure is famously difficult even when annotations are provided. The approach according to various embodiments can perform this generative part without ground-truth information. Experiments have demonstrated that such models learn a reasonable posterior distribution over scene structure and significantly improve over manually designed prior distributions. The approach can optimize both the scene structure and the parameters of the synthetic scene generator to yield satisfactory results.

[0035] As described above, such techniques for generating diverse scene graphs can generate a scene or environment that mimics the real world, or the target world or environment. Information about this world or environment can be learned directly from the pixels of image examples of the real world or target world. Such techniques can be used to attempt an accurate reconstruction, but in many embodiments, they can enable the generation of an infinite variety of worlds and environments that can be at least partially based on these real-world or target worlds. A set of rules or scene grammar can be provided that describes the world at a micro level and defines object-specific relationships. For example, instead of a person having to manually generate at least the layout of each scene or image, the person can specify or select rules that can be used to automatically generate that scene or image. The scene graph in at least one embodiment can provide a detailed description of the layout of a three-dimensional world. In at least one embodiment, it is also possible to generate a string that provides a complete representation or definition of the layout of a three-dimensional scene using a recursive expansion of the rules for generating the scene structure. As described above, a generation model can be used to perform the expansion and generate the scene structure. In at least one embodiment, this scene structure can be stored as a JSON file, or using another such format. This JSON file can then be provided as input to a rendering engine for generating an image or scene. The rendering engine can extract the appropriate asset data to use when rendering individual objects.

[0036] As described above, the rendered data can be used to present a virtual environment, such as for gaming or VR applications. This rendering can be used to generate training data for such applications, as well as for other applications, such as training models for autonomous or semi-autonomous machines, such as vehicle navigation or robot simulation. There may be various libraries of assets that can be selected for these renderings so that the environment can be appropriate for various geographical locations, points in time, etc. Each pixel of the rendered image can be labeled to indicate the type of object that the pixel represents. In at least one example, for each object, a bounding box or other position indicator, as well as a depth determined for each pixel of the 3D scene, a normal at that pixel position, etc., can be generated. This information can be extracted from the rendering engine by utilizing a rendering function suitable for extracting data in at least some examples.

[0037] In at least one embodiment, an artist can provide a set of assets, select a set of rules, and the entire virtual environment can be generated without manual input by the artist. In some embodiments, there may be a library of assets that the artist can select from. For example, to generate an environment that includes appropriate visual objects and layouts and is based on Japanese cities, but does not directly correspond to or represent Japanese cities, the artist can select a scene structure for "Japanese cities" and assets for "Japanese cities". In some embodiments, the artist may have the ability to adjust this environment by indicating what the artist likes or dislikes. For example, the artist may not want cars on the streets for this application. Thus, the artist may indicate that the artist does not want to include cars, or at least does not want to include them in a particular area, or does not want to associate them with a particular object type, and a new scene graph can be generated with the cars removed and the appropriate relationships updated. In some embodiments, the user can provide, obtain, utilize, or generate two or more subgraphs such that the user can indicate what the user likes and dislikes. These subgraphs can then be used to generate a new scene that better meets the user's expectations. By such means, the user can easily generate a virtual environment with a particular look and visual appearance without the need for specialized knowledge of the creation process or the need to manually place, move, or adjust objects within the scene or set of scenes. By such means, an ordinary person can become a 3D artist with a minimum of effort on that person's part.

[0038] Figure 4 shows an example process 400 for generating an image of a scene that can be utilized by various embodiments. For this process and other processes presented herein, unless otherwise specified, additional, fewer, or alternative steps may be present and may be performed in a similar order or in an alternative order, or at least partially in parallel, within the scope of the various embodiments. In this example, a rule set to be used for at least one scene to be generated is determined 402. This may include, among other such options, the user generating these rules or selecting from among several rule sets. The individual rules within the set can define the relationships between the types of objects within the scene. This rule set can be sampled 404 based on the determined probabilities to generate a scene structure that includes the relationships of the objects defined by the rules. In at least one embodiment, this can include a hierarchical scene structure having nodes of a hierarchy corresponding to the types of objects in the scene. The parameters to be used when rendering each of these objects can be determined 406, such as by sampling from an appropriate data set. Next, a scene graph can be generated 408 based on the scene structure, but with appropriate parameters applied to the individual nodes or objects. The scene graph can be provided 410 to a renderer 410 or other target for generating an image or scene, along with an asset library or other source of object content. A rendered image of the scene, rendered based on the determined scene graph, can be received 412. This rendered image may include object labels and retain the scene structure if the image is used as training data as described herein.

[0039] To train the neural networks described herein, various techniques can be used. For example, a generative model can be trained to analyze unlabeled images that can correspond to images captured in a real-world setting. The generative network can then be trained to generate scenes having a similar appearance, layout, and other such looks. Both real and synthetic scenes can exist. When passing these scenes to a deep neural network, a set of features corresponding to positions in a high-dimensional space, such as a 1,000-dimensional space, can always be extracted within that scene. In that case, this scene can be considered to be composed of the points of those features in this high-dimensional space rather than the pixels in the image. The network can be trained so that the features corresponding to the synthetic scene in this feature space match the features corresponding to the real scene. Thus, it can be difficult to distinguish the features of the real and synthetic scenes in the feature space.

[0040] In at least one embodiment, this can be achieved using reinforcement learning. As described above, the goal may be to align the two entire datasets, but since this is an uncorrelated teacherless space, there is no data or correlation regarding the specific feature points to be aligned. Since the goal in many situations is not to generate an exact copy of the scene but rather to generate a similar scene, it may be sufficient to simply align the distribution of feature points within this feature space. Thus, the training procedure can compare the real scene and the synthetic scene globally. To evaluate a scene, the distribution of its feature points within the feature space can be compared to the distribution of feature points for other scenes. In various approaches, it can be difficult to determine whether a scene is realistic or useful, in addition to whether the structure is appropriate, based solely on signal comparison. Thus, the training approach according to at least one embodiment can extract signals from each individual data point itself without the need to look at the entire dataset. In this way, the signal can be evaluated regarding how well a particular scene is aligned with the entire dataset. In that case, the likelihood that a particular scene can be decomposed into all synthetic scenes can be calculated. The calculation can provide the likelihood that this particular scene is synthetic. In at least one embodiment, the calculation can be performed by using kernel density estimation (KDE). Using KDE, the probability that this scene belongs to the distribution of synthetic scenes can be obtained. Using KDE, the probability that this scene belongs to the distribution of real scenes can also be calculated. In at least one embodiment, the ratio of these values can be analyzed and this ratio can be used to optimize the system. Maximizing this ratio (log) as a reward function for the scene provides a signal that can be optimized for each individual scene.

[0041] FIG. 5 shows an example process 500 for training a network to generate realistic images that can be utilized by at least one embodiment. In this example, as described above with respect to FIG. 4, a scene graph and assets are obtained 502. Using the scene graph and assets, a composite image of the scene can be generated 504. For this generated image, the locations of feature points in an n-dimensional feature space can be determined 506, where n can be equal to the number of rules in the set used to generate the scene graph. These feature points for the generated image can be compared 508 with the distribution of feature points for the composite image in that feature space. Based on the comparison, a first probability that this generated image is synthetic can be determined 510. The feature points for the generated image can also be compared 512 with the distribution of feature points for real images in that feature space. Based on the comparison, a second probability that this generated image is realistic can be determined 514. The ratio of these two probabilities can be calculated 516, and one or more weights of the trained network can be adjusted to optimize for this ratio.

[0042] Another embodiment can utilize the discriminator of a GAN. The GAN can be trained to use the discriminator portion to determine whether the generated scene is realistic. The network can then be optimized so that the discriminator can determine with a high probability that the scene is realistic. However, current renderers produce high-quality images, but these images may still be identified as not being real captured images, so such an approach can be difficult because the discriminator may be able to primarily distinguish the differences even when the images are very structurally similar. In such cases, the rendered images do not confuse the discriminator into thinking they are real images, so the discriminator cannot provide useful information and the GAN may break down during training. In at least one embodiment, an image transformation can be performed before providing this image data to the GAN in an attempt to improve the appearance of these synthesized images. Image transformation may help reduce the style gap between real and synthetic images and may help with the low-level visual appearance related to textures or reflections that can make an image look synthetic rather than real. This can be advantageous, for example, for systems that utilize ray tracing to generate reflections and other lighting effects.

[0043] In another embodiment, a process for guaranteeing termination can be used. The neural network can be defined to execute a fixed number of steps, such as 150 steps, for computational reasons. If this number is too small, it may not be possible to fully generate the scene graph based on more rules to be analyzed and developed. Thus, the generated scene graph will be incomplete, resulting in an inaccurate rendering. In at least one embodiment, the network can be made executable up to its limit. If that limit is insufficient for the scene, the remaining features can be determined and a negative reward can be applied to the model to generate those features again. Such an approach may result in a scene that does not contain all the features originally desired, but guarantees that a scene that matches the rendering limit can be rendered.

[0044] In at least one embodiment, the client device 602 can generate content for a session using components of the content application 604 on the client device 602 and data stored locally on that client device. In at least one embodiment, a content application 624 (e.g., an image generation or editing application) executed at the content server 620 can initiate at least a session associated with the client device 602 so as to be able to utilize a session manager and user data stored in the user database 634. If required for this type of content or platform, the content manager 626 determines the content 632, renders the content 632 using a rendering engine, and uses a transmission manager 622 suitable for sending by download, streaming, or another such transmission channel to send the content 632 to the client device 602. In at least one embodiment, this content 632 can include assets that can be used by the rendering engine to render a scene based on a determined scene graph. In at least one embodiment, the client device 602 that receives this content can provide this content to the corresponding content application 604, and this content application 604 may additionally or alternatively include a rendering engine for rendering at least a portion of this content for presentation via the client device 602, such as image or video content via the display 606 and audio such as voice and music via at least one audio playback device 608 such as a speaker or headphones.In at least one embodiment, at least a portion of the content may already be stored on the client device 602, rendered on the client device 602, or otherwise accessible to the client device 602, such that for at least that portion of the content, transmission via the network 640 is not required, for example, if the content has been previously downloaded or may be locally stored on a hard drive or optical disk. In at least one embodiment, a transmission mechanism such as data streaming may be used to transfer this content from the server 620 or the content database 634 to the client device 602. In at least one embodiment, at least a portion of the content may be obtained or streamed from another source, such as a third-party content service 660 that may also include a content application 662 for generating or providing the content. In at least one embodiment, a portion of this functionality may be performed using multiple processors within one or more computing devices, which may include a combination of multiple computing devices or a CPU and GPU.

[0045] In at least one embodiment, the content application 624 includes a content manager 626 that can determine or analyze the content before it is sent to the client device 602. In at least one embodiment, the content manager 626 can also include or operate with other components capable of generating, modifying, or enhancing the content to be provided. In at least one embodiment, this can include a rendering engine for rendering image or video content. In at least one embodiment, a scene graph generation component 628 can be used to generate a scene graph from a set of rules and other such data. In at least one embodiment, an image generation component 630, which may also include a neural network, can generate an image from this scene graph. Then, in at least one embodiment, the content manager 626 can cause this generated image to be sent to the client device 602. In at least one embodiment, the content application 604 on the client device 602 can also include components such as a rendering engine, a scene graph generator 612, and an image generation module 614 so that any or all of this functionality can be additionally or alternatively executed on the client device 602. In at least one embodiment, the content application 662 on the third-party content service system 660 can also include such functionality. In at least one embodiment, the location where at least a portion of this functionality is executed can be configurable or can depend on factors such as, among other such factors, the type of the client device 602 or the availability of a network connection with an appropriate bandwidth. In at least one embodiment, the system for content generation can include any suitable combination of hardware and software in one or more locations.In at least one embodiment, the generated image or video content of one or more resolutions may be provided to other client devices 650 or made available to other client devices 650, for example, for downloading or streaming from a media source that stores a copy of the image or video content. In at least one embodiment, this may include transmitting an image of game content for a multi-player game, and different client devices may display the content at different resolutions, including one or more super-resolutions.

[0046] In this example, these client devices can include any suitable computing device, which can include, among other options, desktop computers, laptop computers, set-top boxes, streaming devices, gaming consoles, smartphones, tablet computers, VR headsets, AR goggles, wearable computers, or smart TVs. Each client device can send requests via at least one wired or wireless network, which can include, among other such options, the Internet, Ethernet®, local area network (LAN), or cellular network. In this example, these requests can be sent to an address associated with a cloud provider, which may operate or control one or more electronic resources within a cloud provider environment that can include a data center or server farm, etc. In at least one embodiment, the requests may be received or processed by at least one edge server located on the network edge and outside of at least one security layer associated with the cloud provider environment. In this way, latency can be reduced by enabling the client device to interact with a closer server while improving the security of the resources within the cloud provider environment.

[0047] In at least one embodiment, such a system can be used to perform graphical rendering operations. In other embodiments, such a system can be used for other purposes, such as to perform simulation operations for testing or validating autonomous machine applications or to perform deep learning operations. In at least one embodiment, such a system can be implemented using edge devices or may incorporate one or more virtual machines (VMs). In at least one embodiment, such a system can be implemented at least partially within a data center or using at least partially cloud computing resources.

[0048] Inference and training logics FIG. 7A shows inference and / or training logic 715 used to perform inference and / or training operations with respect to one or more embodiments. Details regarding the inference and / or training logic 715 are provided below in conjunction with FIGS. 7A and / or 7B.

[0049] In at least one embodiment, the inference and / or training logic 715 may include, without limitation, code and / or data storage 701 for storing forward propagation and / or output weights, and / or input / output data, and / or other parameters that configure neurons or layers of a neural network that are trained and / or used to infer in aspects of one or more embodiments. In at least one embodiment, the training logic 715 may include or be coupled to code and / or data storage 701 for storing graph code or other software for controlling timing and / or order, and the code and / or data storage 701 has weight and / or other parameter information loaded therein to configure logic including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, code such as graph code loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, the code and / or data storage 701 stores the weight parameters and / or input / output data of each layer of a neural network that is trained or used in conjunction with one or more embodiments while propagating the input / output data and / or weight parameters forward during training and / or inference using aspects of one or more embodiments. In at least one embodiment, any portion of the code and / or data storage 701 may be included with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory.

[0050] In at least one embodiment, any portion of the code and / or data storage 701 may be internal or external to one or more processors, or other hardware logic devices or circuits. In at least one embodiment, the code and / or the code and / or data storage 701 may be cache memory, dynamic randomly addressable memory (“DRAM”), static randomly addressable memory (“SRAM”), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the selection of whether the code and / or the code and / or data storage 701 is internal or external to, for example, a processor, or the selection of whether it is composed of DRAM, SRAM, flash, or some other type of storage, may be determined according to on-chip versus off-chip available storage, latency requirements of the training and / or inference functions being executed, the batch size of the data used in neural network inference and / or training, or any combination of these factors.

[0051] In at least one embodiment, the inference and / or training logic 715 may include, without limitation, backward propagation and / or output weights corresponding to neurons or layers of a neural network that are trained and / or used for inference in one or more embodiments, and / or code and / or data storage 705 for storing input / output data, and / or may be coupled thereto. In at least one embodiment, the code and / or data storage 705 stores the weight parameters and / or input / output data of each layer of a neural network that is trained or used in conjunction with one or more embodiments while backpropagating the input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, the training logic 715 may include or be coupled to code and / or data storage 705 for storing graph code or other software for controlling timing and / or order, and the code and / or data storage 705 has weight and / or other parameter information loaded therein to configure logic including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, code such as graph code loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any portion of the code and / or data storage 705 may be included with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory. In at least one embodiment, any portion of the code and / or data storage 705 may be internal or external to one or more processors, or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 705 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage.In at least one embodiment, the choice of whether the code and / or data storage 705 is, for example, internal or external to the processor, or the choice of whether it is composed of DRAM, SRAM, flash, or some other type of storage, may be determined according to the on-chip versus off-chip available storage, the latency requirements of the training and / or inference functions to be executed, the batch size of the data used in the neural network inference and / or training, or any combination of these factors.

[0052] In at least one embodiment, the code and / or data storage 701 and the code and / or data storage 705 may be separate storage structures. In at least one embodiment, the code and / or data storage 701 and the code and / or data storage 705 may be the same storage structure. In at least one embodiment, the code and / or data storage 701 and the code and / or data storage 705 may be partially the same storage structure and partially separate storage structures. In at least one embodiment, any part of the code and / or data storage 701 and the code and / or data storage 705 may be included together with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including the system memory.

[0053] In at least one embodiment, the inference and / or training logic 715 includes, without limitation, one or more arithmetic logic units (“ALUs”) 710 including integer and / or floating point units to perform logical and / or arithmetic operations based at least in part on and / or indicated by training and / or inference code (e.g., graph code), the result of which may generate activations (e.g., output values from a layer or neuron within a neural network) stored in the activation storage 720, which are a function of the code and / or data storage 701, and / or the input / output and / or weight parameter data stored in the code and / or data storage 705. In at least one embodiment, the activations stored in the activation storage 720 are generated in accordance with linear algebra calculations and / or matrix-based calculations performed by the ALU 710 in response to executing instructions or other code, where the weight values stored in the code and / or data storage 705 and / or the code and / or data storage 701 are used as operands along with other values such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in the code and / or data storage 705, or the code and / or data storage 701, or another storage on-chip or off-chip.

[0054] In at least one embodiment, the ALU 710 is included within one or more processors, or other hardware logic devices or circuits, but in another embodiment, the ALU 710 may be external to the processors or other hardware logic devices or circuits that use them (e.g., a coprocessor). In at least one embodiment, the ALU 710 may be included within the execution unit of a processor, or may be distributed among execution units of processors, either within the same processor or different processors of different types (e.g., a central processing unit, a graphics processing unit, a fixed function unit, etc.), such that it is accessible within an ALU bank that may be otherwise included. In at least one embodiment, the code and / or data storage 701, the code and / or data storage 705, and the activation storage 720 may be in the same processor or other hardware logic device or circuit, and in another embodiment, they may be in different processors or other hardware logic devices or circuits, or in any combination of the same processor or other hardware logic device or circuit and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of the activation storage 720 may be included together with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory. Further, the inference and / or training code may be stored together with other code accessible to the processor or other hardware logic or circuit, and may be fetched and / or processed using the fetch, decode, scheduling, execution, retirement, and / or other logic circuits of the processor.

[0055] In at least one embodiment, the activation storage 720 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the activation storage 720 may be fully or partially inside or outside one or more processors or other logic circuits. In at least one embodiment, the selection of whether the activation storage 720 is, for example, inside or outside a processor, or the selection of being composed of DRAM, SRAM, flash, or some other type of storage, may be determined according to on-chip versus off-chip available storage, the latency requirements of the training and / or inference functions being executed, the batch size of the data used in the neural network inference and / or training, or a combination from any of these factors. In at least one embodiment, the inference and / or training logic 715 shown in FIG. 7A may be used in conjunction with an application-specific integrated circuit (ASIC) such as a TensorFlow® processing unit from Google, an inference processing unit (IPU) from Graphcore®, or a Nervana® (e.g., "Lake Crest") processor from Intel Corp. In at least one embodiment, the inference and / or training logic 715 shown in FIG. 7A may be used in conjunction with other hardware such as central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or field programmable gate array (FPGA).

[0056] FIG. 7B shows inference and / or training logic 715 according to at least one or more embodiments. In at least one embodiment, the inference and / or training logic 715 may include, without limiting to hardware logic, in which computing resources are dedicated to one or more layers of neurons in a neural network for weight values or other information, or otherwise used only in conjunction with them. In at least one embodiment, the inference and / or training logic 715 shown in FIG. 7B may be used in conjunction with an application specific integrated circuit (ASIC) such as a TensorFlow® processing unit from Google, an inference processing unit (IPU) from Graphcore®, or a Nervana® (e.g., "Lake Crest") processor from Intel Corporation. In at least one embodiment, the inference and / or training logic 715 shown in FIG. 7B may be used in conjunction with other hardware such as central processing unit (CPU) hardware, graphics processing unit ("GPU") hardware, or field programmable gate array (FPGA). In at least one embodiment, the inference and / or training logic 715 includes, without limitation, code and / or data storage 701, and code and / or data storage 705, and uses these to store code (e.g., graph code), weight values, and / or bias values, gradient information, momentum values, and / or other information including other parameters or hyperparameter information. In at least one embodiment shown in FIG. 7B, each of the code and / or data storage 701 and the code and / or data storage 705 is associated with dedicated computing resources such as computing hardware 702 and computing hardware 706, respectively. In at least one embodiment, each of the computing hardware 702 and the computing hardware 706 includes one or more ALUs that execute mathematical functions such as linear algebra functions only on the information stored in the code and / or data storage 701 and the code and / or data storage 705, respectively, and the results are stored in the activation storage 720.

[0057] In at least one embodiment, each of code and / or data storage 701 and 705, and corresponding computing hardware 702 and 706, respectively corresponds to different layers of a neural network, such that the activation resulting from one "storage / compute pair 701 / 702" of code and / or data storage 701 and computing hardware 702 is provided as an input to a "storage / compute pair 705 / 706" of code and / or data storage 705 and computing hardware 706 in order to reflect the conceptual organization of the neural network. In at least one embodiment, storage / compute pairs 701 / 702 and 705 / 706 may correspond to two or more layers of a neural network. In at least one embodiment, additional storage / compute pairs (not shown) may be included in inference and / or training logic 715 after or in parallel with storage / compute pairs 701 / 702 and 705 / 706.

[0058] Data center FIG. 8 shows an exemplary data center 800 in which at least one embodiment may be used. In at least one embodiment, data center 800 includes a data center infrastructure layer 810, a framework layer 820, a software layer 830, and an application layer 840.

[0059] As shown in FIG. 8, in at least one embodiment, the data center infrastructure layer 810 may include a resource orchestrator 812, grouped computing resources 814, and node computing resources (“node C.R.”) 816(1) to 816(N), where “N” represents any positive integer. In at least one embodiment, the node C.R. 816(1) to 816(N) may include any number of central processing units (“CPU”) or other processors (including accelerators, field programmable gate arrays (FPGA), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., semiconductor drives or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VM”), power modules, and cooling modules, but are not limited thereto. In at least one embodiment, one or more of the node C.R. 816(1) to 816(N) may be a server having one or more of the computing resources described above.

[0060] In at least one embodiment, the grouped computing resources 814 may include separate groups of node C.R.s housed within one or more racks (not shown), or multiple racks housed in a data center at various graphical locations (also not shown). Separate groups of node C.R.s within the grouped computing resources 814 may include grouped computing resources, network resources, memory resources, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, some node C.R.s that include a CPU or processor may be grouped within one or more racks to provide computing resources for supporting one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.

[0061] In at least one embodiment, the resource orchestrator 812 may configure or otherwise control one or more node C.R.s 816(1)-816(N) and / or the grouped computing resources 814. In at least one embodiment, the resource orchestrator 812 may include a software design infrastructure ("SDI") management entity for the data center 800. In at least one embodiment, the resource orchestrator may include hardware, software, or some combination thereof.

[0062] As shown in FIG. 8, in at least one embodiment, the framework layer 820 includes a job scheduler 822, a configuration manager 824, a resource manager 826, and a distributed file system 828. In at least one embodiment, the framework layer 820 may include a framework for supporting software 832 of the software layer 830 and / or one or more applications 842 of the application layer 840. In at least one embodiment, the software 832 or the application 842 may each include web-based service software or an application, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 820 may be a kind of free and open-source software web application framework, such as Apache Spark (registered trademark) (hereinafter referred to as "Spark") that can use the distributed file system 828 for large-scale data processing (e.g., "big data"), but is not limited thereto. In at least one embodiment, the job scheduler 822 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 800. In at least one embodiment, the configuration manager 824 may be able to configure different layers, such as the software layer 830 and the framework layer 820 including Spark and the distributed file system 828 for supporting large-scale data processing. In at least one embodiment, the resource manager 826 may be able to manage clustered or grouped computing resources mapped or allocated to support the distributed file system 828 and the job scheduler 822. In at least one embodiment, the clustered or grouped computing resources may include grouped computing resources 814 in the data center infrastructure layer 810.In at least one embodiment, the resource manager 826 may manage these mapped or allocated computing resources in cooperation with the resource orchestrator 812.

[0063] In at least one embodiment, the software 832 included in the software layer 830 may include software used by at least a portion of the node C.R. 816(1) - 816(N), the grouped computing resources 814, and / or the distributed file system 828 of the framework layer 820. One or more types of software may include, but are not limited to, Internet web page search software, email virus scan software, database software, and streaming video content software.

[0064] In at least one embodiment, the application 842 included in the application layer 840 may include one or more types of applications used by at least a portion of the node C.R. 816(1) - 816(N), the grouped computing resources 814, and / or the distributed file system 828 of the framework layer 820. One or more types of applications may include, but are not limited to, any number of genomics applications, recognition computing, and software for training or inference, machine learning applications including machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0065] In at least one embodiment, any one of configuration manager 824, resource manager 826, and resource orchestrator 812 may implement any number and type of self-corrective measures based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-corrective measures may prevent the data center operator of data center 800 from determining configurations that may be defective and may eliminate parts of the data center that are not being fully utilized and / or have low performance.

[0066] In at least one embodiment, data center 800 may include tools, services, software, or other resources for training one or more machine learning models or for predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by calculating weight parameters according to a neural network architecture using the software and computing resources described above with respect to data center 800. In at least one embodiment, a trained machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 800 by using weight parameters calculated by one or more techniques described herein.

[0067] In at least one embodiment, the data center may use a CPU, an application specific integrated circuit (ASIC), a GPU, an FPGA, or other hardware to perform training and / or inference using the resources described above. Further, the one or more software and / or hardware resources described above may be configured as a service to enable a user to perform training or inference of information, such as image recognition, speech recognition, or other artificial intelligence services.

[0068] Using the inference and / or training logic 715, inference and / or training operations associated with one or more embodiments are performed. Details regarding the inference and / or training logic 715 are provided below in conjunction with FIGS. 7A and / or 7B. In at least one embodiment, the inference and / or training logic 715 may be used in the system of FIG. 8 for inference or prediction operations, at least in part based on the training operations of the neural network described herein, the functions and / or architecture of the neural network, or the weight parameters calculated using the use cases of the neural network.

[0069] Using such components, a variety of scene graphs can be generated from one or more rule sets, and the scene graphs can be used to generate training data or image content representing one or more scenes of a virtual environment.

[0070] Computer system FIG. 9 is a block diagram showing an exemplary computer system, which may be formed with a processor that may include an execution unit for executing instructions, along with interconnected devices and components, a system-on-chip (SoC), or some combination of these 900, according to at least one embodiment. In at least one embodiment, computer system 900 may include, without limitation, components such as processor 902 for using an execution unit that includes logic for executing an algorithm for processing data in accordance with the present disclosure, such as in the embodiments described herein. In at least one embodiment, computer system 900 may include a processor such as a PENTIUM® processor family, Xeon™, Itanium® , XScale™, and / or StrongARM™, Intel® Core™, or Intel® Nervana™ microprocessor available from Intel Corporation of Santa Clara, California, although other systems (including PCs having other microprocessors, engineering workstations, set-top boxes, etc.) may be used. In at least one embodiment, computer system 900 may execute a version of the WINDOWS® operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX® and Linux®), embedded software, and / or graphical user interfaces may be used.

[0071] Embodiments may be used in other devices such as portable devices and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and portable PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor ("DSP"), a system-on-chip, network computers ("NetPCs"), set-top boxes, network hubs, wide area network ("WAN") switches, or any other system capable of executing one or more instructions according to at least one embodiment.

[0072] In at least one embodiment, computer system 900 may include, without limitation, a processor 902, which may include, without limitation, one or more execution units 908 for performing training and / or inference of a machine learning model according to the techniques described herein. In at least one embodiment, computer system 900 is a single-processor desktop or server system, but in another embodiment, computer system 900 may be a multi-processor system. In at least one embodiment, processor 902 may include, without limitation, a complex instruction set computer (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 902 may be coupled to a processor bus 910, which may transmit data signals between processor 902 and other components within computer system 900.

[0073] In at least one embodiment, processor 902 may include, without limitation, a level 1 (L1) internal cache memory (cache) 904. In at least one embodiment, processor 902 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory may be external to processor 902. Other embodiments may also include a combination of both internal and external caches, depending on the particular implementation and requirements. In at least one embodiment, register file 906 may store different types of data in various registers including, without limitation, integer registers, floating point registers, status registers, and instruction pointer registers.

[0074] In at least one embodiment, the execution unit 908, which includes without limitation logic for performing integer and floating point operations, is also in the processor 902. In at least one embodiment, the processor 902 may also include a microcode ( "u-code") read only memory ( "ROM") that stores microcode for certain macro instructions. In at least one embodiment, the execution unit 908 may include logic for handling a packed instruction set 909. In at least one embodiment, by including the packed instruction set 909 in the instruction set of a general purpose processor along with the associated circuitry for executing the instructions, operations used by many multimedia applications can be executed using the packed data of the general purpose processor 902. In one or more embodiments, by performing operations on packed data using the full width of the processor's data bus, many multimedia applications can be accelerated and executed more efficiently, thereby eliminating the need to transfer smaller units of data between the processor's data buses to perform one or more operations on one data element at a time.

[0075] In at least one embodiment, the execution unit 908 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, the computer system 900 may include a memory 920 without limitation. In at least one embodiment, the memory 920 may be implemented as a dynamic random access memory ( "DRAM") device, a static random access memory ( "SRAM") device, a flash memory device, or other memory device. In at least one embodiment, the memory 920 may store instructions 919 and / or data 921 represented by data signals that may be executed by the processor 902.

[0076] In at least one embodiment, a system logic chip may be coupled to a processor bus 910 and a memory 920. In at least one embodiment, the system logic chip may include, without limitation, a memory controller hub (MCH) 916, and the processor 902 may communicate with the MCH 916 via the processor bus 910. In at least one embodiment, the MCH 916 may provide a high-bandwidth memory path 918 to the memory 920 for storing instructions and data and for storing graphics commands, data, and textures. In at least one embodiment, the MCH 916 may direct data signals between the processor 902, the memory 920, and other components of the computer system 900 and may bridge data signals between the processor bus 910, the memory 920, and the system I / O interface 922. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 916 may be coupled to the memory 920 via the high-bandwidth memory path 918, and the graphics / video card 912 may be coupled to the MCH 916 via an Accelerated Graphics Port (AGP) interconnect 914.

[0077] In at least one embodiment, computer system 900 may use a system I / O 922, which is a proprietary hub interface bus for coupling the MCH 916 to an I / O controller hub (“ICH”) 930. In at least one embodiment, the ICH 930 may provide direct connections to several I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, without limitation, a high-speed I / O bus for connecting peripheral devices to the memory 920, chipset, and processor 902. By way of example, it may include, without limitation, a legacy I / O controller 923 including an audio controller 929, a firmware hub (“flash BIOS”) 928, a wireless transceiver 926, data storage 924, a user input and keyboard interface 925, a serial expansion port such as a Universal Serial Bus (“USB”), and a network controller 934. The data storage 924 may comprise a hard disk drive, a floppy (registered trademark) disk drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0078] In at least one embodiment, FIG. 9 shows a system including interconnected hardware devices or “chips,” while in other embodiments, FIG. 9 may show an exemplary system-on-chip (“SoC”). In at least one embodiment, the devices may be interconnected by proprietary interconnects, standard interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of the computer system 900 may be interconnected using a Compute Express Link (CXL) interconnect.

[0079] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 715 is used. Details regarding the inference and / or training logic 715 are provided below in conjunction with FIGS. 7A and / or 7B. In at least one embodiment, the inference and / or training logic 715 may be used in the system of FIG. 9 for inference or prediction operations, at least in part based on weight parameters calculated using the training operations of the neural network, the functionality and / or architecture of the neural network, or the use cases of the neural network described herein.

[0080] Using such components, diverse scene graphs can be generated from one or more rule sets, and the scene graphs can be used to generate training data or image content representing one or more scenes of a virtual environment.

[0081] FIG. 10 is a block diagram showing an electronic device 1000 for utilizing a processor 1010 according to at least one embodiment. In at least one embodiment, the electronic device 1000 may be, for example, without limitation, a notebook, tower server, rack server, blade server, laptop, desktop, tablet, mobile device, phone, embedded computer, or any other suitable electronic device.

[0082] In at least one embodiment, system 1000 may include, without limitation, a processor 1010 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, processor 1010 is coupled using a bus or interface such as a 1°C bus, a system management bus (“SMBus”), a low pin count (“LPC”) bus, a serial peripheral interface (“SPI”), a high definition audio (“HDA”) bus, a serial advance technology attachment (“SATA”) bus, a universal serial bus (“USB”) (versions 1, 2, 3), or a universal asynchronous receiver / transmitter (“UART”) bus. In at least one embodiment, FIG. 10 shows a system including interconnected hardware devices or “chips,” while in other embodiments, FIG. 10 may show an exemplary system on a chip (“SoC”). In at least one embodiment, the devices shown in FIG. 10 may be interconnected using proprietary interconnects, standard interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of FIG. 10 may be interconnected using a Compute Express Link (CXL) interconnect.

[0083] In at least one embodiment, FIG. 10 shows a display 1024, a touch screen 1025, a touch pad 1030, a Near Field Communications unit (NFC) 1045, a sensor hub 1040, a thermal sensor 1046, an Express Chipset (EC) 1035, a Trusted Platform Module (TPM) 1038, a BIOS / firmware / flash memory (BIOS, FW flash) 1022, a DSP 1060, a drive 1020 such as a Solid State Disk (SSD) or a Hard Disk Drive (HDD), a wireless local area network unit (WLAN) 1050, a Bluetooth unit 1052, a Wireless Wide Area Network unit (WWAN) 1056, a Global Positioning System (GPS) 1055, a camera such as a USB3.0 camera (USB3.0 camera) 1054, and / or a Low Power Double Data Rate (LPDDR) memory unit (LPDDR3) 1015 implemented, for example, in accordance with the LPDDR3 standard. These components may each be implemented in any suitable manner.

[0084] In at least one embodiment, other components may be communicatively coupled to the processor 1010 via the components described above. In at least one embodiment, an accelerometer 1041, an ambient light sensor (ALS) 1042, a compass 1043, and a gyroscope 1044 may be communicatively coupled to a sensor hub 1040. In at least one embodiment, a thermal sensor 1039, a fan 1037, a keyboard 1046, and a touch pad 1030 may be communicatively coupled to an EC 1035. In at least one embodiment, a speaker 1063, headphones 1064, and a microphone (mic) 1065 may be communicatively coupled to an audio unit (audio codec and class D amplifier) 1062, and this audio unit may be communicatively coupled to a DSP 1060. In at least one embodiment, the audio unit 1064 may include, for example and without limitation, an audio coder / decoder (codec) and a class D amplifier. In at least one embodiment, a SIM card (SIM) 1057 may be communicatively coupled to a WWAN unit 1056. In at least one embodiment, components such as a WLAN unit 1050 and a Bluetooth unit 1052, as well as the WWAN unit 1056, may be implemented in a next generation form factor (NGFF).

[0085] To perform inference and / or training operations associated with one or more embodiments, an inference and / or training logic 715 is used. Details regarding the inference and / or training logic 715 are provided below in conjunction with FIGS. 7A and / or 7B. In at least one embodiment, the inference and / or training logic 715 may be used in the system of FIG. 10 for inference or prediction operations, based at least in part on weight parameters calculated using the training operations of the neural networks, the functions and / or architectures of the neural networks, or the use cases of the neural networks described herein.

[0086] Using such components, a variety of scene graphs can be generated from one or more rule sets, and the scene graphs can be used to generate training data or image content representing one or more scenes of a virtual environment.

[0087] FIG. 11 is a block diagram of a processing system according to at least one example. In at least one embodiment, system 1100 includes one or more processors 1102 and one or more graphics processors 1108 and may be a single-processor desktop system, a multi-processor workstation system, or a server system having a number of processors 1102 or processor cores 1107. In at least one embodiment, system 1100 is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in a mobile device, portable device, or embedded device.

[0088] In at least one embodiment, system 1100 may include or be incorporated in a server-based gaming platform, a game console including a game and media console, a mobile gaming console, a portable gaming console, or an online gaming console. In at least one embodiment, system 1100 is a mobile phone, smart phone, tablet computing device, or mobile Internet device. In at least one embodiment, processing system 1100 may also include, be coupled to, or be integrated within wearable devices such as smart watch wearable devices, smart eyewear devices, augmented reality devices, or virtual reality devices. In at least one embodiment, processing system 1100 is a television or set-top box device having one or more processors 1102 and a graphical interface generated by one or more graphics processors 1108.

[0089] In at least one embodiment, each of one or more processors 1102 includes one or more processor cores 1107 for processing instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of one or more processor cores 1107 is configured to process a particular instruction set 1109. In at least one embodiment, instruction set 1109 may facilitate computing via a complex instruction set computing (CISC), reduced instruction set computing (RISC), or very long instruction word (VLIW). In at least one embodiment, processor cores 1107 may each process different instruction sets 1109, which may include instructions that facilitate emulation of other instruction sets. In at least one embodiment, processor cores 1107 may also include other processing devices such as a digital signal processor (DSP).

[0090] In at least one embodiment, processor 1102 includes cache memory 1104. In at least one embodiment, processor 1102 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, cache memory is shared among various components of processor 1102. In at least one embodiment, processor 1102 may also use an external cache (e.g., a level 3 (L3) cache or a last level cache (LLC)) (not shown), which may be shared among processor cores 1107 using known cache coherence techniques. In at least one embodiment, a register file 1106 is further included in processor 1102, and this register file may include different types of registers (e.g., integer registers, floating point registers, status registers, and instruction pointer registers) for storing different types of data. In at least one embodiment, register file 1106 may include general-purpose registers or other registers.

[0091] In at least one embodiment, one or more processors 1102 are coupled to one or more interface buses 1110 to transmit communication signals such as address, data, or control signals between the processor 1102 and other components within the system 1100. In at least one embodiment, the interface bus 1110 can be a processor bus such as a version of a Direct Media Interface (DMI) bus in one embodiment. In at least one embodiment, the interface 1110 is not limited to the DMI bus and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 1102 includes an integrated memory controller 1116 and a platform controller hub 1130. In at least one embodiment, the memory controller 1116 facilitates communication between the memory device and other components of the system 1100, while the platform controller hub (PCH) 1130 provides connections to I / O devices via a local I / O bus.

[0092] In at least one embodiment, the memory device 1120 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or any other memory device having suitable performance to serve as a process memory. In at least one embodiment, the memory device 1120 operates as a system memory for the system 1100 and can store data 1122 and instructions 1121 for use when one or more processors 1102 execute an application or process. In at least one embodiment, the memory controller 1116 is also coupled to an optional external graphics processor 1112, which may communicate with one or more graphics processors 1108 within the processor 1102 to perform graphics and media operations. In at least one embodiment, the display device 1111 can be connected to the processor 1102. In at least one embodiment, the display device 1111 can include one or more of an internal display device such as a mobile electronic device or a laptop device, or an external display device attached via a display interface (e.g., a display port, etc.). In at least one embodiment, the display device 1111 can include a head-mounted display (HMD) such as a stereoscopic display device for use in virtual reality (VR) applications or augmented reality (AR) applications.

[0093] In at least one embodiment, the platform controller hub 1130 enables peripheral devices to be connected to the memory device 1120 and the processor 1102 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 1146, a network controller 1134, a firmware interface 1128, a wireless transceiver 1126, a touch sensor 1125, and a data storage device 1124 (e.g., a hard disk drive, a flash memory, etc.). In at least one embodiment, the data storage device 1124 can be connected via a storage interface (e.g., SATA) or a peripheral bus such as a Peripheral Component Interconnect bus (e.g., PCI, PCI Express). In at least one embodiment, the touch sensor 1125 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 1126 can be a WiFi transceiver, a Bluetooth transceiver, or a mobile network transceiver such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 1128 enables communication with system firmware and can be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, the network controller 1134 can enable a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 1110. In at least one embodiment, the audio controller 1146 is a multi-channel high-definition audio controller. In at least one embodiment, the system 1100 includes an optional legacy I / O controller 1140 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system.In at least one embodiment, the platform controller hub 1130 can also be connected to one or more universal serial bus (USB) controller 1142 connected input devices, such as a combination of a keyboard and a mouse 1143, a camera 1144, or other USB input devices.

[0094] In at least one embodiment, instances of the memory controller 1116 and the platform controller hub 1130 may be integrated into a separate external graphics processor, such as the external graphics processor 1112. In at least one embodiment, the platform controller hub 1130 and / or the memory controller 1116 may be external to one or more processors 1102. For example, in at least one embodiment, the system 1100 can include an external memory controller 1116 and a platform controller hub 1130, which may be configured as a memory controller hub and a peripheral device controller hub in a system chipset that communicates with the processor 1102.

[0095] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 715 is used. Details regarding the inference and / or training logic 715 are provided below in conjunction with FIGS. 7A and / or 7B. In at least one embodiment, some or all of the inference and / or training logic 715 may be incorporated into the graphics processor 1500. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs embodied in the graphics processor. Further, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than the logic shown in FIGS. 7A or 7B. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALU of the graphics processor to execute one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0096] Using such components, diverse scene graphs can be generated from one or more rule sets, and the scene graphs can be used to generate training data or image content representing one or more scenes of a virtual environment.

[0097] FIG. 12 is a block diagram of a processor 1200 having at least one embodiment of one or more processor cores 1202A - 1202N, an integrated memory controller 1214, and an integrated graphics processor 1208. In at least one embodiment, the processor 1200 can include a lesser number of additional cores including the additional core 1202N represented by the dashed rectangle. In at least one embodiment, each of the processor cores 1202A - 1202N includes one or more internal cache units 1204A - 1204N. In at least one embodiment, each processor core can also access one or more shared cache units 1206.

[0098] In at least one embodiment, the internal cache units 1204A - 1204N and the shared cache unit 1206 represent the cache memory hierarchy within the processor 1200. In at least one embodiment, the cache memory units 1204A - 1204N may include at least one level of cache for instructions and data within each processor core, as well as one or more levels of shared intermediate - level cache such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, where the highest - level cache before the external memory is classified as the LLC. In at least one embodiment, cache coherence logic maintains coherence among the various cache units 1206 and 1204A - 1204N.

[0099] In at least one embodiment, the processor 1200 may also include a set of one or more bus controller units 1216 and a system agent core 1210. In at least one embodiment, the one or more bus controller units 1216 manage a set of peripheral buses such as one or more PCI or PCI Express buses. In at least one embodiment, the system agent core 1210 provides management functions for the various processor components. In at least one embodiment, the system agent core 1210 includes one or more integrated memory controllers 1214 for managing access to various external memory devices (not shown).

[0100] In at least one embodiment, one or more of the processor cores 1202A - 1202N include support for simultaneous multithreading. In at least one embodiment, the system agent core 1210 includes components for coordinating and operating cores 1202A - 1202N during multithreaded processing. In at least one embodiment, the system agent core 1210 may further include a power control unit (PCU), which includes logic and components for adjusting the power state of one or more of the processor cores 1202A - 1202N and the graphics processor 1208.

[0101] In at least one embodiment, the processor 1200 further includes a graphics processor 1208 for performing graphics processing operations. In at least one embodiment, the graphics processor 1208 is coupled to a shared cache unit 1206 and a system agent core 1210 that includes one or more integrated memory controllers 1214. In at least one embodiment, the system agent core 1210 also includes a display controller 1211 for causing the output of the graphics processor to be provided to one or more attached displays. In at least one embodiment, the display controller 1211 may also be a separate module coupled to the graphics processor 1208 via at least one interconnect, or may be integrated within the graphics processor 1208.

[0102] In at least one embodiment, a ring - based interconnect unit 1212 is used to couple the internal components of the processor 1200. In at least one embodiment, alternative interconnect units such as point - to - point interconnects, switch interconnects, or other techniques may be used. In at least one embodiment, the graphics processor 1208 is coupled to the ring interconnect 1212 via an I / O link 1213.

[0103] In at least one embodiment, the I / O link 1213 represents at least one of a variety of I / O interconnects including an on-package I / O interconnect that facilitates communication between various processor components and a high-performance embedded memory module 1218 such as an eDRAM module. In at least one embodiment, each of the processor cores 1202A-1202N and the graphics processor 1208 uses the embedded memory module 1218 as a shared last-level cache.

[0104] In at least one embodiment, the processor cores 1202A-1202N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, the processor cores 1202A-1202N are heterogeneous from the perspective of an instruction set architecture (ISA), where one or more of the processor cores 1202A-1202N execute a common instruction set, but one or more other cores of the processor cores 1202A-1202N execute a subset of the common instruction set, or a different instruction set. In at least one embodiment, the processor cores 1202A-1202N are heterogeneous from the perspective of a microarchitecture, where one or more cores with a relatively high power consumption are coupled with one or more cores with a lower power consumption. In at least one embodiment, the processor 1200 can be implemented on one or more chips or as a SoC integrated circuit.

[0105] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 715 is used. Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7A and / or 7B. In at least one embodiment, some or all of inference and / or training logic 715 may be incorporated within processor 1200. For example, in at least one embodiment, the training and / or inference techniques described herein may utilize one or more of the ALUs embodied in graphics processor 1512, graphics cores 1202A - 1202N, or other components of FIG. 12. Further, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than that shown in FIGS. 7A or 7B. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALU of graphics processor 1200 to execute one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0106] Using such components, diverse scene graphs can be generated from one or more rule sets, and the scene graphs can be used to generate training data or image content representing one or more scenes of a virtual environment.

[0107] Virtualized computing platform FIG. 13 is an example data flow diagram of a process 1300 for generating and introducing an image processing and inference pipeline according to at least one embodiment. In at least one embodiment, the process 1300 may be introduced for use with imaging devices, processing devices, and / or other types of devices in one or more facilities 1302. The process 1300 may be executed within a training system 1304 and / or within an introduction system 1306. In at least one embodiment, the training system 1304 is used to train, introduce, and implement a machine learning model (e.g., a neural network, an object detection algorithm, a computer vision algorithm, etc.) for use in the introduction system 1306. In at least one embodiment, the introduction system 1306 is configured to offload processing and computing resources across distributed computing environments to reduce infrastructure requirements in the facility 1302. In at least one embodiment, one or more applications within the pipeline may use or call services (e.g., inference, virtualization, computing, AI, etc.) of the introduction system 1306 during execution of the application.

[0108] In at least one embodiment, some of the applications used in the advanced processing and inference pipeline may use a machine learning model or other AI to perform one or more processing steps. In at least one embodiment, the machine learning model may be trained at the facility 1302 using data 1308 (such as imaging data) generated at the facility 1302 and stored in one or more image archive and communication system (PACS) servers at the facility 1302, may be trained using imaging or sequencing data 1308 from one or more other facilities, or may be a combination thereof. In at least one embodiment, the training system 1304 may be used to provide applications, services, and / or other resources for generating a practical and deployable machine learning model for the introduction system 1306.

[0109] In at least one embodiment, the model registry 1324 may be backed up by an object storage that can support version management and object metadata. In at least one embodiment, the object storage may be accessible, for example, from within a cloud platform via a compatibility application programming interface (API) of cloud storage (e.g., cloud 1426 of FIG. 14). In at least one embodiment, the machine learning models within the model registry 1324 may be uploaded, listed, modified, or deleted by a developer or partner of the system interacting with the API. In at least one embodiment, the API may provide access to a way for a user with appropriate credentials to associate a model with an application, thereby enabling the model to be executed as part of running a containerized instance of the application.

[0110] In at least one embodiment, the training pipeline 1404 (FIG. 14) may include a situation where the facility 1302 is training its own machine learning model or has an existing machine learning model that needs to be optimized or updated. In at least one embodiment, imaging data 1308 generated by an imaging device, a sequencing device, and / or other types of devices may be received. In at least one embodiment, when the imaging data 1308 is received, AI-assisted annotation 1310 may be used to assist in generating annotations corresponding to the imaging data 1308 that will be used as ground truth data for the machine learning model. In at least one embodiment, the AI-assisted annotation 1310 may include one or more machine learning models (e.g., a convolutional neural network (CNN)), which may be trained to generate annotations corresponding to a specific type of imaging data 1308 (e.g., from a specific device). In at least one embodiment, the AI-assisted annotation 1310 may then be used directly to generate ground truth data or may be adjusted or fine-tuned using an annotation tool. In at least one embodiment, the AI-assisted annotation 1310, labeled clinical data 1312, or a combination thereof may be used as ground truth data for training the machine learning model. In at least one embodiment, the trained machine learning model may be referred to as the output model 1316 and may be used by the introduction system 1306 described herein.

[0111] In at least one example, the training pipeline 1404 (FIG. 14) may include a situation where the facility 1302 requires a machine learning model to perform one or more processing tasks for one or more applications within the onboarding system 1306, but the facility 1302 may not currently have such a machine learning model (or may not have a model optimized, efficient, or effective for such purposes). In at least one example, an existing machine learning model may be selected from the model registry 1324. In at least one example, the model registry 1324 may include machine learning models trained to perform various different inference tasks on imaging data. In at least one example, the machine learning models of the model registry 1324 may be trained on imaging data from a facility different from the facility 1302 (e.g., a facility in a remote location). In at least one example, the machine learning model may be trained on imaging data from one location, two locations, or any number of locations. In at least one example, when trained on imaging data from a particular location, the training may be performed at that location, or at least in a manner that protects the confidentiality of the imaging data or restricts the transfer of the imaging data outside the facility. In at least one example, when a model is trained or partially trained at one location, the machine learning model may be added to the model registry 1324. In at least one example, the machine learning model may then be retrained or updated at any number of other facilities, and the retrained or updated model may be made available in the model registry 1324. In at least one example, the machine learning model may then be selected from the model registry 1324, may be referred to as the output model 1316, and may be used in the onboarding system 1306 to perform one or more processing tasks for one or more applications of the onboarding system.

[0112] In at least one embodiment, the training pipeline 1404 (FIG. 14) may include a scenario where the facility 1302 needs a machine learning model to execute one or more processing tasks for one or more applications within the onboarding system 1306, but the facility 1302 may not currently have such a machine learning model (or may not have a model optimized, efficient, or effective for such purposes). In at least one embodiment, the machine learning model selected from the model registry 1324 may not be fine-tuned or optimized for the imaging data 1308 generated at the facility 1302 because there may be differences in the population, the robustness of the training data used to train the machine learning model, the diversity of anomalies in the training data, and / or other issues associated with the training data. In at least one embodiment, AI-assisted annotation 1310 may be used to assist in generating annotations corresponding to the imaging data 1308 that will be used as ground truth data for retraining or updating the machine learning model. In at least one embodiment, labeled data 1312 may be used as ground truth data for training the machine learning model. In at least one embodiment, retraining or updating the machine learning model may be referred to as model training 1314. In at least one embodiment, model training 1314, such as AI-assisted annotation 1310, labeled clinic data 1312, or a combination thereof, may be used as ground truth data for retraining or updating the machine learning model. In at least one embodiment, the trained machine learning model may be referred to as the output model 1316 and may be used by the onboarding system 1306 described herein.

[0113] In at least one embodiment, the introduction system 1306 may include software 1318, services 1320, hardware 1322, and / or other components, features, and functions. In at least one embodiment, the introduction system 1306 may include a software “stack,” whereby the software 1318 may be built on top of the services 1320, may use the services 1320 to perform some or all of the processing tasks, and the services 1320 and software 1318 may be built on top of the hardware 1322 and use the hardware 1322 to perform the processing, storage, and / or other computing tasks of the introduction system 1306. In at least one embodiment, the software 1318 may include any number of different containers, where each container may perform an instantiation of an application. In at least one embodiment, each application may perform one or more processing tasks of an advanced processing and inference pipeline (e.g., inference, object detection, feature detection, segmentation, image enhancement, calibration, etc.). In at least one embodiment, the advanced processing and inference pipeline may be defined based on a selection of different containers desired or required to process the imaging data 1308, in addition to the containers that receive and configure the imaging data used by each container and / or used by the facility 1302 after being processed through the pipeline (e.g., to re-convert the output to a usable type of data). In at least one embodiment, a combination of containers within the software 1318 (e.g., that make up the pipeline) may be referred to as a virtual device (described in more detail herein), and the virtual device may use the services 1320 and hardware 1322 to perform some or all of the processing tasks of the applications instantiated in the containers.

[0114] In at least one embodiment, the data processing pipeline may receive input data (e.g., imaging data 1308) in a specific format in response to an inference request (e.g., a request from a user of the introduction system 1306). In at least one embodiment, the input data may represent one or more images, videos, and / or other data representations generated by one or more imaging devices. In at least one embodiment, the data may undergo preprocessing as part of the data processing pipeline and be prepared so that it can be processed by one or more applications. In at least one embodiment, postprocessing may be performed on the output of one or more inference tasks or other processing tasks of the pipeline to prepare the output data for the next application and / or to prepare the output data for transmission and / or use by the user (e.g., in response to an inference request). In at least one embodiment, the inference task may be performed by one or more machine learning models, such as a trained or introduced neural network, which may include the output model 1316 of the training system 1304.

[0115] In at least one embodiment, the tasks of the data processing pipeline may be encapsulated in containers, each container representing an individual fully functional instantiation of an application and a virtualized computing environment capable of referencing a machine learning model. In at least one embodiment, the container or application may be issued to a private (e.g., restricted access) area of a container registry (described in more detail herein), and the trained or introduced model may be stored in the model registry 1324 and associated with one or more applications. In at least one embodiment, an image of an application (e.g., an image of a container) may be available in the container registry and, when selected by a user from the container registry for introduction into the pipeline, the image may be used to generate a container for instantiating the application so that it can be used on the user's system.

[0116] In at least one example, a developer (e.g., a software developer, a clinician, a physician, etc.) may develop, publish, and store an application (e.g., as a container) to perform image processing and / or inference on the provided data. In at least one example, the development, publishing, and / or storage may be performed using a software development kit (SDK) associated with the system (e.g., to ensure that the developed application and / or container complies with or is compatible with the system). In at least one example, the developed application may be tested locally (e.g., at a first facility, for data from the first facility) using an SDK that can support at least a portion of service 1320 as a system (e.g., system 1400 of FIG. 14). In at least one example, a DICOM object can contain anywhere from one to hundreds of images or other types of data and, due to the variations in the data, the developer may be responsible for managing the extraction and preparation of the input data (e.g., setting up the configuration for the application, building pre-processing into the application, etc.). In at least one example, once the application is (e.g., accuracy) verified by system 1400, the application is made available in a container registry for selection and / or implementation by a user and one or more processing tasks may be performed on the data at the user's facility (e.g., a second facility).

[0117] In at least one embodiment, the developer may then share the application or container through a network so that it can be accessed and used by a user of the system (e.g., system 1400 of FIG. 14). In at least one embodiment, the completed and verified application or container may be stored in a container registry, and the associated machine learning model may be stored in a model registry 1324. In at least one embodiment, a requesting entity that issues an inference or image processing request may browse the container registry and / or model registry 1324 to search for applications, containers, datasets, machine learning models, etc., select a desired combination of elements for inclusion in a data processing pipeline, and send an imaging processing request. In at least one embodiment, the request may include input data (and in some instances, associated patient data) required to execute the request, and / or may include a selection of the application and / or machine learning model to be executed when processing the request. In at least one embodiment, the request is then passed to one or more components of an ingress system 1306 (e.g., the cloud) to execute the processing of the data processing pipeline. In at least one embodiment, the processing by the ingress system 1306 may include referring to elements (e.g., applications, containers, models, etc.) selected from the container registry and / or model registry 1324. In at least one embodiment, when a result is generated by the pipeline, the result may be returned to and viewed by the user (e.g., viewed in a viewing application suite running locally, on an in-premises workstation or terminal).

[0118] In at least one embodiment, Service 1320 may be utilized to assist in processing or executing an application or container in a pipeline. In at least one embodiment, Service 1320 may include a computing service, an artificial intelligence (AI) service, a visualization service, and / or other types of services. In at least one embodiment, Service 1320 may provide common functionality to one or more applications of Software 1318, whereby the functionality may be abstracted with respect to services that can be called or utilized by the applications. In at least one embodiment, the functionality provided by Service 1320 may be executed dynamically and more efficiently, and at the same time, may scale well by enabling an application to process data in parallel (e.g., using parallel computing platform 1430 (FIG. 14)). Instead of requiring each application sharing the same functionality provided by Service 1320 to have its own instance of Service 1320, Service 1320 may be shared among various applications. In at least one embodiment, the service may include an inference server or engine that may be used, as a non-limiting example, to perform detection or segmentation tasks. In at least one embodiment, a model training service capable of providing a function for training and / or retraining a machine learning model may be included. In at least one embodiment, a data augmentation service capable of providing extraction, resizing, scaling, and / or other augmentation of GPU-accelerated data (e.g., DICOM, RIS, CIS, REST-compliant, RPC, raw, etc.) may be further included. In at least one embodiment, a visualization service capable of adding image rendering effects such as ray tracing, rasterization, noise removal, sharpening, etc. may be used to add a sense of reality to two-dimensional (2D) and / or three-dimensional (3D) models. In at least one embodiment, a virtual device service capable of realizing beamforming, segmentation, inference, imaging, and / or support for other applications within a pipeline of virtual devices may be included.

[0119] In at least one embodiment, when service 1320 includes an AI service (e.g., an inference service), one or more machine learning models may be executed by calling (as an API call) the inference service (e.g., an inference server) to execute the machine learning model, or its processing, as part of application execution. In at least one embodiment, when another application includes one or more machine learning models for a segmentation task, the application may call the inference service so as to execute a machine learning model for executing one or more of the processing operations associated with the segmentation task. In at least one embodiment, software 1318 implementing an advanced processing and inference pipeline including a segmentation application and an anomaly detection application may be rationalized because each application may call the same inference service to execute one or more inference tasks.

[0120] In at least one embodiment, the hardware 1322 may include a GPU, a CPU, a graphics card, an AI / deep learning system (e.g., an AI supercomputer such as NVIDIA's DGX), a cloud platform, or a combination thereof. In at least one embodiment, different types of hardware 1322 may be used to provide efficient and dedicated support for the software 1318 and services 1320 of the introduction system 1306. In at least one embodiment, the use of GPU processing for local (e.g., at the facility 1302) processing may be implemented within an AI / deep learning system, a cloud system, and / or other processing components of the introduction system 1306 to improve the efficiency, accuracy, and effectiveness of image processing and image generation. In at least one embodiment, the software 1318 and / or services 1320 may be optimized for GPU processing related to deep learning, machine learning, and / or high-performance computing, as non-limiting examples. In at least one embodiment, at least a portion of the computing environment of the introduction system 1306 and / or the training system 1304 may be executed using GPU-optimized software (e.g., a combination of hardware and software of NVIDIA's DGX system) in one or more supercomputers of a data center or a high-performance computing system. In at least one embodiment, the hardware 1322 may include any number of GPUs, and these GPUs may be called to perform parallel processing of data as described herein. In at least one embodiment, the cloud platform may further include GPU processing for GPU-optimized execution of deep learning tasks, machine learning tasks, or other computing tasks. In at least one embodiment, the cloud platform (e.g., NVIDIA's NGC) may be executed using an AI / deep learning supercomputer (e.g., provided by NVIDIA's DGX system) and / or GPU-optimized software as a platform for hardware abstraction and scaling.In at least one embodiment, the cloud platform may integrate an application container clustering system or an orchestration system (e.g., KUBERNETES) for multiple GPUs to enable seamless scaling and load balancing.

[0121] FIG. 14 is a system diagram showing an example system 1400 for generating and introducing an imaging introduction pipeline according to at least one embodiment. In at least one embodiment, system 1400 may be used to implement process 1300 of FIG. 13 and / or other processes including advanced processing and inference pipelines. In at least one embodiment, system 1400 may include a training system 1304 and an introduction system 1306. In at least one embodiment, training system 1304 and introduction system 1306 may be implemented using software 1318, service 1320, and / or hardware 1322 as described herein.

[0122] In at least one embodiment, system 1400 (e.g., training system 1304 and / or introduction system 1306) may be implemented in a cloud computing environment (e.g., cloud 1426). In at least one embodiment, system 1400 may be implemented locally with respect to a healthcare service facility or as a combination of a cloud and local computing resources. In at least one embodiment, access to the API of cloud 1426 may be limited to authorized users via established security measures or protocols. In at least one embodiment, the security protocol may include a web token, which may be signed by an authentication (e.g., AuthN, AuthZ, Gluecon, etc.) service and may have appropriate permissions. In at least one embodiment, the API of the virtual device (described herein) or other instantiations of system 1400 may be limited to a set of public IPs that have been inspected or permitted for the interaction.

[0123] In at least one embodiment, the various components of system 1400 may communicate with each other using any of a variety of different types of networks, including but not limited to local area networks (LANs) and / or wide area networks (WANs) via wired and / or wireless communication protocols. In at least one embodiment, communication between the facility and the components of system 1400 (e.g., to send an inference request, to receive the result of an inference request, etc.) may be communicated via one or more data buses, wireless data protocols (Wi-Fi), wired data protocols (e.g., Ethernet), etc.

[0124] In at least one embodiment, the training system 1304 may execute a training pipeline 1404 similar to that described herein with respect to FIG. 13. In at least one embodiment, when one or more machine learning models are to be used in the introduction pipeline 1410 by the introduction system 1306, the training pipeline 1404 may be used to train or retrain one or more (e.g., pre-trained) models and / or one or more of the pre-trained models 1406 may be implemented (e.g., without the need for retraining or updating). In at least one embodiment, as a result of the training pipeline 1404, an output model 1316 may be generated. In at least one embodiment, the training pipeline 1404 may include any number of processing steps, such as, but not limited to, the conversion or adaptation of imaging data (or other input data). In at least one embodiment, different training pipelines 1404 may be used for different machine learning models used by the introduction system 1306. In at least one embodiment, a training pipeline 1404 similar to the first example described with respect to FIG. 13 may be used for the first machine learning model, a training pipeline 1404 similar to the second example described with respect to FIG. 13 may be used for the second machine learning model, and a training pipeline 1404 similar to the third example described with respect to FIG. 13 may be used for the third machine learning model. In at least one embodiment, any combination of tasks within the training system 1304 may be used depending on what is required for each respective machine learning model. In at least one embodiment, one or more of the machine learning models may already be trained and ready for introduction, such that the machine learning models may not undergo any processing by the training system 1304 and may be implemented by the introduction system 1306.

[0125] In at least one embodiment, the output model 1316 and / or the pre-trained model 1406 may include any type of machine learning model depending on the implementation form or embodiment. In at least one embodiment, without limitation, the machine learning model used by the system 1400 may be a linear regression, a logistic regression, a decision tree, a support vector machine (SVM), a naive Bayes, a k-nearest neighbor (Knn), a k-means clustering, a random forest, a dimensionality reduction algorithm, a gradient boosting algorithm, a neural network (e.g., an auto-encoder, a convolutional, a recurrent, a perceptron, a long / short-term memory (LSTM), a Hopfield, a Boltzmann, a deep belief, a deconvolutional, an adversarial generation, a liquid state machine, etc.), and / or other types of machine learning models.

[0126] In at least one embodiment, the training pipeline 1404 may include AI-assisted annotation, which is described in more detail herein with respect to at least FIG. 15B. In at least one embodiment, the labeled data 1312 (e.g., conventional annotation) may be generated by any number of techniques. In at least one embodiment, the labels or other annotations may be generated within a drawing program (e.g., an annotation program), a computer-aided design (CAD) program, a labeling program, an annotation or label generation program suitable for ground truth, or another type of program, and / or in some instances, may be handwritten. In at least one embodiment, the ground truth data may be generated synthetically (e.g., generated from a computer model or rendering), generated realistically (e.g., designed and generated from real-world data), machine-automated (e.g., using feature analysis and learning to extract features from the data and then generate labels), human-annotated (e.g., a labeler or annotation expert may define the location of the labels), and / or a combination thereof. In at least one embodiment, there may be corresponding ground truth data generated by the training system 1304 for each instance of the imaging data 1308 (or other type of data used by the machine learning model). In at least one embodiment, in addition to or instead of the AI-assisted annotation included in the training pipeline 1404, the AI-assisted annotation may be performed as part of the introduction pipeline 1410. In at least one embodiment, the system 1400 may include a multi-layer platform, which may include a software layer (e.g., software 1318) of a diagnostic application (or other type of application) that can perform one or more medical imaging and diagnostic functions. In at least one embodiment, the system 1400 may be communicatively coupled to the PACS server network of one or more facilities (e.g., via an encrypted link).In at least one embodiment, the system 1400 is configured to access and reference data from a PACS server and may perform operations such as training a machine learning model, introducing a machine learning model, image processing, inference, and / or other operations.

[0127] In at least one embodiment, the software layer may be implemented as a secure, encrypted, and / or authenticated API through which an application or container may be called (e.g., invoked) from an external environment (e.g., facility 1302). In at least one embodiment, the application may then call or execute one or more services 1320 to perform calculation, AI, or visualization tasks associated with each application, and the software 1318 and / or services 1320 may utilize the hardware 1322 to perform processing tasks in an effective and efficient manner.

[0128] In at least one embodiment, the introduction system 1306 may execute an introduction pipeline 1410. In at least one embodiment, the introduction pipeline 1410 may include any number of applications, which may be applied continuously, discontinuously, or otherwise to imaging data (and / or other types of data) generated by imaging devices, sequencing devices, genomics devices, etc., including the AI-assisted annotation described above. In at least one embodiment, as described herein, the introduction pipeline 1410 for an individual device may be referred to as a virtual device for the device (e.g., a virtual ultrasound device, a virtual CT scan device, a virtual sequencing device, etc.). In at least one embodiment, depending on the information required for the data generated by the device, there may be more than one introduction pipeline 1410 for one device. In at least one embodiment, if anomaly detection is required for an MRI machine, a first introduction pipeline 1410 may exist, and if image enhancement is required for the output of the MRI machine, a second introduction pipeline 1410 may exist.

[0129] In at least one embodiment, the image generation application may include processing tasks that include the use of a machine learning model. In at least one embodiment, a user may wish to use their own machine learning model or select a machine learning model from the model registry 1324. In at least one embodiment, the user may implement their own machine learning model or select a machine learning model to include in the application to perform the processing task. In at least one embodiment, the application may be selectable and customizable, and by defining the structure of the application, the introduction and implementation of an application for a specific user is presented as a more seamless user experience. In at least one embodiment, by utilizing other features of the system 1400, such as the service 1320 and the hardware 1322, etc., the onboarding pipeline 1410 can be made even more user-friendly, achieve easier integration, and produce more accurate, efficient, and timely results.

[0130] In at least one embodiment, the onboarding system 1306 may include a user interface 1414 (e.g., a graphical user interface, a web interface, etc.), which may be used to select an application to include in the onboarding pipeline 1410, configure the application, modify or change the application or its parameters or structure, use and interact with the onboarding pipeline 1410 during setup and / or onboarding, and / or interact with the onboarding system 1306 in other ways. Although not shown with respect to the training system 1304, in at least one embodiment, the user interface 1414 (or a different user interface) may be used to select a model to use in the onboarding system 1306, select a model to train or retrain in the training system 1304, and / or interact with the training system 1304 in other ways.

[0131] In at least one embodiment, in addition to the application orchestration system 1428, a pipeline manager 1412 may be used to manage interactions between the applications or containers of the onboarding pipeline 1410 and the services 1320 and / or the hardware 1322. In at least one embodiment, the pipeline manager 1412 may be configured to facilitate interactions from application to application, from an application to the service 1320, and / or from an application or service to the hardware 1322. Although illustrated as being included in the software 1318 in at least one embodiment, this is not intended to be limiting, and in some cases (such as that shown in FIG. 12cc), the pipeline manager 1412 may be included in the service 1320. In at least one embodiment, the application orchestration system 1428 (such as Kubernetes, DOCKER, etc.) may include a container orchestration system that can group applications into containers as logical units for tuning, management, scaling, and onboarding. In at least one embodiment, rather than associating applications (such as rebuilt applications, segmented applications, etc.) from the onboarding pipeline 1410 with individual containers, each application can execute in a self - contained environment (such as at the kernel level) to improve speed and efficiency.

[0132] In at least one embodiment, each application and / or container (or its image) may be developed, modified, and introduced individually (e.g., a first user or developer may develop, modify, and introduce a first application, and a second user or developer may develop, modify, and introduce a second application separately from the first user or developer), whereby it is possible to focus on and pay attention to the tasks of one application and / or container without being interrupted by the tasks of another application or container. In at least one embodiment, communication and cooperation between different containers or applications may be assisted by the pipeline manager 1412 and the application orchestration system 1428. In at least one embodiment, as long as the predicted inputs and / or outputs of each container or application are known to the system (e.g., based on the structure of the application or container), the application orchestration system 1428 and / or the pipeline manager 1412 can facilitate communication between and sharing of resources among the respective applications or containers. In at least one embodiment, since one or more of the applications or containers in the introduction pipeline 1410 can share the same services and resources, the application orchestration system 1428 may orchestrate services or resources, perform load balancing, and determine sharing among different applications or containers. In at least one embodiment, a scheduler may be used to track the resource requirements of applications or containers, the current or planned usage of these resources, and the availability of resources. In at least one embodiment, in this way, the scheduler may allocate resources to different applications and distribute resources among applications considering the system requirements and availability.In some examples, the scheduler (and / or other components of the application orchestration system 1428) may determine resource availability and allocation based on constraints imposed on the system (e.g., user constraints), such as quality of service (QoS), and the urgency of data output required (e.g., to determine whether to perform real-time processing or deferred processing).

[0133] In at least one embodiment, the services 1320 utilized and shared by the applications or containers of the introduction system 1306 may include computing services 1416, AI services 1418, visualization services 1420, and / or other types of services. In at least one embodiment, an application may call (e.g., execute) one or more of the services 1320 to perform processing operations for the application. In at least one embodiment, the computing service 1416 may be utilized by an application to perform supercomputing or other high-performance computing (HPC) tasks. In at least one embodiment, parallel processing may be performed using the computing service 1416 (e.g., using the parallel computing platform 1430) to substantially simultaneously process data through one or more of the applications and / or to substantially simultaneously process one or more tasks of one application. In at least one embodiment, the parallel computing platform 1430 (e.g., NVIDIA's CUDA) may enable general-purpose computing on graphics processing units (GPGPUs) (e.g., GPU 1422). In at least one embodiment, the software layer of the parallel computing platform 1430 may provide a virtual instruction set and access to the parallel computing elements of the GPU to execute compute kernels. In at least one embodiment, the parallel computing platform 1430 may include memory, and in some embodiments, the memory may be shared among multiple containers and / or between different processing tasks within one container. In at least one embodiment, inter-process communication (IPC) calls may be generated for multiple containers and / or multiple processes within a container to use the same data from the shared segment of the memory of the parallel computing platform 1430 (e.g., when multiple different stages of an application or multiple applications process the same information).In at least one embodiment, rather than creating a copy of the data and moving the data to different locations in memory (e.g., read / write operations), the same data at the same location in memory may be used for any number of processing tasks (e.g., at the same time, different times, etc.). In at least one embodiment, when data is used and new data is generated as a result of processing, this information about the new location of the data may be stored in and shared among various applications. In at least one embodiment, the location of the data and the location of the updated or modified data may be part of the definition of how the payload is understood within the container.

[0134] In at least one embodiment, the AI service 1418 may be utilized to execute an inference service for executing a machine learning model associated with an application (e.g., tasked with performing one or more processing tasks of the application). In at least one embodiment, the AI service 1418 may utilize the AI system 1424 to execute a machine learning model (e.g., a neural network such as a CNN) for segmentation, reconstruction, object detection, feature detection, classification, and / or other inference tasks. In at least one embodiment, the application of the ingestion pipeline 1410 may perform inferences on the imaging data using one or more of the output model 1316 from the training system 1304 and / or other models of the application. In at least one embodiment, two or more instances of inferences using an application orchestration system 1428 (e.g., a scheduler) may be available. In at least one embodiment, the first category may include a high-priority / low-latency path that can achieve a higher service level agreement, such as for performing inferences for emergency requests during an emergency or for a radiologist during a diagnosis. In at least one embodiment, the second category may include a standard-priority path that can be used for non-emergency requests or when the analysis may be performed later. In at least one embodiment, the application orchestration system 1428 may allocate resources (e.g., the service 1320 and / or the hardware 1322) based on the priority paths for different inference tasks of the AI service 1418.

[0135] In at least one embodiment, the shared storage may be attached to the AI service 1418 within the system 1400. In at least one embodiment, the shared storage may operate as a cache (or other type of storage device) and may be used to process inference requests from applications. In at least one embodiment, when an inference request is sent, the request may be received by a set of API instances of the ingress system 1306, and one or more instances may be selected (e.g., for best fit, for load balancing, etc.) to process the request. In at least one embodiment, to process the request, the request may be placed in a database, the machine learning model may be identified from the model registry 1324 if it is not yet in the cache, and the verification step may ensure that the appropriate machine learning model is loaded into the cache (e.g., shared storage) and / or a copy of the model may be saved in the cache. In at least one embodiment, if the application has not yet been run or there are not sufficient instances of the application, a scheduler (e.g., the pipeline manager 1412) may be used to start the application referenced in the request. In at least one embodiment, if the inference server for running the model has not yet been started, the inference server may be started. Any number of inference servers may be started per model. In at least one embodiment, in a pull model where the inference servers are clustered, the model may be cached whenever load balancing is advantageous. In at least one embodiment, the inference server may be statically loaded onto the corresponding distributed server.

[0136] In at least one embodiment, the inference may be performed using an inference server that runs within a container. In at least one embodiment, an instance of the inference server may be associated with a model (optionally with multiple versions of the model). In at least one embodiment, when a request to perform an inference on a model is received, if no instance of the inference server exists, a new instance may be loaded. In at least one embodiment, when starting the inference server, the model may be passed to the inference server, such that as long as the inference server is running as a different instance, it may serve different models using the same container.

[0137] In at least one embodiment, during the execution of an application, an inference request may be received for a given application, a container (e.g., hosting an instance of an inference server) may be loaded (if not already loaded), and a start procedure may be called. In at least one embodiment, the container's preprocessing logic may load, decode, and / or execute any additional preprocessing on the input data (e.g., using a CPU and / or GPU). In at least one embodiment, once the data is prepared for inference, the container may execute the inference on the data as needed. In at least one embodiment, this may include a single inference call for one image (e.g., an X-ray of a hand), or may request inferences for hundreds of images (e.g., chest CTs). In at least one embodiment, the application may summarize the results before completion, which may include, without limitation, generating a single confidence score, pixel-level segmentation, voxel-level segmentation, visualization, or text for summarizing the findings. In at least one embodiment, different models or applications may be assigned different priorities. For example, there may be models with real-time (TAT < 1 minute) priority, and models with low priority (e.g., TAT < 10 minutes). In at least one embodiment, the model execution time may be measured from the requesting facility or entity and may include cross-partner network transit time in addition to the execution on the inference service.

[0138] In at least one embodiment, the transfer of requests between the service 1320 and the inference application may be hidden behind the software development kit (SDK), and robust transfer may be provided through a queue. In at least one embodiment, for a combination of individual application / tenant IDs, requests are queued via an API, and the SDK pulls requests from the queue and provides the requests to the application. In at least one embodiment, the name of the queue may be provided in an environment where the SDK picks up requests. In at least one embodiment, asynchronous communication via the queue may be useful because when the communication becomes available, any instance of the application can pick up the work by that communication. The results may be returned via the queue to prevent data loss. In at least one embodiment, the highest-priority work can proceed to the queue with most instances of the application connected to the queue, while the lowest-priority work can proceed to the queue that processes tasks in the order received with one instance connected to the queue, so the queue can also provide a function to segment the work. In at least one embodiment, the application may be executed on a GPU-accelerated instance generated in the cloud 1426, and the inference service may perform inference on the GPU.

[0139] In at least one embodiment, visualization services 1420 may be utilized to generate visualizations for viewing the output of application and / or onboarding pipeline 1410. In at least one embodiment, GPU 1422 may be utilized by visualization services 1420 to generate the visualizations. In at least one embodiment, rendering effects such as ray tracing may be implemented by visualization services 1420 to generate higher quality visualizations. In at least one embodiment, the visualizations may include, without limitation, rendering of 2D images, rendering of 3D volumes, reconstruction of 3D volumes, 2D tomography slices, virtual reality displays, augmented reality displays, and the like. In at least one embodiment, a virtual interactive display or interactive environment (e.g., a virtual environment) for a user of the system to interact with may be generated using a virtualized environment. In at least one embodiment, visualization services 1420 may include an internal visualizer, cinematics, and / or other rendering or image processing capabilities or functions (e.g., ray tracing, rasterization, internal optics, etc.).

[0140] In at least one embodiment, the hardware 1322 may include a GPU 1422, an AI system 1424, a cloud 1426, and / or any other hardware used to execute the training system 1304 and / or the introduction system 1306. In at least one embodiment, the GPU 1422 (e.g., NVIDIA's TESLA and / or QUADRO GPUs) may include any number of GPUs, which may be used to execute processing tasks for the computing service 1416, the AI service 1418, the visualization service 1420, other services, and / or any features or functions of the software 1318. For example, with respect to the AI service 1418, preprocessing may be performed on the imaging data (or other types of data used by the machine learning model) using the GPU 1422, postprocessing may be performed on the output of the machine learning model, and / or inference may be performed (e.g., the machine learning model may be executed). In at least one embodiment, the cloud 1426, the AI system 1424, and / or other components of the system 1400 may use the GPU 1422. In at least one embodiment, the cloud 1426 may include a GPU-optimized platform for deep learning tasks. In at least one embodiment, the AI system 1424 may use a GPU, and at least a portion of the cloud 1426, or the role of deep learning or inference, may be executed using one or more AI systems 1424. Thus, although the hardware 1322 is shown as individual components, this is not intended to be limiting, and any component of the hardware 1322 may be combined with and utilized by any other component of the hardware 1322.

[0141] In at least one embodiment, the AI system 1424 may include a dedicated computing system (e.g., a supercomputer or HPC) configured for inference, deep learning, machine learning, and / or other artificial intelligence tasks. In at least one embodiment, the AI system 1424 (e.g., NVIDIA's DGX) may include GPU-optimized software (e.g., a software stack), which may be executed using a plurality of GPUs 1422 in addition to a CPU, RAM, storage, and / or other components, features, or functions. In at least one embodiment, one or more AI systems 1424 may be implemented in the cloud 1426 (e.g., in a data center) to perform some or all of the AI-based processing tasks of the system 1400.

[0142] In at least one embodiment, cloud 1426 may include a GPU-accelerated infrastructure (e.g., NVIDIA's NGC), which may provide a GPU-optimized platform for executing the processing tasks of system 1400. In at least one embodiment, cloud 1426 may include an AI system 1424 (e.g., as a platform for hardware abstraction and scaling) for executing one or more of the AI-based tasks of system 1400. In at least one embodiment, cloud 1426 may utilize multiple GPUs and be integrated with an application orchestration system 1428 to enable seamless scaling and load balancing between applications and services 1320. In at least one embodiment, cloud 1426 may be tasked with executing at least a portion of the services 1320 of system 1400, including the computing service 1416, AI service 1418, and / or visualization service 1420 described herein. In at least one embodiment, cloud 1426 may perform large and small batch inferences (e.g., execution of NVIDIA's TensorRT), provide an accelerated parallel computing API and platform 1430 (e.g., NVIDIA's CUDA), execute an application orchestration system 1428 (e.g., KUBERNETES), provide a graphics rendering API and platform (e.g., ray tracing for generating high-quality cinematics, 2D graphics, 3D graphics, and / or other rendering techniques), and / or provide other functions for system 1400.

[0143] FIG. 15A shows a data flow diagram of a process 1500 for training, retraining, or updating a machine learning model, according to at least one embodiment. In at least one embodiment, process 1500 may be executed using the system 1400 of FIG. 14 as a non-limiting example. In at least one embodiment, process 1500 may utilize the service 1320 and / or the hardware 1322 of the system 1400 described herein. In at least one embodiment, the refined model 1512 generated by process 1500 may be executed by the introduction system 1306 for one or more containerized applications within the introduction pipeline 1410.

[0144] In at least one embodiment, model training 1314 may include retraining or updating an initial model 1504 (e.g., a pre-trained model) using new training data (e.g., a customer dataset 1506 and / or new input data such as new ground truth data associated with the input data). In at least one embodiment, to retrain or update the initial model 1504, the output or loss layer of the initial model 1504 may be reset, may be deleted, and / or may be replaced with an updated or new output or loss layer. In at least one embodiment, the initial model 1504 may have parameters (e.g., weights and / or biases) remaining from a previous training that were previously fine-tuned, such that training or retraining 1314 does not take as long or require as much processing as training the model from scratch. In at least one embodiment, during model training 1314, by having a reset or replaced output or loss layer of the initial model 1504, the parameters may be updated or readjusted for the new dataset based on a loss calculation associated with the accuracy of the output or loss layer when generating predictions for the new customer dataset 1506 (e.g., the image data 1308 of FIG. 13).

[0145] In at least one embodiment, the pre-trained model 1406 may be stored in a data store or registry (e.g., the model registry 1324 of FIG. 13). In at least one embodiment, the pre-trained model 1406 may be trained at one or more facilities that are at least partially different from the facility that executes process 1500. In at least one embodiment, to protect the privacy and rights of patients, subjects, and customers of different facilities, the pre-trained model 1406 may be trained on-site using customer or patient data generated on-site. In at least one embodiment, the pre-trained model 1406 may be trained using the cloud 1426 and / or other hardware 1322, but privacy-protected confidential patient data cannot be transferred to, used by, or accessed by any component of the cloud 1426 (or other off-site hardware). In at least one embodiment, when the pre-trained model 1406 is trained using patient data from two or more facilities, the pre-trained model 1406 may be trained individually for each facility and then trained on patient or customer data from another facility. In at least one embodiment, if customer or patient data has been released from privacy concerns (e.g., by waiver for experimental use) or if customer or patient data is included in a public dataset, etc., the pre-trained model 1406 may be trained on-site and / or off-site, such as in a data center or other cloud computing infrastructure, using customer or patient data from any number of facilities.

[0146] In at least one embodiment, when selecting an application to use in the import pipeline 1410, the user can also select a machine learning model that will be used with the specific application. In at least one embodiment, the user may not have a model to use, and thus, the user may select a pre-trained model 1406 to use with the application. In at least one embodiment, the pre-trained model 1406 may be optimized to generate accurate results for the customer dataset 1506 of the user's facility (e.g., based on patient diversity, demographics, type of medical imaging device used, etc.). In at least one embodiment, before introducing the pre-trained model 1406 into the import pipeline 1410 for use with an application, the pre-trained model 1406 may be updated, retrained, and / or fine-tuned for use in each facility.

[0147] In at least one embodiment, the user may select a pre-trained model 1406 that will be updated, retrained, and / or fine-tuned, and the pre-trained model 1406 may be referred to as an initial model 1504 for training the system 1304 within the process 1500. In at least one embodiment, using the customer dataset 1506 (e.g., imaging data, genomics data, sequencing data, or other types of data generated by the facility's devices), model training 1314 (which may include, without limitation, transfer learning) may be performed on the initial model 1504 to generate a refined model 1512. In at least one embodiment, ground truth data corresponding to the customer dataset 1506 may be generated by the training system 1304. In at least one embodiment, the ground truth data may be at least partially generated by clinicians, scientists, physicians, and practitioners at the facility (e.g., as the labeled clinic data 1312 of FIG. 13).

[0148] In at least one embodiment, AI-assisted annotation 1310 may be used in some examples to generate ground truth data. In at least one embodiment, AI-assisted annotation 1310 (implemented, for example, using an AI-assisted annotation SDK) may utilize a machine learning model (e.g., a neural network) to generate ground truth data that is suggested or predicted for a customer dataset. In at least one embodiment, user 1510 may use an annotation tool within a user interface (graphical user interface (GUI)) on computing device 1508.

[0149] In at least one embodiment, user 1510 may interact with the GUI via computing device 1508 to edit or fine-tune (automated) annotations. In at least one embodiment, polygon editing features may be used to move polygon vertices to more accurate or fine-tuned locations.

[0150] In at least one embodiment, when customer dataset 1506 obtains associated ground truth data, the ground truth data (from, for example, AI-assisted annotation, manual labeling, etc.) may be used during model training 1314 to generate refined model 1512. In at least one embodiment, customer dataset 1506 may be applied to initial model 1504 any number of times, and the ground truth data may be used to update the parameters of initial model 1504 until an acceptable level of accuracy for refined model 1512 is achieved. In at least one embodiment, when refined model 1512 is generated, refined model 1512 may be introduced into one or more deployment pipelines 1410 in a facility to perform one or more processing tasks on medical imaging data.

[0151] In at least one embodiment, the refinement model 1512 may be uploaded to a pre-trained model 1406 of a model registry 1324 that will be selected by another facility. In at least one embodiment, this process may be completed at any number of facilities, whereby the refinement model 1512 may be further refined any number of times for a new dataset to produce a more general model.

[0152] FIG. 15B is a diagram of an example of a client-server architecture 1532 for enhancing an annotation tool using a pre-trained annotation model according to at least one embodiment. In at least one embodiment, the AI-assisted annotation tool 1536 may be instantiated based on the client-server architecture 1532. In at least one embodiment, the annotation tool 1536 of the imaging application may assist, for example, a radiologist in identifying organs and abnormalities. In at least one embodiment, the imaging application may include a software tool that assists a user 1510 in identifying a few extreme points on a specific target organ in a raw image 1534 (such as, for example, a 3D MRI or CT scan) and receives an automatically annotated result for all 2D slices of a specific organ. In at least one embodiment, the result may be stored in a data store as training data 1538 and may be used as ground truth data for training (such as, without limitation). In at least one embodiment, when a computing device 1508 sends extreme points for AI-assisted annotation 1310, for example, a deep learning model may receive this data as input and may return an inference result of a segmented organ or abnormality. In at least one embodiment, a pre-instantiated annotation tool, such as the AI-assisted annotation tool 1536B in FIG. 15B, may be extended by making an API call (such as, for example, API call 1544) to a server, such as an annotation support server 1540, that can include a set of pre-trained models 1542 stored in an annotation model registry. In at least one embodiment, the annotation model registry may store pre-trained models 1542 (such as, for example, machine learning models such as deep learning models) that are pre-trained to perform AI-assisted annotation for a specific organ or abnormality. These models may be further updated by using the training pipeline 1404.In at least one embodiment, a pre-installed annotation tool may be improved over time as new labeled clinic data 1312 is added.

[0153] Using such components, a variety of scene graphs can be generated from one or more rule sets, and the scene graphs can be used to generate training data or image content representing one or more scenes of a virtual environment.

[0154] Other variations are within the scope of the present disclosure. Accordingly, while the disclosed techniques are capable of various modifications and alternative configurations, certain exemplary embodiments thereof have been shown in the drawings and described in detail above. However, there is no intention to limit the present disclosure to the specific one or more disclosed forms, and on the contrary, it is intended to cover all modifications, alternative configurations, and equivalents that fall within the spirit and scope of the disclosure as defined in the appended claims.

[0155] In the context of describing the disclosed embodiments (in particular, in the context of the following claims), the use of the terms "a", "an", and "the", as well as similar indicators, should be construed to cover both the singular and the plural, unless otherwise stated in this specification or clearly contradicted by the context, and should not be construed as a definition of the terms. The terms "comprising", "having", "including", and "containing" should be construed as open-ended terms (meaning "including but not limited to") unless otherwise stated. The term "connected" is construed as being partially or fully enclosed, attached to, or joined to one another, even if there is something intervening, when it refers to a physical connection without modification. The detailed description of a range of values herein is merely intended to function as a concise way of referring individually to each separate value within the range, unless otherwise stated herein or unless each separate value is incorporated into the specification as if it were individually detailed herein. The use of the term "set" (e.g., "a set of items") or "subset" should be construed as a non-empty set comprising one or more members, unless otherwise stated or negated by the context. Further, unless otherwise stated or negated by the context, the term "subset" of a corresponding set does not necessarily refer to a strict subset of the corresponding set, and the subset and the corresponding set may be equal.

[0156] Conjunctive terms such as "at least one of A, B, and C" or phrases in the form of "at least one of A, B, and C" are understood in the context generally used to indicate that an item, term, etc. is A or B or C, or a non-empty subset of any of the sets of A, B, and C, unless there is a specific description to the contrary or it is not clearly negated by the context. For example, in an illustrative example of a set having three members, the conjunctive phrases "at least one of A, B, and C" and "at least one of A, B, and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such conjunctive terms do not generally imply that a given embodiment requires the presence of at least one of each of A, at least one of B, and at least one of C. Further, unless otherwise specified or not clearly negated by the context, the term "plurality" indicates a plural state (e.g., "a plurality of items" indicates multiple items). A plurality means at least two items, but may be more if explicitly or indicated by the context. Further, unless otherwise specified or not clear from the context to the contrary, the phrase "based on" means "at least partially based on" and does not mean "based only on".

[0157] The operations of the processes described in this specification can be performed in any suitable order, unless otherwise stated in this specification or clearly contradicted by the context. In at least one embodiment, a process such as the process described in this specification (or a variation and / or combination thereof) is executed under the control of one or more computer systems configured with executable instructions and is implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors, by hardware, or by a combination thereof. In at least one embodiment, the code is stored in a computer-readable storage medium in the form of a computer program comprising a plurality of instructions executable by, for example, one or more processors. In at least one embodiment, the computer-readable storage medium excludes transient signals (e.g., propagating transient electrical or electromagnetic transmissions) but includes non-transient computer-readable storage media within a transceiver of a transient signal (e.g., buffers, caches, and queues). In at least one embodiment, the code (e.g., executable code or source code) is stored in a set of one or more non-transient computer-readable storage media, which storage media stores (or has other memory for storing) executable instructions that, when executed by one or more processors of a computer system (i.e., as a result of execution), cause the computer system to perform the operations described in this specification. The set of non-transient computer-readable storage media comprises, in at least one embodiment, a plurality of non-transient computer-readable storage media, and one or more of the individual non-transient storage media of the plurality of non-transient computer-readable storage media do not have all the code, but the plurality of non-transient computer-readable storage media collectively store all the code.In at least one embodiment, the executable instructions are executed such that different instructions are executed by different processors. For example, a non-transitory computer-readable storage medium stores the instructions, a main central processing unit ("CPU") executes some of the instructions, and a graphics processing unit ("GPU") executes other instructions. In at least one embodiment, different components of a computer system have separate processors, and the different processors execute different subsets of instructions.

[0158] Accordingly, in at least one embodiment, a computer system is configured to implement one or more services that perform the operations of the processes described herein, singly or in combination, and such computer system is composed of applicable hardware and / or software that enables the execution of the operations. Further, a computer system implementing at least one embodiment of the present disclosure is a single device, and in another embodiment, is a distributed computer system comprising multiple devices operating in different ways, such that the distributed computer system performs the operations described herein without a single device performing all the operations.

[0159] Any examples provided herein, or use of exemplary language (e.g., "such as") are intended merely to clarify embodiments of the present disclosure and do not limit the scope of the present disclosure unless otherwise claimed. No language in this specification should be construed as indicating any non-claimed element as essential to the practice of the present disclosure.

[0160] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference had been individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.

[0161] In the specification and claims, the terms "coupled" and "connected" may be used along with their derivatives. It should be understood that these terms may not be intended as synonyms for each other. Rather, in certain instances, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. Also, "coupled" may mean that two or more elements are not in direct contact with each other but still co - act or interact with each other.

[0162] Unless otherwise specifically stated, throughout the specification, terms such as "process", "compute", "calculate", or "determine" refer to the act and / or process of a computer or computing system, or similar electronic computing device, that manipulates and / or transforms data represented as physical quantities, such as electronic quantities, within the registers and / or memory of the computing system into other data similarly represented as physical quantities within the memory, registers, or other such information storage devices, transmission devices, or display devices of the computing system.

[0163] Similarly, the term "processor" may refer to any device, or portion of a device, that processes electronic data from registers and / or memory and can transform that electronic data into other electronic data that can be stored in registers and / or memory. By way of non-limiting example, a "processor" may be a CPU or a GPU. A "computing platform" may comprise one or more processors. A "software" process as used herein may include software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Also, each process may refer to a plurality of processes for executing instructions serially or in parallel, continuously or intermittently. The terms "system" and "method" are used interchangeably herein only insofar as a system can embody one or more methods and a method can be considered a system.

[0164] In this specification, reference can be made to obtaining, acquiring, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. Obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished in various ways, such as receiving the data as a parameter of a function call or a call to an application programming interface. In some implementations, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be realized by transferring data via a serial or parallel interface. In another implementation, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be realized by transferring data via a computer network from the entity providing the data to the entity acquiring it. Reference can also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data can be realized by transferring the data as a parameter of an input or output of a function call, an application programming interface, or an inter-process communication mechanism.

[0165] The above discussion describes exemplary implementations of the techniques described, but other architectures may be used to implement the described functions, and this other architecture is intended to be within the scope of this disclosure. Further, for purposes of discussion, specific assignments of roles are defined above, but the various functions and roles may be assigned and divided in different ways depending on the situation.

[0166] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as exemplary forms of implementing the claims.

Claims

1. Generating a scene structure using at least a subset of a plurality of rules from a rule set, wherein the plurality of rules specify relationships between types of objects in a scene; Determining one or more parameter values for one or more of the objects represented by the scene structure; Generating a scene graph including the parameter values based on the scene structure; Providing the scene graph to a rendering engine to render an image of the scene; Causing the rendering engine to include at least one of an object label or scene structure information with the image of the scene; comprising; A computer-implemented method, wherein the image can be used as training data for training a neural network.

2. The computer-implemented method according to claim 1, further comprising providing a library of object-related content for use in rendering the objects represented by the scene structure using the one or more parameter values for the objects.

3. Receiving an input indicating that one of the rules or objects is not a target to be used in the scene; Updating the scene graph to reflect the input; Providing the updated scene graph to the rendering engine to render an updated image of the scene. The computer-implemented method according to claim 1, further comprising.

4. Generating a variety of sets of scene graphs using the rule set; Generating a virtual environment using the variety of sets of scene graphs. The computer-implemented method according to claim 1, further comprising.

5. The computer-implemented method according to claim 4, wherein the virtual environment is a gaming environment, and the rendering engine uses the variety of sets of scene graphs to render one or more images of the gaming environment for one or more players of a game corresponding to the gaming environment.

6. The computer-implemented method of claim 4, further comprising utilizing the virtual environment for one or more test simulations for one or more autonomous or semi-autonomous machines.

7. The computer-implemented method of claim 1, wherein the scene structure is generated from the rule set in an unsupervised manner without data annotation on one or more input images.

8. Assigning probabilities to the rules of the rule set; Selecting the subset of the plurality of rules by sampling the rules according to the probabilities; The computer-implemented method of claim 1, further comprising.

9. The computer-implemented method of claim 8, further comprising determining a scene mask for indicating which of one or more rules of the rule set are selectable through the sampling.

10. Generating the scene structure using an iterative process in which, for one or more of a plurality of time steps, a determination is made as to whether to perform an evolution on one or more object types within the scene structure, wherein the object types are arranged at various levels of a hierarchical scene structure. The computer-implemented method of claim 1, further comprising the generating step.

11. The computer-implemented method of claim 1, wherein a generation model is used to generate the scene structure using the rule set.

12. The computer-implemented method of claim 11, further comprising generating a latent vector as an input to the generation model, wherein the latent vector has a length equal to the number of rules in the rule set.

13. A system, At least one processor; When executed by the at least one processor, cause the system to, Generate a scene structure using at least a subset of a plurality of rules from a rule set, wherein the plurality of rules specify relationships between types of objects in a scene; Determine one or more parameter values for one or more of the objects represented by the scene structure; Generate a scene graph including the parameter values based on the scene structure; Providing the scene graph to a rendering engine to render an image of the scene, and causing the rendering engine to include at least one of an object label or scene structure information with the image of the scene a memory including instructions to cause including a system, wherein the image can be used as training data for training a neural network.

14. When the instructions are executed, causing the system to receive an input indicating that at least one of the rule or object is not the target used in the scene, update the scene graph to reflect the input, and provide the updated scene graph to the rendering engine to render an updated image of the scene The system according to claim 13, further causing.

15. When the instructions are executed, causing the system to generate a variety of sets of scene graphs using the rule set, and generate a virtual environment using the variety of sets of scene graphs The system according to claim 13, further causing.

16. The system is a system for performing graphical rendering operations, a system for performing simulation operations, a system for performing simulation operations to test or verify an autonomous machine application, a system for performing deep learning operations, a system implemented using an edge device, a system incorporating one or more virtual machines (VMs), a system implemented at least partially in a data center, a system implemented at least partially using cloud computing resources The system according to claim 13, including at least one of.

17. A method for synthesizing training data, comprising: obtaining a rule set including a plurality of rules specifying relationships between types of objects in a scene; generating a plurality of scene structures based on the plurality of rules; determining one or more parameter values for each of the objects represented by the plurality of scene structures; generating a plurality of scene graphs including the parameter values based on the plurality of scene structures; A step of providing the plurality of scene graphs to a rendering engine to render a set of images including data about the object represented by the images; A step of causing the rendering engine to include at least one of an object label or scene structure information together with the set of images; comprising; A method, wherein the set of images represents training data used when training one or more neural networks.

18. The method according to claim 17, further comprising a step of providing a library of object-related content used to render the object represented by the scene structure using the one or more parameter values about the object to the rendering engine.

19. A step of receiving an input indicating that one of the rules or the object is not the target to be utilized; A step of updating the plurality of scene graphs to reflect the input; A step of providing the updated scene graph to the rendering engine to render the updated set of images; The method according to claim 17, further comprising.

20. The method according to claim 17, wherein the scene structure is generated from the set of rules in an unsupervised manner without data annotation on one or more input images.

21. A step of assigning probabilities to the rules of the set of rules; A step of selecting a subset of the plurality of rules for each scene structure by sampling the rules according to the probabilities; The method according to claim 17, further comprising.

Citation Information

Patent Citations

  • Image editing device

    JP1993189531A

  • Image processing method and image processor

    JP1998255081A

  • Navigation system

    JP2004213663A

  • A scene graph for defining three-dimensional graphic objects.

    JP2014516432A

  • Hierarchical system and method for on-demand loading of data in a navigation system

    US20080091732A1