Neural rendering for inverse graphics generation

By combining the StyleGAN generator and the differentiable graph renderer, an inverse graph network was trained, solving the problem of converting 2D images to 3D models. This enabled efficient and accurate 3D reconstruction and viewpoint control, reducing the reliance on manual annotation.

CN115151915BActive Publication Date: 2026-01-20NVIDIA CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202180012865.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-03-05
Filing Date
2021-03-06
Publication Date
2026-01-20
Estimated Expiration
2041-03-06

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently generate accurate 3D models from 2D images, particularly due to a lack of sufficient labeled training data and camera pose information, resulting in inadequate model training and an inability to generate realistic 3D reconstructions.

Method used

We employ a combination of StyleGAN generator and differentiable graph renderer. By generating a multi-view dataset, we train the inverse graph network, use the differentiable renderer to decipher the underlying code of the generator, generate accurate 3D information, and optimize the network through a cycle consistency loss function to reduce the reliance on manual annotation.

Benefits of technology

It significantly improves 3D reconstruction performance, reduces the need for manual annotation, generates realistic 3D models, and allows control over the object's viewpoint and background, reducing training time and resource costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115151915B_ABST
    Figure CN115151915B_ABST
Patent Text Reader

Abstract

The present disclosure proposes a method of training an inverse graphics network. An image synthesis network can generate training data for the inverse graphics network. In turn, the inverse graphics network can teach the synthesis network about physical three-dimensional (3D) controls. This method can provide accurate 3D reconstruction of objects out of 2D images using a trained inverse graphics network while requiring little to no annotation of the provided training data. This method can extract and untangle 3D knowledge learned from generative models by leveraging differentiable renderers, making the untangled generative models act as controllable 3D "neural renderers" that supplement traditional graphics renderers.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross Reference to Related Applications

[0002] This application claims priority to U.S. Non-Provisional Patent Application Serial No. 17 / 193,405, titled “Neural Rendering for Inverse Graphics Generation,” filed March 5, 2021, which claims benefit of Provisional Patent Application Serial No. 62 / 986,618, titled “3D Neural Rendering and Inverse Graphics with StyleGAN Renderer,” filed March 6, 2020, all of which are hereby incorporated by reference in their entirety and for all purposes.

[0003] This application is also related to co-pending U.S. Patent Application Serial No. 17 / 019,120, titled “Labeling Images Using a Neural Network,” filed September 11, 2020, and co-pending U.S. Patent Application Serial No. 17 / 020,649, titled “Generating Labels for Synthetic Images Using One or More Neural Networks,” filed September 14, 2020, each of which is hereby incorporated by reference in its entirety and for all purposes. BACKGROUND

[0004] A wide variety of industries rely on three-dimensional (3D) modeling for a variety of purposes, including those that need to generate 3D environment representations. To provide realistic complex environments, there needs to be a variety of different types of objects, or similar objects with different appearances, to avoid unrealistic repetition or omission. Unfortunately, obtaining a large number of three-dimensional models can be a complex, expensive, and time (and resource) intensive process. It can be desirable to generate 3D environments from a large amount of available two-dimensional (2D) data, but many existing approaches do not provide adequate 3D model generation based on 2D data. For promising approaches that can involve machine learning, for example, there still needs to be a sufficient number and variety of labeled training data to train the machine learning. An insufficient number and variety of annotated training data instances can hinder the model from being adequately trained to produce acceptably accurate results. BRIEF DESCRIPTION OF DRAWINGS

[0005] Various embodiments according to the present disclosure will be described with reference to the drawings, in which:

[0006] Figure 1A 、 1B and 1C illustrate image data that can be utilized according to at least one embodiment;

[0007] Figure 2 illustrates a neural rendering and inverse graphics pipeline according to at least one embodiment;

[0008] Figure 3 illustrates an image generation system according to at least one embodiment;

[0009] Figure 4 illustrates a representative point for determining view or pose information for an object according to at least one embodiment;

[0010] Figure 5 illustrates a process for training an inverse graphics network according to at least one embodiment;

[0011] Figure 6 illustrates components of a system for training and utilizing an inverse graphics network according to at least one embodiment;

[0012] Figure 7A illustrates inference and / or training logic according to at least one embodiment;

[0013] Figure 7B illustrates inference and / or training logic according to at least one embodiment;

[0014] Figure 8 illustrates an example data center system according to at least one embodiment;

[0015] Figure 9 illustrates a computer system according to at least one embodiment;

[0016] Figure 10 illustrates a computer system according to at least one embodiment;

[0017] Figure 11 illustrates at least a portion of a graphics processor according to one or more embodiments;

[0018] Figure 12 illustrates at least a portion of a graphics processor according to one or more embodiments;

[0019] Figure 13 is an example dataflow graph of a high-level compute pipeline according to at least one embodiment;

[0020] Figure 14is a system diagram of an example system for training, adapting, instantiating, and deploying machine learning models in an advanced computing pipeline, in accordance with at least one embodiment; and

[0021] Figure 15A and Figure 15B A dataflow graph of a process for training a machine learning model is shown, in accordance with at least one embodiment, as well as a client-server architecture that leverages a pre-trained annotation model to augment an annotation tool. DETAILED DESCRIPTION

[0022] Methods in accordance with various embodiments can provide for the generation of three-dimensional (3D) models, or the inference of one or more 3D properties (e.g., shape, texture, or lighting) from two-dimensional (2D) inputs (e.g., images or video frames). In particular, various embodiments provide methods that train inverse graphics models to generate accurate 3D models or representations from 2D images. In at least some embodiments, the generated models can be used to generate multiple views of an object from different perspectives or in different poses, while other image features or aspects remain fixed. The generated images can include pose, view, or camera information used to generate each such image. These images can be used as training images for an inverse graphics network. The inverse graphics network can utilize such a set of input images of a single object to generate and refine a model of that object. In at least some embodiments, these networks can be trained together, where the 3D model output by the inverse graphics network can be fed as input training data to the generator network for improving the accuracy of the generator network. A combined loss function can be used with terms for both the inverse graphics network and the generator network in order to optimize both networks together using a common training dataset. Relative to previous approaches, such methods can provide improved 3D reconstruction performance, can provide for class generalization (i.e., the model need not be trained for a specific type of object classification), and can significantly reduce the need for manual annotation (e.g., from hours to a minute or less).

[0023] In at least one embodiment, an inverse graphics network can be trained to generate a 3D model or representation from a 2D image of an object (e.g., an input image 100 showing a front view of a vehicle Figure 1A As can be appreciated, the input image presents at least three challenges for reconstructing a 3D model or representation of the vehicle. The first challenge is that there is only image data of the front of the vehicle in this example, so image data for other portions of the vehicle (e.g., the back, underneath, or sides) need to be completely inferred and generated. For example, Figure 1B Other view images 140 are shown in FIG. 1 Figure 1BDifferent views of the vehicle's side are shown, representing portions of the vehicle's side not included in the original input image. Another challenge is the lack of depth information in this 2D image. Any depth or shape information must be inferred solely from this single two-dimensional view. Yet another challenge is the absence of camera or pose information or annotations provided with the input image 100 to serve as a reference when instructed to generate other views or portions of a 3D model.

[0024] To train a model, network, or algorithm (e.g., an inverse graphical model) to generate or infer such 3D information, including semantic information, methods according to various embodiments can provide view images 140 (e.g., those in...). Figure 1B (As shown) are used as training images for the network. However, as mentioned earlier, it is difficult to obtain a sufficient number of views for a given object, and even if a sufficient number are obtained, these images must be annotated with enough information (such as pose, camera, or view data) so that those images can be used as ground-based training data.

[0025] The methods according to various embodiments can utilize the ability to acquire input images (e.g., Figure 1A A generator (image 100) generates a set of images of the object in different views. In at least one embodiment, this can be a generative neural network, such as a generative adversarial network (GAN). A GAN can be provided with an input image and pose, view, or camera data, and can generate an image representation of the object in the input image that has or corresponds to the provided pose, view, or camera data. To generate a set of training images for this object, a set of pose, view, or camera information (manually, randomly, or according to a specified pattern or rule) can be provided, and an image can be generated for each unique pose, view, or camera input. Another advantage of this process is that the GAN can provide annotations without any additional processing, since the pose, view, or camera information is known for generating images. An example GAN that can be used for such purposes is the Style Generative Adversarial Network, also known simply as StyleGAN, developed by NVIDIA. According to one or more embodiments, the StyleGAN model extends the general GAN ​​architecture to include a mapping network to map points in the latent space to intermediate latent spaces, which can control the "style" (e.g., pose, view, or camera information) at each point in the generator model, and can also introduce noise as a source of variation at each point. In one or more embodiments, the StyleGAN model can be used to generate a large number of accurate multi-view datasets in a short time with relatively low processing resources. The StyleGAN model can take an input image 140 and generate multiple different view images 140 for training the inverse graph network. The inverse graph network can then be trained to take 2D input images and generate or infer relatively accurate 3D information 180, such as... Figure 1CAs shown, the input can involve, for example, the shape of the object (which can be represented by a 3D mesh, for example), the lighting of the object (e.g., the direction, intensity, and color of one or more light sources), and the texture of the object (e.g., a complete image dataset representing all relevant parts of the object). As is known in applications such as computer graphics, a view of the object can be generated by projecting the texture onto the mesh, lighting the mesh using appropriate lighting information, and then rendering an image of the object from a determined perspective of a virtual camera. Other groups or types of 3D information can also be used for different use cases, applications, or embodiments. In turn, the 3D information generated by the inverse graphics network can then be used as training data for the StyleGAN model. In various embodiments, these networks can be trained separately using separate loss functions or together using a common loss function.

[0026] Figure 2 An example training pipeline 200 that can be used in accordance with various embodiments is shown. This example pipeline 200 includes two different renderers. The first renderer is a generator network, such as a GAN (e.g., StyleGAN), and the second renderer in this example is a differentiable graphics renderer, such as a differentiable interpolation-based renderer (DIB-R). DIB-R is a differentiable rendering framework that allows the computation of gradients for all pixels in an image. This framework can treat foreground rasterization as a weighted interpolation of local attributes and background rasterization as a distance-based global geometric aggregation, allowing for exact optimization of vertex positions, colors, normals, light directions, and texture coordinates through various light models. In this example, the generator 204 is used as a synthetic data generator with effective annotations of a multi-view dataset 206. This dataset can then be used to train an inverse graphics network 212 that predicts 3D attributes from 2D images. This network can be used to disentangle the latent code of the generator through a carefully designed mapping network.

[0027] In this example, the input images of the multi-view dataset can be provided as input to the inverse graphics network 208, which can utilize an inference network to infer 3D information, such as the shape, lighting, and texture of the object in the image. This 3D information can be fed as input to the differentiable graphics renderer along with the camera information from the input image for training purposes. This renderer can utilize this 3D information to generate shape information, 3D models, or one or more images for the given camera information. These renderings can then be compared to the relevant ground truth data using an appropriate loss function to determine a loss value. This loss can then be used to adjust one or more network weights or parameters. As previously mentioned, the output of the inverse graphics network 208 can also be used to further train or adjust the generator 204 or StyleGAN to generate more accurate images.

[0028] Figure 3 A portion 300 of such a conduit is shown, which can be used according to at least one embodiment, as illustrated with respect to Figure 1. Figure 2 As shown, 3D information (or "code") about objects in an image can be inferred, such as mesh 304, texture 306, and background information 308. Camera information extracted from annotations of one or more input images can also be utilized. In this example, a machine learning processor (MLP) is used to process mesh 304 to extract one or more dimensions or latent features. Texture 306 and background data 308 can be processed with corresponding encoders 314, 316 (e.g., one or more convolutional neural networks (CNNs)) to further extract relevant dimensions or features. These features or dimensions can then be passed to a disentangling module 318 that includes a mapping network. The mapping network will attempt to map features from various codes into a single latent code 320, which is provided to a generator 322 (or renderer) to generate an output image 322. In at least one embodiment, a portion of the latent code will correspond to camera information, and the remainder will correspond to the mesh, texture, and background. Before generating the latent code 320, it may be possible to merge the features into a single set of features instead of three sets of features for the mesh, texture, and background. These features, along with camera features, can be selected and merged using a selection matrix S into a latent code of 320. This information can then be fed into a generator (e.g., StyleGAN) to render the image. The latent code can also take other forms, such as a latent space or feature vectors with defined dimensions, for example, approximately 500 features. Each dimension's entry may contain information about the image, but may not know which feature or dimension controls or contains which information or aspects of the image. To generate different image views of the same object, the features corresponding to the camera (e.g., the first 100 known features) can be changed while keeping other features unchanged. This approach ensures that only the view or pose changes, while other aspects of the object or image remain constant between rendered images.

[0029] In at least some embodiments, at least some type of camera or object view or pose information may be required, as provided as a note. Figure 4One example of an annotation point 402 provided in proximity to a class of objects in a collection of image views 400 is shown. Weak camera information containing only a subset of potential features can be used, rather than labeling all key points of an object (which can require a lot of effort). In this example, the camera pose can be divided into many different ranges, such as twelve ranges. Given these feature points and pose ranges, this is enough to get a rough estimate of the object location for annotating images to use as training data. The process can then be initialized by this less accurate camera pose information. This process can also handle multiple types of objects, such as people, animals, vehicles, etc.

[0030] Methods in accordance with various embodiments can thus leverage differentiable rendering to help train one or more neural networks to perform inverse graphics-related tasks, as can include (but are not limited to) such as from monocular (e.g., 2D) photographs. To train high-performance models, many traditional approaches rely on multi-view images that are not readily available in practice. In contrast, recent generative image synthesis GANs appear to implicitly acquire 3D knowledge during training: object viewpoints can be manipulated by simple manipulation of latent codes. However, these latent codes often lack further physical interpretation, so GANs cannot be easily inverted to perform explicit 3D reasoning. The 3D knowledge learned by generative models can be at least partially extracted and untangled by using differentiable renderers. In at least one embodiment, one or more generative adversarial networks (GANs) can be utilized as multi-view data generators to train inverse graphics networks. This can be performed using off-the-shelf differentiable renderers and trained inverse graphics networks as teachers to untangle the latent codes of GANs into interpretable 3D attributes. In various approaches, the entire architecture can be trained iteratively by using cycle-consistency loss. This approach can provide significantly improved performance over traditional inverse graphics networks trained on existing datasets, both quantitatively and through user studies. This untangled GAN can also be used as a controllable 3D “neural renderer” that can be used to complement traditional graphics renderers.

[0031] The ability to infer 3D properties (e.g., geometry, texture, materials, and lighting) from photos is key in many domains (e.g., AR / VR, robotics, architecture, and computer vision). Interest in this problem is exploding, evidenced by the large body of published work and several released 3D datasets over the past few years. The process of going from images to 3D is often referred to as “inverse graphics” because the problem is the inverse of the rendering process in graphics, where a 3D scene is projected to a 2D image by considering the geometry and material properties of the objects, and the light sources present in the scene. Most work in inverse graphics assumes that 3D labels are available during training, and trains a neural network to predict these labels. To ensure high quality 3D ground truth, synthetic datasets such as ShapeNet are often used. However, models trained on synthetic datasets often struggle on real photos due to the domain gap with synthetic images.

[0032] To circumvent at least some of these issues, another alternative method of training inverse graphics networks can be used to circumvent the need for 3D ground truth during training. A graphics renderer can be made differentiable, which allows direct inference of 3D properties from images using gradient-based optimization. At least some of these methods can predict the geometry, texture, and lighting of an image with a neural network by minimizing the difference between the input image and an image rendered from these properties. While impressive results have been achieved in certain methods, many of these methods still require some form of implicit 3D supervision, such as multi-view images of the same object with known cameras. On the other hand, generative models of images seem to implicitly learn 3D information, where manipulation of the latent code can generate images of the same scene from different perspectives. However, the learned latent space often lacks physical interpretation and is often un Factored, where properties such as the 3D shape and color of an object cannot be manipulated independently.

[0033] Methods in accordance with at least one embodiment can extract and factor the 3D knowledge learned by generative models by leveraging a differentiable graphics renderer. In at least one embodiment, a generator such as a GAN can be used as a generator of multi-view images to train an inverse graphics neural network using a differentiable renderer. In turn, the inverse graphics network can be used to inform the generator image formation process with knowledge in graphics, effectively factoring the latent space of the GAN. In at least one embodiment, a GAN (e.g., StyleGAN) can be connected with an inverse graphics network to form a single architecture that can be trained iteratively using cycle-consistency loss. This approach can yield a trained network that can significantly outperform inverse graphics networks on existing datasets, and can provide controllable 3D generation and image processing using the factored generative model.

[0034] A pipeline that can be used for this approach is outlined in the preceding with respect toFigure 2 A method is described. This method can combine two types of renderers: a GAN-based neural "renderer" and a differentiable graphics renderer. In at least one embodiment, this method can exploit the fact that GANs can learn to generate highly realistic object images and allow reliable control over the virtual cameras used to generate views of these objects. A set of camera views can be selected manually or otherwise, e.g., using coarse angle annotations. A GAN, e.g., StyleGAN, can then be used to generate a large number of examples view-by-view. Such a dataset can be used to train an inverse graphics network with a differentiable renderer, e.g., DIB-R. The trained inverse graphics network can be used to disentangle the latent code of the GAN and turn the GAN into a 3D neural renderer, where explicit 3D properties can be controlled.

[0035] In at least one embodiment, a multi-view image can be generated with a generator such as a StyleGAN model. An example StyleGAN model is a 16-layer neural network that maps a latent code z e Z drawn from a normal distribution to a real image. The code z is first mapped to an intermediate latent code w e W, which is converted to W * is called the transformed latent space to separate it from the intermediate latent space W. The transformed latent code w * is then injected as style information into the StyleGAN synthesis network.

[0036] Different layers can control different image properties in the generator. The styles in early layers adjust the camera view, while the styles of the middle and higher layers affect shape, texture, and background. It is empirically determined that the latent code in the first four layers controls the camera view in at least one StyleGAN model. That is, if a process samples a new code but keeps the remaining dimensions of w * fixed (called the content code), then it can generate images of the same object depicted in different views. It can be further observed that the sampled code actually represents a fixed camera view. That is, if is kept fixed but the remaining dimensions of w * are sampled, then the generator can generate images of different objects in the same camera view. The objects in each view will be aligned because this generator is used as a multi-view data generator.

[0037] In an example method, a number of views can be manually selected to cover all common viewing angles of the object, ranging from 0-360 degrees in azimuth and approximately 0-30 degrees in elevation. This approach can take care to select viewing angles where the object appears most consistent. Since inverse graphics techniques typically utilize camera pose information, the selected view codes can be annotated with coarse absolute camera poses. Specifically, each view code can be categorized as one of, for example, twelve azimuthal increments uniformly sampled along 360 degrees. A fixed elevation (e.g., 0°) and camera distance can be assigned to each code. These camera poses can provide very coarse annotations of the actual poses, since they act as an initialization of the camera that will be optimized during training. This approach provides annotations for all views and the entire dataset in a relatively short period of time (e.g., one minute or less). Such results can make the annotation effort practically negligible. For each view, a large number of content codes can be sampled to synthesize different objects in these views. Since differentiable renderers such as DIB-R can also utilize segmentation masks during training, networks such as MaskRCNN can be further applied on the generated dataset to obtain instance segmentation. Since the generator can sometimes generate unrealistic images or images with multiple objects, images with multiple instances or small masks (less than 10% of the entire image area) can be filtered out in at least one embodiment.

[0038] A method according to at least one embodiment can aim to train a 3D prediction network f parameterized by θ to infer 3D shapes (as can be represented as meshes) as well as textures from images. Let I v denote an image in a view V from the generator dataset, and M denote its corresponding object mask. The inverse graphics network makes the following prediction: {S, T} f θ = (I V ), where S denotes the predicted shape, and T denotes the texture map. The shape S is deformed from a sphere. While DIB-R also supports ray prediction, its performance can not be sufficient to provide realistic images, so ray estimation is omitted in this discussion.

[0039] To train the network, a renderer such as DIB-R can be employed as a differentiable graphics renderer that takes {S, T} and V as input and generates a rendered image I V = r(S, T, V) and a rendered mask M'. Following DIB-R, the loss function takes the following form:

[0040] L(I, S, T, V; θ) = λ col L col (I, I') + λ percept L percept (I, I') + L IOU(M, M') + λ sm L sm (S) + λ lap L lap (S) + λ mov L mov (S)

[0041] Here, L col is the standard L1 image reconstruction loss defined in the RGB color space, while L percept is a perceptual loss that helps predict more realistic looking textures. Note that the rendered image has no background, so L col and L percept are computed by using the mask. L IOU computes the intersection over union between the ground truth mask and the rendered mask. Regularization losses such as Laplacian loss L lap and flattening loss L sm are typically used to ensure that the shape behaves well. Finally, L mov adjusts the shape deformation to be small and uniform.

[0042] Since multiple view images of each object are also accessible, a multi-view consistency loss can be included. Specifically, the loss for each object k can be given by:

[0043]

[0044] where While more views provide more constraints, empirically two views have been shown to be sufficient. The view pair (i, j) can be sampled randomly for efficiency. The above loss function can be used to jointly train the neural network f and optimize the viewing camera V. It can be assumed that different images generated from the same object correspond to the same view V. Jointly weighting network optimization of the camera can let this approach effectively handle noisy initial camera annotations.

[0045] Inverse graphics models allow 3D meshes and textures to be inferred from a given image. These 3D properties can then be used to disentangle the generator’s latent space and turn the generator into a fully controllable 3D neural renderer, which can be referred to as StyleGAN-R, for example. It can be noted that StyleGAN actually synthesizes not only objects, it also generates backgrounds to form complete scenes. The approach in at least some embodiments can also provide control over the background, enabling the neural renderer to render 3D objects into a desired scene. To obtain the background from a given image, the object can be masked out in at least one embodiment.

[0046] A mapping network can be trained and used to map the view, shape (e.g. mesh), texture, and background into the latent code of the generator. Since the generator can not be fully disentangled, the entire generative model can be fine-tuned while keeping the inverse graphics network fixed. The mapping network, e.g. the example shown in Figure 3 , can map the view to the first four layers, and the shape, texture, background to the last twelve layers of W * . For simplicity, the first four layers can be denoted as Wv*and the last twelve layers as WS*TB, where Wv*∈ R2048and WS*TB∈ R3008. Note that there can be different number of layers in other models or networks. In this example, the mapping network for view V and shape S are separate MLPs, while the texture T and background B are CNN layers:

[0047] z view = g v (v; θ v ), z shape = g s (s; θ s ), z txt = g t (t; θ t ), z bck = g b (b; θ b ),

[0048] where z view ∈ R, z shape , z txt , z bck ∈ R 3008 and θ v , θ s , θ t , θ b are network parameters. The shape, texture, and background codes can be soft combined into the final latent code as follows:

[0049]

[0050] where denotes element-wise product, and s m , s t , s b ∈ R 3008Shared across all samples. To enable disentanglement, each dimension of the final code can only be explained by one attribute (e.g., shape, texture, or background). Thus, the process according to at least one embodiment can use SoftMax to normalize each dimension of s. In practice, it is determined that mapping V to a high-dimensional code can be challenging because the dataset can only contain a limited number of views, and V can be limited to azimuth, elevation, and scale. One approach is to map V to a subset of , where the method empirically selects a number, e.g., 144 out of 2048, or the dimension with the highest correlation to the annotated view. Thus, in this example z view ∈ R 144 ∈ R144.

[0051] In at least one embodiment, the mapping network can be trained and the StyleGAN model fine-tuned in two separate stages. In one example, the weights of the StyleGAN model are frozen and only the mapping network is trained. In one or more embodiments, this helps or improves the ability of the mapping network to output reasonable latent codes for the StyleGAN model. The process can then fine-tune both the StyleGAN model and the mapping network to better disentangle different attributes. In the warm-up stage, the view code can be sampled in selected views, and the remaining dimensions of w * ∈ W * can be sampled. One can attempt to minimize the L2 difference between the mapping code and the StyleGAN model code w * . To encourage disentanglement in the latent space, one can penalize the entropy of each dimension i of s. An example overall loss function for this mapping network can then be given by:

[0052]

[0053] By training the mapping network, one can disentangle views, shapes, and textures in the original StyleGAN model, but the background can remain entangled. Thus, one can fine-tune the model to achieve better disentanglement. One can fine-tune the StyleGAN model in conjunction with a cycle-consistency loss. In particular, by inputting the sampled shape, texture, and background into the StyleGAN model, one can obtain a synthetic image. This approach can encourage consistency between the original sampled attributes and the shape, texture, and background predicted from the StyleGAN synthetic image by the inverse graphics network. The same background B can be filled with two different {S, T} pairs to generate two images I1and I2. One can then encourage the re-synthesized backgrounds B1and B2to become similar. This loss attempts to disentangle the background from the foreground object. During training, imposing a consistency loss on B in image space can cause the image to blur, so it can be constrained in code space. An example of the fine-tuning loss takes the following form:

[0054]

[0055] In one example, the inverse graphics model based on DIB-R was trained using Adam with a learning rate of 1e 4 , λ IOU , λ col , λ lap , λ sm , and λ mov were set to 3, 20, 5, 5, and 2.5, respectively. The model was first trained for 3,000 iterations using the L col loss, and then fine-tuned to make the texture more realistic by adding the L percept loss. This process set λ 5 perception to 0.5. The model converged in 200,000 iterations with a batch size of 16. Training on four V100 GPUs took approximately 120 hours. The training produced high-quality 3D reconstruction results, including the quality of the predicted shape and texture, and the diversity of the obtained 3D shapes. This approach also works for more challenging (e.g., jointed) courses, such as animals.

[0056] In at least one embodiment, Adam is used with a learning rate of 1e 5StyleGAN-R model trained with a learning rate of 0.001 and a batch size of 16. The warm-up phase was performed for 700 iterations and joint fine-tuning was performed for another 2500 iterations. Using the provided input image, the process first predicts the mesh and texture using the trained inverse graphics model, and then inputs these 3D attributes into StyleGAN-R to generate a new image. For comparison, the same 3D attributes are provided to the DIB-R graphics renderer (i.e., the OpenGL renderer). It can be noted that DIB-R can only render the predicted object, while StyleGAN-R also has the ability to render the object into a desired background. It was found that StyleGAN-R produced a relatively consistent image compared to the input image. The shape and texture were well preserved, while only the background had a slight content shift.

[0057] This approach was tested when processing StyleGAN synthetic images from the test set and real images. Specifically, given an input image, the approach predicts 3D attributes using the inverse graphics network and extracts the background by making object masks using Mask-RCNN. Then, the approach manipulates these attributes and feeds them into StyleGAN-R to synthesize new views.

[0058] To control the viewpoint, the process first freezes the shape, texture, and background, changing only the camera viewpoint. Meaningful results were obtained, especially in terms of shape and texture. For comparison, another approach that has been explored is to directly optimize the latent code of the GAN (in the example, the code of the original StyleGAN) through an L2 image reconstruction loss. However, in at least some embodiments, such an approach can fail to generate reasonable images, demonstrating the importance of the mapping network and fine-tuning the entire architecture in a loop using the 3D inverse graphics network.

[0059] To control the shape, texture, and background, this approach can attempt to manipulate these or other 3D attributes while keeping the camera viewpoint fixed. In one example, the shape of all cars can be changed to one picked shape, and a neural rendering performed using StyleGAN-R is performed. Such a process successfully swaps the shape of the cars while keeping other characteristics. The process is also able to modify minor parts of the car, such as the trunk and headlight. The same experiment can be performed but swapping the texture and background. In some embodiments, swapping the texture can also slightly modify the background, which suggests that further improvements can be sought when disentangling the two. This framework also works well when provided with real images, as the images of StyleGAN are very realistic.

[0060] The StyleGAN codebase provides models for different object classes at different resolutions. Here, we take the car model at 512x384 as an example. The model contains 16 layers, where every two consecutive layers form a block. Each block has a different number of channels. In the last block, the model generates a 32-channel feature map at 512x384 resolution. Finally, a learned RGB transformation function is applied to convert the feature map to an RGB image. The feature maps of each block can be visualized through the learned RGB transformation function. Specifically, for a feature map of size h x w x c in each block, the process can sum along the feature dimension first, forming a h x w x 1 tensor. The process can repeat this 32 times and generate a new feature map of size h x w x 32. This allows the information of all channels to be preserved and directly applied in the last block to convert it to an RGB image. In this example, blocks 1 and 2 do not exhibit interpretable structure, while the car shape starts to appear in blocks 3-5. Block 4 has a rough outline of the car, which becomes further refined in block 5. From blocks 6 to 8, the shape of the car becomes increasingly refined, and the background scene also appears. This supports the view angle to be controlled in blocks 1 and 2 (e.g., the first 4 layers), while in this example, the shape, texture, and background exist in the last 12 layers. The car shape and texture, as well as the background scene at different view angles, all have high consistency. Note that for articulated objects such as horses and birds, the StyleGAN model can not perfectly preserve the target joints at different view angles, which can lead to challenges in training high-precision models using multi-view consistency loss.

[0061] As mentioned earlier, inverse graphics tasks require camera pose information during training, which can be challenging to obtain for real images. The pose is usually obtained by annotating the key points of each object and running structure from motion (SFM) techniques to compute the camera parameters. However, key point annotation is very time-consuming - it takes about 3-5 minutes per object in one experiment. StyleGAN models can be used to significantly reduce the annotation work, as samples with the same share the same view angle. Therefore, this process only needs to select a few The poses are assigned to the camera poses. Specifically, the poses can be assigned to several bins sufficient to train the inverse graph network, where the camera uses these bins as initialization for joint optimization during training, along with the network parameters. In one example, each view is annotated with a coarse absolute camera pose (which can be further optimized during training). Specifically, one example could start by selecting 12 azimuth angles: [0°, 30°, 60°, 90°, 120°, 150°, 180°, 210°, 240°, 270°, 300°, 330°]. Given a StyleGAN view, the process could include manually classifying which azimuth angle it is closer to and assigning it a corresponding label with a fixed elevation angle (0°) and camera distance.

[0062] To demonstrate the effectiveness of this camera initialization, it can be compared with another inverse graph network trained with more accurate camera initialization. This initialization is performed on each selected view of a single car example. The annotation of object keypoints was done manually, which took approximately 3-4 hours (about 200 minutes for 39 views). Note that this is still a significantly reduced annotation workload compared to the 200-350 hours required to annotate keypoints for each object in the Pascal3D dataset. Camera parameters can then be computed using Structure of Motion (SfM). Two inverse graph networks, initialized with different cameras and trained as the view model and keypoint model respectively, can be referenced separately. While training takes the same amount of time, the view model saves annotation time. The performance of the view model and the keypoint model is comparable to almost the same 2D IOU reprojection scores on the StyleGAN test set. Furthermore, during training, the two camera systems converge to the same location. This can be evaluated by converting all views to quaternions and comparing the differences between rotation axes and rotation angles. Across all views, the average difference in rotation axes is only 1.43°, and the difference in rotation angles is 0.42°. The maximum difference in rotation axes is only 2.95°, and the difference in rotation angles is 1.11°. Both qualitative and quantitative comparisons show that the view camera initialization is sufficient to train an accurate inverse graph network, requiring no additional annotations. This demonstrates a scalable approach to creating multi-view datasets using StyleGAN, requiring approximately one minute of annotation time per class.

[0063] Figure 5An example process 500 for training an inverse graphics network that can be used in accordance with various embodiments is shown. It will be appreciated that for this and other processes presented herein, additional, fewer, or alternative steps can be performed in similar or alternative order, or at least partially in parallel, within the scope of various embodiments, unless otherwise specifically noted. Further, while discussed with respect to streaming video content, it will be appreciated that such augmentations can be provided to individual images or sequences of images, stored video files, augmented or virtual reality streams, or other such content. In this example, a two-dimensional (2D) training image including a representation of an object is received. A set of camera poses is also received 504, or otherwise determined. Using the 2D image information with a generator network (e.g., StyleGAN), a set of view images is generated, including representations of the object with views according to the set of camera poses. Here, each generated image will include or be associated with corresponding camera pose information.

[0064] In this example, the set of generated view images can then be provided 508 as input to the inverse graphics network. A set of three-dimensional (3D) information for the object in the image is inferred 510, which can include shape, lighting, and texture, for example. One or more representations of the object can then be rendered 512 using a differentiable renderer using the 3D information from the input image and camera information. The rendered representations can then be compared 514 to corresponding ground truth data to determine one or more loss values. One or more network parameters or weights can then be adjusted 516 in an attempt to minimize the loss. A determination can be made 518 as to whether an end condition has been met, such as network convergence, reaching a maximum number of training passes, or all training data being processed, among other such options. If not, the process can continue with the next 2D training image. If the end criteria has been met, the optimized network parameters can be provided 520 for inference. At least some rendered output from the inverse graphics network can also be provided 524 as training data for further training or fine-tuning of the generator network. In some embodiments, the generator network and the inverse graphics network can be trained together using a common loss function.

[0065] As an example, Figure 6A network configuration 600 that can be used to provide or enhance content is shown. In at least one embodiment, a client device 602 can generate content for a session using components of a content application 604 on the client device 602 and data stored locally on the client device. In at least one embodiment, a content application 624 (e.g., an image generation or editing application) executing on a content server 620 (e.g., a cloud server or edge server) can initiate a session associated with at least the client device 602, as can utilize a session manager and user data stored in a user database 634, and can cause content 632 to be determined by a content manager 626. This content 632 can be transmitted to the client device 602 using an appropriate transport manager 622, and sent through a download, stream, or other such transport channel, if needed for this type of content or platform. In at least one embodiment, this content 632 can include 2D or 3D assets that can be used by a rendering engine to render a scene based on a determined scene graph. In at least one embodiment, a client device 602 receiving this content can provide the content to a corresponding content application 604, which can also or alternatively include a rendering engine (if necessary) for rendering at least some of the content for presentation by the client device 602, such as presenting image or video content through a display 606, and presenting audio such as sound and music through at least one audio playback device 608 such as speakers or headphones. For example, for live video content captured by one or more cameras, such a rendering engine can not be needed unless to enhance the video content in some manner. In at least one embodiment, at least some of the content can already be stored on the client device 602, rendered on the client device 602, or accessible to the client device 602, such that at least this portion of the content does not need to be transmitted over the network 640, such as in the case where the content can have been previously downloaded or stored locally on a hard drive or optical disc. In at least one embodiment, a transport mechanism such as a data stream can be used to transmit the content from the server 620 or content database 634 to the client device 602. In at least one embodiment, at least a portion of the content can be obtained or streamed from another source such as a third party content service 660, which can also include a content application 662 for generating or providing the content. In at least one embodiment, portions of this functionality can be performed using multiple computing devices or multiple processors within one or more computing devices such as a combination of CPUs and GPUs.

[0066] In at least one embodiment, content application 624 includes a content manager 626 that can determine or analyze content prior to transmission to client device 602. In at least one embodiment, content manager 626 can also include or work in cooperation with other components that can generate, modify, or enhance content to be provided. In at least one embodiment, this can include a rendering engine for rendering image or video content. In at least one embodiment, this rendering engine is part of an inverse graphics network. In at least one embodiment, an image, video, or scene generation component 628 can be used to generate image, video, or other media content. In at least one embodiment, an inverse graphics component 630, which can also include a neural network, can generate representations based on inferred 3D information, as discussed and suggested herein. In at least one embodiment, content manager 626 can cause this content (enhanced or unenhanced) to be transmitted to client device 602. In at least one embodiment, content application 604 on client device 602 can also include components such as a rendering engine, image or video generator 612, and inverse graphics module 614, such that any or all of this functionality can additionally or alternatively be performed on client device 602. In at least one embodiment, content application 662 on third party content service system 660 can also include such functionality. In at least one embodiment, a location where at least some of this functionality is performed can be configurable, or can depend on factors such as a type of client device 602 or availability of a network connection with sufficient bandwidth. In at least one embodiment, a system for content generation can include any suitable combination of hardware and software in one or more locations. In at least one embodiment, generated image or video content at one or more resolutions can also be provided to or made available to other client devices 650, such as for download or streaming from a media source that stores a copy of this image or video content. In at least one embodiment, this can include transmission of game content for a multiplayer game, where different client devices can display this content at different resolutions including one or more super resolutions.

[0067] In this example, the client devices can include any appropriate computing devices, such as can include desktop computers, notebook computers, set-top boxes, streaming devices, game consoles, smartphones, tablet computers, VR headsets, AR eyewear, wearable computers, or smart televisions. Each client device can submit requests across at least one wired or wireless network, which can include the Internet, an Ethernet network, a local area network (LAN), or a cellular network, among other such options. In this example, the requests can be submitted to an address associated with a cloud provider, which can operate or control one or more electronic resources under a cloud provider environment, such as can include a data center or a server farm. In at least one embodiment, the requests can be received or processed by at least one edge server that is located at the edge of the network and outside of at least one security layer associated with the cloud provider environment. In this way, latency can be reduced by enabling the client devices to interact with servers that are closer in distance, while also improving security of resources in the cloud provider environment.

[0068] In at least one embodiment, such a system can be used to perform graphics rendering operations. In other embodiments, such a system can be used for other purposes, such as to provide image or video content for testing or validating autonomous machine applications, or to perform deep learning operations. In at least one embodiment, such a system can be implemented using edge devices, or can incorporate one or more virtual machines (VMs). In at least one embodiment, such a system can be implemented at least partially in a data center or at least partially using cloud computing resources.

[0069] Inference and training logic

[0070] Figure 7A Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. In at least one embodiment, as shown in FIG. 7, inference and / or training logic 715 can include, among other things, neural processing units (NPUs) 720, a non-transitory computer-readable medium 725 that stores code including instructions that, when executed by one or more processors, cause the one or more processors to perform functions described herein, and a processor 730 that reads and executes the instructions. Figure 7A and / or Figure 7B Details regarding inference and / or training logic 715 are provided with respect to FIG. 7.

[0071] In at least one embodiment, inference and / or training logic 715 can include, without limitation, code and / or data storage 701 for storing forward and / or output weights and / or input / output data, and / or other parameters of neurons or layers of a neural network configured in aspects of one or more embodiments that are trained and / or used for inferencing. In at least one embodiment, training logic 715 can include or be coupled to code and / or data storage 701 for storing graph code or other software to control timing and / or order, where weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code, such as graph code, loads weight or other parameter information into processor ALUs based on an architecture of a neural network to which that code corresponds. In at least one embodiment, code and / or data storage 701 stores weight parameters and / or input / output data of each layer of a neural network trained or used in conjunction with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inferencing using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 701 can be included with other on-chip or off-chip data storage, including a processor’s Ll, L2, or L3 cache or system memory.

[0072] In at least one embodiment, any portion of code and / or data storage 701 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 701 can be cache memory, dynamic random addressable memory (“DRAM”), static random addressable memory (“SRAM”), non-volatile memory (such as flash memory), or other storage. In at least one embodiment, a choice of whether code and / or data storage 701 is internal or external to a processor, e.g., or made up of DRAM, SRAM, flash or some other storage type, can depend on available storage space on-chip or off-chip, latency requirements of training and / or inferencing functions being performed, batch size of data used in inferencing and / or training of a neural network, or some combination of these factors.

[0073] In at least one embodiment, inference and / or training logic 715 can include, without limitation, code and / or data storage 705 to store backward and / or output weights and / or input / output data for neurons or layers of a neural network trained and / or used for inferencing in aspects of one or more embodiments. In at least one embodiment, code and / or data storage 705 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with one or more embodiments during backward propagation of input / output data and / or weight parameters during training and / or inferencing using aspects of one or more embodiments. In at least one embodiment, training logic 715 can include or be coupled to code and / or data storage 705 to store graph code or other software to control timing and / or order, wherein weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic unit(s) (ALUs)). In at least one embodiment, code, such as graph code, loads weight or other parameter information into processor ALUs based on an architecture for a neural network to which that code corresponds. In at least one embodiment, any portion of code and / or data storage 705 can be included with other on-chip or off-chip data storage, including a processor’s L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data storage 705 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 705 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., Flash memory), or other storage. In at least one embodiment, a choice of whether code and / or data storage 705 is internal or external to a processor, e.g., whether it is made up of DRAM, SRAM, Flash memory, or some other storage type, depends on whether available storage is on-chip or off-chip, latency requirements of training and / or inferencing functions being performed, batch size of data being used in inferencing and / or training of a neural network, or some combination of these factors.

[0074] In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 can be separate storage structures. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 can be the same storage structure. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 can be partially the same storage structure and partially separate storage structures. In at least one embodiment, any portion of code and / or data storage 701 and code and / or data storage 705 can be included with other on-chip or off-chip data storage, including a processor’s L1, L2, or L3 cache or system memory.

[0075] In at least one embodiment, inference and / or training logic 715 can include, without limitation, one or more arithmetic logic units (“ALUs”) 710 (including integer and / or floating point units) for performing logical and / or mathematical operations based, at least in part, on training and / or inference code (e.g., graph code) or instructions therefrom. In at least one embodiment, results of operations performed by ALUs 710 can produce activations (e.g., output values from layers or neurons within a neural network) stored in activation storage 720 that are functions of input / output and / or weight parameter data stored in code and / or data storage 701 and / or code and / or data storage 705. In at least one embodiment, activations stored in activation storage 720 are generated by linear algebraic and / or matrix-based mathematics performed by ALUs 710 in response to executing instructions or other code, where weight values stored in code and / or data storage 705 and / or code and / or data storage 701 are used as operands along with other values such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which can be stored in code and / or data storage 705 or code and / or data storage 701 or other on-chip or off-chip storage.

[0076] In at least one embodiment, one or more ALUs 710 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment one or more ALUs 710 can be external to a processor or other hardware logic device or circuit using them (e.g., a co-processor). In at least one embodiment, one or more ALUs 710 can be included within execution units of a processor, or otherwise included in a group of ALUs accessible by execution units of a processor, which can be within a same processor or distributed between different types of processors (e.g., central processing units, graphics processing units, fixed function units, etc.). In at least one embodiment, code and / or data storage 701, code and / or data storage 705, and activation storage 720 can be on a same processor or other hardware logic device or circuit, while in another embodiment they can be in different processors or other hardware logic devices or circuits or some combination of same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 720 can be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory. Moreover, inference and / or training code can be stored with other code accessible to a processor or other hardware logic or circuit, and can be fetched and / or processed using fetch, decode, schedule, execute, exit, and / or other logic circuits of a processor.

[0077] In at least one embodiment, the active memory 720 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the active memory 720 may be wholly or partially located inside or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the active memory 720 is internal to or external to the processor may depend on the available on-chip or off-chip storage, the latency requirements for training and / or inference functions, the batch size of data used in inference and / or training the neural network, or some combination of these factors. For example, it may include DRAM, SRAM, flash memory, or other memory types. In at least one embodiment, Figure 7A The inference and / or training logic 715 shown can be used in conjunction with an application-specific integrated circuit (“ASIC”), such as those from Google. Processing unit, from Graphcore TM Inference processing units (IPUs) or from Intel Corp. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 7A The inference and / or training logic 715 shown can be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as field programmable gate array (“FPGA”)

[0078] Figure 7B Inference and / or training logic 715 according to at least one or more embodiments is illustrated. In at least one embodiment, the inference and / or training logic 715 may include, but is not limited to, hardware logic, wherein computational resources are dedicated or otherwise uniquely used in conjunction with weight values ​​or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, Figure 7B The inference and / or training logic 715 shown can be used in conjunction with an application-specific integrated circuit (ASIC), such as those from Google. Processing unit, from Graphcore TM Inference processing units (IPUs) or from Intel Corp. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 7BThe inference and / or training logic 715, as shown in FIG. 7B, can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware or other hardware, such as field programmable gate arrays (FPGAs). In at least one embodiment, the inference and / or training logic 715 includes, without limitation, code and / or data storage 701 and code and / or data storage 705, which can be used to store code (e.g., graph code), weight values and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In Figure 7B In at least one embodiment, each of code and / or data storage 701 and code and / or data storage 705 is associated with a dedicated computing resource, such as computing hardware 702 and computing hardware 706, respectively. In at least one embodiment, each of computing hardware 702 and computing hardware 706 includes one or more ALUs that perform mathematical functions (e.g., linear algebraic functions) only on information stored in code and / or data storage 701 and code and / or data storage 705, respectively, the results of which are stored in activation storage 720.

[0079] In at least one embodiment, each of code and / or data storage 701 and 705 and corresponding computing hardware 702 and 706, respectively, correspond to different layers of a neural network, such that activations resulting from one “storage / computing pair 701 / 702” of code and / or data storage 701 and computing hardware 702 are provided as input to the next “storage / computing pair 705 / 706” of code and / or data storage 705 and computing hardware 706 in order to reflect the conceptual organization of a neural network. In at least one embodiment, each storage / computing pair 701 / 702 and 705 / 706 can correspond to more than one neural network layer. In at least one embodiment, additional storage / computing pairs (not shown) can be included in inference and / or training logic 715 after or in parallel with storage / computing pairs 701 / 702 and 705 / 706.

[0080] Data Center

[0081] Figure 8 An example data center 800 that can use at least one embodiment is shown. In at least one embodiment, data center 800 includes a data center infrastructure layer 810, a framework layer 820, a software layer 830 and an application layer 840.

[0082] In at least one embodiment, as Figure 8As shown, the data center infrastructure layer 810 can include a resource orchestrator 812, grouped computing resources 814, and node computing resources (“node C.R.s”) 816(1)-816(N), where “N” represents any positive integer. In at least one embodiment, node C.R.s 816(1)-816(N) can include, but are not limited to, any number of central processing units (“CPUs” or “processors”), including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc., memory devices (e.g., dynamic random access memory), storage devices (e.g., solid state or disk drives), network input / output (“NWI / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more node C.R.s of node C.R.s 816(1)-816(N) can be a server having one or more of the above-described computing resources.

[0083] In at least one embodiment, grouped computing resources 814 can include individual groups of node C.R.s housed within one or more racks (not shown), or housed within a number of racks (also not shown) within various geographic locations. Individual groups of node C.R.s within grouped computing resources 814 can include groups of computing, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several node C.R.s including CPUs or processors can be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks can also include any number of power modules, cooling modules, and network switches, in any combination.

[0084] In at least one embodiment, resource orchestrator 812 can configure or otherwise control one or more node C.R.s 816(1)-816(N) and / or grouped computing resources 814. In at least one embodiment, resource orchestrator 812 can include a software design infrastructure (“SDI”) management entity for data center 800. In at least one embodiment, resource orchestrator 108 can comprise hardware, software, or some combination thereof.

[0085] In at least one embodiment, as Figure 8As shown, framework layer 820 includes a job scheduler 822, a configuration manager 824, a resource manager 826, and a distributed file system 828. In at least one embodiment, framework layer 820 may include a framework of software 832 supporting software layer 830 and / or one or more applications 842 supporting application layer 840. In at least one embodiment, software 832 or application 842 may respectively include web-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, framework layer 820 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark, which can utilize distributed file system 828 for large-scale data processing (e.g., "big data"). TM (Hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 832 may include a Spark driver to facilitate the scheduling of workloads supported by various layers of the data center 800. In at least one embodiment, the configuration manager 824 may be able to configure different layers, such as the software layer 830 and the framework layer 820, which includes Spark and a distributed file system 828 for supporting large-scale data processing. In at least one embodiment, the resource manager 826 is able to manage cluster or group computing resources mapped to or allocated to support the distributed file system 828 and the job scheduler 822. In at least one embodiment, the cluster or group computing resources may include group computing resources 814 on the data center infrastructure layer 810. In at least one embodiment, the resource manager 826 may coordinate with the resource coordinator 812 to manage these mapped or allocated computing resources.

[0086] In at least one embodiment, the software 832 included in the software layer 830 may include software used by at least a portion of the nodes CR816(1)-816(N), the grouped computing resources 814, and / or the distributed file system 828 of the framework layer 820. One or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.

[0087] In at least one embodiment, one or more applications 842 included in application layer 840 can include one or more types of applications used by at least portions of node C.R.s 816(1)-816(N), grouped computing resources 814, and / or distributed file system 828 of framework layer 820. One or more types of applications can include, but are not limited to, any number and / or type of genomics applications, cognitive computing and machine learning applications including training or inferencing software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0088] In at least one embodiment, any of configuration manager 824, resource manager 826, and resource orchestrator 812 can implement any number and type of self-modifying actions based on any amount and type of data acquired in any technically feasible fashion. In at least one embodiment, self-modifying actions can alleviate data center operators of data center 800 from making possibly poor configuration decisions and can avoid underutilization and / or poorly performing portions of a data center.

[0089] In at least one embodiment, data center 800 can include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information in accordance with one or more embodiments described herein. For example, in at least one embodiment, a machine learning model can be trained by computing weight parameters according to a neural network architecture using software and computing resources described above with respect to data center 800. In at least one embodiment, using weight parameters computed by one or more training techniques described herein, a trained machine learning model corresponding to one or more neural networks can be used to infer or predict information using resources described above with respect to data center 800.

[0090] In at least one embodiment, a data center can use CPUs, application specific integrated circuits (ASICs), GPUs, FPGAs, or other hardware to perform training and / or inferencing using resources described above. Moreover, one or more software and / or hardware resources described above can be configured as a service to allow users to train or perform information inference such as image recognition, speech recognition, or other artificial intelligence services.

[0091] Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments described herein using one or more examples of a neural network, of a support Figure 7A and / or Figure 7BDetails are provided regarding the inference and / or training logic 715. In at least one embodiment, the inference and / or training logic 715 can be implemented in the system. Figure 8 Used in systems for reasoning or predicting operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0092] Such components can be used to train an inverse graph network using a set of images generated by a generator network, where aspects of an object remain fixed while pose or view information varies across the set of images.

[0093] Computer System

[0094] Figure 9 This is a block diagram illustrating an exemplary computer system according to at least one embodiment. The exemplary computer system may be a system of interconnected devices and components, a system-on-a-chip (SoC), or some combination thereof formed with a processor, which may include an execution unit to execute instructions. In at least one embodiment, according to this disclosure, such as the embodiments described herein, computer system 900 may include, but is not limited to, components such as processor 902, whose execution unit includes logic to execute algorithms for process data. In at least one embodiment, computer system 900 may include a processor, such as those available from Intel Corporation of Santa Clara, California. Processor family, Xeon™ XScale™ and / or StrongARM™ Core TM or Nervana TM A microprocessor may be used, although other systems (including PCs, engineering workstations, set-top boxes, etc.) with other microprocessors may also be used. In at least one embodiment, computer system 900 may execute a version of the Windows operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (such as UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.

[0095] Embodiments can be used in other devices such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, embedded applications can include a microcontroller, a digital signal processor ("DSP"), a system on a chip, a network computer ("NetPC"), a set-top box, a network hub, a wide area

[0096] In at least one embodiment, computer system 900 can include, but is not limited to, processor 902, which can include, but is not limited to, one or more execution units 908 to perform, e.g., machine learning model training and / or inferencing, in accordance with techniques described herein. In at least one embodiment, computer system 900 is a single processor desktop or server system, but in another embodiment, computer system 900 can be a multiprocessor system. In at least one embodiment, processor 902 can include, but is not limited to, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor implementing a combo of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 902 can be coupled to a processor bus 910 that can transmit data signals between processor 902 and other components in computer system 900.

[0097] In at least one embodiment, processor 902 can include, but is not limited to, level 1 ("Ll") internal cache memory ("cache") 904. In at least one embodiment, processor 902 can have a single -level internal cache or multi-level internal cache. In at least one embodiment, cache memory can reside in the processor 902's external. Other embodiments can include a combination of internal and external caches based on specific implementation and requirements. In at least one embodiment, register file 906 can store different types of data within various registers including, but not limited to, integer registers, floating point registers, status registers, and instruction pointer registers.

[0098] In at least one embodiment, execution unit 908 includes, without limitation, logic to perform integer and floating-point operations, including bit- wide operations. In at least one embodiment, processor 902 can also include a microcode (“ucode”) read only memory (“ROM”) that stores microcode for certain macro instructions. In at least one embodiment, execution unit 908 can also include logic to handle a packed instruction set 909. In at least one embodiment, by including packed instruction set 909 in a general-purpose processor, many multimedia applications can be accelerated by using full width of data bus of processor 902. In one or more embodiments, by using full width of data bus of processor for operations on packed data, many multimedia applications can be executed more efficiently and speedier.

[0099] In at least one embodiment, execution unit 908 can also be used in a microcontroller, embedded processor, graphics device, DSP, and other types of logic circuits. In at least one embodiment, computer system 900 can include, without limitation, memory 920. In at least one embodiment, memory 920 can be implemented as a Dynamic Random Access Memory (“DRAM”) device, a Static Random Access Memory (“SRAM”) device, a flash memory device, or other memory device. In at least one embodiment, memory 920 can store instruction(s) 919 and / or data 921 represented by data signals that can be executed by processor 902.

[0100] In at least one embodiment, a system logic chip can be coupled to processor bus 910 and memory 920. In at least one embodiment, system logic chip can include, without limitation, a memory controller hub (“MCH”) 916 and processor 902 can communicate with MCH 916 via processor bus 910. In at least one embodiment, MCH 916 can provide a high bandwidth memory path 918 to memory 920 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, MCH 916 can direct data signals between processor 902, memory 920, and other components in computer system 900, and

[0101] In at least one embodiment, computer system 900 can use system I / O 922, which is a proprietary hub interface bus to couple MCH 916 to I / O controller hub (“ICH”) 930. In at least one embodiment, ICH 930 can provide a direct connection to some I / O devices and indirectly through a high-speed I / O bus. In at least one embodiment, the high-speed I / O bus can include, without limitation, a

[0102] In at least one embodiment, Figure 9 A system including interconnected hardware devices or “chips” is shown, while in other embodiments, Figure 9 An exemplary system on a chip (SoC) can be shown. In at least one embodiment, devices can be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 900 are interconnected using a compute express link (CXL) interconnect.

[0103] Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7L and / or 7M. Figure 7A and / or Figure 7B Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7L and / or 7M. Figure 9 Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7L and / or 7M.

[0104] Such components can be used to train the inverse graphics network using a set of images generated by the generator network, where aspects of the object remain fixed while pose or view information varies between images of the set.

[0105] Figure 10 FIG. 10 is a block diagram illustrating an electronic device 1000 for use with a processor 1010, in accordance with at least one embodiment. In at least one embodiment, electronic device 1000 can be, for example and without limitation, a laptop, a tower server, a rack server, a blade server, a laptop computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device.

[0106] In at least one embodiment, system 1000 can include, without limitation, a processor 1010 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 1010 is coupled using a bus or interface, such as an 1C bus, a System Management Bus (“SMBus”), a Low Pin Count (LPC) bus, a Serial Peripheral Interface (“SPI”), a High Definition Audio (“HDA”) bus, a Serial Advanced Technology Attachment (“SATA”) bus, a Universal Serial Bus (“USB”) (versions 1, 2, 3), or a Universal Asynchronous Receiver / Transmitter (“UART”) bus. In at least one embodiment, processor 1010 is a complex instruction set computer (“CISC”) or reduced instruction set computer (“RISC”) processor, multiprocessor, microcontroller, digital signal processor, embedded processor, graphics processing unit, or any other microprocessor or central processing unit (“CPU”) or controller. Figure 10 In at least one embodiment, system 1000 is shown to include an interconnect of hardware devices or “chips,” while in other embodiments, system 1000 can be shown to be a single chip with one or more processors 1010 and interconnects as shown. Figure 10 In at least one embodiment, system 1000 can be shown to be an exemplary system on a chip (“SoC”). In at least one embodiment, devices shown in FIG. 10 can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of system 1000 are interconnected using compute express link (“CXL”) interconnects. Figure 10 In at least one embodiment, devices shown in FIG. 10 can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of system 1000 are interconnected using compute express link (“CXL”) interconnects. Figure 10 In at least one embodiment, devices shown in FIG. 10 can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of system 1000 are interconnected using compute express link (“CXL”) interconnects.

[0107] In at least one embodiment, devices shown in FIG. 10 can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of system 1000 are interconnected using compute express link (“CXL”) interconnects. Figure 10The display 1024, touch screen 1025, touch pad 1030, near field communication unit ("NFC") 1045, sensor hub 1040, thermal sensor 1046, Express Chipset ("EC") 1035, Trusted Platform Module ("TPM") 1038, BIOS / firmware / flash memory ("BIOS, FW Flash") 1022, DSP 1060, drive 1020 (such as a solid state disk ("SSD") or a hard disk drive ("HDD")), wireless local area network unit ("WLAN") 1050, Bluetooth unit 1052, wireless wide area network unit ("WWAN") 1056, Global Positioning System ("GPS") 1055, camera ("USB 3.0 camera") 1054 (such as a USB 3.0 camera), and / or Low Power Double Data Rate ("LPDDR") memory unit ("LPDDR3") 1015 implemented in, for example, LPDDR3 standard can each be implemented in any suitable manner.

[0108] In at least one embodiment, other components can be communicatively coupled to processor 1010 by components described above. In at least one embodiment, accelerometer 1041, ambient light sensor ("ALS") 1042, compass 1043, and gyroscope 1044 can be communicatively coupled to sensor hub 1040. In at least one embodiment, thermal sensor 1039, fan 1037, keyboard 1036, and touch pad 1030 can be communicatively coupled to EC 1035. In at least one embodiment, speaker 1063, earpiece 1064, and microphone ("mic") 1065 can be communicatively coupled to audio unit ("audio codec and class D amplifier") 1062, which in turn can be communicatively coupled to DSP 1060. In at least one embodiment, audio unit 1062 can include, for example and without limitation, an audio coder / decoder ("codec") and a class D amplifier. In at least one embodiment, SIM card ("SIM") 1057 can be communicatively coupled to WWAN unit 1056. In at least one embodiment, components such as WLAN unit 1050 and Bluetooth unit 1052, as well as WWAN unit 1056, can be implemented as a Next Generation Form Factor ("NGFF").

[0109] Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7 and 8. Figure 7A and / or Figure 7B Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7 and 8. In at least one embodiment, inference and / or training logic 715 can be used in Figure 10Inference and / or prediction operations in the system can be based, at least in part, on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0110] Such components can be used to train an inverse graphics network using a set of images generated by a generator network, where aspects of the object remain fixed while pose or view information varies between images of the set.

[0111] Figure 11 is a block diagram of a processing system in accordance with at least one embodiment. In at least one embodiment, system 1100 includes one or more processors 1102 and one or more graphics processors 1108, and can be a single processor desktop system, a multiprocessor workstation system, or a server system having many processors 1102 or processor cores 1107. In at least one embodiment, system 1100 is a processing platform incorporated within a system on a chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices.

[0112] In at least one embodiment, system 1100 can include or be incorporated within a server-based gaming platform, a game console, a media console, a mobile gaming console, a handheld game console, or an online game console that includes game and media processing consoles. In at least one embodiment, system 1100 is a mobile phone, a smart phone, a tablet device, or a mobile internet device. In at least one embodiment, processing system 1100 can also include or be incorporated within a wearable device such as a smart watch wearable device, a smart eyewear device, an augmented reality device, or a virtual reality device. In at least one embodiment, processing system 1100 is a television or set-top box device having one or more processors 1102 and a graphical interface generated by one or more graphics processors 1108.

[0113] In at least one embodiment, one or more processors 1102 each include one or more processor cores 1107 to process instructions which, when executed, implement the operations for system and user software. In at least one embodiment, each of the one or more processor cores 1107 is configured to process a specific instruction set 1109. In at least one embodiment, instruction set 1109 can facilitate Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or computing via a Very Long Instruction Word (VLIW). In at least one embodiment, processor cores 1107 can each process a different instruction set 1109, which can include instructions to facilitate emulation of other instruction sets. In at least one embodiment, processor core 1107 can also include other processing devices, such as a digital signal processor (DSP).

[0114] In at least one embodiment, processor 1102 includes cache memory 1104. In at least one embodiment, processor 1102 can have single level cache or multiple levels of cache. In at least one embodiment, cache memory is shared among various components of processor 1102. In at least one embodiment, processor 1102 also uses an external cache (e.g., a level three (L3) cache or last level cache (LLC)) (not shown), which can be shared between processor cores 1107 using known cache coherency techniques. In at least one embodiment, register file 1106 is additionally included in processor 1102, which processor can include different types of registers to store different types of data (e.g., integer registers, floating point registers, status registers, and instruction pointer registers). In at least one embodiment, register file 1106 can include general registers or other registers.

[0115] In at least one embodiment, one or more processor(s) 1102 are coupled with one or more interface bus(es) 1110 to transmit communication signals between processor 1102 and other components in system 1100, such as address, data, or control signals. In at least one embodiment, interface bus 1110 can be a version of a processor bus, such as a direct media interface (DMI) bus, in at least one embodiment. In at least one embodiment, interface bus 1110 is not limited to DMI bus, and can include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In at least one embodiment, processor 1102 includes integrated memory controller 1116 and platform controller hub 1130. In at least one embodiment, memory controller 1116 facilitates communication between memory devices and other components of processing system 1100, while platform controller hub 1130 provides connections to input / output (I / O) devices via local I / O bus.

[0116] In at least one embodiment, memory device 1120 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase-change memory device, or a device with suitable performance for use as processor memory. In at least one embodiment, memory device 1120 may be used as system memory of processing system 1100 to store data 1122 and instructions 1121 for use when one or more processors 1102 execute an application or process. In at least one embodiment, memory controller 1116 is also coupled to an optional external graphics processor 1112, which may communicate with one or more graphics processors 1108 of processor 1102 to perform graphics and media operations. In at least one embodiment, display device 1111 may be connected to processor 1102. In at least one embodiment, display device 1111 may include one or more internal display devices, such as in mobile electronic devices or laptop devices, or external display devices connected via a display interface (e.g., DisplayPort). In at least one embodiment, the display device 1111 may include a head-mounted display (HMD), such as a stereoscopic display device for virtual reality (VR) or augmented reality (AR) applications.

[0117] In at least one embodiment, platform controller hub 1130 enables peripherals coupled to bridge 1122 to interact with a processor coupled to memory controller hub 1110. In at least one embodiment, platform controller hub 1130 can mediate communication between a processor coupled to memory controller hub 1110 and devices coupled to bridge 1122 via an I / O bus (e.g., a PCI bus). In at least one embodiment, platform controller hub 1130 can provide a plurality of different I / O buses for coupling to devices. In at least one embodiment, platform controller hub 1130 can enable devices to be connected to memory controller hub 1110 via an I / O bus (e.g., via an on-chip I / O bus such as a PCI bus).

[0118] In at least one embodiment, memory controller 1116 and instances of platform controller hub 1130 can be integrated into a discrete external graphics processor, such as external graphics processor 1112. In at least one embodiment, platform controller hub 1130 and / or memory controller 1116 can be external to one or more processor(s) 1102. For example, in at least one embodiment, system 1100 can include an external memory controller 1116 and platform controller hub 1130, which can be configured as a memory controller hub and a peripheral controller hub within a system chipset that is discrete from processor(s) 1102.

[0119] Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are described in more detail below in conjunction with FIGS. 7 and 8. Figure 7A and / or Figure 7BDetails regarding the inference and / or training logic 715 are provided. In at least one embodiment, some or all of inference and / or training logic 715 can be incorporated with graphics processor 1100. For example, in at least one embodiment, the training and / or inference techniques described herein can use one or more ALUs embodied in a graphics processor. Further, in at least one embodiment, the inference and / or training operations described herein can be accomplished with logic other than that shown. Figure 7A or Figure 7B In at least one embodiment, weight parameters can be stored in on-chip or off-chip memory and / or registers (shown or not) that configure ALUs of a graphics processor to perform one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0120] Such components can be used to train an inverse graphics network using a set of images generated by a generator network, where aspects of the object remain fixed while pose or view information varies between images of the set.

[0121] Figure 12 is a block diagram of a processor 1200 having one or more processor cores 1202A-1202N, an integrated memory controller 1214, and an integrated graphics processor 1208, according to at least one embodiment. In at least one embodiment, processor 1200 can include additional cores, up to and including an additional core 1202N represented by a dashed lined in at least one embodiment. In at least one embodiment, each processor core 1202A-1202N includes one or more internal cache units 1204A-1204N. In at least one embodiment, each processor core can also include access to one or more shared cache units 1206.

[0122] In at least one embodiment, internal cache units 1204A-1204N and shared cache unit 1206 represent a cache memory hierarchy within processor 1200. In at least one embodiment, cache memory units 1204A-1204N can include at least one level of cache memory such as level one (LI), level two (L2), level three (L3), level four (L4), or other levels of cache, within each processor core 1202A-1202N and shared level cache(s) 1206, where the highest level of cache memory prior to main memory is referred to as the LLC. In at least one embodiment, cache coherence logic maintains coherency for the various cache units 1206 and 1204A-1204N.

[0123] In at least one embodiment, processor 1200 also includes a set of one or more bus controller units 1216 and a system agent core 1210. In at least one embodiment, one or more bus controller units 1216 manage a set of peripheral buses, such as one or more PCI or PCIe buses. In at least one embodiment, system agent core 1210 provides management functionality for various processor components. In at least one embodiment, system agent core 1210 includes one or more integrated memory controllers 1214 to manage access to various external memory devices (not shown), including support for data bus protocols such as DDR SDRAM.

[0124] In at least one embodiment, one or more processor cores 1202A-1202N include support to run in multiple threads simultaneously. In at least one embodiment, system agent core 1210 includes components for coordination and operation of cores 1202A-1202N during multi-threaded processing. In at least one embodiment, system agent core 1210 can additionally include a power control unit (PCU), including logic and components to govern one or more power states of processor cores 1202A-1202N and graphics processor 1208.

[0125] In at least one embodiment, processor 1200 also includes graphics processor 1208, which can be configured to perform a graphics processing operations. In at least one embodiment, graphics processor 1208 couples with shared cache unit 1206, and system agent core 1210, including one or more integrated memory controllers 1214. In at least one embodiment, system agent core 1210 also includes a display controller 1211 for driving one or more coupled displays to present graphics processor output. In at least one embodiment, display controller 1211 can also be a separate module coupled with graphics processor 1208 via at least one interconnect, or can be integrated within graphics processor 1208.

[0126] In at least one embodiment, ring based interconnect unit 1212 is used to couple the internal components of processor 1200. In at least one embodiment, an alternative interconnect unit can be used, such as a point-to-point interconnect, a switched interconnect, or other technology. In at least one embodiment, graphics processor 1208 couples with ring interconnect 1212 via I / O link 1213.

[0127] In at least one embodiment, I / O link 1213 represents at least one of a variety of I / O interconnects, including packaged I / O interconnects that facilitate communication between various processor components and high-performance embedded memory module 1218 (e.g., eDRAM module). In at least one embodiment, each of processor cores 1202A-1202N and graphics processor 1208 uses embedded memory module 1218 as a shared last-level cache.

[0128] In at least one embodiment, processor cores 1202A-1202N are homogeneous cores executing a common instruction set architecture. In at least one embodiment, processor cores 1202A-1202N are heterogeneous in terms of instruction set architecture (ISA), with one or more processor cores 1202A-1202N executing a common instruction set, while one or more other processor cores 1202A-1202N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, processor cores 1202A-1202N are heterogeneous in terms of microarchitecture, with one or more cores having relatively high power consumption coupled to one or more power cores having lower power consumption. In at least one embodiment, processor 1200 may be implemented on one or more chips or implemented as a SoC integrated circuit.

[0129] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. The following is combined with... Figure 7A and / or Figure 7B Details regarding the inference and / or training logic 715 are provided. In at least one embodiment, some or all of the inference and / or training logic 715 may be incorporated into the processor 1200. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in... Figure 12 The graphics processor 1512, graphics core 1202A-1202N, or other components are used. Furthermore, in at least one embodiment, the inference and / or training operations described herein can use, except... Figure 7A or Figure 7B The logic is performed using logic other than that shown. In at least one embodiment, weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of the graphics processor 1200 to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein. Such components can be used to generate enhanced content, such as image or video content with upgraded resolution, reduced artifact presence, and enhanced visual quality.

[0130] Such components can be used to train an inverse graphics network using a set of images generated by a generator network, where aspects of the object remain fixed while pose or view information varies between images of the set.

[0131] Virtualized computing platform

[0132] Figure 13 is an example data flow diagram of a process 1300 of generating and deploying image processing and inference pipelines, in accordance with at least one embodiment. In at least one embodiment, process 1300 can be deployed for use with imaging devices, processing devices, and / or other device types at one or more facilities 1302. Process 1300 can be executed within a training system 1304 and / or a deployment system 1306. In at least one embodiment, training system 1304 can be used to perform training, deployment, and implementation of machine learning models (e.g., neural networks, object detection algorithms, computer vision algorithms, etc.) for deployment system 1306. In at least one embodiment, deployment system 1306 can be configured to offload processing and computing resources in a distributed computing environment to reduce infrastructure requirements of facilities 1302. In at least one embodiment, one or more applications in a pipeline can use or call services (e.g., inference, visualization, computation, AI, etc.) of deployment system 1306 during application execution.

[0133] In at least one embodiment, some applications used in an advanced processing and inference pipeline can use machine learning models or other AI to perform one or more processing steps. In at least one embodiment, machine learning models can be trained at facilities 1302 using data 1308 (e.g., imaging data) generated at facilities 1302 (and stored on one or more picture archiving and communication systems (PACS) servers at facilities 1302), can be trained using imaging or sequencing data 1308 from another facility or facilities, or a combination thereof. In at least one embodiment, training system 1304 can be used to provide applications, services, and / or other resources to generate working, deployable machine learning models for deployment system 1306.

[0134] In at least one embodiment, model registry 1324 can be supported by object storage, which can support versioning and object metadata. In at least one embodiment, model registry 1324 can be accessed from within a cloud platform through, for example, cloud storage (e.g., Amazon S3), cloud object storage (e.g., Amazon S3 Glacier), and / or other cloud storage services. Figure 14compatible application programming interface (API) to access object storage. In at least one embodiment, machine learning models within model registry 1324 can be uploaded, listed, modified, or deleted by developers or partners of systems that interact with API. In at least one embodiment, API can provide access to methods that allow users with appropriate credentials to associate models with applications such that models can be executed as part of execution of containerized instantiations of applications.

[0135] In at least one embodiment, training pipeline 1404 Figure 14 ) can include scenarios in which facility 1302 is training their own machine learning models, or has existing machine learning models that need to be optimized or updated. In at least one embodiment, imaging data 1308 generated by imaging devices, sequencing devices, and / or other types of devices can be received. In at least one embodiment, once imaging data 1308 is received, AI assisted annotation 1310 can be used to help generate annotations corresponding to imaging data 1308 to be used as ground truth data for machine learning models. In at least one embodiment, AI assisted annotation 1310 can include one or more machine learning models (e.g., convolutional neural networks (CNNs)) that can be trained to generate annotations corresponding to certain types of imaging data 1308 (e.g., from certain devices). In at least one embodiment, AI assisted annotation 1310 can then be used directly, or can be adjusted or fine-tuned using annotation tools to generate ground truth data. In at least one embodiment, AI assisted annotation 1310, labeled clinical data 1312, or a combination thereof can be used as ground truth data to train machine learning models. In at least one embodiment, trained machine learning models can be referred to as output models 1316, and can be used by deployment system 1306, as described herein.

[0136] In at least one embodiment, training pipeline 1404 Figure 14) can include situations in which facility 1302 needs a machine learning model for performing one or more processing tasks for one or more applications in deployment system 1306, but facility 1302 can not currently have such a machine learning model (or can not have a model that is optimized, efficient, or effective for this purpose). In at least one embodiment, an existing machine learning model can be selected from model registry 1324. In at least one embodiment, model registry 1324 can include machine learning models that are trained to perform a variety of different inferencing tasks on imaging data. In at least one embodiment, the machine learning models in model registry 1324 can have been trained on imaging data from different facilities (e.g., facilities located remotely from facility 1302). In at least one embodiment, a machine learning model can have been trained on imaging data from one location, two locations, or any number of locations. In at least one embodiment, when training on imaging data from a particular location, the training can be performed at that location, or at least in a manner that protects the confidentiality of the imaging data or limits the transfer of the imaging data offsite. In at least one embodiment, once a model is trained, or partially trained, at a location, the machine learning model can be added to model registry 1324. In at least one embodiment, the machine learning model can then be retrained or updated at any number of other facilities, and the retrained or updated model can be used in model registry 1324. In at least one embodiment, a machine learning model can then be selected from model registry 1324 (and referred to as output model 1316), and can be used in deployment system 1306 to perform one or more processing tasks for one or more applications of the deployment system.

[0137] In at least one embodiment, training pipeline 1404( Figure 14In at least one embodiment, scenario can include facility 1302 that requires a machine learning model for performing one or more processing tasks for deploying one or more applications in deployment system 1306, but facility 1302 can not currently have such a machine learning model (or can not have an optimized, efficient, or effective model). In at least one embodiment, due to population differences, robustness of training data used to train a machine learning model, diversity of training data anomalies, and / or other issues with training data, a machine learning model selected from model registry 1324 can not be fine-tuned or optimized for imaging data 1308 generated at facility 1302. In at least one embodiment, AI assisted annotation 1310 can be used to help generate annotations corresponding to imaging data 1308 for use as ground truth data to train or update a machine learning model. In at least one embodiment, labeled clinical data 1312 can be used as ground truth data to train a machine learning model. In at least one embodiment, retraining or updating a machine learning model can be referred to as model training 1314. In at least one embodiment, model training 1314 (e.g., AI assisted annotation 1310, labeled clinical data 1312, or a combination thereof) can be used as ground truth data to retrain or update a machine learning model. In at least one embodiment, a trained machine learning model can be referred to as output model 1316 and can be used by deployment system 1306, as described herein.

[0138] In at least one embodiment, deployment system 1306 can include software 1318, services 1320, hardware 1322, and / or other components, features, and functionality. In at least one embodiment, deployment system 1306 can include a software “stack” such that software 1318 can be built on top of services 1320, and can use services 1320 to perform some or all processing tasks, and services 1320 and software 1318 can be built on top of hardware 1322 and use hardware 1322 to perform processing, storage, and / or other computing tasks of deployment system. In at least one embodiment, software 1318 can include any number of different containers, where each container can execute an instantiation of an application. In at least one embodiment, each application can perform one or more processing tasks in a high-level processing and inference pipeline (e.g., inference, object detection, feature detection, segmentation, image enhancement, calibration, etc.). In at least one embodiment, a high-level processing and inference pipeline can be defined based on a selection of different containers desired or required to process imaging data 1308 (e.g., to convert output back to a usable data type, in addition to receiving and configuring containers for use by each container with imaging data for use and / or use by facility 1302 after processing through the pipeline. In at least one embodiment, a combination of containers within software 1318 (e.g., which make up a pipeline) can be referred to as a virtual instrument (as described in greater detail herein), and a virtual instrument can utilize services 1320 and hardware 1322 to perform some or all processing tasks of applications instantiated in containers.

[0139] In at least one embodiment, a data processing pipeline can receive input data (e.g., imaging data 1308) in a specific format in response to an inference request (e.g., a request from a user of deployment system 1306). In at least one embodiment, input data can represent one or more images, videos, and / or other data representations generated by one or more imaging devices. In at least one embodiment, data can be pre-processed as part of a data processing pipeline to prepare data for processing by one or more applications. In at least one embodiment, post-processing can be performed on output of one or more inference tasks or other processing tasks of a pipeline to prepare output data for a next application and / or to prepare output data for transmission and / or use by a user (e.g., in response to an inference request). In at least one embodiment, inference tasks can be performed by one or more machine learning models, such as trained or deployed neural networks, which can include output models 1316 of training system 1304.

[0140] In at least one embodiment, the tasks of the data processing pipeline can be encapsulated in containers, each container representing a discrete, fully functional instantiation of an application and a virtualized computing environment capable of referencing a machine learning model. In at least one embodiment, containers or applications can be published to a private (e.g., limited access) area of ​​a container registry (described in more detail herein), and trained or deployed models can be stored in a model registry 1324 and associated with one or more applications. In at least one embodiment, an image of an application (e.g., a container image) can be used in the container registry, and once a user selects an image from the container registry for deployment in the pipeline, that image can be used to generate containers for instantiation of the application for use by the user's system.

[0141] In at least one embodiment, a developer (e.g., a software developer, clinician, physician, etc.) can develop, publish, and store an application (e.g., as a container) for performing image processing and / or inference on provided data. In at least one embodiment, a software development kit (SDK) associated with the system can be used to perform development, publication, and / or storage (e.g., to ensure that the developed application and / or container conforms to or is compatible with the system). In at least one embodiment, the developed application can be tested locally using the SDK (e.g., at a first facility, testing data from a first facility), the SDK serving as a system (e.g.,...). Figure 14 System 1400 may support at least some services 1320. In at least one embodiment, since DICOM objects may contain one to hundreds of images or other data types, and due to variations in the data, the developer may be responsible for managing (e.g., setting up constructs for preprocessing built into the application, etc.) the extraction and preparation of incoming data. In at least one embodiment, once verified by system 1400 (e.g., for accuracy), the application becomes available in the container registry for user selection and / or implementation to perform one or more processing tasks on data at the user's facility (e.g., a second facility).

[0142] In at least one embodiment, the developer can then share the application or container over a network for the system (e.g., Figure 14of the system 1400) by users. In at least one embodiment, completed and validated applications or containers can be stored in a container registry, and related machine learning models can be stored in a model registry 1324. In at least one embodiment, a requesting entity (which provides an inference or image processing request) can browse the container registry and / or model registry 1324 to obtain applications, containers, datasets, machine learning models, etc., select a desired combination of elements to include in a data processing pipeline, and submit an image processing request. In at least one embodiment, a request can include input data necessary to perform the request (and, in some examples, data related to a patient), and / or can include a selection of applications and / or machine learning models to be executed in processing the request. In at least one embodiment, a request can then be passed to one or more components of the deployment system 1306 (e.g., a cloud) to perform processing of the data processing pipeline. In at least one embodiment, processing by the deployment system 1306 can include referencing elements (e.g., applications, containers, models, etc.) selected from the container registry and / or model registry 1324. In at least one embodiment, once results are generated through the pipeline, the results can be returned to a user for review (e.g., for review in a viewing application suite executed on a local, on-premises workstation or terminal).

[0143] In at least one embodiment, to help process or execute applications or containers in a pipeline, services 1320 can be utilized. In at least one embodiment, services 1320 can include computing services, artificial intelligence (Al) services, visualization services, and / or other service types. In at least one embodiment, services 1320 can provide functionality that is common to one or more applications in software 1318, and thus functionality can be abstracted as a service that can be called or utilized by applications. In at least one embodiment, functionality provided by services 1320 can run dynamically and more efficiently, while also allowing applications to process data in parallel (e.g., using Figure 14the parallel computing platform 1430) to scale well. In at least one embodiment, rather than requiring each application that requires the same functionality provided by a shared service 1320 to have a respective instance of the service 1320, the service 1320 can be shared among and between various applications. In at least one embodiment, as a non-limiting example, a service can include an inference server or engine that can be used to perform detection or segmentation tasks. In at least one embodiment, a model training service can be included, which can provide machine learning model training and / or retraining capabilities. In at least one embodiment, a data augmentation service can be further included, which can provide GPU-accelerated data (e.g., DICOM, RIS, CIS, REST-compliant, RPC, raw, etc.) extraction, resizing, scaling, and / or other augmentations. In at least one embodiment, a visualization service can be used, which can add image rendering effects (e.g., ray tracing, rasterization, de-noising, sharpening, etc.) to add realism to two-dimensional (2D) and / or three-dimensional (3D) models. In at least one embodiment, a virtual instrument service can be included, which provides beamforming, segmentation, inference, imaging, and / or support to other applications within a pipeline of a virtual instrument.

[0144] In at least one embodiment, where the services 1320 include an AI service (e.g., an inference service), as part of an application’s execution, one or more machine learning models can be executed by invoking (e.g., as an API call) the inference service (e.g., inference server) to execute one or more machine learning models or processing thereof. In at least one embodiment, where another application includes one or more machine learning models for a segmentation task, the application can invoke the inference service to execute the machine learning model for performing one or more processing operations associated with the segmentation task. In at least one embodiment, software 1318 implementing an advanced processing and inference pipeline, which includes a segmentation application and an anomaly detection application, can be pipelined, as each application can invoke the same inference service to perform one or more inference tasks. In at least one embodiment, hardware 1322 can include a GPU, CPU, graphics card, AI / deep learning system (e.g., an AI supercomputer such as NVIDIA’s DGX), a cloud platform, or a combination thereof.

[0145] In at least one embodiment, different types of hardware 1322 can be used to provide efficient, specially-built support for software 1318 and services 1320 in deployment system 1306. In at least one embodiment, GPU processing can be implemented for local processing within an AI / deep learning system, in a cloud system, and / or in other processing components of deployment system 1306 (e.g., at facility 1302) to improve efficiency, accuracy, and performance of image processing and generation. In at least one embodiment, software 1318 and / or services 1320 can be optimized for GPU processing, by way of non-limiting example with respect to deep learning, machine learning, and / or high performance computing. In at least one embodiment, at least some of computing environments of deployment system 1306 and / or training system 1304 can be executed in a data center, one or more supercomputer or high performance computer systems with GPU-optimized software (e.g., a combination of hardware and software of NVIDIA DGX systems). In at least one embodiment, hardware 1322 can include any number of GPUs that can be called upon to perform data processing in parallel, as described herein. In at least one embodiment, a cloud platform can also include GPU processing for GPU-optimized execution of deep learning tasks, machine learning tasks, or other computing tasks. In at least one embodiment, an AI / deep learning supercomputer and / or GPU-optimized software (e.g., as provided on NVIDIA’s DGX systems) can be used as a hardware abstraction and scaling platform to execute a cloud platform (e.g., NVIDIA’s NGC), in at least one embodiment. In at least one embodiment, a cloud platform can integrate an application container clustering system or orchestration system (e.g., KUBERNETES) across multiple GPUs to enable seamless scaling and load balancing.

[0146] Figure 14 is a system diagram of an example system 1400 for generating and deploying imaging deployment pipelines, in accordance with at least one embodiment. In at least one embodiment, system 1400 can be used to implement processes 1300 and / or other processes of FIG. 13, including advanced processing and inference pipelines. In at least one embodiment, system 1400 can include training system 1304 and deployment system 1306. In at least one embodiment, training system 1304 and deployment system 1306 can be implemented using software 1318, services 1320, and / or hardware 1322, as described herein. Figure 13

[0147] ​In at least one embodiment, system 1400 (e.g., training system 1304 and / or deployment system 1306) can be implemented in a cloud computing environment (e.g., using cloud 1426). In at least one embodiment, system 1400 can be implemented locally (with respect to a medical service facility), or as a combination of cloud computing resources and local computing resources. In at least one embodiment, access to APIs in cloud 1426 can be limited to authorized users by instituting security measures or protocols. In at least one embodiment, security protocols can include network tokens that can be signed by an authentication (e.g., AuthN, AuthZ, Gluecon, etc.) service and can carry appropriate authorization. In at least one embodiment, APIs (described herein) of a virtual instrument or other instances of system 1400 can be limited to a set of public IPs that have been vetted or authorized for interaction.

[0148] In at least one embodiment, various components of system 1400 can communicate information between each other using any of a plurality of different network types, including but not limited to local area networks (LANs) and / or wide area networks (WANs) via wired and / or wireless communication protocols. In at least one embodiment, communication between facilities and components of system 1400 (e.g., for sending inference requests, for receiving results of inference requests, etc.) can be communicated through one or more data buses, wireless data protocols (Wi-Fi), wired data protocols (e.g., Ethernet), etc.

[0149] In at least one embodiment, similar to training pipelines 1302 described herein with respect to Figure 13 In at least one embodiment, training pipeline 1404 can be used to train or retrain one or more (e.g., pre-trained) models, and / or implement one or more pre-trained models 1406 (e.g., without retraining or updating), where deployment system 1306 will use the one or more machine learning models in deployment pipeline 1410. In at least one embodiment, as a result of training pipeline 1404, output models 1316 can be generated. In at least one embodiment, training pipeline 1404 can include any number of processing steps, such as but not limited to conversion or adaptation of imaging data (or other input data). In at least one embodiment, different training pipelines 1404 can be used for different machine learning models used by deployment system 1306. In at least one embodiment, training pipeline 1404 of a first example described with respect to Figure 13 In at least one embodiment, training pipeline 1404 of a first example described with respect to Figure 13 In at least one embodiment, training pipeline 1404 of a first example described with respect to Figure 13The training pipeline 1404 of the third example described can be used for the third machine learning model. In at least one embodiment, any combination of tasks within the training system 1304 can be used depending on the requirements of each respective machine learning model. In at least one embodiment, one or more machine learning models can already be trained and ready for deployment, so the training system 1304 can not perform any processing on the machine learning model and the one or more machine learning models can be implemented by the deployment system 1306.

[0150] In at least one embodiment, the output model 1316 and / or the pre-trained model 1406 can include any type of machine learning model, depending on implementation or embodiment. In at least one embodiment and without limitation thereto, machine learning models used by the system 1400 can include using linear regression, logistic regression, decision trees, support vector machines (SVM), Naive Bayes, k- nearest neighbors (Knn), k-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutional, recurrent, perceptrons, long / short-term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolutional, generative adversarial, liquid state machines, etc.), and / or other types of machine learning models.

[0151] In at least one embodiment, the training pipeline 1404 can include AI-assisted annotation, as described herein with respect to at least Figure 15BIn at least one embodiment, labeled clinical data 1312 (e.g., traditional annotations) can be generated by any number of techniques. In at least one embodiment, labels or other annotations can be generated in a drawing program (e.g., annotation program), a computer aided design (CAD) program, a labeling program, another type of application suitable for generating annotations or labels for ground truth, and / or can be hand drawn, in some examples. In at least one embodiment, ground truth data can be synthetically generated (e.g., from computer models or renderings), realistically generated (e.g., designed and generated from real world data), automatically generated by a machine (e.g., using feature analysis and learning to extract features from data and then generate labels), manually annotated (e.g., by a marker or annotation specialist defining locations of labels), and / or combinations thereof. In at least one embodiment, for each instance of imaging data 1308 (or other data types used by a machine learning model), there can be corresponding ground truth data generated by training system 1304. In at least one embodiment, AI-assisted annotation can be performed as part of deployment pipeline 1410; in addition to or instead of AI-assisted annotation included in training pipeline 1404. In at least one embodiment, system 1400 can include a multi-tiered platform that can include a software tier of diagnostic applications (or other application types) (e.g., software 1318) that can perform one or more medical imaging and diagnostic functions. In at least one embodiment, system 1400 can be communicatively coupled to (e.g., via encrypted links) a PACS server network of one or more facilities. In at least one embodiment, system 1400 can be configured to access and reference data from a PACS server to perform operations such as training machine learning models, deploying machine learning models, image processing, inferencing, and / or other operations.

[0152] In at least one embodiment, a software tier can be implemented as a secure, encrypted, and / or authenticated API through which applications or containers can be invoked (e.g., called) from an external environment (e.g., facility 1302). In at least one embodiment, an application can then call or execute one or more services 1320 to perform computing, AI, or visualization tasks associated with a respective application, and software 1318 and / or services 1320 can utilize hardware 1322 to perform processing tasks in an efficient and effective manner.

[0153] In at least one embodiment, deployment system 1306 can execute deployment pipelines 1410. In at least one embodiment, deployment pipelines 1410 can include any number of applications that can be sequential, non-sequential, or otherwise applied to imaging data (and / or other data types) - including AI assisted annotation - generated by imaging devices, sequencing devices, genomics devices, etc., as described above. In at least one embodiment, a deployment pipeline 1410 for an individual device can be referred to as a virtual instrument for a device (e.g., a virtual ultrasound instrument, a virtual CT scan instrument, a virtual sequencing instrument, etc.), as described herein. In at least one embodiment, there can be more than one deployment pipeline 1410 for a single device, depending on information desired from data generated by a device. In at least one embodiment, where anomalies are desired to be detected from MRI machines, there can be a first deployment pipeline 1410, and where image enhancement is desired from output of MRI machines, there can be a second deployment pipeline 1410.

[0154] In at least one embodiment, image generation applications can include processing tasks that include use of machine learning models. In at least one embodiment, a user can wish to use their own machine learning model, or select a machine learning model from model registry 1324. In at least one embodiment, a user can implement their own machine learning model or select a machine learning model for inclusion in an application that performs a processing task. In at least one embodiment, applications can be selectable and customizable, and by defining a construction of an application, deployment and implementation of an application for a particular user is presented as a more seamless user experience. In at least one embodiment, by leveraging other features of system 1400 (e.g., services 1320 and hardware 1322), deployment pipelines 1410 can be more user friendly, provide easier integration, and produce more accurate, efficient, and timely results.

[0155] In at least one embodiment, deployment system 1306 can include a user interface 1414 (e.g., a graphical user interface, a web interface, etc.) that can be used to select applications to include in deployment pipelines 1410, arrange applications, modify or change applications or parameters or constructions thereof, use and interact with deployment pipelines 1410 during setup and / or deployment, and / or otherwise interact with deployment system 1306. In at least one embodiment, although not shown with respect to training system 1304, user interface 1414 (or a different user interface) can be used to select models for use in deployment system 1306, for selecting models for training or retraining in training system 1304, and / or for otherwise interacting with training system 1304.

[0156] In at least one embodiment, in addition to application orchestration system 1428, pipeline manager 1412 can be used to manage interactions between applications or containers of deployment pipeline 1410 and services 1320 and / or hardware 1322. In at least one embodiment, pipeline manager 1412 can be configured to facilitate interactions from application to application, from application to service 1320, and / or from application or service to hardware 1322. In at least one embodiment, although illustrated as included in software 1318, this is not intended to be limiting, and in some examples, pipeline manager 1412 can be included in services 1320. In at least one embodiment, application orchestration system 1428 (e.g., Kubernetes, DOCKER, etc.) can include a container orchestration system that can group applications into containers as logical units for orchestration, management, scaling, and deployment. In at least one embodiment, by associating applications (e.g., rebuilt applications, split applications, etc.) from deployment pipeline 1410 with individual containers, each application can execute in a self-contained environment (e.g., at kernel level) to improve speed and efficiency.

[0157] In at least one embodiment, each application and / or container (or image thereof) can be separately developed, modified, and deployed (e.g., a first user or developer can develop, modify, and deploy a first application, a second user or developer can develop, modify, and deploy a second application separate from first user or developer), which can allow for focus and attention to tasks of a single application and / or container without being impeded by tasks of another application or container. In at least one embodiment, pipeline manager 1412 and application coordination system 1428 can facilitate communication and collaboration between different containers or applications. In at least one embodiment, as long as intended inputs and / or outputs of each container or application are known to system (e.g., based on construction of application or container), application coordination system 1428 and / or pipeline manager 1412 can facilitate communication between and among each application or container and sharing of resources. In at least one embodiment, as one or more applications or containers in deployment pipeline 1410 can share same services and resources, application coordination system 1428 can coordinate, load balance, and determine sharing of services or resources between and among various applications or containers. In at least one embodiment, a scheduler can be used to track resource needs of applications or containers, current or planned use of these resources, and resource availability. Accordingly, in at least one embodiment, a scheduler can allocate resources to different applications and among and between applications, taking into account needs and availability of system. In some examples, a scheduler (and / or other components of application coordination system 1428) can determine resource availability and distribution based on constraints imposed on system (e.g., user constraints), such as quality of service (QoS), urgency of data output (e.g., to determine whether to perform real-time processing or delayed processing), etc.

[0158] In at least one embodiment, services 1320 utilized by and shared by applications or containers in deployment system 1306 can include compute services 1416, AI services 1418, visualization services 1420, and / or other service types. In at least one embodiment, applications can invoke (e.g., execute) one or more services 1320 to perform processing operations for applications. In at least one embodiment, applications can utilize compute services 1416 to perform supercomputing or other high performance computing (HPC) tasks. In at least one embodiment, one or more compute services 1416 can be utilized to perform parallel processing (e.g., using parallel computing platform 1430) to process data substantially simultaneously by one or more applications and / or one or more tasks of a single application. In at least one embodiment, parallel computing platform 1430 (e.g., NVIDIA’s CUDA) can implement general purpose computing on GPUs (GPGPU) (e.g., GPU 1422). In at least one embodiment, a software layer of parallel computing platform 1430 can provide access to a virtual instruction set and parallel computing elements of a GPU to execute compute kernels. In at least one embodiment, parallel computing platform 1430 can include memory, and in some embodiments, can share memory between and among multiple containers, and / or between and among different processing tasks within a single container. In at least one embodiment, inter-process communication (IPC) calls can be generated for multiple containers and / or multiple processes within a container to use the same data from a shared memory segment of parallel computing platform 1430 (e.g., where multiple different stages of an application or applications are processing the same information). In at least one embodiment, rather than copying data and moving data to different locations in memory (e.g., read / write operations), the same data in the same location in memory can be used for any number of processing tasks (e.g., at same time, at different times, etc.). In at least one embodiment, as data is used to generate new data as a result of processing, this information of new locations of data can be stored and shared between various applications. In at least one embodiment, locations of data, as well as locations of updated or modified data, can be part of a definition of how to understand a payload in a container.

[0159] In at least one embodiment, AI services 1418 can be utilized to perform inferencing services for executing machine learning models associated with applications (e.g., task is to perform one or more processing tasks for an application). In at least one embodiment, AI services 1418 can utilize AI system 1424 to execute machine learning models (e.g., neural networks such as CNNs) for segmentation, reconstruction, object detection, feature detection, classification, and / or other inferencing tasks. In at least one embodiment, an application of deployment pipeline 1410 can use one or more output models 1316 from training system 1304 and / or other models of an application to perform inferencing on imaging data. In at least one embodiment, two or more examples of inferencing can be available using application coordination system 1428 (e.g., a scheduler). In at least one embodiment, a first category can include a high priority / low latency path, which can implement a higher service level agreement, such as for performing inferencing on urgent requests in emergency situations, or for radiologists during a diagnosis process. In at least one embodiment, a second category can include a standard priority path, which can be used for requests that can not be urgent or can have analysis performed at a later time. In at least one embodiment, application coordination system 1428 can allocate resources (e.g., services 1320 and / or hardware 1322) based on a priority path for different inferencing tasks of AI services 1418.

[0160] In at least one embodiment, shared storage can be installed to AI services 1418 in system 1400. In at least one embodiment, shared storage can operate as a cache (or other storage device type) and can be used to process inference requests from applications. In at least one embodiment, when an inference request is submitted, a set of API instances of deployment system 1306 can receive the request and can select one or more instances (e.g., for best fit, for load balancing, etc.) to process the request. In at least one embodiment, to process the request, the request can be input into a database, a machine learning model can be located from model registry 1324 if not already in cache, a validation step can ensure that appropriate machine learning model is loaded into cache (e.g., shared storage), and / or a copy of the model can be saved to cache. In at least one embodiment, if an application has not already been running or there are not enough instances of an application, a scheduler (e.g., of pipeline manager 1412) can be used to start the application referenced in the request. In at least one embodiment, if an inference server has not already been started to execute the model, an inference server can be started. Each model can start any number of inference servers. In at least one embodiment, in a pull model of clustering inference servers, a model can be cached whenever load balancing is favorable. In at least one embodiment, inference servers can be statically loaded into respective distributed servers.

[0161] In at least one embodiment, inference can be performed using inference servers running in containers. In at least one embodiment, an instance of an inference server can be associated with a model (and optionally multiple versions of a model). In at least one embodiment, if an instance of an inference server does not exist when a request to perform inference on a model is received, a new instance can be loaded. In at least one embodiment, when an inference server is started, a model can be passed to the inference server so that the same container can be used to service different models as long as the inference server is run as a different instance.

[0162] In at least one embodiment, during application execution, an inference request for a given application can be received and a container (e.g., an instance hosting an inference server) can be loaded (if not already loaded) and a launcher can be invoked. In at least one embodiment, pre-processing logic in a container can load, decode, and / or perform any additional pre-processing on incoming data (e.g., using CPU and / or GPU). In at least one embodiment, once data is ready for inference, a container can infer on data as needed. In at least one embodiment, this can include a single inference call on one image (e.g., a hand X-ray), or can require inference on hundreds of images (e.g., a chest CT). In at least one embodiment, an application can summarize results before completion, which can include, without limitation, a single confidence score, a pixel-level segmentation, a voxel-level segmentation, generating a visualization, or generating text to summarize results. In at least one embodiment, different priorities can be assigned for different models or applications. For example, some models can have real-time (TAT less than 1 minute) priority, while other models can have lower priority (e.g., TAT less than 10 minutes). In at least one embodiment, model execution time can be measured from a requesting authority or entity, and can include cooperative network traversal time as well as execution time of an inference service.

[0163] In at least one embodiment, transfer of requests between service 1320 and inference applications can be hidden behind a software development kit (SDK), and robust transfer can be provided through queues. In at least one embodiment, requests will be placed in queues through an API for individual application / tenant ID combinations, and SDK will pull requests from queues and provide requests to applications. In at least one embodiment, a name of a queue can be provided in an environment from which SDK will pick up queues. In at least one embodiment, asynchronous communication through queues can be useful because it can allow any instance of an application to pick up work when it is available. Results can be transferred back through queues to ensure no data loss. In at least one embodiment, queues can also provide an ability to split work, because highest priority work can go into a queue that connects to most instances of an application, while lowest priority work can go into a queue that connects to a single instance that processes tasks in order of receipt. In at least one embodiment, an application can run on GPU-accelerated instances that are spawned in cloud 1426, and inference service can perform inference on GPUs.

[0164] In at least one embodiment, visualization service 1420 can be utilized to generate visualizations for viewing application and / or deployment pipeline 1410 output. In at least one embodiment, visualization service 1420 can utilize GPU 1422 to generate visualizations. In at least one embodiment, visualization service 1420 can implement rendering effects such as ray tracing to generate higher quality visualizations. In at least one embodiment, visualizations can include, without limitation, 2D image rendering, 3D volume rendering, 3D volume reconstruction, 2D tomographic slices, virtual reality displays, augmented reality displays, etc. In at least one embodiment, virtualization environment can be used to generate virtual interactive displays or environments (e.g., virtual environments) for interaction by system users (e.g., doctors, nurses, radiologists, etc.). In at least one embodiment, visualization service 1420 can include internal visualizers, movie and / or other rendering or image processing capabilities or functionality (e.g., ray tracing, rasterization, internal optics, etc.).

[0165] In at least one embodiment, hardware 1322 can include GPU 1422, AI system 1424, cloud 1426, and / or any other hardware used to execute training system 1304 and / or deployment system 1306. In at least one embodiment, GPU 1422 (e.g., NVIDIA’s TESLA and / or QUADRO GPUs) can include any number of GPUs that can be used to perform processing tasks for any features or functionality of computing service 1416, AI service 1418, visualization service 1420, other services, and / or software 1318. For example, for AI service 1418, GPU 1422 can be used to perform pre-processing on imaging data (or other data types used by machine learning models), perform post-processing on outputs of machine learning models, and / or perform inferencing (e.g., to execute machine learning models). In at least one embodiment, cloud 1426, AI system 1424, and / or other components of system 1400 can use GPU 1422. In at least one embodiment, cloud 1426 can include a GPU-optimized platform for deep learning tasks. In at least one embodiment, AI system 1424 can use GPUs, and one or more AI systems 1424 can be used to perform cloud 1426 (or at least portions of tasks that are deep learning or inferencing). Likewise, although hardware 1322 is shown as discrete components, this is not intended to be limiting, and any component of hardware 1322 can be combined with, or utilized by, any other component of hardware 1322.

[0166] In at least one embodiment, AI system 1424 can include a purpose-built computing system (e.g., a supercomputer or HPC) configured for inferencing, deep learning, machine learning, and / or other artificial intelligence tasks. In at least one embodiment, AI system 1424 (e.g., NVIDIA’s DGX) can include software (e.g., a software stack) that can use multiple GPUs 1422 to perform split-GPU optimizations in addition to CPUs, RAM, storage, and / or other components, features, or functions. In at least one embodiment, one or more AI systems 1424 can be implemented in cloud 1426 (e.g., in a data center) to perform some or all of AI-based processing tasks of system 1400.

[0167] In at least one embodiment, cloud 1426 can include a GPU-accelerated infrastructure (e.g., NVIDIA’s NGC) that can provide a GPU-optimized platform for performing processing tasks of system 1400. In at least one embodiment, cloud 1426 can include AI system 1424 for performing one or more AI-based tasks of system 1400 (e.g., as a hardware abstraction and scaling platform). In at least one embodiment, cloud 1426 can integrate with application orchestration system 1428 that utilizes multiple GPUs to enable seamless scaling and load balancing between and among applications and services 1320. In at least one embodiment, cloud 1426 can be responsible for performing at least some services 1320 of system 1400, including compute services 1416, AI services 1418, and / or visualization services 1420, as described herein. In at least one embodiment, cloud 1426 can perform batched inferencing (e.g., perform NVIDIA’s TENSORRT), provide an accelerated parallel computing API and platform 1430 (e.g., NVIDIA’s CUDA), perform application orchestration system 1428 (e.g., KUBERNETES), provide a graphics rendering API and platform (e.g., for ray tracing, 2D graphics, 3D graphics, and / or other rendering techniques to produce higher quality cinematic effects), and / or can provide other functionality for system 1400.

[0168] Figure 15A A dataflow graph for process 1500 for training, retraining, or updating a machine learning model is shown, in accordance with at least one embodiment. In at least one embodiment, process 1500 can be performed using, as a non-limiting example, NVIDIA’s Figure 14The process 1500 can be performed by the system 1400. In at least one embodiment, the process 1500 can utilize the services 1320 and / or hardware 1322 of the system 1400, as described herein. In at least one embodiment, the refined model 1512 generated by the process 1500 can be executed by the deployment system 1306 for one or more containerized applications in the deployment pipeline 1410.

[0169] In at least one embodiment, the model training 1314 can include retraining or updating the initial model 1504 (e.g., a pre-trained model) using new training data (e.g., new input data (such as the customer dataset 1506), and / or new ground truth data associated with the input data). In at least one embodiment, to retrain or update the initial model 1504, the output or loss layers of the initial model 1504 can be reset or deleted, and / or replaced with updated or new output or loss layers. In at least one embodiment, the initial model 1504 can have previously fine-tuned parameters (e.g., weights and / or biases) that are retained from previous training, so the training or retraining 1314 can not take as long or require as much processing as training a model from scratch. In at least one embodiment, during the model training 1314, by resetting or replacing the output or loss layers of the initial model 1504, the parameters of the new dataset can be updated and re-tuned as predictions are generated on the new customer dataset 1506 (e.g., image data 1308) based on loss calculations associated with the accuracy of the output or loss layers. Figure 13 In at least one embodiment, the model training 1314 can include retraining or updating the initial model 1504 (e.g., a pre-trained model) using new training data (e.g., new input data (such as the customer dataset 1506), and / or new ground truth data associated with the input data). In at least one embodiment, to retrain or update the initial model 1504, the output or loss layers of the initial model 1504 can be reset or deleted, and / or replaced with updated or new output or loss layers. In at least one embodiment, the initial model 1504 can have previously fine-tuned parameters (e.g., weights and / or biases) that are retained from previous training, so the training or retraining 1314 can not take as long or require as much processing as training a model from scratch. In at least one embodiment, during the model training 1314, by resetting or replacing the output or loss layers of the initial model 1504, the parameters of the new dataset can be updated and re-tuned as predictions are generated on the new customer dataset 1506 (e.g., image data 1308) based on loss calculations associated with the accuracy of the output or loss layers.

[0170] In at least one embodiment, the pre-trained model 1406 can be stored in a data store or registry (e.g., the data store 1302) for use by the deployment system 1306 to deploy the model 1406 to one or more containerized applications in the deployment pipeline 1410. Figure 13model registry 1324). In at least one embodiment, pre-trained models 1406 can have been trained, at least in part, at one or more facilities other than the facility executing process 1500. In at least one embodiment, to protect the privacy and rights of patients, subjects, or customers of different facilities, pre-trained models 1406 can have been trained locally using locally generated customer or patient data. In at least one embodiment, pre-trained models 1406 can be trained using cloud 1426 and / or other hardware 1322, but confidential, privacy protected patient data can not be transferred to, used by, or accessed by any component of cloud 1426 (or other non-local hardware). In at least one embodiment, if pre-trained models 1406 are trained using patient data from more than one facility, pre-trained models 1406 can have been individually trained for each facility before training on patient or customer data from another facility. In at least one embodiment, customer or patient data from any number of facilities can be used to train pre-trained models 1406 locally and / or externally, such as in a data center or other cloud computing infrastructure, for example, in cases where customer or patient data has been de-identified (e.g., by waiver, for experimental use, etc.), or where customer or patient data is included in a public dataset.

[0171] In at least one embodiment, when selecting an application to use in deployment pipeline 1410, a user can also select a machine learning model for use with the particular application. In at least one embodiment, a user can not have a model to use, so the user can select a pre-trained model 1406 to use with the application. In at least one embodiment, pre-trained models 1406 can not be optimized for generating accurate results on a customer dataset 1506 of a user’s facility (e.g., based on patient diversity, demographics, types of medical imaging devices used, etc.). In at least one embodiment, pre-trained models 1406 can be updated, retrained, and / or fine-tuned for use at individual facilities before being deployed into deployment pipeline 1410 for use with one or more applications.

[0172] In at least one embodiment, a user can select a pre-trained model 1406 to update, retrain, and / or fine-tune, and the pre-trained model 1406 can be referred to as an initial model 1504 for training system 1304 in process 1500. In at least one embodiment, a customer dataset 1506 (e.g., imaging data, genomic data, sequencing data, or other data types generated by equipment at a facility) can be used to perform model training 1314 (which can include, without limitation, transfer learning) on initial model 1504 to generate a refined model 1512. In at least one embodiment, ground truth data corresponding to customer dataset 1506 can be generated by training system 1304. In at least one embodiment, ground truth data can be generated at least in part by a clinician, scientist, physician, practitioner at a facility (e.g., as labeled clinical data 1312 in FIG. 13B). Figure 13

[0173] In at least one embodiment, AI-assisted annotation 1310 can be used in some examples to generate ground truth data. In at least one embodiment, AI-assisted annotation 1310 (e.g., implemented using an AI-assisted annotation SDK) can utilize a machine learning model (e.g., a neural network) to generate suggested or predicted ground truth data for a customer dataset. In at least one embodiment, a user 1510 can use annotation tools within a user interface (graphical user interface (GUI)) on computing device 1508.

[0174] In at least one embodiment, user 1510 can interact with GUI via computing device 1508 to edit or fine-tune annotations or automated annotations. In at least one embodiment, a polygon editing feature can be used to move vertices of a polygon to more precise or fine-tuned locations.

[0175] In at least one embodiment, once customer dataset 1506 has associated ground truth data, the ground truth data (e.g., from AI-assisted annotation, manual labeling, etc.) can be used during model training 1314 to generate refined model 1512. In at least one embodiment, customer dataset 1506 can be applied to initial model 1504 any number of times, and ground truth data can be used to update parameters of initial model 1504 until an acceptable level of accuracy is reached for refined model 1512. In at least one embodiment, once refined model 1512 is generated, refined model 1512 can be deployed within one or more deployment pipelines 1410 at a facility for performing one or more processing tasks with respect to medical imaging data.

[0176] ​In at least one embodiment, a refined model 1512 can be uploaded to pre-trained models 1406 in model registry 1324 for selection by another facility. In at least one embodiment, his process can be completed at any number of facilities such that refined model 1512 can be further refined any number of times on new data sets to generate more general purpose models.

[0177] Figure 15B is an example illustration of a client-server architecture 1532 for augmenting annotation tools with pre-trained annotation models, in accordance with at least one embodiment. In at least one embodiment, an AI-assisted annotation tool 1536 can be instantiated based on client-server architecture 1532. In at least one embodiment, annotation tools 1536 in an imaging application can assist radiologists, for example, in identifying organs and abnormalities. In at least one embodiment, an imaging application can include software tools that help a user 1510 identify a few extreme points on a particular organ of interest in a raw image 1534 (e.g., in a 3D MRI or CT scan), for example, and receive automatic annotation results for all 2D slices of a particular organ. In at least one embodiment, results can be stored as training data 1538 in a data store and used as ground truth data for training, for example, but not by way of limitation. In at least one embodiment, when a computing device 1508 sends extreme points for AI-assisted annotation 1310, a deep learning model, for example, can receive that data as input and return an inference result that segments an organ or abnormality. In at least one embodiment, a pre-instantiated annotation tool (e.g., AI-assisted annotation tool 1536B in Figure 15B AI-assisted annotation tools 1536B in an imaging application) can be augmented by making API calls (e.g., API call 1544) to a server, such as an annotation helper server 1540, which can include a set of pre-trained models 1542 stored in an annotation model registry, for example. In at least one embodiment, an annotation model registry can store pre-trained models 1542 (e.g., machine learning models, such as deep learning models) that are pre-trained to perform AI-assisted annotation on a particular organ or abnormality. In at least one embodiment, these models can be further updated by using a training pipeline 1404. In at least one embodiment, as new labeled clinical data 1312 is added, a pre-installed annotation tool can be improved over time.

[0178] Such components can be used to train an inverse graphics network using a set of images generated by a generator network, where aspects of the object remain fixed while pose or view information varies between images of the set.

[0179] Other variations are within the spirit of the present disclosure. Thus, while the disclosed technology is susceptible to various modifications and alternative constructions, certain illustrated embodiments thereof are shown in the drawings and have been described above in detail. It should, however, be understood that there is no intention to limit the disclosure to the specific form or forms disclosed, but on the contrary, the intention is to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of the disclosure, as defined in the appended claims.

[0180] Unless otherwise stated or contradicted by context, the use of the terms "a" and "an" and "the" and similar referents in the context of describing the disclosed embodiments (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated by context. The terms "comprising," "having," "including," and "containing" are to be construed as open-ended terms (meaning "including, but not limited to") unless otherwise noted by context. The term "connected" (when used without modification) is to be construed as partly or wholly encompassed in, attached to, or joined together with, even if there are some intervening matters. Unless otherwise indicated herein, the reference herein to a range of values is intended merely as a shorthand method of referring individually to each separate value falling within the range, and each separate value is incorporated in the specification as if it were individually recited herein. Unless otherwise indicated or contradicted by context, the use of the term "set" (e.g., "set of items") or "subset" is to be construed as a non-empty set of one or more members. Furthermore, unless otherwise indicated or contradicted by context, the term "subset" of a corresponding set does not necessarily denote a proper subset of the corresponding set, but rather the subset and the corresponding set can be equal.

[0181] Unless explicitly indicated otherwise or contradicted by context, conjunction language such as phrases in the form "at least one of A, B, and C" or "at least one of A, B, and C" is to be construed in context as generally used to indicate items, clauses, etc., that can be A or B or C, or any non-empty subset of the set of A and B and C. For example, in the illustrative example of a set having three members, the conjunction phrases "at least one of A, B, and C" and "at least one of A, B, and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such conjunction language is generally not intended to imply that certain embodiments require the existence of at least one of A, at least one of B, and at least one of C. Additionally, unless otherwise indicated or contradicted by context, the term "plurality" denotes a state of plurality (e.g., "a plurality of items" denotes a plurality of items). The number of items in a plurality of items is at least two, but can be more if explicitly indicated or indicated by context. Furthermore, unless otherwise indicated or clear from context, the phrase "based on" means "based at least in part on" rather than "based only on."

[0182] Unless otherwise indicated herein, or otherwise clearly contradicted by context, the operations of a process described herein can be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations and / or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions to perform the operations of the processes, and are implemented as code (e.g., executable instructions, one or more computer programs or one or more applications) executing collectively on one or more processing units, by hardware or combinations thereof. In at least one embodiment, the code is stored on a computer-readable storage medium, such as a computer program product, which is readable by a computer system including one or more processors. The code is executed by the computer system to perform the operations of the processes described herein. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transitory signals (e.g., a propagating transient electric or electromagnetic transmission). In at least one embodiment, the code (e.g., executable instructions or source code) is stored on a set of one or more non-transitory computer-readable storage media (or other memory storage) having stored thereon executable instructions that, as a result of being executed by one or more processors of a computer system (i.e., as a result of being executed), cause the computer system to perform operations as described herein. In at least one embodiment, a set of non-transitory computer-readable storage media includes multiple non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media of the multiple non-transitory computer-readable storage media lack all of the code, with the multiple non-transitory computer-readable storage media collectively storing the entire code. In at least one embodiment, executable instructions are executed to cause different instructions to be executed by different processors, e.g., a non-transitory computer-readable storage medium stores instructions and a main central processing unit (“CPU”) executes some instructions, while a graphics processing unit (“GPU”) executes other instructions. In at least one embodiment, different components of a computer system have separate processors and different processors execute different subsets of the instructions.

[0183] Accordingly, in at least one embodiment, a computer system is configured to implement one or more services that individually or collectively perform operations of processes described herein, and such a computer system is configured with applicable hardware and / or software to enable implementation of the operations. Moreover, a computer system implementing at least one embodiment of the present disclosure is a single device, and in another embodiment is a distributed computer system including multiple devices operating in different manners such that the distributed computer system performs operations described herein, and such that a single device does not perform all of the operations.

[0184] The use of any and all examples, or exemplary language (e.g., "such as") provided herein, is intended merely to better illuminate embodiments of the disclosure and does not pose a limitation on the scope of the disclosure unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.

[0185] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.

[0186] In the description and claims, the terms "coupled" and "connected," along with derivatives thereof, can be used. It should be understood that these terms are not intended as synonyms for each other. Rather, in particular embodiments, "connected" or "coupled" can be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" can also mean that two or more elements are not in direct contact with each other, but yet still co-operate or interact with each other.

[0187] Unless specifically stated otherwise, it can be appreciated that throughout the specification terms such as "processing," "computing," "calculating," "determining," or the like, refer to the action and / or processes of a computer or computing system, or similar electronic

[0188] In a similar manner, the term "processor" can refer to any device or portion of a device that processes electronic data from registers and / or memory to transform that electronic data into other electronic data that can be stored in registers and / or memory. As a non-limiting example, a "processor" can be a CPU or GPU. A "computing platform" can include one or more processors. As used herein, a "software" process can include, for example, software and / or hardware entities such as tasks, threads, and intelligent agents that perform work over time. Likewise, each process can refer to multiple processes to sequentially or concurrently execute instructions, either continuously or intermittently. The terms "system" and "method" can be used interchangeably herein so long as a system can embody one or more methods and a method can be considered a system.

[0189] In this document, obtaining, accessing, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine can be referenced. Analog and digital data can be obtained, accessed, received, or inputted in a variety of ways, such as by receiving data as a parameter of a function call or a call to an application programming interface. In some implementations, the process of obtaining, accessing, receiving, or inputting analog or digital data can be accomplished by transmitting data via a serial or parallel interface. In another implementation, the process of obtaining, accessing, receiving, or inputting analog or digital data can be accomplished by transmitting data from a providing entity to an accessing entity via a computer network. Providing, outputting, transmitting, sending, or presenting analog or digital data can also be referenced. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data can be accomplished by transmitting data as an input or output parameter of a function call, a parameter of an application programming interface, or an interprocess communication mechanism.

[0190] Although the above discussion discloses example implementations of the described technology, other architectures can be utilized and are intended to fall within the scope of the present disclosure. Moreover, although a specific division of responsibilities has been defined above for purposes of discussion, various functions and responsibilities can be distributed and divided in different ways depending on circumstances.

[0191] Further, although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claims.

Claims

1. A computer-implemented method, comprising: Provide a two-dimensional image of the object as input to generate a network; The generative network is used to generate a set of view images of the object from different view representations; The set of view images and the information of the different views are provided as input to the inverse graph network; For each view image in the set, the inverse graph network is used to determine the set of 3D information; For each view image of the set, the representation of the object is rendered using the set of 3D information and the corresponding view information; The rendered representation is compared with the corresponding ground reality data to determine at least one loss value, the ground reality data being at least partially based on the annotation data; as well as One or more network parameters of the inverse graph network are adjusted, at least in part, based on the at least one loss value.

2. The computer-implemented method as described in claim 1, further comprising: At least a subset of the representation of the object rendered by the inverse graph network is provided as training data to further train the generative network.

3. The computer-implemented method as described in claim 2, further comprising: The inverse graph network and the generative network are trained together using a common loss function.

4. The computer-implemented method of claim 1, wherein the generative network is a style generative adversarial network that enables features only related to the camera view to be adjusted for generating a set of images of the view.

5. The computer-implemented method as described in claim 1, further comprising: A selection matrix is ​​used to reduce the dimensionality of the image features to be included in the underlying code for rendering the representation of the object.

6. The computer-implemented method as described in claim 5, further comprising: The representation of the object is rendered, at least in part, based on the underlying code and using a differentiable renderer.

7. The computer-implemented method of claim 5, wherein the underlying code includes camera features corresponding to the view.

8. The computer-implemented method of claim 1, wherein the three-dimensional information of the object includes at least one of the object's shape, texture, light, or background.

9. The computer-implemented method of claim 1, wherein the two-dimensional image is input into the generative network and annotated with weakly precise camera information corresponding to a subset of object features.

10. A system comprising: At least one processor; as well as Memory, comprising instructions that, when executed by at least one processor, cause the system to: Provide a two-dimensional image of the object as input to generate a network; The generative network is used to generate a set of view images of the object from different view representations; The set of view images and the information of the different views are provided as input to the inverse graph network; For each view image in the set, the inverse graph network is used to determine the set of 3D information; For each view image of the set, the representation of the object is rendered using the set of 3D information and the corresponding view information; The rendered representation is compared with the corresponding ground reality data to determine at least one loss value, the ground reality data being at least partially based on the annotation data; as well as One or more network parameters of the inverse graph network are adjusted, at least in part, based on the at least one loss value.

11. The system of claim 10, wherein when the instructions are executed, the system is further configured to: At least a subset of the representation of the object rendered by the inverse graph network is provided as training data to further train the generative network.

12. The system of claim 11, wherein when the instructions are executed, the system further causes: The inverse graph network and the generative network are trained together using a common loss function.

13. The system of claim 10, wherein the generative network is a style generative adversarial network that enables features only related to the camera view to be adjusted for generating a set of images of the view.

14. The system of claim 10, wherein when the instructions are executed, the system further causes: A selection matrix is ​​used to reduce the dimensionality of the image features to be included in the underlying code for rendering the representation of the object; and The representation of the object is rendered based at least in part on the underlying code using a differentiable renderer.

15. The system of claim 10, wherein the system comprises at least one of the following: A system used to perform graphics rendering operations; A system used to perform simulation operations; A system used to perform simulations to test or validate autonomous machine applications; A system used to perform deep learning operations; Systems implemented using edge devices; A system that merges one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.

16. A computer-implemented method, comprising: Receive two-dimensional images; as well as A three-dimensional representation of the two-dimensional image is generated using an inverse graph network, which is trained at least in part by the following operations: Using a generative network and a two-dimensional input image of the object, a set of view images of the object represented from different views is generated; For each view image in the set, the inverse graph network is used to determine the set of 3D information; For each view image of the set, the representation of the object is rendered using the set of 3D information and the information of the corresponding view; The rendered representation is compared with the corresponding ground reality data to determine at least one loss value, the ground reality data being at least partially based on the annotation data; as well as One or more network parameters of the inverse graph network are adjusted, at least in part, based on the at least one loss value.

17. The computer-implemented method of claim 16, further comprising: At least a subset of the representation of the object rendered by the inverse graph network is provided as training data to further train the generative network.

18. The computer-implemented method of claim 17, further comprising: The inverse graph network and the generative network are trained together using a common loss function.

19. The computer-implemented method of claim 16, wherein the generative network is a style generative adversarial network that enables features only related to the camera view to be adjusted for generating a set of images of the view.

20. The computer-implemented method of claim 16, wherein the two-dimensional input image is annotated with weakly precise camera information corresponding to a subset of object features.

Citation Information

Patent Citations

  • Generating labels for synthetic images using one or more neural networks

    US20220083807A1

  • Labeling images using a neural network

    US20220084204A1

  • Differentiable rendering pipeline for inverse graphics

    US20190087985A1

  • Three-Dimensional Segmentation from Two-Dimensional Intracardiac Echocardiography Imaging

    US20190261945A1

  • Neural Networks for Discovering Latent Factors from Data

    US20190354806A1