Ray tracing volumetric particle for real-time novel view synthesis
Volumetric particle representations with ray tracing and bounding volume hierarchies address the challenge of real-time high-resolution rendering of 3D scenes from novel views, supporting distorted cameras and complex lighting effects.
Patent Information
- Application Number
- JP2025065548
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-17
- Filing Date
- 2025-04-11
- Publication Date
- 2025-12-23
AI Technical Summary
Traditional methods for generating novel views of 3D models from multiple captured images struggle with real-time performance at high resolution and quality, particularly when dealing with non-pinhole cameras and higher-order lighting effects like shadows and reflections.
Utilizing volumetric particle representations aligned with scene geometry, combined with ray tracing and bounding volume hierarchies, to efficiently render images by performing hit testing on fewer volumetric particles, which support distorted cameras and lighting effects.
Enables high-quality, real-time rendering of 3D scenes from arbitrary viewpoints, including support for distorted cameras and complex lighting effects, with reduced resource requirements and latency.
Smart Images

Figure 2025186159000001_ABST
Abstract
Description
[Technical Field]
[0001] There are a variety of operations, such as computer animation and environment simulation, that may require generating an image of at least one three-dimensional (3D) model of a scene. [Background technology]
[0002] 3D models useful for such purposes can be generated by combining data from multiple captured images of a physical object. Often, it is necessary to generate an image of the model from a novel perspective, different from the captured images of the physical object. Traditional approaches to generating such novel views typically cannot achieve real-time performance at high resolution and quality. More recent approaches use rasterization with radiance field representation (NeRF), which can achieve acceptable performance at interactive speeds, but these approaches accept the drawbacks of rasterization. In particular, these approaches provide significant support for any non-pinhole cameras (e.g., cameras with fisheye lenses or other types of distortion) and rolling shutters, and do not provide support for higher-order lighting effects such as shadows and reflections. Summary of the Invention
[0003] Various embodiments according to the present disclosure will now be described with reference to the drawings. [Brief explanation of the drawings]
[0004] [Figure 1A] FIG. 1 illustrates a digital representation of an object, according to at least one embodiment. [Figure 1B] FIG. 1 illustrates a digital representation of an object, according to at least one embodiment. [Figure 1C] FIG. 1 illustrates a digital representation of an object, according to at least one embodiment. [Figure 2A] FIG. 1 illustrates a geometric mesh approximation for one or more volumetric representations, according to at least one embodiment. [Figure 2B] FIG. 1 illustrates a geometric mesh approximation for one or more volumetric representations, according to at least one embodiment. [Figure 2C] FIG. 1 illustrates a geometric mesh approximation for one or more volumetric representations, according to at least one embodiment. [Figure 3A] FIG. 1 illustrates a ray tracing for a representation of an object and values determined along the traced ray, according to at least one embodiment. [Figure 3B] FIG. 1 illustrates a ray tracing for a representation of an object and values determined along the traced ray, according to at least one embodiment. [Figure 3C] FIG. 1 illustrates a ray tracing for a representation of an object and values determined along the traced ray, according to at least one embodiment. [Figure 3D] FIG. 1 illustrates a ray tracing for a representation of an object and values determined along the traced ray, according to at least one embodiment. [Figure 3E] FIG. 1 illustrates a ray tracing for a representation of an object and values determined along the traced ray, according to at least one embodiment. [Figure 4A] FIG. 1 illustrates components of an example content generation system, according to at least one embodiment. [Figure 4B] FIG. 1 illustrates components of an example rendering pipeline, according to at least one embodiment. [Figure 5] FIG. 1 illustrates an example process for generating an image of an object or scene, possibly from a novel view, according to at least one embodiment. [Figure 6]FIG. 1 illustrates components of a distributed system that can be utilized to generate and serve content, according to at least one embodiment. [Figure 7A] FIG. 1 illustrates inference and / or training logic, according to at least one embodiment. [Figure 7B] FIG. 1 illustrates inference and / or training logic, according to at least one embodiment. [Figure 8] FIG. 1 illustrates an example data center system, according to at least one embodiment. [Figure 9] FIG. 1 illustrates a computer system according to at least one embodiment. [Figure 10] FIG. 1 illustrates a computer system according to at least one embodiment. [Figure 11] FIG. 1 illustrates at least a portion of a graphics processor according to one or more embodiments. [Figure 12] FIG. 1 illustrates at least a portion of a graphics processor according to one or more embodiments. [Figure 13] FIG. 1 is an example data flow diagram for an advanced computing pipeline, according to at least one embodiment. [Figure 14] FIG. 1 is a system diagram of an example system for training, adapting, instantiating, and deploying machine learning models in an advanced computing pipeline, according to at least one embodiment. [Figure 15A] FIG. 1 illustrates a data flow diagram for a process of training a machine learning model, according to at least one embodiment, and a client-server architecture for powering an annotation tool with a pre-trained annotation model. [Figure 15B] FIG. 1 illustrates a data flow diagram for a process of training a machine learning model, according to at least one embodiment, and a client-server architecture for powering an annotation tool with a pre-trained annotation model. DETAILED DESCRIPTION OF THE INVENTION
[0005] In the following description, various embodiments are described. For purposes of explanation, specific configurations and details are set forth to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that the embodiments may be practiced without the specific details. Furthermore, well-known features may be omitted or simplified so as not to obscure the described embodiments.
[0006] The systems and methods described herein may be used by, but are not limited to, non-autonomous vehicles (e.g., in one or more advanced driver assistance systems (ADAS)), semi-autonomous vehicles, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles hitched to one or more trailers, airships, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, trains, underwater vehicles, remotely operated vehicles such as drones, and / or other types of vehicles. Additionally, the systems and methods described herein may be used for a variety of purposes, including, by way of illustration and not limitation, machine control, machine movement, machine driving, synthetic data generation, model training or updating, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environment simulation, object or agent simulation and / or digital twinning, data center processing, conversational artificial intelligence (AI), generative AI, manipulation using one or more large language models (LLMs) or one or more vision language models (VLMs), light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, and / or other suitable uses.
[0007] The disclosed embodiments may comprise a variety of different systems, such as: automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, marine systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems incorporating one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems for performing generative AI operations, systems for performing operations using one or more LLMs or VLMs, systems for performing light transport simulations, systems for performing collaborative content creation of 3D assets, systems implemented at least partially using cloud computing resources, and / or other types of systems.
[0008] Techniques according to various illustrative embodiments provide for efficient rendering of high-quality images of a three-dimensional (3D) object or scene from various views. These views may include any suitable views, including novel views not previously represented in previously acquired or created data for the object. Rendering can be achieved in part through the use of volumetric particle representations using ray tracing. Object models can be represented using a set of volumetric particles (e.g., 2D / 3D Gaussian distributions or Lagrangian representations of color and / or other similar information) that are aligned with the underlying structure or geometry (e.g., thin structures) of the scene being rendered. Volumetric particles can be encapsulated in bounding meshes (or other proxy geometry), which can be used to efficiently construct bounding volume hierarchies (BVHs). Such techniques enable significant acceleration of graphics hardware and efficient hit detection. For a view to be rendered (e.g., a novel view), ray tracing can be performed to determine the intersection of rays with the bounding mesh of the volumetric particles (e.g., the geometric envelope around a 3D Gaussian) or proxy geometry corresponding to that view. For a given volumetric particle, once a hit with the proxy geometry is determined, the precise intersection location with the volumetric particle can be calculated (if there is a true intersection), and a distribution value (e.g., the maximum response of the Gaussian along the ray) can be calculated and returned for that ray. If the ray passes through one or more semi-transparent volumetric particles, a color value can be determined based on the values returned by such particles.In at least one embodiment, samples extracted from intersecting particles (one or more samples per particle) can be volume-rendered until a transmittance threshold (or other similar criterion) is reached. These color values can then be used to render a specified view of the scene. Such a process provides high-quality rendered images and offers numerous improvements over traditional rasterization-based approaches, including high efficiency and support for distorted cameras. Such an approach can also support evaluating gradients for the backward pass, allowing backpropagation to fit the parameters of the best-rendering set of particles to a set of training images with ground-truth poses.
[0009] Variations of this and other similar functionality may be used within the scope of various embodiments, as will be apparent to those skilled in the art in light of the teachings and suggestions contained herein.
[0010] When an image of a scene is rendered, as previously described, the rendering process may involve generating an image representation of one or more objects from a specified viewpoint. There are many ways to digitally represent an object or object model, such as using a geometric mesh or a particle cloud with color information. Other information, such as information related to material properties, may also be stored for such a representation. In some cases, a complete 3D model may be synthetically generated by a digital artist or a generative modeler. In other cases, a 3D object model may be reconstructed from a set of captured 2D images of a physical object. FIG. 1A illustrates an example view 100 showing the location of a set of 2D images 104 of a physical object 102. The images may include any suitable number of camera images (e.g., around 250), depending in part on the desired level of detail. Each of these 2D images 104 may capture the physical object 102 from a different viewpoint and from a different location. In some cases, the images may also be captured with different camera settings or under different lighting conditions. It may be desirable to capture a sufficiently large number of images from a wide variety of views to generate a sufficiently accurate 3D (or 4D) digital model or representation of the physical object 102. However, it should be understood that, if desired, a 3D model can be inferred from as few as one image (e.g., if previously encoded in the data).
[0011] The collection of 2D images 104 can then be analyzed to attempt to generate an accurate 3D digital representation. This may include preprocessing such as aligning the images, adjusting for varying camera parameters or lighting conditions, and performing noise reduction. A neural network or modeling algorithm can then analyze the data from the various images, such as attempting to extract and correlate various image features. This may include correlating the positions of extracted (or otherwise determined) particles 132 or "fitting" these particles with respect to a common coordinate system or reference frame, as illustrated in the example view 130 of FIG. 1B. In this example, the representation is a set of volumetric particles (e.g., a particle cloud) formed from multiple particles 132 with associated color values (as well as other types of values, such as surface properties, as discussed elsewhere herein). Other representations, such as meshes, can also be generated.
[0012] In at least one embodiment, a light transport simulation process, such as ray tracing, can be used with such a model to generate an image of the object model from at least one specified viewpoint. As previously discussed, this may differ from a captured or previously generated view of the corresponding object. When using a volumetric particle representation such as that shown in FIG. 1B, there may be a large number of particles to perform ray tracing and hit testing on, which may require significant amounts of time and resources. Even in the case of mesh or other representations, the amount of data to be processed may prohibit real-time performance. Therefore, techniques according to various embodiments may use a different type of object representation that can be processed much faster, such as when performing ray tracing and hit testing. One such representation involves the use of a set of volumetric particles. In at least one embodiment, volumetric particles are three-dimensional representations that may be ellipsoidal in shape. An object representation such as that shown in sample view image 160 of FIG. 1C may be composed of a set of volumetric particles 162 of different shapes and / or dimensions. These volumetric particles can be selected and oriented to align with the underlying structure or geometry of one or more objects for a scene. Each volumetric particle can contain color information in the form of a 2D Gaussian distribution, a Lagrangian distribution, or other similar representation. When a ray intersects (or passes through) a volumetric particle, its color can change based on the position and direction of the ray and can return a color similar to the color that would be returned if the ray were cast against the particle cloud of FIG. 1B.
[0013] Volumetric particles can offer several advantages over traditional point-based, mesh-based, or other similar approaches. In a first example, hit testing can be performed much more quickly because there are far fewer volumetric particles, which are the underlying particles or geometric instances (e.g., triangles) of a mesh. Volumetric particles can represent large portions of an object model, and if a cast ray does not intersect the boundary of a volumetric particle, there is no need to sample any particles within that volumetric particle for that ray. Another advantage of volumetric particles is that individual particles can contain continuous distributions, which can potentially provide fairly reliable data for any sample particle within a volumetric particle. Furthermore, using a continuous distribution representation can also reduce the presence of noise and questionable data.
[0014] In at least one embodiment, ray tracing can be performed directly on these volumetric particles. However, at least some ray tracing hardware can achieve acceleration and / or performance improvements by using a geometric representation of these volumetric particles for hit testing. The geometric representation can be defined by a small number of particles in space, thereby reducing resource requirements and the time required for hit testing. FIG. 2A shows an image view 200 of an example volumetric particle. The shading variation indicates that the color values of the interior distribution can change based on location and direction, and that the distribution can take on many different shapes or forms. A geometric representation 204 can be generated that serves as a type of bounding volume for the volumetric particle. While the geometric representation 204 includes particles outside the volumetric particle 202, the geometric representation 204 can be more lightweight and faster to use for performing hit testing or analysis. Any suitable shape can be used to represent a volumetric particle, but since volumetric particles may be substantially ellipsoidal in nature, it may be advantageous for the representative geometry to take the form of a rhombohedron or other similar geometry, which may have as many as six sides to represent the entire bounding volume of the ellipsoid.
[0015] These geometric representations 232 can be used to represent objects, as shown in view 230 of Figure 2B. Processes such as ray tracing and hit testing can be performed on such geometric representations to quickly determine regions of the object model where sampling should (or should not) be performed. It can be seen that the number of particles required to define the geometric representation 232 is significantly less than that of the particle cloud representation of Figure 1B, and is also significantly less complex than a volumetric particle representation such as that illustrated in Figure 1C.
[0016] There may be additional optimizations or representations that may be useful for specific ray tracing or processing hardware. For example, FIG. 2C shows a view 260 of the geometric representation of FIG. 2B, where each geometric representation defines a rectangular bounding volume 262 (or proxy geometry). These rectangular bounding volumes are also all aligned to a common reference frame so that their edges in the image are all either horizontal or vertical. These rectangular bounding volumes may be part of a bounding volume hierarchy (BVH). Such a representation may be advantageously used as part of a BVH ray tracing acceleration structure that can be optimized for specific ray tracing hardware, such as the RTX hardware commercially available from NVIDIA Corporation. Other similar representations may be used as appropriate.
[0017] Once an appropriate set of geometric proxies or bounding volumes has been determined, ray tracing can be performed using a configuration 300 such as that illustrated in FIG. 3A. In such a configuration, a virtual camera 302 can be positioned at a specified location with a specified orientation, thereby giving the camera a specified perspective of an object representation, such as the set of geometric proxies 302. Rays 304 can be cast with respect to this camera position to determine the colors to be used for various pixel locations in a pixel grid 306 corresponding to the rendered image. A given ray may have an intersection with, or "hit," one or more of the geometric proxies. As shown in the example view 320 of FIG. 3B, at most one point of intersection of the cast ray 304 with the geometric proxy representation 302 can be determined, and this point can be the starting point 322 along the edge of the representation where the ray intersected. As shown in Figure 3B, the top ray 304 is determined to intersect with four geometric proxies, while the bottom ray 324 is shown to intersect with three different geometric proxies. Such an approach can be used to quickly narrow down the portion of the object model for which sampling is performed for a given ray. If no geometric proxies are intersected for a given ray, then no sampling needs to be performed for that ray.
[0018] After a ray intersection is determined, sampling can be performed on volumetric particles within the intersected geometric proxy. As shown in the example view 340 of FIG. 3C, there may be multiple points 342 sampled for a given ray within the identified volumetric particle. If any of the points correspond to an opaque surface, sampling of further points along that ray is not necessary. For a given ray, as long as the previously sampled point is at least partially transparent, further points can be sampled (and further sampling for reflections, etc.). In some scenarios where a ray intersects a geometric proxy, even if there is no actual intersection with the corresponding volumetric particle, such an approach still significantly and quickly reduces the search space.
[0019] FIG. 3D shows a more detailed view 360 of an example sampling process, according to at least one embodiment. Once a volumetric particle 366 for sampling has been identified using a geometric proxy or bounding volume 364, ray tracing can be performed to analyze various sample points of the volumetric particle intersected by a cast ray 362. As shown, one or more sample points can be determined for a given ray, which may depend on the transmission characteristics of the hit point, as discussed above. The color (or other pixel value) to return for a given sample or hit point can be determined by analyzing the distribution (e.g., Gaussian, Lagrangian, linear, or other) at that point. A cross section 370 through such a representation shows the shape of the distribution 372 in terms of a range of color values. This distribution can represent the distribution of colors at different feature locations in space corresponding to the volumetric particle. For the same volumetric particle, the returned color value can depend on the location and direction of the incident ray. Thus, different angles or views may return different color values from a single volumetric particle. This gives a reasonable approximation of the number of individual feature points used to generate volumetric particles and determine an appropriate distribution. Once sampled, these values can be used for tasks such as rendering images from object or scene representations. These values can also support evaluating the gradient of a backward pass through a reconstruction or generative model, allowing backpropagation to fit the parameters of the best-rendering set of particles to a set of training images taken from ground truth poses.
[0020] Figure 3E shows an example curve for a 3D anisotropic Gaussian. In the first view 380, four rays are cast through different points in the Gaussian. The second view 382 shows a plot of the corresponding density values (as 1D Gaussians) for each ray. As shown, the density values and location of the response values vary for each ray. The third view 384 illustrates the transmittance curve for each ray cast. The transmittance gives an indication of the transparency of the surface at the corresponding hit point and is useful for understanding not only the contribution but also whether further hits need to be considered for the ray. It can also be seen that the shape of the transmittance curve, or the dropoff in transmittance, varies from location to location. The transmittance values can be used to generate a shadow map, at least in one embodiment. As mentioned previously, different directions can have similarly different curves for the same Gaussian distribution or other similar distributions. Using a Gaussian field model, this can be equated to a 1D Gaussian sum, from which analytical integrals can be calculated. The amount of obscuration experienced by these rays can be taken to be equal to the sum of the integrals of the respective 1D Gaussians over the rays, with the transmittance value corresponding to the negative exponent of the integral from the start of the corresponding ray.
[0021] Such techniques can be used to represent potentially large and complex 3D scenes using a set of volumetric particles, which can represent 3D Gaussian or Lagrangian distributions, among others. These volumetric particles can be used to quickly generate images of such scenes from arbitrary viewpoints, and potentially even novel viewpoints. Volumetric particles can also be generated using algorithms that can reduce resource requirements and latency in some cases. Such algorithms can also be used to fit these volumetric particles, such as to build such representations from captured images or other similar data of a scene. The use of ray tracing also has an advantage over other techniques in that it can support distorted and / or curved cameras (e.g., cameras with fisheye lenses or rolling shutters), which can be important for operations related to automotive applications and robotics. Ray tracing also allows for the evaluation of light along individual rays, which is important for photorealistic rendering and relighting, such as through the use of path tracing renderers. In at least one embodiment, the system can evaluate the piecewise penetration of light along rays, enabling the simulation of environmental effects (e.g., fog and smoke). The system can also represent secondary effects such as shadows, reflections, refraction, and depth of field. Incorporating these effects is important for photorealistic rendering, as well as interactions such as relighting the scene. Such a process can be scalable to large scenes, at least in part due to the availability of spatial acceleration structures, such as using bounding volume hierarchies to quickly identify intersections between rays and volumetric particles, as discussed above.
[0022] In at least one embodiment, no assumptions are made regarding the camera model used; only the use of camera parameters can be used to generate the ray to be cast. Rays can be traced against a single BVH representing the entire scene, or a combination of BVHs representing individual objects. BVHs can be constructed from the volumetric particles discussed above. For Gaussian-based volumetric particles, the response of a Gaussian kernel (or Gabor kernel, etc.) can fall off rapidly away from its center. In one or more example embodiments, the response of a generalized Gaussian kernel p(x) can be expressed as:
number
number
number
number
[0023] For example, when casting millions of rays into a scene with millions of volumetric particles, it is beneficial to efficiently determine which Gaussian τ-volumes intersect with which rays. To do this using accelerated ray tracing hardware, we can construct tight proxy bounding triangular meshes around each particle, as shown in Figure 2A. These triangular meshes can be processed using existing optimized ray-mesh intersection routines that leverage hardware-accelerated ray tracing frameworks. In at least one embodiment, the proxy geometry that envelops the Gaussian τ-envelope as tightly as possible is computed as a regular polyhedron (e.g., tetrahedron, octahedron, or icosahedron) transformed by the Gaussian translation μ, rotation R, and scaling S. Ray tracing the proxy geometry for volumetric particles allows us to discard most Gaussians whose sample response along the ray is less than τ. In contrast to conventional approaches, such approaches can tightly fit highly elongated isotropic volumetric particles, which may dominate certain operations and would otherwise incur significant computational costs.
[0024] If we can identify the volumetric particles (Gaussians in this example) that contribute to the ray, it may be appropriate to sample the values of each and integrate their contributions sequentially along the ray. A first example sampling strategy involves integrating a single sample per Gaussian. This sample may correspond to the point on the ray with the largest Gaussian response. In other words, L can be approximated as:
number
number
number
number
number
number
[0025] Another example approach may involve estimating L using multiple significant samplings, including:
number
number
[0026] Such a sampler can be computed iteratively by tracing the Gaussian from front to back. In at least one embodiment, samples can be rejected based on ωρ(o+vt) and N samples can be generated for the currently hit Gaussian. The transmittance sampling term is taken into account by considering only the closest sample along the ray.
[0027] Ray-tracing programming models allow for constraints on how ray-mesh intersections are evaluated and where calculations are performed. Therefore, ensuring an algorithm meets these constraints is crucial for high-performance processing. Specifically, this means structuring the algorithm as a combination of shader programs, such as ray-generation, closest-hit, or any-hit shaders, which can be evaluated at different times as rays are fired and intersect with primitives. A simple approach would be to use closest-hit ray casting to find all intersections along the ray in order of intersection. However, this approach results in many redundant calculations for every ray. Previous work has proposed structuring the traversal as slabs, as shown in Figure 3D, where the any-hit program collects all intersections within a fixed-width subregion of the ray. The collected intersections are then sorted and integrated in the ray-generation program. This process is repeated for each slab. This approach is limited to a fixed number of hits per slab, which can lead to inaccurate results. In contrast to previous approaches, at least one embodiment presented here consists of collecting hits and sorting them with an any-hit program. The hits are stored in a fixed-size array in the ray payload. When the array is full, the traversal is stopped by reporting the most distant hit. Integration is then performed in the ray-generation program, and subsequent cast rays collect further hits along the ray.
[0028] The technique proposed here supports situations where particles are extremely densely packed on hard surfaces, which, depending on the parameter selection, can make various conventional techniques inefficient or inaccurate. A volumetric tracing algorithm may be used, which involves tracing dynamic ray slabs from a ray generation shader. An any hit shader can be used to store and sort the K closest hits in the ray-payload buffer. Once the any hit shader determines that the K closest samples have been collected, the technique can return to ray generation to process the contributions from these samples. Tracing the next slab can resume from the end distance of the previous slab or the distance to the Kth closest sample, whichever is closer. Such a technique is important to avoid missing densely packed particles, which can be relatively common in certain scenes.
[0029] In cases where rays correspond to pixels in an image, further performance gains can be made. Rather than casting a ray for each pixel individually, it is possible to cast rays corresponding to a small tile of pixels (for example, a 2x2 tile of pixels). Evaluation can still be performed individually per pixel in the ray generation shader, and only the ray-casting Gaussian intersection in the closet-hit shader is shared across all pixels in the tile. Such an approach can result in up to a 50% performance improvement with only a negligible loss of quality for 2x2 fragment tiles.
[0030] Such an algorithm can also be used for multiple samples per Gaussian. In at least one embodiment, a sorted cache buffer of samples can be maintained within the ray-generation shader. Specifically, for each Gaussian in the K closest hitting ray payload buffer, N samples can be generated. Samples closer than the next hit can be used to update the integral. Samples farther than the next hit can be cached in the sorted buffer of samples. The cached samples can be examined before each hit evaluation. Samples closer than the next hit can be used to update the integral and then removed from the cache. When the cache buffer becomes full, the farthest samples can be discarded.
[0031] At least for Gaussian particles, operations such as pruning, cloning, and splitting can be applied to the Gaussian particles. These properties may be desirable to ensure that the model distributes particle capacity to better represent the learned scene. In one or more embodiments, since trace functions may occur in 3D space, cloning and splitting criteria that use 3D gradients instead of 2D gradients may be applied. Finally, the BVH can be reconstructed for each training iteration. This operation does not incur significant overhead and can be used to handle variations in particle volume.
[0032] FIG. 4A illustrates an example system for rendering images, video frames, or other instances of image-related content, according to at least one embodiment. Such a system can include or incorporate functionality presented herein and can generate 3D representations of objects or scenes, such as by using a sparse voxel hierarchy. In this example, images are rendered for objects and / or scenes (or other views, portions, or regions) within a virtual environment 400, but such a system can also be used to render images for semi-virtual or real environments. The virtual environment 400 may include geometry and other data representing shapes or objects within the environment, such as three-dimensional (3D) objects representing or contained in a scene occurring within the environment, including foreground objects such as people and vehicles, or background objects such as roads and buildings, among others. In at least some embodiments, at least a portion of the inserted content may be obtained from a source such as an asset repository 402 or other similar location, which may include content such as geometry, texture, and density data that can be used to render one or more objects positioned in a view of the scene. At least some of the assets may be generated using a sparse voxel architecture as discussed herein. In at least some embodiments or instances, there may be a user device 404 running a content generation or management application that allows a user to generate and / or select assets 402 to be rendered within the virtual environment 400 or to generate and / or select assets 402 to be rendered within the virtual environment 400. The user device 404 may also allow a user to control aspects of the rendered images, such as the location or pose of objects in the scene, as well as the viewpoint and other parameters of a virtual camera used in rendering the images of the virtual environment 400. After the images are rendered, they may be stored in an image repository 422 and / or provided for display on the user device or display device 424, among others.
[0033] In this example, at least one computational resource 406 is used to perform rendering or other image generation. The resource may correspond to one or more servers, which may be located, for example, locally or across at least one network, among others. Alternatively, in some embodiments, rendering may be performed at least partially on the user device 404. The computational resource 406 may obtain or receive data used for rendering, which may include geometry, attribute, texture, and / or density data about the virtual environment, objects, scenes, or assets, as well as information about the location and pose of those objects in the scene and the parameters of a virtual camera used to determine the view of the rendered scene. This information may be received by a content application 408, which may be running on the computational resource's central processing unit (CPU) 410, responsible for tasks such as gathering data, rendering images, and formatting or encoding the resulting images, among other operations. The content application may cooperate with, for example, a rendering manager 412, which may be responsible for coordinating the operation of a rendering pipeline running on the computational resources 406, which may include modules 414 or processes responsible for tasks such as geometry-related tasks (including lighting and shading tasks) or other similar tasks. Offset determinations used in an attempt to avoid self-intersections may account for error and may be implemented in these modules. In at least some embodiments, at least some rendering tasks may be performed using one or more graphics processing units (GPUs) 420A-D of the computational resources, as well as possibly one or more processors or compute instances (physical or virtual) of one or more other computational resources.
[0034] Tasks such as light transport simulation (e.g., ray tracing, path tracing, ray marching, etc.) or volumetric sampling may be performed using a single processor, such as a single GPU, or may have operations distributed across multiple GPUs 420A-D. In this example, there may be a pool or set of GPUs 420A-D, and the resource manager 418 may be responsible for at least partially allocating GPUs to perform processing for an operation. If it is desirable or beneficial to use more than one GPU, the resource manager 418 may allocate one or more GPUs with appropriate capacity or capabilities. This may include allocating the number of GPUs specified in the request or determining the number of GPUs to allocate based in part on the request. In some embodiments, the resource manager may monitor available bandwidth or memory to determine how many and which GPUs to allocate; for example, if the bandwidth impact due to the transfer of ray information is not significant, high bandwidth capacity may allow the operation to be spread across more GPUs, while in a bandwidth-constrained system, the resource manager may attempt to allocate as few GPUs as possible to reduce the number of transfer messages required.
[0035] In at least one embodiment, data partitioning may be performed, for example, by the rendering manager 412, and allocation of data to different processors may be performed by the system's resource manager 418. The resource manager may receive information from the rendering component and may select an appropriate processor from a pool of available processors 420 or processor capacity. In some embodiments, the rendering application may select the partitioning, but in other embodiments, the renderer may not have control over data partitioning, and data partitioning may be performed by a separate management component (not shown in FIG. 4A ).
[0036] FIG. 4B shows an example image generation pipeline 450 that may be used in a virtual environment 400 such as that illustrated in FIG. 4A to render one or more images, such as video frames in a sequence. In this example, pixel data 452 (which may include G-buffer data for major surfaces) for the current frame being rendered may be received as input to a surface interaction component 454 of the rendering system. The surface interaction component 454 may use this data to attempt to determine data for specific types of surface interactions (e.g., reflection, transmission, diffraction, and / or refraction) in the pixel data and provide this data to a backprojection and G-buffer patching component 456, which may perform the backpropagation discussed herein to locate corresponding points for those surface interactions and use this data to patch a G-buffer 468, thereby providing updated input for subsequent frames being rendered. The data is then provided to a light sample generation component 458 for performing light sampling, a ray tracing lighting component 460 for performing ray tracing lighting, and one or more shaders 462 that can set pixel colors for various pixels of the frame based at least in part on the determined lighting information (along with other information such as color, texture, etc.). As previously mentioned, errors can be determined from the ray traced lighting 460 and / or shader 462 components and used to determine offset values for the spawn points of secondary light rays. The results can be accumulated by an accumulation module 464 or component to generate an output frame 466 of a desired size, resolution, or format.
[0037] In at least one embodiment, shader 462 can perform the backprojection step. Once the backprojection pass is complete and the gradient surface parameters have been patched into the current G-buffer, the renderer can perform the lighting pass. Using information from the lighting pass and the lighting results from the previous frame, gradients can be calculated and filtered and used for history rejection. Such techniques can be used to calculate robust temporal gradients between the current and previous frames using a temporal denoiser for ray-traced renderers. Such backprojection-based techniques can also work through surface interactions and can work with a rasterized G-buffer. Traditional techniques for backprojection omit G-buffer patching and instead rely on raw current G-buffer samples, which also give false-positive gradients. Patching the surface parameters can eliminate false positives in the majority of cases, making the denoised image very stable, yet still responsive to lighting changes. Once the backprojection pass is complete and the gradient surface parameters have been patched into the current G-buffer, the renderer can perform the lighting pass. Using information from the illumination path and illumination results from previous frames, gradients are calculated and filtered and used for history rejection.
[0038] In at least some embodiments, components of the rendering pipeline may employ one or more machine learning (ML) models or deep neural networks (DNNs), which may include, for example, generative networks that generate image content. Machine learning may also be used to avoid self-intersections with traced paths or rays, for example, where appropriate offsets or spawn locations are inferred based on multiple sources of error discussed herein, or to avoid self-intersections while attempting to use the smallest possible offsets (to provide accurate color and lighting information) or that would otherwise introduce image artifacts.
[0039] FIG. 5 illustrates an example process 500 that can perform efficient rendering of an image of an object from a specified view, such as a novel view, according to at least one embodiment. It should be understood that for this and other processes presented herein, there may be additional, fewer, or alternative steps performed, or steps in a similar or alternative order, or at least partially parallel steps, within the scope of various embodiments, unless otherwise specified. Furthermore, while this example is discussed with respect to an object generated from multiple captured images of a physical object, it may also be used within the scope of various embodiments for other types of object representations (e.g., scenes) used to generate content that is not limited to 2D images. In this example, multiple images of at least one physical object may be obtained 502, where each image may be captured from a different viewpoint. Feature points (or other similar representation data) may be extracted from the images, and these extracted feature points may be fitted 504 to a common frame of reference to generate a point-based representation of the object. These points can be used to generate a representation of the object, i.e., composed of a set of volumetric particles, where each volumetric particle can use a 3D function, such as a Gaussian function, a Lagrangian function, or a distribution function, to represent the value of the corresponding feature point. A geometric mesh or a set of proxy geometries can be used to represent the object 506, at least for purposes of efficient hit testing and hardware acceleration. Ray tracing can be performed to determine appropriate color values (or other relevant values, including, for example, without limitation, instance or identity values and / or semantic information) to use to render an image of the object from a particular viewpoint. For a given ray, the intersection point of that ray can be determined 508 with respect to the proxy geometry (or geometric mesh) corresponding to at least one volumetric particle. Such an approach enables efficient hit testing.Based on the intersections with the proxy geometry, actual intersections of the cast ray with one or more corresponding volumetric particles can be determined 510. Response values can be determined for such actual hits with the volumetric particles. The response values can be used to determine at least one pixel value of an image of the object from the specified viewpoint 512. If further rays are determined to be cast 514, the process continues with the next ray. If no more rays are cast for this image, the color (and / or identification information, semantics) and / or pixel values from the cast ray can be provided 516 for use in generating an image of the object from the selected viewpoint. As discussed, in at least one embodiment, color values of semi-transparent points can be combined at least until a transparency threshold or other similar criterion is met.
[0040] In at least one embodiment, the volumetric particle representation is not limited to a single image but can be used to render content that may include or correspond to various types of representations of one or more objects in a scene or environment. For example, the rendered content may include video frames, streaming media, or multi-dimensional object representations that may be useful for various operations, including, but not limited to, operations related to games, animations, simulations, autonomous navigation, or virtual reality (VR) / augmented reality (AR) / enhanced reality (ER) applications, among others.
[0041] Aspects of the various techniques presented herein can be lightweight enough to be performed in real time at various locations, such as on a client device, including a personal computer or gaming console. Such processing can be performed on or for content generated on the client device or received by the client device from an external source, such as streaming data or other content received over at least one network from a cloud server 620 or a third-party service 660, among others. In some cases, at least a portion of this content processing, generation, composition, and / or determination may be performed by one of these other devices, systems, or entities and then provided to the client device (or another similar recipient) for presentation or other similar use.
[0042] 6 illustrates an example network configuration 600 that can be used to provide, generate, modify, encode, process, and / or transmit image data or other similar content. In at least one embodiment, a client device 602 can generate or receive data for a session using components of a content application 604 on the client device 602 and data stored locally on the client device. In at least one embodiment, a content application 624 executing on a server 620 (e.g., a cloud server or an edge server) can initiate a session in association with at least one client device 602, utilize user data stored in a session manager and user database 636, and have the content manager 626 determine content such as one or more digital assets (e.g., implicit and / or explicit object representations including sparse voxel grid representations, meshes, and textures) from an asset repository 634. The content manager 626, in cooperation with the rendering module 628, can generate or select objects, digital assets, or other content to be placed in a scene or other virtual environment. Views of these objects can be rendered by the rendering module 628 and provided for presentation via the client device 602. In at least one embodiment, the rendering module 628 can cooperate with a content generator 630, which can determine the image content (or other content representation) to be rendered by the rendering module 628 as part of the content provision or generated by, among other things, the sparse voxel hierarchical VAE discussed herein. A training manager 632 can be used to train any or all of the generative models used.At least a portion of the rendered content (or a representation used to render the content) is transmitted to the client device 602 using an appropriate transmission manager 622, and may be transmitted by download, streaming, or another similar transmission channel. An encoder may be used to encode and / or compress at least a portion of this data before transmitting it to the client device 602. In at least one embodiment, a client device 602 receiving such content may provide the content to a corresponding content application 604, which may additionally or alternatively include a graphical user interface 610, a content manager 612, and a rendering module 614 for providing, compositing, rendering, compositing, modifying, or using the content for presentation (or otherwise) on or by the client device 602. A decoder may also be used to decode data received over the network 640 for presentation via the client device 602, such as images and video content through a display 606 and audio, such as voice or music, through at least one audio playback device 608, such as speakers or headphones. In at least one embodiment, at least some of this content is already stored on, rendered on, or accessible to the client device 602, such that transmission over the network 640 is not required for at least some of the content, such as if the content was previously downloaded locally or stored locally on a hard drive or optical disk. In at least one embodiment, a transmission mechanism, such as data streaming, can be used to transfer this content from the server 620 or user database 636 to the client device 602.In at least one embodiment, at least some of this content may be obtained, enhanced, and / or streamed from another source, such as a third party service 660 or other client device 650, which may also include a content application 662 for generating, enhancing, or providing content. In at least one embodiment, some of this functionality may be performed using multiple computing devices or multiple processors within one or more computing devices, e.g., a combination of CPUs and GPUs.
[0043] In this example, these client devices may include any suitable computing device, which may include a desktop computer, a notebook computer, a set-top box, a streaming device, a gaming console, a smartphone, a tablet computer, a virtual reality headset, an augmented reality goggles, a wearable computer, or a smart television. Each client device may submit requests over at least one wired or wireless network, which may include the Internet, an Ethernet, a local area network (LAN), or a cellular network, among others. In this example, these requests may be submitted to an address associated with a cloud provider, which may operate or control one or more electronic resources in a cloud provider environment, which may include a data center or server farm. In at least one embodiment, the requests may be received or processed by at least one edge server located at the network edge and outside at least one security layer associated with the cloud provider environment. In this manner, allowing client devices to interact with servers in close proximity can reduce latency while improving security of resources within the cloud provider environment.
[0044] In at least one embodiment, such a system may be used to perform graphical rendering operations. In other embodiments, such a system may be used for other purposes, such as providing image or video content for testing or validating autonomous machine applications or performing deep learning operations. In at least one embodiment, such a system may be implemented using edge devices or may incorporate one or more virtual machines (VMs). In at least one embodiment, such a system may be implemented at least in part in a data center or at least in part using cloud computing resources.
[0045] Inference and Training Logic Figure 7A illustrates inference and / or training logic 715 used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 715 are provided below in conjunction with Figures 7A and / or 7B.
[0046] In at least one embodiment, the inference and / or training logic 715 may include, without limitation, code and / or data storage 701 for storing forward and / or output weights, and / or input / output data, and / or other parameters for configuring neurons or layers of a neural network that is trained and / or used to infer in one or more embodiments. In at least one embodiment, the training logic 715 may include or be coupled to code and / or data storage 701 for storing graph code or other software for controlling the timing and / or sequence of logic loaded with weights and / or other parameter information, including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code, such as graph code, loads weights or other parameter information into a processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, code and / or data storage 701 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with one or more embodiments during forward propagation of the input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 701 may be included with other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache, or system memory.
[0047] In at least one embodiment, any portion of code and / or data storage 701 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 701 may be cache memory, dynamic randomly addressable memory (“DRAM”), static randomly addressable memory (“SRAM”), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether code and / or data storage 701 is internal or external to a processor, or whether it is comprised of DRAM, SRAM, flash, or some other type of storage, may depend on the available storage on-chip versus off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of data used in neural network inference and / or training, or any combination of these factors.
[0048] In at least one embodiment, the inference and / or training logic 715 may include, without limitation, code and / or data storage 705 for storing back and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used to infer in accordance with one or more aspects of the embodiment. In at least one embodiment, the code and / or data storage 705 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with one or more aspects of the embodiment while backpropagating input / output data and / or weight parameters during training and / or inference using one or more aspects of the embodiment. In at least one embodiment, training logic 715 may include or be coupled to code and / or data storage 705 for storing graph code or other software for timing and / or sequencing control, and code and / or data storage 705 may be loaded with weights and / or other parameter information to configure logic including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code, such as graph code, loads weights or other parameter information into processor ALUs based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any portion of code and / or data storage 705 may be included with other on-chip or off-chip data storage, including processor L1, L2, or L3 caches or system memory. In at least one embodiment, any portion of code and / or data storage 705 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 705 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage.In at least one embodiment, the choice of whether code and / or data storage 705 is internal or external to the processor, for example, or whether it is comprised of DRAM, SRAM, flash, or some other type of storage, may depend on the storage available on-chip versus off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of data used in neural network inference and / or training, or any combination of these factors.
[0049] In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 may be separate storage structures. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 may be the same storage structure. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 may be partially the same storage structure or partially separate storage structures. In at least one embodiment, any portion of code and / or data storage 701 and code and / or data storage 705 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.
[0050] In at least one embodiment, the inference and / or training logic 715 may include one or more arithmetic logic units (“ALUs”) 710, including, without limitation, integer and / or floating point units, for performing logical and / or arithmetic operations based at least in part on or indicated by the training and / or inference code (e.g., graph code), the results of which may generate activations (e.g., output values from layers or neurons in a neural network) stored in activation storage 720, which are functions of input / output and / or weight parameter data stored in code and / or data storage 701 and / or code and / or data storage 705. In at least one embodiment, the activations stored in activation storage 720 are generated according to linear algebra and / or matrix-based calculations performed by ALU 710 in response to executing instructions or other code, where weight values stored in code and / or data storage 705 and / or code and / or data storage 701 are used as operands along with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data storage 705, or code and / or data storage 701, or in another storage, on-chip or off-chip.
[0051] In at least one embodiment, ALU 710 is included within one or more processors or other hardware logic devices or circuits, while in other embodiments, ALU 710 may be external to the processors or other hardware logic devices or circuits that use them (e.g., a coprocessor). In at least one embodiment, ALU 710 may be included within an execution unit of a processor, or may be included within an ALU bank that is otherwise accessible by execution units of a processor, either within the same processor or distributed among different processors of different types (e.g., central processing units, graphics processing units, fixed function units, etc.). In at least one embodiment, code and / or data storage 701, code and / or data storage 705, and activation storage 720 may be in the same processor or other hardware logic devices or circuits, while in other embodiments, they may be in different processors or other hardware logic devices or circuits, or some combination of the same processor or other hardware logic devices or circuits and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 720 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Additionally, inference and / or training code may be stored with other code accessible to the processor or other hardware logic or circuitry, and may be fetched and / or processed using the processor's fetch, decode, schedule, execute, retire, and / or other logic.
[0052] In at least one embodiment, activation storage 720 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, activation storage 720 may be completely or partially internal to or external to one or more processors or other logic circuits. In at least one embodiment, the choice of whether activation storage 720 is internal or external to a processor, or whether it is comprised of DRAM, SRAM, flash memory, or some other type of storage, for example, may depend on available on-chip versus off-chip storage, latency requirements of the training and / or inference functions being performed, batch sizes of data used in inference and / or training of neural networks, or any combination of these factors. In at least one embodiment, the inference and / or training logic 715 shown in Figure 7A may be used in conjunction with an application-specific integrated circuit ("ASIC"), such as Google's TensorFlow® processing unit, Graphcore™'s inference processing unit (IPU), or Intel Corp.'s Nervana® (e.g., "Lake Crest") processor. In at least one embodiment, the inference and / or training logic 715 shown in Figure 7A may also be used in conjunction with other hardware, such as central processing unit ("CPU") hardware, graphics processing unit ("GPU") hardware, or field programmable gate arrays ("FPGAs").
[0053] FIG. 7B illustrates inference and / or training logic 715, according to at least one or more embodiments. In at least one embodiment, the inference and / or training logic 715 may include, without limitation, hardware logic in which computational resources are dedicated to, or otherwise used only in conjunction with, weight values or other information corresponding to one or more layers of neurons in a neural network. In at least one embodiment, the inference and / or training logic 715 illustrated in FIG. 7B may be used in conjunction with an application-specific integrated circuit (ASIC), such as Google's TensorFlow® processing unit, Graphcore™'s inference processing unit (IPU), or Intel Corp.'s Nervana® (e.g., "Lake Crest") processor. In at least one embodiment, the inference and / or training logic 715 illustrated in FIG. 7B may be used in conjunction with other hardware, such as central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or field programmable gate arrays (FPGAs). In at least one embodiment, inference and / or training logic 715 includes, without limitation, code and / or data storage 701 and code and / or data storage 705, which may be used to store code (e.g., graph code), weight and / or bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In at least one embodiment illustrated in FIG. 7B , code and / or data storage 701 and code and / or data storage 705 are each associated with dedicated computational resources, such as computation hardware 702 and computation hardware 706, respectively. In at least one embodiment, computation hardware 702 and computation hardware 706 each include one or more ALUs that perform mathematical functions, such as linear algebraic functions, solely on the information stored in code and / or data storage 701 and code and / or data storage 705, respectively, with the results stored in activation storage 720.
[0054] In at least one embodiment, each of code and / or data storage 701 and 705 and corresponding computational hardware 702 and 706 corresponds to a different layer of a neural network, such that activations resulting from one storage / computation pair 701 / 702 of code and / or data storage 701 and computational hardware 702 are provided as input to a storage / computation pair 705 / 706 of code and / or data storage 705 and computational hardware 706 to reflect the conceptual organization of the neural network. In at least one embodiment, each of storage / computation pairs 701 / 702 and 705 / 706 may correspond to two or more layers of the neural network. In at least one embodiment, additional storage / computation pairs (not shown) may be included in the inference and / or training logic 715 after or in parallel with storage and computation pairs 701 / 702 and 705 / 706.
[0055] Data Center 8 illustrates an example data center 800 in which at least one embodiment may be used. In at least one embodiment, data center 800 includes a data center infrastructure layer 810, a framework layer 820, a software layer 830, and an application layer 840.
[0056] 8, in at least one embodiment, data center infrastructure layer 810 may include a resource orchestrator 812, grouped computing resources 814, and node computing resources (“node CRs”) 816(1) through 816(N), where “N” represents a positive integer. In at least one embodiment, node CRs 816(1) through 816(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid state drives or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power supply modules, cooling modules, etc. In at least one embodiment, one or more of the nodes CR 816(1)-816(N) may be a server having one or more of the computing resources described above.
[0057] In at least one embodiment, grouped computing resources 814 may include separate groups of node CRs housed within one or more racks (not shown), or multiple racks housed in a data center at various geographic locations (also not shown). Separate groups of node CRs within grouped computing resources 814 may include grouped compute resources, network resources, memory resources, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, several node CRs, including CPUs or processors, may be grouped within one or more racks to provide compute resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power supply modules, cooling modules, and network switches in any combination.
[0058] In at least one embodiment, resource orchestrator 812 may configure or otherwise control one or more nodes CR 816(1)-816(N) and / or grouped computing resources 814. In at least one embodiment, resource orchestrator 812 may include a software design infrastructure (“SDI”) management entity for data center 800. In at least one embodiment, resource orchestrator 812 may include hardware, software, or some combination thereof.
[0059] As shown in FIG. 8 , in at least one embodiment, framework layer 820 includes a job scheduler 822, a configuration manager 824, a resource manager 826, and a distributed file system 828. In at least one embodiment, framework layer 820 may include frameworks to support software 832 in software layer 830 and / or one or more applications 842 in application layer 840. In at least one embodiment, software 832 or applications 842 may each include web-based service software or applications, such as those offered by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, framework layer 820 may be a type of free and open-source software web application framework, such as, but not limited to, Apache Spark® (hereinafter “Spark”), which can use distributed file system 828 for large-scale data processing (e.g., “big data”). In at least one embodiment, job scheduler 822 may include a Spark driver to facilitate scheduling of workloads supported by various tiers of data center 800. In at least one embodiment, configuration manager 824 may be capable of configuring different tiers, such as software tier 830, as well as framework tier 820, which includes Spark and distributed file system 828 to support large-scale data processing. In at least one embodiment, resource manager 826 may be capable of managing clustered or grouped computing resources that are mapped or allocated to support distributed file system 828 and job scheduler 822. In at least one embodiment, the clustered or grouped computing resources may include grouped computing resources 814 in data center infrastructure tier 810.In at least one embodiment, resource manager 826 may manage these mappings or allocated computing resources in conjunction with resource orchestrator 812.
[0060] In at least one embodiment, software 832 included in software layer 830 may include software used by nodes CR 816(1)-816(N), grouped computing resources 814, and / or at least a portion of distributed file system 828 of framework layer 820. The one or more types of software may include, but are not limited to, internet web page searching software, email virus scanning software, database software, and streaming video content software.
[0061] In at least one embodiment, applications 842 included in application layer 840 may include one or more types of applications used by at least a portion of nodes CR 816(1)-816(N), grouped computing resources 814, and / or distributed file system 828 of framework layer 820. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive compute, and machine learning applications including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0062] In at least one embodiment, any of configuration manager 824, resource manager 826, and resource orchestrator 812 may implement any number and types of self-correcting actions based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-correcting actions may prevent data center operators of data center 800 from making potentially poor configuration decisions and from potentially avoiding underutilized and / or poorly performing portions of the data center.
[0063] In at least one embodiment, data center 800 may include tools, services, software, or other resources for training one or more machine learning models or for predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, in at least one embodiment, machine learning models may be trained by calculating weight parameters according to a neural network architecture using the software and computing resources described above with respect to data center 800. In at least one embodiment, trained machine learning models corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 800 by using weight parameters calculated by one or more techniques described herein.
[0064] In at least one embodiment, the data center may use a CPU, application specific integrated circuit (ASIC), GPU, FPGA, or other hardware to perform training and / or inference using the resources described above. Additionally, one or more of the software and / or hardware resources described above may be configured as a service to enable a user to train or perform inference on information, such as image recognition, speech recognition, or other artificial intelligence services.
[0065] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding the inference and / or training logic 715 are provided below in conjunction with Figures 7A and / or 7B. In at least one embodiment, the inference and / or training logic 715 may be used in the system Figure 8 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.
[0066] Such components can be used to generate sparse voxel-grid representations of 3D objects, including large-scale scenes.
[0067] Computer Systems 9 is a block diagram illustrating an exemplary computer system, which may be a system having interconnected devices and components formed with a processor that may include an execution unit for executing instructions, a system-on-a-chip (SoC), or some combination thereof 900, according to at least one embodiment. In at least one embodiment, computer system 900 may include components such as, without limitation, a processor 902 for using an execution unit that includes logic for executing algorithms for processing data in accordance with the present disclosure, such as in the embodiments described herein. In at least one embodiment, computer system 900 may include a processor such as the PENTIUM® processor family, Xeon™, Itanium®, XScale™ and / or StrongARM™, Intel® Core™, or Intel® Nervana™ microprocessors available from Intel Corporation of Santa Clara, California, although other systems may be used (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.). In at least one embodiment, computer system 900 may run a version of the WINDOWS® operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX® and Linux®), embedded software, and / or graphical user interfaces may also be used.
[0068] Embodiments may be used in other devices, such as portable devices and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and portable PCs. In at least one embodiment, embedded applications may include microcontrollers, digital signal processors ("DSPs"), systems-on-chips, network computers ("NetPCs"), set-top boxes, network hubs, wide area network ("WAN") switches, or any other system capable of executing one or more instructions according to at least one embodiment.
[0069] In at least one embodiment, computer system 900 may include, without limitation, a processor 902, which may include one or more execution units 908 for performing machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, computer system 900 is a single-processor desktop or server system, while in other embodiments, computer system 900 may be a multiprocessor system. In at least one embodiment, processor 902 may include, without limitation, a complex instruction set computing ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") computing microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 902 may be coupled to a processor bus 910, which may transmit digital signals between processor 902 and other components within computer system 900.
[0070] In at least one embodiment, processor 902 may include, without limitation, level 1 ("L1") internal cache memory ("cache") 904. In at least one embodiment, processor 902 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, cache memory may be external to processor 902. Other embodiments may include a combination of both internal and external cache, depending on the particular implementation and needs. In at least one embodiment, register file 906 may store different types of data in various registers, including, without limitation, integer registers, floating-point registers, status registers, and an instruction pointer register.
[0071] In at least one embodiment, processor 902 also includes an execution unit 908, including, without limitation, logic for performing integer and floating-point operations. In at least one embodiment, processor 902 may also include microcode (“u-code”) read-only memory (“ROM”) that stores microcode for certain macroinstructions. In at least one embodiment, execution unit 908 may include logic for a packed instruction set 909. In at least one embodiment, including a packed instruction set 909, along with associated circuitry for executing the instructions, in the instruction set of general-purpose processor 902 allows operations used by many multimedia applications to be performed using packed data in general-purpose processor 902. In one or more embodiments, many multimedia applications can be accelerated and run more efficiently by performing operations on packed data using the full width of the processor's data bus, thereby eliminating the need to transfer smaller units of data between the processor's data bus to perform one or more operations on one data element at a time.
[0072] In at least one embodiment, execution unit 908 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 900 may include, without limitation, memory 920. In at least one embodiment, memory 920 may be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other memory device. In at least one embodiment, memory 920 may store instructions 919 and / or data 921 represented by data signals that may be executed by processor 902.
[0073] In at least one embodiment, a system logic chip may be coupled to the processor bus 910 and the memory 920. In at least one embodiment, the system logic chip may include, without limitation, a memory controller hub (“MCH”) 916, and the processor 902 may communicate with the MCH 916 via the processor bus 910. In at least one embodiment, the MCH 916 may provide a high-bandwidth memory path 918 to the memory 920 for instruction and data storage, and for storing graphics commands, data, and textures. In at least one embodiment, the MCH 916 may route data signals between the processor 902, the memory 920, and other components of the computer system 900, and may bridge data signals between the processor bus 910, the memory 920, and the system I / O 922. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 916 may be coupled to memory 920 via a high-bandwidth memory path 918, and the graphics / video card 912 may be coupled to the MCH 916 via an Accelerated Graphics Port (“AGP”) interconnect 914.
[0074] In at least one embodiment, computer system 900 may use system I / O 922, a proprietary hub interface bus for coupling MCH 916 to I / O controller hub (“ICH”) 930. In at least one embodiment, ICH 930 may provide direct connection to several I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus for connecting peripherals to memory 920, a chipset, and processor 902. Examples may include, but are not limited to, an audio controller 929, a firmware hub (“flash BIOS”) 928, a wireless transceiver 926, data storage 924, a legacy I / O controller 923 including a user input and keyboard interface 925, a serial expansion port 927 such as a Universal Serial Bus (“USB”), and a network controller 934. Data storage 924 may comprise a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0075] In at least one embodiment, FIG. 9 illustrates a system including interconnected hardware devices or "chips," while in other embodiments, FIG. 9 may illustrate an exemplary system on a chip ("SoC"). In at least one embodiment, the devices may be interconnected with a proprietary interconnect, a standard interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 900 may be interconnected using a compute express link (CXL) interconnect.
[0076] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding the inference and / or training logic 715 are provided below in conjunction with Figures 7A and / or 7B. In at least one embodiment, the inference and / or training logic 715 may be used in system Figure 9 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.
[0077] Such components can be used to generate sparse voxel-grid representations of 3D objects, including large-scale scenes.
[0078] 10 is a block diagram illustrating an electronic device 1000 for utilizing a processor 1010, according to at least one embodiment. In at least one embodiment, electronic device 1000 may be, for example, without limitation, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop, a tablet, a mobile device, a phone, an embedded computer, or any other suitable electronic device.
[0079] In at least one embodiment, system 1000 may include a processor 1010 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices, including, without limitation, an I / C bus, a System Management Bus (“SMBus”), a Low Pin Count (“LPC”) bus, a Serial Peripheral Interface (“SPI”), a High Definition Audio (“HDA”) bus, a Serial Advance Technology Attachment (“SATA”) bus, a Universal Serial Bus (“USB”) (versions 1, 2, 3, etc.), or a Universal Asynchronous Receiver / Transmitter (“UART”) bus, or other bus or interface. In at least one embodiment, FIG. 10 depicts a system including interconnected hardware devices or "chips," although in other embodiments, FIG. 10 may depict an exemplary system-on-a-chip ("SoC"). In at least one embodiment, the devices depicted in FIG. 10 may be interconnected using proprietary interconnects, standard interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of FIG. 10 may be interconnected using a Compute Express Link (CXL) interconnect.
[0080] In at least one embodiment, FIG. 10 illustrates a display 1024, a touch screen 1025, a touch pad 1030, a Near Field Communications unit ("NFC") 1045, a sensor hub 1040, a thermal sensor 1046, an Express Chipset ("EC") 1035, a Trusted Platform Module ("TPM") 1038, a BIOS / firmware / flash memory ("BIOS,FW flash") 1022, a DSP 1060, a drive 1020, such as a solid state disk ("SSD") or hard disk drive ("HDD"), a wireless local area network unit ("WLAN") 1050, a Bluetooth unit 1052, a wireless wide area network unit ("WWAN") 1054, a Bluetooth module 1056, a Bluetooth-enabled device ("Device") 1058 ... The memory may include a GPS (Global Positioning System) 1056, a Global Positioning System (GPS) 1055, a camera such as a USB 3.0 camera ("USB 3.0 Camera") 1054, and / or a Low Power Double Data Rate ("LPDDR") memory unit ("LPDDR3") 1015, implemented, for example, to the LPDDR3 standard. Each of these components may be implemented in any suitable manner.
[0081] In at least one embodiment, other components may be communicatively coupled to the processor 1010 via the components discussed above. In at least one embodiment, an accelerometer 1041, an ambient light sensor (“ALS”) 1042, a compass 1043, and a gyroscope 1044 may be communicatively coupled to the sensor hub 1040. In at least one embodiment, a thermal sensor 1039, a fan 1037, a keyboard 1036, and a touchpad 1030 may be communicatively coupled to the EC 1035. In at least one embodiment, a speaker 1063, headphones 1064, and a microphone (“mic”) 1065 may be communicatively coupled to an audio unit (“audio codec and class d amplifier”) 1062, which may be communicatively coupled to the DSP 1060. In at least one embodiment, audio unit 1062 may include, for example, without limitation, an audio coder / decoder ("codec") and a Class D amplifier. In at least one embodiment, SIM card ("SIM") 1057 may be communicatively coupled to WWAN unit 1056. In at least one embodiment, components such as WLAN unit 1050 and Bluetooth unit 1052, as well as WWAN unit 1056, may be implemented in a Next Generation Form Factor ("NGFF").
[0082] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding the inference and / or training logic 715 are provided below in conjunction with Figures 7A and / or 7B. In at least one embodiment, the inference and / or training logic 715 may be used in system Figure 10 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.
[0083] Such components can be used to generate sparse voxel-grid representations of 3D objects, including large-scale scenes.
[0084] 11 is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, system 1100 includes one or more processors 1102 and one or more graphics processors 1108 and may be a single-processor desktop system, a multi-processor workstation system, or a server system with multiple processors 1102 or processor cores 1107. In at least one embodiment, system 1100 is a processing platform integrated into a system-on-chip (SoC) integrated circuit for use in a mobile, handheld, or embedded device.
[0085] In at least one embodiment, system 1100 may include or be incorporated into a server-based gaming platform, a game console including a game and media console, a mobile gaming console, a handheld game console, or an online game console. In at least one embodiment, system 1100 is a mobile phone, a smart phone, a tablet computing device, or a mobile internet device. In at least one embodiment, processing system 1100 may also include, be coupled to, or be integrated into a wearable device, such as a smart watch wearable device, a smart eyewear device, an augmented reality device, or a virtual reality device. In at least one embodiment, processing system 1100 is a television or set-top box device having one or more processors 1102 and a graphical interface generated by one or more graphics processors 1108.
[0086] In at least one embodiment, the one or more processors 1102 each include one or more processor cores 1107 for processing instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 1107 is configured to process a particular instruction set 1109. In at least one embodiment, the instruction set 1109 may facilitate computing via complex instruction set computing (CISC), reduced instruction set computing (RISC), or very long instruction word (VLIW). In at least one embodiment, each of the processor cores 1107 may process a different instruction set 1109, which may include instructions that facilitate emulation of other instruction sets. In at least one embodiment, the processor cores 1107 may also include other processing devices, such as a digital signal processor (DSP).
[0087] In at least one embodiment, processor 1102 includes cache memory 1104. In at least one embodiment, processor 1102 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, cache memory is shared among various components of processor 1102. In at least one embodiment, processor 1102 also uses an external cache (e.g., a level 3 (L3) cache or last level cache (LLC)) (not shown), which may be shared among processor cores 1107 using known cache coherence techniques. In at least one embodiment, processor 1102 further includes register file 1106, which may include different types of registers (e.g., integer registers, floating-point registers, status registers, and instruction pointer registers) for storing different types of data. In at least one embodiment, register file 1106 may include general-purpose registers or other registers.
[0088] In at least one embodiment, one or more processors 1102 are coupled to one or more interface buses 1110 to transmit communication signals, such as address, data, or control signals, between the processors 1102 and other components in the system 1100. In at least one embodiment, the interface bus 1110 may be a processor bus, such as, in one embodiment, a version of a Direct Media Interface (DMI) bus. In at least one embodiment, the interface bus 1110 is not limited to a DMI bus, but may include one or more Peripheral Component Interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 1102 includes an integrated memory controller 1116 and a platform controller hub 1130. In at least one embodiment, memory controller 1116 facilitates communication between memory devices and other components of system 1100, while platform controller hub (PCH) 1130 provides connectivity to I / O devices via a local I / O bus.
[0089] In at least one embodiment, memory device 1120 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase-change memory device, or any other memory device with performance suitable for serving as process memory. In at least one embodiment, memory device 1120 may operate as system memory for system 1100, storing data 1122 and instructions 1121 for use by one or more processors 1102 when executing applications or processes. In at least one embodiment, memory controller 1116 also couples to an optional external graphics processor 1112, which may communicate with one or more graphics processors 1108 within processor 1102 to perform graphics and media operations. In at least one embodiment, a display device 1111 may be connected to processor 1102. In at least one embodiment, display device 1111 may include one or more of an internal display device, such as a mobile electronic device or laptop device, or an external display device attached via a display interface (e.g., a display port, etc.). In at least one embodiment, display device 1111 may include a head-mounted display (HMD), such as a stereoscopic display device for use in virtual reality (VR) or augmented reality (AR) applications.
[0090] In at least one embodiment, platform controller hub 1130 allows peripherals to be connected to memory device 1120 and processor 1102 via a high-speed I / O bus. In at least one embodiment, the I / O peripherals include, but are not limited to, an audio controller 1146, a network controller 1134, a firmware interface 1128, a wireless transceiver 1126, a touch sensor 1125, and a data storage device 1124 (e.g., hard disk drive, flash memory, etc.). In at least one embodiment, the data storage device 1124 can be connected via a storage interface (e.g., SATA) or via a peripheral bus such as a Peripheral Component Interconnect bus (e.g., PCI, PCI Express). In at least one embodiment, the touch sensor 1125 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, wireless transceiver 1126 may be a WiFi transceiver, a Bluetooth transceiver, or a mobile network transceiver such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, firmware interface 1128 enables communication with system firmware and may be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, network controller 1134 may enable network connectivity to a wired network. In at least one embodiment, a high-performance network controller (not shown) couples to interface bus 1110. In at least one embodiment, audio controller 1146 is a multi-channel high-definition audio controller. In at least one embodiment, system 1100 includes an optional legacy I / O controller 1140 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system.In at least one embodiment, platform controller hub 1130 can also connect to one or more universal serial bus (USB) controller 1142 connected input devices, such as a keyboard and mouse 1143 combination, a camera 1144, or other USB input devices.
[0091] In at least one embodiment, instances of memory controller 1116 and platform controller hub 1130 may be integrated into a separate external graphics processor, such as external graphics processor 1112. In at least one embodiment, platform controller hub 1130 and / or memory controller 1116 may be external to one or more processors 1102. For example, in at least one embodiment, system 1100 may include external memory controller 1116 and platform controller hub 1130, which may be configured as a memory controller hub and a peripheral controller hub within a system chipset that communicates with processor 1102.
[0092] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding the inference and / or training logic 715 are provided below in conjunction with FIG. 7A and / or FIG. 7B . In at least one embodiment, some or all of the inference and / or training logic 715 may be incorporated into graphics processor 1500. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs embodied in the graphics processor. Furthermore, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than that illustrated in FIG. 7A and / or FIG. 7B . In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of the graphics processor for executing one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0093] Such components can be used to generate sparse voxel-grid representations of 3D objects, including large-scale scenes.
[0094] 12 is a block diagram of a processor 1200 having one or more processor cores 1202A-1202N, an integrated memory controller 1214, and an integrated graphics processor 1208, according to at least one embodiment. In at least one embodiment, processor 1200 may include a smaller number of additional cores, including additional core 1202N, represented by a dashed box. In at least one embodiment, each of processor cores 1202A-1202N includes one or more internal cache units 1204A-1204N. In at least one embodiment, each processor core also has access to one or more shared cache units 1206.
[0095] In at least one embodiment, internal cache units 1204A-1204N and shared cache unit 1206 represent a cache memory hierarchy within processor 1200. In at least one embodiment, cache units 1204A-1204N may include at least one level of instruction and data cache within each processor core, as well as one or more levels of shared mid-level cache, such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, where the highest level of cache before external memory is classified as LLC. In at least one embodiment, cache coherence logic maintains coherency between the various cache units 1206 and 1204A-1204N.
[0096] In at least one embodiment, processor 1200 may also include a set of one or more bus controller units 1216 and a system agent core 1210. In at least one embodiment, one or more bus controller units 1216 manage a set of peripheral buses, such as one or more PCI or PCI Express buses. In at least one embodiment, system agent core 1210 provides management functions for various processor components. In at least one embodiment, system agent core 1210 includes one or more integrated memory controllers 1214 for managing access to various external memory devices (not shown).
[0097] In at least one embodiment, one or more of processor cores 1202A-1202N include support for simultaneous multithreading. In at least one embodiment, system agent core 1210 includes components for coordinating processor cores 1202A-1202N during multithreaded processing. In at least one embodiment, system agent core 1210 may further include a power control unit (PCU), which includes logic and components for coordinating one or more power states of processor cores 1202A-1202N and graphics processor 1208.
[0098] In at least one embodiment, processor 1200 further includes a graphics processor 1208 for performing graphics processing operations. In at least one embodiment, graphics processor 1208 couples to a shared cache unit 1206 and to a system agent core 1210 that includes one or more integrated memory controllers 1214. In at least one embodiment, system agent core 1210 also includes a display controller 1211 for driving the output of the graphics processor to one or more coupled displays. In at least one embodiment, display controller 1211 may also be a separate module coupled to graphics processor 1208 via at least one interconnect or may be integrated within graphics processor 1208.
[0099] In at least one embodiment, a ring-based interconnect unit 1212 is used to couple the internal components of processor 1200. In at least one embodiment, alternative interconnect units, such as a point-to-point interconnect, a switched interconnect, or other techniques, may be used. In at least one embodiment, graphics processor 1208 couples to ring-based interconnect unit 1212 via I / O link 1213.
[0100] In at least one embodiment, I / O link 1213 represents at least one of a variety of I / O interconnects, including an on-package I / O interconnect that facilitates communication between various processor components and a high-performance embedded memory module 1218, such as an eDRAM module. In at least one embodiment, each of processor cores 1202A-1202N and graphics processor 1208 use embedded memory module 1218 as a shared last-level cache.
[0101] In at least one embodiment, processor cores 1202A-1202N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, processor cores 1202A-1202N are heterogeneous in terms of instruction set architecture (ISA), where one or more of processor cores 1202A-1202N execute a common instruction set, while one or more other of processor cores 1202A-1202N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, processor cores 1202A-1202N are heterogeneous in terms of microarchitecture, where one or more cores with relatively high power consumption are combined with one or more cores with lower power consumption. In at least one embodiment, processor 1200 can be implemented on one or more chips or as an SoC integrated circuit.
[0102] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 715 are provided below in conjunction with FIG. 7A and / or 7B. In at least one embodiment, some or all of the inference and / or training logic 715 may be incorporated into processor 1200. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs embodied in graphics processor 1208, graphics cores 1202A-1202N, or other components of FIG. 12. Furthermore, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than that illustrated in FIG. 7A and / or 7B. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALU of graphics processor 1200 for executing one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0103] Such components can be used to generate sparse voxel-grid representations of 3D objects, including large-scale scenes.
[0104] Virtualized Computing Platform FIG. 13 is an example data flow diagram of a process 1300 for generating and deploying an image processing and inference pipeline, according to at least one embodiment. In at least one embodiment, the process 1300 is deployed for use with imaging devices, processing devices, and / or other device types at one or more facilities 1302. The process 1300 may be performed within a training system 1304 and / or within a deployment system 1306. In at least one embodiment, the training system 1304 may be used to train, deploy, and implement machine learning models (e.g., neural networks, object detection algorithms, computer vision algorithms, etc.) for use in the deployment system 1306. In at least one embodiment, the deployment system 1306 may be configured to offload processing and computational resources between distributed computing environments to reduce infrastructure requirements at the facility 1302. In at least one embodiment, one or more applications in the pipeline may use or call services (e.g., inference, virtualization, computation, AI, etc.) of the deployment system 1306 during application execution.
[0105] In at least one embodiment, some of the applications used in the advanced processing and inference pipeline may use machine learning models or other AI to perform one or more processing steps. In at least one embodiment, the machine learning models may be trained at the facility 1302 using data (e.g., imaging data) 1308 generated at the facility 1302 (and stored in one or more picture archiving and communication system (PACS) servers at the facility 1302), may be trained using imaging or sequencing data 1308 from another facility, or a combination thereof. In at least one embodiment, the training system 1304 may be used to provide applications, services, and / or other resources for generating practical, deployable machine learning models for the deployment system 1306.
[0106] In at least one embodiment, the model registry 1324 may be backed by object storage that can support versioning and object metadata. In at least one embodiment, the object storage may be accessible, for example, from within a cloud platform, via a cloud storage compatibility application programming interface (API). In at least one embodiment, machine learning models in the model registry 1324 may be uploaded, listed, modified, or deleted by a system developer or partner by interacting with the API. In at least one embodiment, the API may provide access to methods that allow users with appropriate credentials to associate models with applications, thereby enabling the models to be executed as part of running a containerized instance of the application.
[0107] In at least one embodiment, training system 1304 ( FIG. 13 ) may include a scenario in which facility 1302 is training its own machine learning model or has an existing machine learning model that needs to be optimized or updated. In at least one embodiment, imaging data 1308 generated by an imaging device, a sequencing device, and / or another type of device may be received. In at least one embodiment, once imaging data 1308 is received, AI-assisted annotation 1310 may be used to assist in generating annotations corresponding to imaging data 1308, which will be used as ground truth data for the machine learning model. In at least one embodiment, AI-assisted annotation 1310 may include one or more machine learning models (e.g., convolutional neural networks (CNNs)), which may be trained to generate annotations corresponding to a particular type of imaging data 1308 (e.g., from a particular device). In at least one embodiment, AI-assisted annotation 1310 may then be used directly to generate ground truth data or may be adjusted or fine-tuned using an annotation tool. In at least one embodiment, the AI-assisted annotations 1310, the labeled data 1312, or a combination thereof may be used as ground truth data for training a machine learning model. In at least one embodiment, the trained machine learning model may be referred to as an output model 1316 and may be used by the deployment system 1306 described herein.
[0108] In at least one embodiment, the training pipeline may include a situation where facility 1302 needs a machine learning model to use in performing one or more processing tasks for one or more applications in installation system 1306, but facility 1302 may not currently have such a machine learning model (or may not have a model optimized, efficient, or effective for such purposes). In at least one embodiment, an existing machine learning model may be selected from model registry 1324. In at least one embodiment, model registry 1324 may include machine learning models trained to perform a variety of different inference tasks on imaging data. In at least one embodiment, the machine learning models in model registry 1324 may have been trained on imaging data from a facility different from facility 1302 (e.g., a facility in a remote location). In at least one embodiment, the machine learning models may have been trained on imaging data from one location, two locations, or any number of locations. In at least one embodiment, when training on imaging data from a particular location, the training may occur at that location, or at least in a manner that protects the confidentiality of the imaging data or limits its transfer off-premise. In at least one embodiment, once a model is trained or partially trained at one location, the machine learning model may be added to model registry 1324. In at least one embodiment, the machine learning model may then be retrained or updated at any number of other facilities, and the retrained or updated model may be made available in model registry 1324. In at least one embodiment, the machine learning model may then be selected from model registry 1324, referred to as output model 1316, and used in installation system 1306 to perform one or more processing tasks for one or more applications of the installation system.
[0109] In at least one embodiment, a scenario may include facility 1302 requesting a machine learning model to use in performing one or more processing tasks for one or more applications in installation system 1306, but facility 1302 may not currently have such a machine learning model (or may not have a model optimized, efficient, or effective for such purposes). In at least one embodiment, the machine learning model selected from model registry 1324 may not be fine-tuned or optimized for the imaging data 1308 generated at facility 1302 due to differences in the population, robustness of the training data used to train the machine learning model, the variety of anomalies in the training data, and / or other issues with the training data. In at least one embodiment, AI-assisted annotation 1310 may be used to assist in generating annotations corresponding to imaging data 1308 to be used as ground truth data for retraining or updating the machine learning model. In at least one embodiment, labeled data 1312 may be used as ground truth data for training the machine learning model. In at least one embodiment, retraining or updating the machine learning model may be referred to as model training 1314. In at least one embodiment, model training 1314, e.g., AI-assisted annotations 1310, labeled data 1312, or a combination thereof, may be used as ground truth data to retrain or update a machine learning model. In at least one embodiment, the trained machine learning model may be referred to as an output model 1316 and may be used by the deployment system 1306 described herein.
[0110] In at least one embodiment, deployment system 1306 may include software 1318, services 1320, hardware 1322, and / or other components, features, and functions. In at least one embodiment, deployment system 1306 may include a software “stack” whereby software 1318 may be built on top of services 1320 and may use services 1320 to perform some or all processing tasks, and services 1320 and software 1318 may be built on top of hardware 1322 and may use hardware 1322 to perform processing, storage, and / or other computational tasks of deployment system 1306. In at least one embodiment, software 1318 may include any number of different containers, where each container may perform an instantiation of an application. In at least one embodiment, each application may perform one or more processing tasks (e.g., inference, object detection, feature detection, segmentation, image enhancement, calibration, etc.) of an advanced processing and inference pipeline. In at least one embodiment, an advanced processing and inference pipeline may be defined based on a selection of different containers desired or required for processing of the imaging data, in addition to the containers that receive and configure the imaging data 1308 for use by each container and / or for use by facility 1302 after processing through the pipeline (e.g., converting output back to a usable data type). In at least one embodiment, the combination of containers in software 1318 (e.g., configuring a pipeline) may be referred to as a virtual device (described in more detail herein), which may utilize services 1320 and hardware 1322 to perform some or all of the processing tasks of the applications instantiated in the containers.
[0111] In at least one embodiment, the data processing pipeline may receive input data (e.g., imaging data 1308) in a particular format in response to an inference request (e.g., a request from a user of the deployment system 1306). In at least one embodiment, the input data may represent one or more images, videos, and / or other data representations generated by one or more imaging devices. In at least one embodiment, the data may undergo pre-processing as part of the data processing pipeline to prepare the data for processing by one or more applications. In at least one embodiment, post-processing may be performed on the output of one or more inference or other processing tasks of the pipeline to prepare output data for a subsequent application and / or for transmission and / or use by a user (e.g., in response to an inference request). In at least one embodiment, the inference tasks may be performed by one or more machine learning models, such as trained or deployed neural networks, which may include the output model 1316 of the training system 1304.
[0112] In at least one embodiment, tasks in a data processing pipeline may be encapsulated in containers, each representing a separate, fully functional instantiation of an application and a virtualized computing environment that can reference a machine learning model. In at least one embodiment, containers or applications may be published to a private (e.g., restricted access) area of a container registry (described in more detail herein), and trained or deployed models may be stored in a model registry 1324 and associated with one or more applications. In at least one embodiment, images of applications (e.g., container images) may be available in the container registry, and when selected from the container registry by a user for deployment into the pipeline, the images may be used to generate a container for instantiating the application for use on the user's system.
[0113] In at least one embodiment, a developer (e.g., a software developer, clinician, physician, etc.) may develop, publish, and store an application (e.g., as a container) to perform image processing and / or inference on the provided data. In at least one embodiment, the development, publishing, and / or storage may be performed using a software development kit (SDK) associated with the system (e.g., to ensure that the developed application and / or container conforms to or is compatible with the system). In at least one embodiment, the developed application may be tested locally (e.g., at a first facility and on data from the first facility) using the SDK that can support at least a portion of the services 1320 as a system (e.g., system 1200 of FIG. 12 ). In at least one embodiment, because a DICOM object can contain anywhere from one to hundreds of images or other types of data and there is variation in the data, the developer may be responsible for managing the extraction and preparation of the incoming data (e.g., setting up configurations for the application, building pre-processing into the application, etc.). In at least one embodiment, once verified by system 1300 (e.g., for accuracy, etc.), the application may be made available in a container registry for selection and / or implementation by a user, and one or more processing tasks may be performed on the data at the user's facility (e.g., a second facility).
[0114] In at least one embodiment, the developer may then share the application or container over a network for access and use by users of the system (e.g., system 1300 of FIG. 13 ). In at least one embodiment, the completed and validated application or container may be stored in a container registry, and the associated machine learning model may be stored in a model registry 1324. In at least one embodiment, a requesting entity issuing an inference or image processing request may browse the container registry and / or the model registry 1324 to select a desired combination of elements for inclusion in the data processing pipeline, such as an application, container, dataset, machine learning model, etc., and submit the image processing request. In at least one embodiment, the request may include input data (and in some instances, associated patient data) required to execute the request and / or may include a selection of an application and / or machine learning model to be executed in processing the request. In at least one embodiment, the request may then be passed to one or more components of the deployment system 1306 (e.g., the cloud) to execute the processing of the data processing pipeline. In at least one embodiment, processing by the deployment system 1306 may include referencing selected elements (e.g., applications, containers, models, etc.) from the container registry and / or the model registry 1324. In at least one embodiment, once results are produced by the pipeline, the results may be returned to the user for viewing (e.g., viewed locally in a viewing application suite running on a local workstation or terminal).
[0115] In at least one embodiment, services 1320 may be utilized to assist in the processing or execution of applications or containers in the pipeline. In at least one embodiment, services 1320 may include computational services, artificial intelligence (AI) services, visualization services, and / or other types of services. In at least one embodiment, services 1320 may provide functionality common to one or more applications of software 1318, whereby functionality may be abstracted into services that can be called or utilized by the applications. In at least one embodiment, the functionality provided by services 1320 may be performed dynamically and more efficiently, while also scaling well by allowing applications to process data in parallel (e.g., using parallel computing platform 1230 (FIG. 12)). In at least one embodiment, services 1320 may be shared among various applications, rather than requiring each application that shares the same functionality provided by service 1320 to have its own instance of service 1320. In at least one embodiment, services may include, by way of non-limiting example, an inference server or engine that may be used to perform detection or segmentation tasks. In at least one embodiment, a model training service may be included that can provide training and / or retraining of machine learning models. In at least one embodiment, a data augmentation service may further be included that can provide GPU-accelerated data (e.g., DICOM, RIS, CIS, REST-compliant, RPC, raw, etc.) extraction, resizing, scaling, and / or other enhancements. In at least one embodiment, a visualization service may be used that can add image rendering effects such as ray tracing, rasterization, denoising, sharpening, etc. to add realism to two-dimensional (2D) and / or three-dimensional (3D) models. In at least one embodiment, a virtual device service may be included that provides beamforming, segmentation, inference, imaging, and / or support for other applications in the virtual device pipeline.
[0116] In at least one embodiment, if services 1320 include an AI service (e.g., an inference service), the one or more machine learning models may be executed by calling (e.g., as an API call) an inference service (e.g., an inference server) to execute the machine learning models, or their processing, as part of the application's execution. In at least one embodiment, if another application includes one or more machine learning models for a segmentation task, the application may call the inference service to execute the machine learning models to perform one or more of the processing operations associated with the segmentation task. In at least one embodiment, software 1318 implementing advanced processing and inference pipelines, including a segmentation application and an anomaly detection application, may be streamlined because each application may call the same inference service to perform one or more inference tasks.
[0117] In at least one embodiment, hardware 1322 may include a GPU, a CPU, a graphics card, an AI / deep learning system (e.g., an AI supercomputer such as NVIDIA's DGX), a cloud platform, or a combination thereof. In at least one embodiment, different types of hardware 1322 may be used to provide efficient and dedicated support for software 1318 and services 1320 of deployment system 1306. In at least one embodiment, the use of GPU processing may be implemented to perform processing locally (e.g., at facility 1302), within the AI / deep learning system, in a cloud system, and / or in other processing components of deployment system 1306 to improve the efficiency, accuracy, and effectiveness of image processing and generation, etc. In at least one embodiment, software 1318 and / or services 1320 may be optimized for GPU processing related to, by way of non-limiting example, deep learning, machine learning, and / or high-performance computing. In at least one embodiment, at least a portion of the computing environment of deployment system 1306 and / or training system 1304 may be executed using GPU-optimized software (e.g., a hardware and software combination of NVIDIA's DGX system) in one or more supercomputers or high-performance computing systems in a data center. In at least one embodiment, hardware 1322 may include any number of GPUs, which may be called upon to perform data parallel processing as described herein. In at least one embodiment, the cloud platform may further include GPU processing for GPU-optimized execution of deep learning tasks, machine learning tasks, or other computing tasks. In at least one embodiment, the cloud platform (e.g., NVIDIA's NGC) may be executed using an AI / deep learning supercomputer (e.g., provided by NVIDIA's DGX system) and / or GPU-optimized software as a hardware abstraction and scaling platform.In at least one embodiment, the cloud platform may integrate an application container clustering system or orchestration system (e.g., Kubernetes) across multiple GPUs to enable seamless scaling and load balancing.
[0118] 14 is a system diagram for an example system 1400 for generating and deploying an imaging deployment pipeline according to at least one embodiment. In at least one embodiment, system 1400 may be used to implement process 1300 of FIG. 13 and / or other processes, including advanced processing and inference pipelines. In at least one embodiment, system 1400 may include training system 1304 and deployment system 1306. In at least one embodiment, training system 1304 and deployment system 1306 may be implemented using software 1318, services 1320, and / or hardware 1322, as described herein.
[0119] In at least one embodiment, system 1400 (e.g., training system 1304 and / or deployment system 1306) may be implemented in a cloud computing environment (e.g., using cloud 1426). In at least one embodiment, system 1400 may be implemented locally with respect to a healthcare service facility or as a combination of both cloud and local computing resources. In at least one embodiment, access to the cloud 1426 API may be limited to authorized users via enacted security measures or protocols. In at least one embodiment, the security protocol may include a web token, which may be signed by an authentication service (e.g., AuthN, AuthZ, Gluecon, etc.) and may have appropriate permissions. In at least one embodiment, the API of a virtual device (described herein) or other instantiation of system 1400 may be limited to a set of verified or authorized public IPs for interaction.
[0120] In at least one embodiment, the various components of system 1400 may communicate with one another using any of a variety of different types of networks, including, but not limited to, a local area network (LAN) and / or a wide area network (WAN), via wired and / or wireless communication protocols. In at least one embodiment, communications between the facility and the components of system 1400 (e.g., to send inference requests, receive results of inference requests, etc.) may be communicated via one or more data buses, wireless data protocols (Wi-Fi), wired data protocols (e.g., Ethernet), etc.
[0121] In at least one embodiment, the training system 1304 may execute a training pipeline 1404 similar to that described herein with respect to FIG. 13 . In at least one embodiment, if one or more machine learning models are to be used by the deployment system 1306 in the deployment pipeline 1410, the training pipeline 1404 may be used to train or retrain one or more (e.g., pre-trained) models and / or to implement one or more of the pre-trained models 1406 (e.g., without the need for retraining or updating). In at least one embodiment, the training pipeline 1404 may result in an output model 1316. In at least one embodiment, the training pipeline 1404 may include any number of processing steps, including, but not limited to, transforming or adapting imaging data (or other input data). In at least one embodiment, different training pipelines 1404 may be used for different machine learning models used by the deployment system 1306. In at least one embodiment, a training pipeline 1404 similar to the first example described with respect to Figure 13 may be used for a first machine learning model, a training pipeline 1404 similar to the second example described with respect to Figure 13 may be used for a second machine learning model, and a training pipeline 1404 similar to the third example described with respect to Figure 13 may be used for a third machine learning model. In at least one embodiment, any combination of tasks within training system 1304 may be used depending on what is required for each respective machine learning model. In at least one embodiment, one or more of the machine learning models may already be trained and ready for deployment, such that they do not need to undergo any processing by training system 1304 and may be implemented by deployment system 1306.
[0122] In at least one embodiment, output model 1316 and / or pre-trained model 1406 may include any type of machine learning model, depending on the implementation or embodiment. In at least one embodiment, without limitation, the machine learning models used by system 1400 may include machine learning models using linear regression, logistic regression, decision trees, support vector machines (SVMs), naive Bayes, k-nearest neighbor (Knn), k-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutional, recurrent, perceptron, long / short-term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolution, generative adversarial, liquid state machines, etc.), and / or other types of machine learning models.
[0123] In at least one embodiment, the training pipeline 1404 may include AI-assisted annotation, as described in more detail herein with respect to at least FIG. 14B . In at least one embodiment, the labeled data 1312 (e.g., traditional annotations) may be generated by any number of techniques. In at least one embodiment, the labels or other annotations may be generated within a drawing program (e.g., an annotation program), a computer-aided design (CAD) program, a labeling program, another type of program suitable for generating ground truth annotations or labels, and / or may be handwritten in some instances. In at least one embodiment, the ground truth data may be synthetically generated (e.g., generated from a computer model or rendering), realistically generated (e.g., designed and generated from real-world data), machine-automated (e.g., using feature analysis and learning to extract features from data and then generate labels), human-annotated (e.g., a labeler or annotation expert may define label locations), and / or a combination thereof. In at least one embodiment, for each instance of imaging data 1308 (or other type of data used by a machine learning model), there may be corresponding ground truth data generated by training system 1304. In at least one embodiment, AI-assisted annotation may be performed as part of deployment pipeline 1410 in addition to or instead of AI-assisted annotation included in training pipeline 1404. In at least one embodiment, system 1400 may include a multi-tiered platform, which may include a software layer (e.g., software 1318) of a diagnostic application (or other type of application) that may perform one or more medical imaging and diagnostic functions. In at least one embodiment, system 1400 may be communicatively coupled (e.g., via an encrypted link) to one or more institution's PACS server networks.In at least one embodiment, the system 1400 can be configured to access referenced data from a PACS server to perform operations such as training machine learning models, deploying machine learning models, image processing, inference, and / or other operations.
[0124] In at least one embodiment, the software layer may be implemented as a secure, encrypted, and / or authenticated API through which applications or containers may be invoked (e.g., called) from an external environment (e.g., facility 1302). In at least one embodiment, the applications may then call or execute one or more services 1320 to perform computational, AI, or visualization tasks associated with the respective application, and the software 1318 and / or services 1320 may utilize hardware 1322 to perform processing tasks in an effective and efficient manner. In at least one embodiment, communications sent to or received by the training system 1304 and the deployment system 1306 may occur using a pair of DICOM adapters 1402A, 1402B.
[0125] In at least one embodiment, the deployment system 1306 may execute an deployment pipeline 1410. In at least one embodiment, the deployment pipeline 1410 may include any number of applications, including the AI-assisted annotations described above, that may be applied sequentially, non-sequentially, or otherwise to imaging data (and / or other types of data) generated by an imaging device, a sequencing device, a genomics device, etc. In at least one embodiment, as described herein, the deployment pipeline 1410 for an individual device may be referred to as a virtual instrument for the device (e.g., a virtual ultrasound instrument, a virtual CT scan instrument, a virtual sequencing instrument, etc.). In at least one embodiment, there may be more than one deployment pipeline 1410 per device, depending on the information needed for the data generated by the device. In at least one embodiment, a first deployment pipeline 1410 may be present if anomaly detection is required for the MRI machine, and a second deployment pipeline 1410 may be present if image enhancement is required for the output of the MRI machine.
[0126] In at least one embodiment, the image generation application may include processing tasks that involve the use of machine learning models. In at least one embodiment, a user may desire to use their own machine learning model or select a machine learning model from the model registry 1324. In at least one embodiment, a user may implement their own machine learning model or select a machine learning model to include in the application to perform the processing task. In at least one embodiment, applications may be selectable and customizable, and by defining the structure of the application, the deployment and implementation of the application for a particular user may be presented as a more seamless user experience. In at least one embodiment, by utilizing other features of the system 1400, such as the services 1320 and the hardware 1322, the deployment pipeline 1410 may become even more user-friendly, provide easier integration, and produce more accurate, efficient, and timely results.
[0127] In at least one embodiment, the deployment system 1306 may include a user interface (“UI”) 1414 (e.g., a graphical user interface, a web interface, etc.), which may be used to select applications for inclusion in the deployment pipeline 1410, configure applications, modify or change applications or their parameters or structure, use and interact with the deployment pipeline 1410 during setup and / or deployment, and / or otherwise interact with the deployment system 1306. In at least one embodiment, although not shown with respect to the training system 1304, the UI 1414 (or a different user interface) may be used to select models for use in the deployment system 1306, to select models to train or retrain in the training system 1304, and / or to otherwise interact with the training system 1304.
[0128] In at least one embodiment, a pipeline manager 1412 may be used in addition to an application orchestration system 1428 to manage interactions between applications or containers in the deployment pipeline 1410 and services 1320 and / or hardware 1322. In at least one embodiment, the pipeline manager 1412 may be configured to facilitate application-to-application interactions, application-to-service interactions, and / or application or service-to-hardware interactions. While shown in at least one embodiment as being included in software 1318, this is not intended to be limiting, and in some instances, the pipeline manager 1412 may be included in services 1320. In at least one embodiment, the application orchestration system 1428 (e.g., Kubernetes, DOCKER, etc.) may include a container orchestration system that can group applications into containers as logical units for coordination, management, scaling, and deployment. In at least one embodiment, by associating applications from the deployment pipeline 1410 (e.g., a reassembly application, a segmentation application, etc.) with individual containers, each application can run in a self-contained environment (e.g., at the kernel level) for increased speed and efficiency.
[0129] In at least one embodiment, each application and / or container (or image thereof) may be developed, modified, and deployed individually (e.g., a first user or developer may develop, modify, and deploy a first application, and a second user or developer may develop, modify, and deploy a second application separately from the first user or developer), thereby allowing focused attention to be paid to the tasks of one application and / or container without being distracted by the tasks of another application or container. In at least one embodiment, communication and coordination between different containers or applications may be facilitated by pipeline manager 1412 and application orchestration system 1428. In at least one embodiment, as long as the expected inputs and / or outputs of each container or application are known by the system (e.g., based on the application's or container's structure), application orchestration system 1428 and / or pipeline manager 1412 can facilitate communication between, and sharing of resources between, each of the applications or containers. In at least one embodiment, one or more of the applications or containers in the deployment pipeline 1410 may share the same services and resources, and therefore the application orchestration system 1428 may orchestrate, load balance, and determine sharing of services or resources among the various applications or containers. In at least one embodiment, a scheduler may be used to track the resource requirements of the applications or containers, the current or planned usage of those resources, and the availability of the resources. In at least one embodiment, the scheduler may thus allocate resources to different applications and distribute them among the applications taking into account the requirements and availability of the system.In some instances, the scheduler (and / or other components of the application orchestration system 1428) may determine resource availability and allocation based on constraints imposed on the system (e.g., user constraints), such as quality of service (QoS), the urgency of needing data output (e.g., to determine whether to perform real-time or delayed processing), etc.
[0130] In at least one embodiment, services 1320 utilized and shared by applications or containers of deployment system 1306 may include compute services 1416, AI services 1418, visualization services 1420, and / or other types of services. In at least one embodiment, applications may call (e.g., execute) one or more of services 1320 to perform processing operations for the application. In at least one embodiment, compute services 1416 may be utilized by applications to perform supercomputing or other high-performance computing (HPC) tasks. In at least one embodiment, parallel processing may be performed utilizing compute services 1416 (e.g., using parallel computing platform 1430) to substantially simultaneously process data via one or more of the applications and / or to substantially simultaneously process one or more tasks of an application. In at least one embodiment, parallel computing platform 1430 (e.g., NVIDIA's CUDA) may enable general-purpose computing (GPGPU) on a GPU (e.g., GPU / Graphics 1422). In at least one embodiment, the software layer of the parallel computing platform 1430 may provide access to virtual instruction sets and parallel computing elements of a GPU to execute computational kernels. In at least one embodiment, the parallel computing platform 1430 may include memory, and in some embodiments, the memory may be shared among multiple containers and / or among different processing tasks within a container. In at least one embodiment, inter-process communication (IPC) calls may be generated so that multiple containers and / or multiple processes within a container use the same data from a shared segment of memory in the parallel computing platform 1430 (e.g., when different stages of an application or multiple applications process the same information).In at least one embodiment, rather than making copies of the data and moving the data to different locations in memory (e.g., read / write operations), the same data in the same location in memory may be used for any number of processing tasks (e.g., at the same time, different times, etc.). In at least one embodiment, as data is used and new data is generated as a result of processing, this information of the new location of the data may be stored in and shared between various applications. In at least one embodiment, the location of the data and the location of updated or modified data may be part of the definition of how the payload is understood within the container.
[0131] In at least one embodiment, AI services 1418 may be utilized to perform inference services for executing machine learning models associated with an application (e.g., tasked with performing one or more processing tasks for the application). In at least one embodiment, AI services 1418 may utilize AI systems 1424 to execute machine learning models (e.g., neural networks such as CNNs) for segmentation, reconstruction, object detection, feature detection, classification, and / or other inference tasks. In at least one embodiment, applications in deployment pipeline 1410 may perform inference on imaging data using output models 1316 from training system 1304 and / or one or more of the application's other models. In at least one embodiment, two or more instances of inference using application orchestration system 1428 (e.g., a scheduler) may be available. In at least one embodiment, a first category may include a high-priority / low-latency path that can achieve higher service level agreements, such as for performing inference on urgent requests during an emergency or for radiologists during a diagnosis. In at least one embodiment, the second category may include a standard priority path that may be used for non-urgent requests or when analysis may be performed at a later time. In at least one embodiment, the application orchestration system 1428 may allocate resources (e.g., services 1320 and / or hardware 1322) based on priority paths for different inference tasks of the AI service 1418.
[0132] In at least one embodiment, shared storage may be attached to AI services 1418 in system 1400. In at least one embodiment, the shared storage may act as a cache (or other type of storage device) and may be used to process inference requests from applications. In at least one embodiment, when an inference request is submitted, the request may be received by a set of API instances in installation system 1306, and one or more instances may be selected (e.g., for best fit, for load balancing, etc.) to process the request. In at least one embodiment, to process the request, the request may be entered into a database, a machine learning model may be identified from model registry 1324 if not already in the cache, and a validation step may ensure that an appropriate machine learning model is loaded into the cache (e.g., shared storage), and / or a copy of the model may be saved in the cache. In at least one embodiment, if the application is not already running or if sufficient instances of the application do not exist, a scheduler (e.g., pipeline manager 1412) may be used to launch the application referenced in the request. In at least one embodiment, an inference server may be started to run the model if it is not already started. Any number of inference servers may be started per model. In at least one embodiment, in a pull model where inference servers are clustered, models may be cached whenever load balancing is advantageous. In at least one embodiment, inference servers may be statically loaded onto corresponding distributed servers.
[0133] In at least one embodiment, inference may be performed using an inference server running within a container. In at least one embodiment, an instance of an inference server may be associated with a model (optionally with multiple versions of the model). In at least one embodiment, when a request to perform inference on a model is received, if an instance of the inference server does not exist, a new instance may be loaded. In at least one embodiment, when the inference server is started, the model may be passed to the inference server, such that the same container may be used to serve different models as long as the inference servers are running as different instances.
[0134] In at least one embodiment, while an application is running, an inference request may be received for a given application, a container (e.g., hosting an instance of an inference server) may be loaded (if not already loaded), and a start procedure may be called. In at least one embodiment, pre-processing logic in the container may load, decode, and / or perform any additional pre-processing on the input data (e.g., using a CPU and / or GPU). In at least one embodiment, once the data is prepared for inference, the container may perform inference on the data as needed. In at least one embodiment, this may include a single inference call for one image (e.g., a hand X-ray) or may request inference for hundreds of images (e.g., a chest CT). In at least one embodiment, the application may summarize results before completion, which may include, without limitation, a single confidence score, pixel-level segmentation, voxel-level segmentation, generation of a visualization, or generation of text to summarize the findings. In at least one embodiment, different models or applications may be assigned different priorities. For example, some models may have real-time priority (TAT<1 minute) while other models may have low priority (e.g., TAT<10 minutes). In at least one embodiment, model execution time may be measured from the requesting facility or entity and may include execution to the inference service as well as partner network traversal time.
[0135] In at least one embodiment, the transition of requests between the service 1320 and the inference application may be hidden behind a software development kit (SDK), and robust transport may be provided through queues. In at least one embodiment, requests are queued via an API for individual application / tenant ID combinations, and the SDK pulls requests from the queue and provides them to the application. In at least one embodiment, a queue name may be provided in the environment where the SDK picks up the request. In at least one embodiment, asynchronous communication via queues may be useful because it allows any instance of the application to pick up work when it becomes available. Results may be sent back via queues to ensure data is not lost. In at least one embodiment, queues may also provide the ability to segment work, as highest priority work may go to a queue with most instances of the application connected to the queue, while lowest priority work may go to a queue with one instance connected to the queue that processes tasks in the order they are received. In at least one embodiment, applications may run on GPU-accelerated instances created in the cloud 1426, and the inference service may perform inference on the GPU.
[0136] In at least one embodiment, a visualization service 1420 may be utilized to generate visualizations for viewing the output of the application and / or deployment pipeline 1410. In at least one embodiment, a GPU / Graphics 1422 may be utilized by the visualization service 1420 to generate the visualizations. In at least one embodiment, rendering effects such as ray tracing may be implemented by the visualization service 1420 to generate higher quality visualizations. In at least one embodiment, visualizations include, but are not limited to, 2D image rendering, 3D volume rendering, 3D volume reconstruction, 2D tomographic slices, virtual reality displays, augmented reality displays, etc. In at least one embodiment, a virtualized environment may be used to generate a virtual interactive display or environment (e.g., a virtual environment) for interaction by a user of the system (e.g., a doctor, nurse, radiologist, etc.). In at least one embodiment, the visualization service 1420 may include internal visualizers, cinematics, and / or other rendering or image processing capabilities or functions (e.g., ray tracing, rasterization, internal optics, etc.).
[0137] In at least one embodiment, hardware 1322 may include GPU / graphics 1422, AI system 1424, cloud 1426, and / or any other hardware used to run training system 1304 and / or deployment system 1306. In at least one embodiment, GPU / graphics 1422 (e.g., NVIDIA TESLA and / or QUADRO GPUs) may include any number of GPUs, which may be used to perform processing tasks for compute services 1416, AI services 1418, visualization services 1420, other services, and / or any features or functionality of software 1318. For example, with respect to AI services 1418, GPU / graphics 1422 may be used to perform pre-processing on imaging data (or other types of data used by machine learning models), post-processing on the output of machine learning models, and / or to perform inference (e.g., machine learning models may be run). In at least one embodiment, cloud 1426, AI system 1424, and / or other components of system 1400 may use GPU / graphics 1422. In at least one embodiment, cloud 1426 may include a GPU-optimized platform for deep learning tasks. In at least one embodiment, AI system 1424 may use GPUs, and cloud 1426, or at least a portion tasked with deep learning or inference, may be executed using one or more AI systems 1424. Thus, while hardware 1322 is illustrated as separate components, this is not intended to be limiting, and any component of hardware 1322 may be combined with or utilized by any other component of hardware 1322.
[0138] In at least one embodiment, AI system 1424 may include a specialized computing system (e.g., a supercomputer or HPC) configured for inference, deep learning, machine learning, and / or other artificial intelligence tasks. In at least one embodiment, AI system 1424 (e.g., NVIDIA's DGX) may include GPU-optimized software (e.g., a software stack) that may be executed using multiple GPUs / Graphics 1422 in addition to CPUs, RAM, storage, and / or other components, features, or functionality. In at least one embodiment, one or more AI systems 1424 may be implemented in cloud 1426 (e.g., in a data center) to perform some or all of the AI-based processing tasks of system 1400.
[0139] In at least one embodiment, cloud 1426 may include a GPU-accelerated infrastructure (e.g., NVIDIA's NGC), which may provide a GPU-optimized platform for executing processing tasks of system 1400. In at least one embodiment, cloud 1426 may include an AI system 1424 (e.g., as a hardware abstraction and scaling platform) for executing one or more of the AI-based tasks of system 1400. In at least one embodiment, cloud 1426 may utilize multiple GPUs and be integrated with application orchestration system 1428 to enable seamless scaling and load balancing among applications and services 1320. In at least one embodiment, cloud 1426 may be tasked with executing at least some of the services 1320 of system 1400, including the compute services 1416, AI services 1418, and / or visualization services 1420 described herein. In at least one embodiment, cloud 1426 may perform large and small batch inference (e.g., NVIDIA's Tensor RT implementation), may provide an accelerated parallel computing API and platform 1430 (e.g., NVIDIA's CUDA), may run an application orchestration system 1428 (e.g., KUBERNETES), may provide a graphics rendering API and platform (e.g., ray tracing, 2D graphics, 3D graphics, and / or other rendering techniques for generating high-quality cinematics), and / or may provide other functionality for system 1400.
[0140] 15A illustrates a data flow diagram of a process 1500 for training, retraining, or updating a machine learning model, according to at least one embodiment. In at least one embodiment, process 1500 may be performed using system 1400 of FIG. 14, as a non-limiting example. In at least one embodiment, process 1500 may utilize services and / or hardware described herein. In at least one embodiment, refined model 1512 generated by process 1500 may be executed by a deployment system for one or more containerized applications in a deployment pipeline.
[0141] In at least one embodiment, model training 1514 may include retraining or updating the initial model 1504 (e.g., a pre-trained model) using new training data (e.g., new input data, such as the customer dataset 1506 and / or new ground truth data associated with the input data). In at least one embodiment, to retrain or update the initial model 1504, an output or loss layer of the initial model 1504 may be reset, removed, and / or replaced with an updated or new output or loss layer. In at least one embodiment, the initial model 1504 may have previously fine-tuned parameters (e.g., weights and / or biases) remaining from previous training, such that training or retraining 1514 does not take as long or require as much processing as training a model from scratch. In at least one embodiment, during model training 1514, by having the output or loss layer of the initial model 1504 reset or replaced, parameters may be updated or retuned for the new data set based on a loss calculation associated with the accuracy of the output or loss layer in generating predictions for the new customer data set 1506.
[0142] In at least one embodiment, the pre-trained model 1506 may be stored in a data store or registry. In at least one embodiment, the pre-trained model 1506 may be trained, at least in part, at one or more facilities different from the facility performing the process 1500. In at least one embodiment, to protect the privacy and rights of patients, subjects, or customers at the different facilities, the pre-trained model 1506 may be trained on-site using customer or patient data generated on-site. In at least one embodiment, the pre-trained model 1306 may be trained using a cloud and / or other hardware, but privacy-protected sensitive patient data may not be transferred to, used by, or accessible by any components of the cloud (or other off-site hardware). In at least one embodiment, if the pre-trained model 1506 is trained using patient data from more than one facility, the pre-trained model 1506 may be trained separately for each facility and then trained on patient or customer data from another facility. In at least one embodiment, the pre-trained model 1506 may be trained on-premise and / or off-premise, such as in a data center or other cloud computing infrastructure, using customer or patient data from any number of facilities, such as if the customer or patient data is free from privacy concerns (e.g., via a waiver for experimental use) or is included in a public dataset.
[0143] In at least one embodiment, when selecting an application for use in the deployment pipeline, a user may also select the machine learning model that will be used with the particular application. In at least one embodiment, a user may not have a model to use, and therefore the user may select a pre-trained model to use with the application. In at least one embodiment, the pre-trained model may not be optimized to produce accurate results for the user's facility's customer dataset 1506 (e.g., based on patient diversity, demographics, type of medical imaging device used, etc.). In at least one embodiment, before introducing the pre-trained model into the deployment pipeline for use with the application, the pre-trained model may be updated, retrained, and / or fine-tuned for use at the respective facility.
[0144] In at least one embodiment, a user may select a pre-trained model to be updated, retrained, and / or fine-tuned, which may be referred to as an initial model 1504 for training the system in process 1500. In at least one embodiment, model training (which may include, without limitation, transfer learning) may be performed on the initial model 1504 using a customer dataset 1506 (e.g., imaging data, genomics data, sequencing data, or other types of data generated by devices at the facility) to generate a refined model 1512. In at least one embodiment, ground truth data corresponding to the customer dataset 1506 may be generated by the training system 1304. In at least one embodiment, the ground truth data may be generated, at least in part, by clinicians, scientists, physicians, or practitioners at the facility.
[0145] In at least one embodiment, AI-assisted annotation may be used in some instances to generate ground truth data. In at least one embodiment, AI-assisted annotation (e.g., implemented using the AI-assisted annotation SDK) may utilize machine learning models (e.g., neural networks) to generate suggested or predicted ground truth data for a customer dataset. In at least one embodiment, a user may use the annotation tool within a user interface (graphical user interface (GUI)) on a computing device.
[0146] In at least one embodiment, a user 1510 may interact with a GUI via computing device 1508 to edit or fine-tune the (automatic) annotation. In at least one embodiment, polygon editing features may be used to move polygon vertices to more precise or fine-tuned locations.
[0147] In at least one embodiment, once the customer dataset 1506 has associated ground truth data, the ground truth data (e.g., from AI-assisted annotation, manual labeling, etc.) may be used during model training to generate the refined model 1512. In at least one embodiment, the customer dataset 1506 may be applied to the initial model 1504 any number of times, and the ground truth data may be used to update the parameters of the initial model 1504 until an acceptable level of accuracy is achieved for the refined model 1512. In at least one embodiment, once the refined model 1512 is generated, the refined model 1512 may be deployed into one or more deployment pipelines at a facility to perform one or more processing tasks on the medical imaging data.
[0148] In at least one embodiment, the refined model 1512 may be uploaded to a model registry of pre-trained models to be selected by another facility. In at least one embodiment, this process may be completed at any number of facilities, whereby the refined model 1512 may be further refined any number of times on new datasets to generate a more general model.
[0149] FIG. 15B is a diagram of an example client-server architecture 1532 for enhancing an annotation tool with a pre-trained annotation model, according to at least one embodiment. In at least one embodiment, an AI-assisted annotation tool 1536 may be instantiated based on the client-server architecture 1532. In at least one embodiment, the AI-assisted annotation tool 1536 in an imaging application may, for example, assist a radiologist in identifying organs and abnormalities. In at least one embodiment, the imaging application may include, by way of non-limiting example, a software tool that helps a user 1510 identify a few extreme points on a particular target organ in raw images 1534 (e.g., of a 3D MRI or CT scan) and receive automatically annotated results for all 2D slices of the particular organ. In at least one embodiment, the results may be stored in a data store as training data 1538 and may be used (for example, without limitation) as ground truth data for training. In at least one embodiment, when computing device 1508 sends extremum points for AI-assisted annotation, a deep learning model, for example, may receive this data as input and return inferences of segmented organs or anomalies. In at least one embodiment, a pre-instantiated annotation tool, such as AI-assisted annotation tool 1536 of FIG. 15B , may be extended by making API calls (e.g., API call 1544) to a server, such as annotation-assisted server 1540, which may include a set of pre-trained models 1542 stored in an annotation model registry. In at least one embodiment, annotation model registry may store pre-trained models 1542 (e.g., machine learning models, such as deep learning models) that have been pre-trained to perform AI-assisted annotation for specific organs or anomalies. These models may be further updated using a training pipeline.In at least one embodiment, the pre-installed annotation tools may be improved over time as new labeled data is added.
[0150] The following clauses can be used to describe various embodiments: 1. A computer-implemented method comprising: Representing one or more objects in a scene using a geometric mesh that approximates a plurality of volumetric particles; determining an intersection of a ray cast with respect to a selected view of the scene with at least a portion of a geometric mesh corresponding to at least one of the volumetric particles; determining a response value of at least one volumetric particle corresponding to an intersection of the rays; using the response values to determine pixel values for an image of the scene rendered from the selected view; 11. A computer-implemented method comprising: 2. The computer-implemented method of clause 1, wherein the volumetric particles are two-dimensional, or three-dimensional, or greater dimensional particles having anisotropy factors along different dimensions. 3. Generating a plurality of volumetric particles based in part on a plurality of two-dimensional images acquired for a plurality of views of the scene. 3. The computer-implemented method of claim 2, further comprising: 4. The computer-implemented method of clause 3, wherein the selected view is different from any of the multiple views from which the multiple two-dimensional images are acquired. 5. The computer-implemented method of clause 1, wherein the volumetric particles appear different colors in different view directions. 6. The computer-implemented method of clause 1, wherein the volumetric particles correspond to a local three-dimensional function, including at least one of a linear function, a Lagrangian function, a Gaussian distribution function, a Gaussian kernel, or a Gabor kernel. 7. Determining that the ray intersects a plurality of translucent volumetric particles; determining pixel values corresponding to the light ray up to at least the transmittance threshold based in part on response values from one or more of the intersected translucent volumetric particles; 2. The computer-implemented method of claim 1, further comprising: 8. The computer-implemented method of clause 1, wherein the view corresponds to a distorted camera with a rolling shutter or a moving virtual camera. 9. The computer-implemented method of clause 1, wherein determining ray intersections is accelerated using hardware acceleration. 10. Generating an image of a scene to be used in at least one of the following operations: robotics, automotive navigation, photorealistic synthetic image generation, or synthetic image relighting. 2. The computer-implemented method of claim 1, further comprising: 11. At least one processor, generating a geometric mesh approximating a plurality of volumetric particles for one or more objects in a scene; determining an intersection of a ray cast with respect to a selected view of the scene with a geometric mesh associated with at least one of the volumetric particles; determining a value of a local three-dimensional function represented by at least one volumetric particle corresponding to an intersection of the rays; using the determined values to determine pixel values for an image of the scene rendered from the selected view; Processing logic for At least one processor comprising: 12. At least one processor according to clause 11, wherein the volumetric particles are three-dimensional particles having anisotropy factors along different dimensions. 13. At least one processor according to clause 11, wherein the volumetric particles correspond to a local three-dimensional function, including at least one of a linear function, a Lagrangian function, or a Gaussian distribution function. 14. The processing logic: determining that a ray intersects a plurality of translucent volumetric particles; determining pixel values corresponding to the light ray up to at least the transmittance threshold based in part on response values from one or more of the intersected translucent volumetric particles; 12. At least one processor according to clause 11, 15. At least one processor: a system for performing simulation operations; A system for performing simulation operations to test or validate an autonomous machine application; A system for running digital twin operations, a system for performing light transport simulations; a system for rendering graphical output; a system for performing deep learning operations; Systems implemented using edge devices, A system for generating or presenting virtual reality (VR) content, A system for generating or presenting augmented reality (AR) content; A system for generating or presenting mixed reality (MR) content, A system incorporating one or more virtual machines (VMs), a system implemented at least in part in a data center; A system for performing hardware testing using simulation; A system for synthetic data generation, a system for performing generative AI operations; A system for performing one or more operations using a large language model (LLM), A system for performing one or more operations using a vision language model (VLM), A collaborative content creation platform for 3D assets, and at least one processor according to clause 11, comprising at least one of the following systems: 16. System 1. A system comprising one or more processors that determine pixel values for an image of a scene rendered from a specified view, in part, by casting a plurality of rays corresponding to the specified view and determining intersection points between the plurality of rays and a mesh of volumetric particles representing one or more objects in the scene, wherein the intersection points correspond to pixel values corresponding to a given ray calculated using response values of one or more volumetric particles that the ray intersects. 17. The system of clause 16, wherein the specified view corresponds to a distorted virtual camera. 18. The system of clause 16, wherein casting multiple rays is accelerated using hardware acceleration. 19. The system of clause 16, wherein the volumetric particles are three-dimensional particles having anisotropy factors along different dimensions. 20. The system a system for performing simulation operations; A system for performing simulation operations to test or validate an autonomous machine application; A system for running digital twin operations, a system for performing light transport simulations; a system for rendering graphical output; a system for performing deep learning operations; a system for performing generative AI operations; A system for performing one or more operations using a large language model (LLM), A system for performing one or more operations using a vision language model (VLM), Systems implemented using edge devices, A system for generating or presenting virtual reality (VR) content, A system for generating or presenting augmented reality (AR) content; A system for generating or presenting mixed reality (MR) content, A system incorporating one or more virtual machines (VMs), a system implemented at least in part in a data center; A system for performing hardware testing using simulation; A system for synthetic data generation, A collaborative content creation platform for 3D assets, and 17. The system of claim 16, including at least one of: a system implemented at least in part using cloud computing resources.
[0151] Other variations are within the scope of the present disclosure. Thus, while the disclosed techniques are susceptible to various modifications and alternative constructions, certain illustrative embodiments thereof have been shown in the drawings and have been described above in detail. It should be understood, however, that there is no intention to limit the disclosure to the particular form or forms disclosed, but on the contrary, the intention is to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of the disclosure as defined by the appended claims.
[0152] The use of the terms "a," "an," and "the," and similar referents in the context of describing the disclosed embodiments (particularly in the context of the claims that follow) should be construed to cover both the singular and the plural and not as defining terms, unless otherwise stated herein or clearly contradicted by context. The terms "comprising," "having," "including," and "containing" are to be construed as open-ended terms (meaning "including, but not limited to") unless otherwise indicated. The term "connected," when unmodified, referring to a physical connection, is to be construed as being partially or completely contained within, attached to, or joined to one another, even if there is something intervening. The recitation of ranges of values herein is merely intended to serve as a shorthand method of individually referring to each separate value falling within the range, unless otherwise stated herein and unless each separate value is incorporated into the specification as if it were individually recited herein. Use of the term "set" (e.g., "set of items") or "subset" should be construed as a non-empty collection having one or more members, unless otherwise stated or constrained by context. Furthermore, unless otherwise stated or constrained by context, the term "subset" of a corresponding set does not necessarily refer to a proper subset of the corresponding set, and a subset and a corresponding set may be equivalent.
[0153] Conjunctive language, such as phrases of the form "at least one of A, B, and C" or "at least one of A, B, and C," is understood in the context in which it is generally used to indicate that an item, term, etc. is A, B, or C, or a non-empty subset of any of the sets A, B, and C, unless specifically stated otherwise or clearly contradicted by context. For example, in the illustrative example of a set having three members, the conjunctive language "at least one of A, B, and C" and "at least one of A, B, and C" refer to any of the following sets: {A}, {B}, {C}, {A,B}, {A,C}, {B,C}, {A,B,C}. Thus, such conjunctive language does not generally imply that a given embodiment requires the presence of at least one A, at least one B, and at least one C, respectively. Further, unless stated otherwise or negated by context, the term "plurality" refers to a plurality (e.g., "a plurality of items" refers to multiple items). A plurality is at least two items, but may be more if indicated explicitly or by context. Further, unless stated otherwise or clear from context, the phrase "based on" means "based at least in part on," and not "based only on."
[0154] The operations of processes described herein may be performed in any suitable order unless otherwise stated herein or clearly contradicted by context. In at least one embodiment, processes such as those described herein (or variations and / or combinations thereof) are implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) performed under the control of one or more computer systems, the code being comprised of executable instructions and collectively executed by one or more processors, by hardware, or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium, e.g., in the form of a computer program comprising instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium excludes transient signals (e.g., propagating transient electrical or electromagnetic transmissions), but includes non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transient signal transceiver. In at least one embodiment, code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media that store executable instructions (or have other memory for storing executable instructions) that, when executed by (i.e., as a result of being executed by) one or more processors of a computer system, cause the computer system to perform the operations described herein. The set of non-transitory computer-readable storage media, in at least one embodiment, comprises a plurality of non-transitory computer-readable storage media, wherein one or more individual non-transitory storage media of the plurality of non-transitory computer-readable storage media do not contain all of the code, but the plurality of non-transitory computer-readable storage media collectively store all of the code.In at least one embodiment, the executable instructions are executed such that different instructions are executed by different processors, e.g., a non-transitory computer-readable storage medium stores the instructions, a main central processing unit ("CPU") executes some instructions, and a graphics processing unit ("GPU") executes other instructions. In at least one embodiment, different components of a computer system have separate processors, and the different processors execute different subsets of the instructions.
[0155] Thus, in at least one embodiment, a computer system is configured to implement one or more services that, singly or collectively, perform the operations of the processes described herein, and such a computer system is configured with applicable hardware and / or software that enables the operations to be performed. Further, a computer system that implements at least one embodiment of the present disclosure may be a single device, or in another embodiment, a distributed computer system comprising multiple devices that operate in different ways, such that the distributed computer system performs the operations described herein such that no single device performs all of the operations.
[0156] Any examples provided herein, or the use of exemplary language (e.g., "such as"), are intended only to further clarify embodiments of the disclosure and do not limit the scope of the disclosure unless otherwise stated. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.
[0157] All references, including publications, patent applications, and patents, cited in this specification are hereby incorporated by reference to the same extent as if each reference was individually indicated to be incorporated by reference and was set forth in its entirety herein.
[0158] In the specification and claims, the terms "coupled" and "connected," along with their derivatives, may be used. It should be understood that these terms may not be intended as synonyms for each other. Rather, in particular instances, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other, but yet still cooperate or interact with each other.
[0159] Unless specifically stated otherwise, terms such as "processing," "computing," "calculating," or "determining" throughout the specification will be understood to refer to the acts and / or processes of a computer or computing system or similar electronic computing device that manipulates and / or transforms data represented as physical quantities, such as electronic quantities, in the registers and / or memory of the computing system into other data similarly represented as physical quantities in the memory, registers, or other such information storage, transmission, or display device of the computing system.
[0160] Similarly, the term "processor" may refer to any device, or portion of a device, that processes electronic data from registers and / or memory and converts the electronic data into other electronic data that can be stored in registers and / or memory. As a non-limiting example, a "processor" may be a CPU or GPU. A "computing platform" may include one or more processors. As used herein, a "software" process may include software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Also, each process may refer to multiple processes for serially or in parallel, continuously or intermittently executing instructions. The terms "system" and "method" are used interchangeably herein, provided that a system may embody one or more methods and a method may be considered a system.
[0161] As used herein, references may be made to obtaining, acquiring, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. Obtaining, acquiring, receiving, or inputting analog and digital data may be implemented in various ways, such as receiving data as parameters of a function call or a call to an application programming interface. In some implementations, the process of obtaining, acquiring, receiving, or inputting analog or digital data may be implemented by transferring data over a serial or parallel interface. In other implementations, the process of obtaining, acquiring, receiving, or inputting analog or digital data may be implemented by transferring data over a computer network from a providing entity to an acquiring entity. References may also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various instances, the process of providing, outputting, transmitting, sending, or presenting analog or digital data may be implemented by transferring data as input or output parameters of a function call, an application programming interface, or an inter-process communication mechanism.
[0162] While the above discussion describes example implementations of the described techniques, other architectures may be used to implement the described functionality, and such other architectures are intended to be within the scope of this disclosure. Furthermore, while a specific distribution of roles is defined for purposes of discussion, various functions and roles may be distributed and divided differently depending on the circumstances.
[0163] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claims.
Claims
1. 1. A computer-implemented method comprising: Representing one or more objects in a scene using a geometric mesh that approximates a plurality of volumetric particles; determining an intersection of a ray cast for a selected view of the scene with at least a portion of the geometric mesh corresponding to at least one of the volumetric particles; determining a response value of at least one of the volumetric particles corresponding to the intersection of the light rays; using the response values to determine pixel values for an image of the scene rendered from the selected view; and 11. A computer-implemented method comprising:
2. The computer-implemented method of claim 1 , wherein the volumetric particles are two-dimensional, three-dimensional, or higher-dimensional particles having anisotropy factors along different dimensions.
3. generating the plurality of volumetric particles based in part on a plurality of two-dimensional images acquired of a plurality of views of the scene; The computer-implemented method of claim 2 further comprising:
4. The computer-implemented method of claim 3 , wherein the selected view is different from any of the plurality of views from which the plurality of two-dimensional images are acquired.
5. The computer-implemented method of claim 1 , wherein the volumetric particles appear different colors in different view directions.
6. The computer-implemented method of claim 1 , wherein the volumetric particles correspond to a local three-dimensional function comprising at least one of a linear function, a Lagrangian function, a Gaussian distribution function, a Gaussian kernel, or a Gabor kernel.
7. determining that the ray intersects a plurality of translucent volumetric particles; determining the pixel values corresponding to the light ray up to at least a transmittance threshold based in part on response values from one or more of the intersected translucent volumetric particles; The computer-implemented method of claim 1 , further comprising:
8. The computer-implemented method of claim 1 , wherein the view corresponds to a distorted camera with a rolling shutter or a moving virtual camera.
9. The computer-implemented method of claim 1 , wherein determining the ray intersection is accelerated using hardware acceleration.
10. 10. The computer-implemented method of claim 1, further comprising generating the image of the scene to be provided to operations related to at least one of robotics, automotive navigation, photorealistic synthetic image generation, or synthetic image relighting.
11. at least one processor, generating a geometric mesh approximating a plurality of volumetric particles for one or more objects in a scene; determining an intersection of a ray cast for a selected view of the scene with the geometric mesh associated with at least one of the volumetric particles; determining a value of a local three-dimensional function represented by at least one of the volumetric particles corresponding to the intersection of the ray; using the determined values to determine pixel values for an image of the scene rendered from the selected view; and Processing logic for At least one processor comprising:
12. 12. The at least one processor of claim 11, wherein the volumetric particles are three-dimensional particles having anisotropy factors along different dimensions.
13. The at least one processor of claim 11 , wherein the volumetric particles correspond to a local three-dimensional function comprising at least one of a linear function, a Lagrangian function, or a Gaussian distribution function.
14. the processing logic: determining that the ray intersects a plurality of translucent volumetric particles; determining the pixel values corresponding to the light ray up to at least a transmittance threshold based in part on response values from one or more of the intersected translucent volumetric particles; 12. At least one processor according to claim 11, for:
15. the at least one processor: a system for performing simulation operations; A system for performing simulation operations to test or validate an autonomous machine application; a system for performing digital twin operations; a system for performing light transport simulations; a system for rendering graphical output; a system for performing deep learning operations; a system implemented using edge devices; a system for generating or presenting virtual reality (VR) content; a system for generating or presenting augmented reality (AR) content; A system for generating or presenting mixed reality (MR) content; a system incorporating one or more virtual machines (VMs); a system implemented at least in part in a data center; A system for performing hardware testing using simulation; A system for synthetic data generation, a system for performing generative AI operations; A system for performing one or more operations using a large language model (LLM), A system for performing one or more operations using a vision language model (VLM), A collaborative content creation platform for 3D assets, and 12. The at least one processor of claim 11, wherein the at least one processor comprises at least one of: a system implemented at least in part using cloud computing resources.
16. It is a system 1. A system comprising: one or more processors that determine pixel values for an image of a scene rendered from a specified view, in part, by casting a plurality of rays corresponding to the specified view and determining intersection points of the plurality of rays with a mesh of volumetric particles representing one or more objects in the scene, wherein the pixel values correspond to a given ray calculated using response values of one or more volumetric particles intersected by the ray.
17. The system of claim 16 , wherein the specified view corresponds to a distorted virtual camera.
18. The system of claim 16 , wherein the casting of the multiple rays is accelerated using hardware acceleration.
19. The system of claim 16 , wherein the volumetric particles are three-dimensional particles having anisotropy factors along different dimensions.
20. The system comprises: a system for performing simulation operations; A system for performing simulation operations to test or validate an autonomous machine application; a system for performing digital twin operations; a system for performing light transport simulations; a system for rendering graphical output; a system for performing deep learning operations; a system for performing generative AI operations; A system for performing one or more operations using a large language model (LLM), A system for performing one or more operations using a vision language model (VLM), a system implemented using edge devices; a system for generating or presenting virtual reality (VR) content; a system for generating or presenting augmented reality (AR) content; A system for generating or presenting mixed reality (MR) content; a system incorporating one or more virtual machines (VMs); a system implemented at least in part in a data center; A system for performing hardware testing using simulation; A system for synthetic data generation, A collaborative content creation platform for 3D assets, and 17. The system of claim 16, comprising at least one of: a system implemented at least in part using cloud computing resources.