Ray tracing volume particles for real-time novel view synthesis

By combining volume particle representation and ray tracing with bounding volume hierarchy, the problem of high resource consumption in existing technologies is solved, achieving efficient rendering of high-quality 3D model images.

CN120976398APending Publication Date: 2025-11-18NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510631198.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-17
Filing Date
2025-05-16
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing methods struggle to generate high-resolution and high-quality 3D model images in real time, especially for arbitrary non-pinhole cameras and complex lighting effects, and are also resource-intensive.

Method used

We employ a combination of volume particle representation and ray tracing. We use volume particle sets to represent 3D models and leverage bounding volume hierarchies to accelerate ray tracing and reduce the resource and time consumption of hit testing.

Benefits of technology

It achieves efficient and high-speed rendering of high-quality 3D model images, supports distortion effects and complex lighting effects, and reduces resource requirements and latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976398A_ABST
    Figure CN120976398A_ABST
Patent Text Reader

Abstract

The method presented herein provides for efficient rendering of a high quality, novel view of a scene, in which case is achieved by a combination of volume particle representation and ray tracing. The object may be represented using a set of body particles (e.g., a 3D distribution) aligned to the infrastructure or geometry of the object. The volume particles may be encapsulated in a bounding grid or proxy geometry that may be used to efficiently compute ray-particle intersections. For a view to render, ray tracing may be performed to determine an intersection of the ray with the proxy geometry. When a hit is determined, an accurate intersection point position with the body particle is calculated and a distribution value of the light is returned. If light passes through a plurality of translucent bulk particles, the color value is determined based on values returned from the particles.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] There are a variety of operations, such as computer animation or environment simulation, in which it can be necessary to generate an image of at least one three-dimensional (3D) model in a scene. The 3D model that can be used for this purpose can be generated by combining data from multiple images of a captured physical object. An image of the model will typically have to be generated from a novel viewpoint that is different from any image of the captured physical object. Prior methods for generating such novel views have generally been unable to achieve real-time performance with higher resolution and quality. Recent methods use rasterization with a representation of radiance field (NeRF), which can achieve acceptable performance at interactive rates, but this approach incorporates the drawbacks of rasterization. In particular, this approach provides important support for arbitrary non-pinhole cameras (e.g., fisheye or other types of cameras with distortion) and rolling shutters, and does not support higher-order lighting effects such as shadows or reflections. BRIEF DESCRIPTION OF DRAWINGS

[0002] Various embodiments according to the present disclosure will be described in reference to the drawings, in which:

[0003] Figures 1A-1C A digital representation of an object is shown in accordance with at least one embodiment;

[0004] Figures 2A-2C One or more geometric mesh approximations of a body are shown in accordance with at least one embodiment;

[0005] Figures 3A-3E Ray tracing for a representation of an object is shown in accordance with at least one embodiment, along with values determined along the traced rays;

[0006] Figure 4A Example components of a content generation system are shown in accordance with at least one embodiment;

[0007] Figure 4B Components of an example rendering pipeline are shown in accordance with at least one embodiment;

[0008] Figure 5 An example process for potentially generating an image of an object or scene from a novel view is shown in accordance with at least one embodiment;

[0009] Figure 6 Components of a distributed system that can be used to generate and provide content are shown in accordance with at least one embodiment;

[0010] Figure 7A Inference and / or training logic is shown in accordance with at least one embodiment;

[0011] Figure 7B Inference and / or training logic is shown in accordance with at least one embodiment;

[0012] Figure 8 An example data center system according to at least one embodiment is shown;

[0013] Figure 9 A computer system according to at least one embodiment is shown;

[0014] Figure 10 A computer system according to at least one embodiment is shown;

[0015] Figure 11 At least a portion of a graphics processor according to one or more embodiments is shown;

[0016] Figure 12 At least a portion of a graphics processor according to one or more embodiments is shown;

[0017] Figure 13 This is an example data flow diagram of an advanced computing pipeline according to at least one embodiment;

[0018] Figure 14 This is a system diagram of an example system for training, adapting, instantiating, and deploying machine learning models in an advanced computing pipeline, according to at least one embodiment; and

[0019] Figure 15A and Figure 15B A data flow diagram of the process for training a machine learning model according to at least one embodiment is shown, as well as a client-server architecture for enhancing annotation tools using a pre-trained annotation model. Detailed Implementation

[0020] In the following description, various embodiments will be described. Specific configurations and details are set forth for illustrative purposes in order to provide a thorough understanding of the embodiments. However, it will also be apparent to those skilled in the art that the embodiments can be practiced without specific details. Furthermore, well-known features may be omitted or simplified so as not to obscure the described embodiments.

[0021] The systems and methods described herein can be used by, but are not limited to, non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more advanced driver assistance systems (ADAS)), manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled with one or more trailers, aircraft, watercraft, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, engineering vehicles, trains, underwater vehicles, remote-controlled vehicles such as drones, and / or other vehicle types. Further, the systems and methods described herein can be used for various purposes such as, but not limited to, machine control, machine motion, machine driving, synthetic data generation, model training or updating, perception, augmented reality, virtual reality, mixed reality, robotics, safety and supervision, simulation and digital twin, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twin, data center processing, conversational AI, generative AI, operations using one or more large language models (LLMs) or one or more visual language models (VLMs), light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, and / or any other suitable application.

[0022] The disclosed embodiments can be included in various different systems such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in data centers, systems for performing conversational AI operations, systems for performing generative AI operations, systems for performing operations using one or more LLMs or one or more VLMs, systems for performing light transport simulation, systems for performing collaborative content creation of 3D assets, systems implemented at least partially using cloud computing resources, and / or other types of systems.

[0023] Methods according to various illustrative embodiments provide for efficiently rendering high quality images of three-dimensional (3D) objects or scenes from a variety of views. These views can include any appropriate view, such as novel views that are not represented in any previously obtained or generated data of the object. The rendering can be implemented in part by using a volume particle representation that utilizes ray tracing. An object model can be represented using a set of volume particles (e.g., 2D / 3D Gaussian distributions or other Lagrangian representations of color and / or other such information) that are aligned with the underlying structure or geometry of the scene to be rendered (e.g., as a thin structure). The volume particles can be encapsulated in bounding meshes (or other proxy geometries) that can be used to efficiently construct a bounding volume hierarchy (BVH). Such a method can enable significant graphics hardware acceleration and efficient hit determination. For a view to be rendered (e.g., a novel view), ray tracing can be performed to determine intersections of a ray with the bounding meshes or proxy geometries of the volume particles (such as the geometric envelope around a 3D Gaussian) corresponding to that view. When a hit is determined with respect to the proxy geometry of a given volume particle, the exact intersection position with the volume particle (if there is a true intersection) can be computed, and the distribution value for that ray (e.g., the maximum Gaussian volume response along the ray) can be computed and returned. If the ray passes through one or more translucent volume particles, color values can be determined based on the values returned from these particles. In at least one embodiment, volume rendering can be performed on samples (one or more samples per particle) extracted from the intersecting particles until a transmittance threshold (or other such criterion) has been reached. These color values can then be used to render the specified view of the scene. Such a process provides for high quality rendered images and improves upon previous rasterization-based methods in many respects, including providing higher efficiency and support for distorted cameras. Such a method can also support evaluation of gradients for backpropagation, allowing backpropagation to fit parameters of the best rendered set of particles to a training set of images of ground truth poses.

[0024] Those of ordinary skill in the art will appreciate, in light of the teachings and suggestions contained herein, that variations of this functionality and other such functionality can also be used within the scope of various embodiments.

[0025] As noted above, when an image of a scene is to be rendered, the rendering process can include generating an image representation of one or more objects from a specified viewpoint. There are many ways in which an object or object model can be represented in digital form, such as by using a geometric mesh with color information or a cloud of particles. Other information can also be stored for such a representation, which can relate to material properties, etc. In some cases, a complete 3D model can be synthetically generated, such as by a digital artist or a generative model. In other cases, a 3D object model can be reconstructed from a set of 2D images of a physical object that are captured. Figure 1A An example view 100 showing the locations of a set of 2D images 104 of a captured physical object 102 is shown. This can include any appropriate number (e.g., on the order of 250) of camera images, which can depend in part on the level of detail desired. Each of these 2D images 104 can be captured from a different location with a different viewpoint of the physical object 102. In some cases, images can also be captured using different camera settings or under different lighting conditions. In order to generate a sufficiently accurate 3D (or 4D) digital model or representation of the physical object 102, it can be desirable to capture a sufficiently large number of images from a wide variety of views. However, it will be appreciated that a 3D model can be inferred from as few as a single image (e.g., with prior encoded in the data), if desired.

[0026] The set of 2D images 104 can then be analyzed to attempt to generate an accurate 3D digital representation. This can include pre-processing, such as aligning the images, adjusting for different camera parameters or lighting conditions, performing noise reduction, etc. A neural network or modeling algorithm can then analyze the data from the various images, such as to attempt to extract and correlate various features of the images. This can include correlating the positions of extracted (or otherwise determined) particles 132, or “fitting” these particles to a common coordinate system or frame of reference, as shown in the example view 130. Figure 1B In this example, the representation is a volumetric particle set (e.g., a cloud of particles) formed from a plurality of particles 132 having associated color values (and other types of values, such as surface properties, as described elsewhere herein). Other representations can also be generated, which can include meshes, etc.

[0027] In at least one embodiment, a light transport simulation process, such as ray tracing, can then be used on such a model to generate an image of the object model from at least one specified viewpoint. As noted above, this can be different from any view of the corresponding object that was captured or previously generated. When using a ray tracing process as described above, for example, the resulting image can be generated by simulating the paths of a plurality of rays 134 from the specified viewpoint, as shown in the example view 140. Figure 1BWhen represented as volume particles, there can be many particles for which ray tracing and hit testing needs to be performed, which can require a large amount of time and resources. Even for meshes or other representations, the amount of data to process can impact real-time performance. Accordingly, methods in accordance with various embodiments can use different types of object representations that are faster to process, such as when performing ray tracing or hit testing. One such representation involves the use of a set of volume particles. In at least one embodiment, a volume particle is a three-dimensional representation that can be elliptical in shape. As Figure 1C The object representation shown in the sample view image 160 can be composed of a set of volume particles 162 of different shapes and / or sizes. These volume particles can be selected and oriented to align them with the underlying structure or geometry of one or more objects in the scene. Each volume particle can contain color information in the form of a 2D Gaussian distribution, a Lagrangian distribution, or other such representation. When a ray intersects (or passes through) a volume particle, the color can vary based on the position and direction of the ray, and can return a color similar to what would be returned if the ray were cast onto Figure 1B the cloud of particles shown above.

[0028] Volume particles can provide several advantages over previous point-based, mesh-based, or other such methods. In a first example, hit testing can be performed more quickly since there are far fewer volume particles aligned with the underlying particles or geometry instances (e.g., triangles) of a mesh. The volume particles can represent important portions of an object model, and if a cast ray does not intersect the boundary of a volume particle, none of the particles in that volume particle need to be sampled for that ray. Another advantage of volume particles is that individual particles can contain a continuous distribution, so there can be reasonably reliable data for any sample particle within the volume particle. Further, the use of a continuous distribution representation can also reduce the presence of noise and spurious data.

[0029] In at least one embodiment, ray tracing can be performed directly on these volume particles. However, for at least some ray tracing hardware, accelerated and / or improved performance can be achieved by using a geometric representation of these volume particles for hit testing. The geometric representation can be defined by a small number of particles in space, which can reduce the resources and time required for hit testing. Figure 2AAn image view 200 of an example volume particle is shown. The variation in shading indicates that the internal distribution of color values can vary based on position and orientation, and that the distribution can take on many different shapes or forms. A geometry representation 204 can be generated, which serves as a type of bounding volume for the volume particle. While the geometry representation 204 will include particles outside of the volume particle 202, the geometry representation 204 can be more lightweight and faster to use to perform hit tests or analysis. Any suitable shape can be used to represent a volume particle, but since volume particles can be roughly ellipsoidal in nature, a representative geometry can advantageously take the form of a rhombic dodecahedron or other such geometry, which can have as few as six edges to represent a bounding volume for the entire ellipsoid.

[0030] These geometry representations 232 can be used to represent objects as shown in view 230 of Figure 2B . Processes such as ray tracing or hit testing can be performed on these geometry representations to quickly determine areas of the object model that should (or should not) be sampled. As can be seen, the number of particles needed to define the geometry representations 232 is far fewer than the particle cloud representation of Figure 1B , and the complexity can also be far lower than the representation of volume particles shown in Figure 1C .

[0031] There can be additional optimizations or representations that are useful for particular ray tracing or processing hardware. For example, Figure 2C shows a view 260 of the geometry representations of Figure 2B , but with each geometry representation having a determined rectangular bounding volume 262 (or proxy geometry shape). These rectangular bounding volumes are also all aligned to a common reference frame, such that the edges in the image are all horizontal or all vertical in orientation. These rectangular bounding volumes can be part of a bounding volume hierarchy (BVH). This representation can advantageously be used as part of a BVH ray tracing acceleration structure, which can be optimized for particular ray tracing hardware such as RTX hardware available from NVIDIA Corporation. Other such representations can be used as appropriate.

[0032] Once an appropriate set of geometry proxies or bounding volumes is determined, ray tracing can be performed using a configuration 300 such as shown in Figure 3A . In this configuration, a virtual camera 302 can be positioned at a specified location with a specified orientation, which can provide a particular viewpoint for the camera of the object representation such as the set of geometry proxies 302. To determine the color that will be used for the various pixel locations of a pixel grid 306 corresponding to an image to be rendered, rays 304 can be cast with respect to this camera location. Any given ray can intersect with one or more geometry proxies or "hit" one or more geometry proxies. As Figure 3BAs shown in example view 320, at most a single intersection of the projected ray 304 with respect to the geometric proxy representation 302 can be determined, which can be an initial point 322 along an edge of the representation at which there is an intersection with the ray. As shown, the top ray 304 is determined to intersect with four geometric proxies, while the bottom ray 324 is shown to intersect with three different geometric proxies. This approach can be used to quickly narrow down the portion or portions of the object model for which a given ray is sampled. If a given ray does not intersect with a geometric proxy, then no sampling need be performed for that ray. Figure 3B As shown, the top ray 304 is determined to intersect with four geometric proxies, while the bottom ray 324 is shown to intersect with three different geometric proxies. This approach can be used to quickly narrow down the portion or portions of the object model for which a given ray is sampled. If a given ray does not intersect with a geometric proxy, then no sampling need be performed for that ray.

[0033] After determining the ray intersection, the volume particles within the intersecting geometric proxy can be sampled. As shown in example view 340, for a given ray within the identified volume particle, there can be multiple points 342 that are sampled. If any of the points correspond to an opaque surface, then no further points along that ray need be sampled. Additional points can be sampled as long as the previously sampled points for the ray are at least partially transmissive (and further sampling for reflection, etc. is performed). This approach can still significantly and quickly reduce the search space even in some scenarios where there can not be an actual intersection with the corresponding volume particle for the ray intersecting with the geometric proxy. Figure 3C

[0034] Figure 3D A more detailed view 360 of an example sampling process is shown, in accordance with at least one embodiment. Once a volume particle 366 is identified for sampling using a geometric proxy or bounding volume 364, ray tracing can be performed and the individual sampled points of the volume particle that intersect with the projected ray 362 are analyzed. As shown, one or more sampled points can be determined for a given ray, which can depend on the transmissive properties of the hit point as discussed previously. The color (or other pixel value) to be returned for a given sampled point or hit point can be determined by analyzing the distribution (e.g., Gaussian, Lagrange, linear, or other) at that point. A cross-sectional view 370 through this representation shows the shape of the distribution 372 with respect to a range of color values. This distribution can represent the color at different feature locations within the space corresponding to the volume particle. For the same volume particle, the color value returned can depend on the position and orientation of the incident ray. Thus, the color value returned from a single volume particle can be different from different angles or perspectives. This provides a reasonable approximation of the number of individual feature points for generating a volume particle and determining an appropriate distribution. Once sampled, these values can be used for tasks such as rendering an image from this object or scene representation. These values can also support evaluating gradients backpropagated through the reconstructed or generative model, enabling backpropagation to fit parameters of the particle set that best render to a training image set of ground truth poses.

[0035] Figure 3E ​Example curves for a 3D anisotropic Gaussian body are shown. In the first view 380, there are four rays projected through different points in the Gaussian body. The second view 382 shows the curve of the corresponding density value (as a ID Gaussian body) for each respective ray. As shown, the location of the density value and the response value for each ray are different. The third view 384 shows the transmittance curve for each projected ray. The transmittance gives an indication of the transparency of the surface at the corresponding hit point, so not only can the contribution be figured out, but it can also be figured out whether additional hits for the ray need to be determined. It can also be seen that the shape or transmittance decay of the transmittance curve is different for each location. In at least one embodiment, the transmittance values can be used to generate a shadow map. As noted above, for the same Gaussian body or other such distribution, different directions can similarly have different curves. With the Gaussian body model, this can be equivalent to a sum of ID Gaussian bodies that can be computed analytically. The amount of occlusion experienced by the ray can be equal to the integral sum of each ID Gaussian body on the ray. The transmittance value corresponds to the exponential of the negative integral from the start of the corresponding ray.

[0036] This approach can be used to represent potentially large and complex 3D scenes using a set of volume particles, where these particles can represent 3D Gaussian distributions or Lagrangian distributions, among other such options. These volume particles can be used to quickly generate images of such scenes from arbitrary and potentially novel viewpoints. The volume particles can also be generated using algorithms that can reduce resource requirements and latency in some cases. Such algorithms can also be used to fit these volume particles, such as constructing such a representation from captured scene images or other such data. The use of ray tracing also has advantages over other approaches, as it can support distorted and / or warped cameras (e.g., cameras with fisheye lenses or rolling shutters), which can be important for operations related to automotive applications and robotics. Ray tracing also allows for the evaluation of light along individual rays, which is important for realistic rendering and relighting, such as by using path tracing renderers. In at least one embodiment, the system can evaluate the piece-wise transmittance of light along a ray, allowing for the simulation of environmental effects (e.g., fog or smoke). The system can also represent secondary effects, such as shadows, reflections, refractions, and depth of field. Fusing these effects is important for realistic rendering, as well as for interactions, such as relighting a scene. This process can also be extended to large scenes, at least in part, to the availability of spatial acceleration structures, such as using a bounding volume hierarchy to quickly identify intersections between rays and volume particles as described above.

[0037] In at least one embodiment, camera parameters alone can be used to generate rays to be cast, even without any assumptions about the camera model to be used. Rays can be traced against a single BVH representing the entire scene or a combination of BVHs representing individual objects. The BVHs can be constructed from volume particles as described above. For volume particles based on a Gaussian distribution, the response of a Gaussian kernel (or Gabor kernel, etc.) can fall off quickly from its center. In one or more example embodiments, the response of a generalized Gaussian kernel p(x) is represented as:

[0038]

[0039] where β is a kernel parameter that controls the fall off (e.g., 1 for a typical Gaussian volume, 2 for a more uniform response). For any given ray cast into a large scene, the sampled response ρ(o + vt s ) of almost all Gaussian volumes along the ray will be close to 0. Although technically even the farthest Gaussian volume will give some very small contribution (the Gaussian support is infinite), by considering only Gaussian volumes whose sample response is above a given threshold (such as can be given by ρ(o + vt s )>τ (typically τ = 0.01), one can practically further approximate In practice, this means that one can use a method to find only those Gaussian volumes that the cast ray passes near, and sample only those Gaussian volumes that the cast ray passes within a specified distance. For Gaussian particles, a Gaussian τ-volume can correspond to an ellipsoid containing every point such that ρ(x) > τ, and a Gaussian τ-envelope is formed by the surface of every point such that ρ(x) = τ.

[0040] For example, when casting millions of rays into a scene with millions of volume particles, it can be beneficial to efficiently determine which Gaussian volumes’ τ-volumes intersect with which rays. To implement this using accelerated ray tracing hardware, one can construct a tight proxy enclosing triangle mesh around each particle, as Figure 2AThese triangular meshes can be processed with existing optimized ray-mesh intersection routines that leverage hardware-accelerated ray tracing frameworks. In at least one embodiment, the proxy geometry that encloses the Gaussian volume τ-envelope as tightly as possible is computed as a regular polyhedron (e.g., tetrahedron, octahedron, or icosahedron) that is transformed by the Gaussian translation μ, rotation R, and scaling S. Ray tracing the proxy geometry for a volume particle allows discarding most Gaussian volumes along a ray whose sample response is less than τ. In contrast to previous methods, this approach can tightly accommodate extremely long and thin isotropic volume particles, which can be common in certain operations and can incur a significant computational cost.

[0041] Once the volume particles that contribute to the ray (in this case, Gaussian volumes) can be identified, the corresponding values can be sampled and their contributions integrated along the ray in order. A first example sampling strategy includes accumulating a single sample per Gaussian volume. This sample can correspond to the point on the ray with the maximum Gaussian response. In other words, L can be approximated as:

[0042]

[0043] and T is:

[0044]

[0045] where is defined as:

[0046]

[0047] and where can be computed as:

[0048]

[0049] where

[0050] o g = S -1 R T (o - μ)

[0051] and

[0052] v g = S -1 R T v

[0053] This approach is efficient, but can produce some amount of aliasing if Gaussian volumes are highly overlapping.

[0054] A method according to another embodiment can include estimating L using multiple importance sampling. This can include using independent biased distributions, such as one per Gaussian volume, which can be given by:

[0055]

[0056] In this case, the Monte Carlo integration of L simplifies to the empirical expectation value of N s samples extracted along the ray:

[0057]

[0058] Such a sampler can be computed iteratively by tracing the Gaussian volume from front to back. In at least one embodiment, N s samples can be generated for the Gaussian volume of the current hit, where samples are rejected based on ω i ρ i (o + vt). The transmittance sampling term is considered by only considering the most recent sample along the ray.

[0059] A ray tracing programming model can impose constraints on how ray- mesh intersections are evaluated and where computations can be performed. Thus, it is important for an algorithm to adapt to these constraints for high performance processing. Specifically, this means constructing the algorithm as a combination of shader programs, such as a ray generation shader, a closest hit shader, or an arbitrary hit shader, which can be evaluated at different times of a ray emission and a ray intersection with a primitive. A naive approach is to use a closest hit ray casting to sequentially find each intersection along a ray. However, this approach can perform a large amount of redundant computations for each ray. As Figure 3D shown, previous work proposes to construct a traversal in slabs, collecting all intersections within a fixed width sub-region of a ray in an arbitrary hit program. The collected intersections are then sorted and integrated into a ray generation program. This process is repeated for each slab. This approach is limited to a fixed number of hits per slab; thus, the results can be inaccurate. At least one embodiment presented herein differs from previous approaches in that it collects hits and sorts them in an arbitrary hit program. The hits are stored in a fixed size array of the ray’s payload. Once the array is full, the traversal is interrupted by reporting the farthest hit. Integration is then performed in the ray generation program, with subsequent ray casts further collecting hits along the ray.

[0060] The approach presented herein can support cases where particles are extremely densely clustered on a hard surface, which would make various previous approaches inefficient or incorrect depending on the choice of parameters. A volume tracing algorithm can be used that involves tracing a dynamic light slab from a ray generation shader. An arbitrary hit shader can be used to store the K closest hits in a ray payload buffer and sort them. Once it is determined that K closest samples have been collected in the arbitrary hit shader, this approach can return to ray generation to process contributions from these samples. Tracing for the next slab can then be resumed from the end distance of the previous slab or the distance to its Kth closest sample, whichever is closer. This approach is important for not missing densely clustered particles, which can be relatively common in some scenes.

[0061] Additional performance can be obtained in cases where the rays correspond to pixels in an image. Rather than projecting a ray for each pixel individually, a ray is projected that corresponds to a small block of pixels (e.g., a 2x2 block of pixels). The evaluation can still be performed in the ray generation shader for each pixel individually, only the ray projection Gaussian intersection in the closest hit shader is shared by all pixels in the block. For a 2x2 fragment block, this approach can result in up to 50% performance improvement with little quality loss.

[0062] This algorithm can also be used for multiple samples per Gaussian volume. In at least one embodiment, a sorted cache buffer of samples can be maintained in the ray generation shader. Specifically, N samples can be generated for each Gaussian volume in the K closest hit ray payload buffer. Samples closer than the next hit can be used to update the integral. Samples farther than the next hit can be cached in a sorted sample buffer. The cached samples can be tested before each hit evaluation: samples closer than the next hit can be used to update the integral and removed from the cache. When the cache buffer is full, the farthest sample can be discarded.

[0063] At least for Gaussian particles, processing such as pruning, cloning, and splitting can be applied to the Gaussian particles. These properties can be desired to ensure that the model allocates its particle capacity to better represent the learned scene. In one or more embodiments, the standard for cloning and splitting using 3D gradients rather than 2D gradients can be applied, as the tracing function can occur in 3D space. Finally, the BVH can be rebuilt at each training iteration. This operation does not incur any significant overhead and can be used to handle changes in the number of particles.

[0064] Figure 4AAn example system for rendering images, video frames, or other instances of image-related content is shown in accordance with at least one embodiment. Such a system can include or incorporate functionality described herein to generate a 3D representation of an object or scene, such as by using a sparse voxel hierarchy. In this example, an image will be rendered for an object and / or scene (or other view, portion, or region) in a virtual environment 400, but such a system can also be used to render images for a semi-virtual or real environment. The virtual environment 400 can include geometry and other data representing shapes or objects in the environment, such as representing three-dimensional (3D) objects that are present or to be included in a scene in the environment, which can include foreground objects (such as people or vehicles) or background objects (such as roads and buildings), among other such options. In at least some embodiments, at least some content to be inserted can be obtained from a source such as an asset repository 402 or other such location, which can contain content (such as geometry, texture, and density data) that can be used to render one or more objects placed in a view of a scene. At least some assets can have been generated using the sparse voxel architecture described herein. In at least some embodiments or instances, there can be a user device 404 running a content generation or management application, which can allow a user to generate and / or select assets 402 to be rendered in or of a virtual environment 400 to be rendered. The user device 404 can also allow a user to control various aspects of an image to be rendered, such as the position or pose of objects in a scene, as well as the viewpoint and other parameters of a virtual camera to be used to render an image of the virtual environment 400. Once rendered, an image can be stored to an image repository 422 and / or provided for display on a user device or display device 424, among other such options.

[0065] In this example, at least one computing resource 406 is used to perform rendering or other image generation. These resources can correspond to one or more servers, for example, which can be located locally or on at least one network, among other such options. In some embodiments, rendering can be performed at least partially on the user device 404. The computing resources 406 can obtain or receive data to be used for rendering, which can include geometry, properties, textures, and / or density data for virtual environments, objects, scenes, or assets, as well as information about the positions and poses of these objects in the scene and parameters of a virtual camera to be used to determine a view of the scene to be rendered. This information can be received to a content application 408, for example, which can execute on a central processing unit (CPU) 410 of the computing resources, which is responsible for tasks such as gathering data, causing rendering images, and performing any formatting or encoding on the resulting images, among other such tasks. The content application can work with a render manager 412, for example, which can be responsible for coordinating the operations of a rendering pipeline executing on the computing resources 406, which can include modules 414 or processes responsible for tasks such as geometry-related tasks (including lighting and shading tasks) or other such tasks. Offset determinations to attempt to avoid self-intersections can take into account errors, and be implemented in these modules. In at least some embodiments, at least some rendering tasks can be performed using one or more GPUs 420A-420D of the computing resources, and potentially using one or more processors or computing instances (physical or virtual) of one or more other computing resources.

[0066] Tasks such as light transport simulation (e.g., ray tracing, path tracing, ray marching, etc.) or volume sampling can be performed using a single processor, such as a single GPU, or can distribute operations across multiple GPUs 420A-420D. In this example, there can be a pool or collection of GPUs 420A-420D, and a resource manager 418 can be at least partially responsible for allocating GPUs to perform processing of operations. If using more than one GPU is desirable or beneficial, the resource manager 418 can allocate one or more GPUs with appropriate capacity or capability. This can include allocating the number of GPUs indicated in a request, or determining the number of GPUs to allocate based in part on the request. In some embodiments, the resource manager is also able to monitor available bandwidth or memory to determine which and how many GPUs to allocate, such as which have high bandwidth capacity can allow operations to be spread across a larger number of GPUs where the impact on bandwidth due to forwarding light ray information is not as severe, while systems with bandwidth-limited can cause the resource manager to attempt to allocate as few GPUs as possible to attempt to reduce the number of forwarding messages required.

[0067] In at least one embodiment, for example, data partitioning can be performed by the rendering manager 412, and data assignment to different processors can be performed by the system's resource manager 418. The resource manager can receive information from the rendering component and can select an appropriate processor from the pool of available processors 420 or processor capacity. In some embodiments, the rendering application can choose the partitioning, while in other embodiments the renderer may not be able to control the data partitioning, which can be handled by a separate management component (…). Figure 4A (Not shown in the image) to complete.

[0068] Figure 4B An example image generation pipeline 450 is shown, which can be used in a virtual environment 400 (such as...). Figure 4A In the example shown, one or more images (such as video frames in a sequence) are rendered. In this example, pixel data 452 of the current frame to be rendered (which may include G-buffer data of the main surface) can be received as input to the surface interaction component 454 of the rendering system. The surface interaction component 454 can use this data to attempt to determine data in the pixel data for any particular type of surface interaction (e.g., reflection, transmission, diffraction, and / or refraction, etc.), and can provide this data to a backpropagation and G-buffer patching component 456, which can perform backpropagation as discussed herein to locate corresponding points of these surface interactions, and use the data to patch a G-buffer 468, which can provide updated input for subsequent frames to be rendered. This data can then be provided to a light sample generation component 458 to perform light sampling, to a ray tracing lighting component 460 to perform ray tracing lighting, and to one or more shaders 462, which can set pixel colors for individual pixels of the frame based at least in part on the determined lighting information (and other information such as color, texture, etc.). As described above, errors can be determined from the ray tracing lighting component 460 and / or the shader component 462, which can be used to determine the offset value of the secondary ray spawn point. The results can be accumulated by the accumulation module 464 or by the component used to generate the output frame 466 of the desired size, resolution, or format.

[0069] In at least one embodiment, shader 462 can perform a reverse projection step. Once the reverse projection pass is complete and the gradient surface parameters have been patched into the current G-buffer, the renderer can perform a lighting pass. Using information from the lighting pass and lighting results from previous frames, gradients can be computed, then filtered and used for historical rejection. This approach can be used to compute a robust temporal gradient between the current frame and previous frames in a temporal denoiser for a ray-tracing renderer. This reverse projection based approach can also work with surface interactions and can work with rasterized G-buffers. Previous reverse projection approaches omitted any G-buffer patching and instead relied on raw current G-buffer samples, which also resulted in false positive gradients. Patching surface parameters can eliminate false positives in most cases, resulting in a very stable denoised image while still reacting quickly to lighting changes. Once the reverse projection pass is complete and the gradient surface parameters have been patched into the current G-buffer, the renderer can perform a lighting pass. Using information from the lighting pass and lighting results from previous frames, gradients can be computed, then filtered and used for historical rejection.

[0070] In at least some embodiments, components of a rendering pipeline can use one or more machine learning (ML) models or deep neural networks (DNNs). For example, this can include generative networks that generate image content. Machine learning can also be used in approaches that avoid self-intersections with traced paths or rays, such as inferring appropriate offsets or generation positions based on multiple error sources to attempt to use the smallest possible offsets (to provide accurate color and lighting information) while avoiding self-intersections or otherwise introducing image artifacts.

[0071] Figure 5An example process 500 that can be performed to efficiently render an image of an object from a specified view, such as a novel view, is shown in accordance with at least one embodiment. It will be appreciated that for this and other processes presented herein, additional, fewer, or alternative steps can be performed, or performed in similar or alternative orders, or at least partially in parallel, within the scope of various embodiments, unless otherwise specifically stated. Moreover, although this example will be discussed with respect to objects generated from multiple images of a physical object, other types of object representations (e.g., scenes) can also be used to generate content that is not limited to 2D images, within the scope of various embodiments. In this example, multiple images of at least one physical object can be obtained 502, where each image can be captured from a different viewpoint. Feature points (or other such representative data) can be extracted from the images, and these extracted feature points can be fitted 504 to a common frame of reference to generate a point-based representation of the object. These points can be used to generate a representation of the object composed of a set of volume particles, where each volume particle can represent values for a corresponding feature point using a 3D function, such as a Gaussian or Lagrangian function or distribution. A set of proxy geometries or a geometry mesh can be used to represent 506 the object, at least for the purposes of efficient hit testing and hardware acceleration. Ray tracing can be performed to determine appropriate color values (or other related values, including but not limited to instance or identity values and / or semantic information) for rendering an image of the object from a particular viewpoint. For a given ray, an intersection of the ray with respect to a proxy geometry (or geometry mesh) corresponding to at least one volume particle can be determined 508. This approach can enable efficient hit testing. Based on the intersection with the proxy geometry, an actual intersection of the cast ray with one or more corresponding volume particles can be determined 510. Response values can be determined for these actual hits with the volume particles. The response values can be used 512 to determine at least pixel values for an image of the object from the specified viewpoint. If it is determined 514 that there are more rays to cast, the process can continue with the next ray. If there are no more rays to cast for this image, the color (and / or identity, semantic) and / or pixel values from the cast rays can be provided 516 for use in generating an image of the object from the selected viewpoint. As discussed, in at least one embodiment, color values for semi-transparent points can be combined until at least a transmissive threshold or other such criteria is met.

[0072] In at least one embodiment, volumetric particles represent content that can be used to render content that is not limited to a single image, but can include or correspond to various types of representations of one or more objects in a scene or environment. For example, rendered content can include video frames, streaming media, or multi-dimensional object representations, such as can be used for various operations, including but not limited to operations related to gaming, animation, simulation, autonomous navigation, or virtual reality (VR) / augmented reality (AR) / enhanced reality (ER) applications, among other such options.

[0073] Various aspects of the various methods presented herein can be sufficiently lightweight to be performed in real-time on a device such as a client device including a personal computer or game console. Such processing can be performed on content generated on or received by the client device or content received from an external source, such as streaming data or other content received from a cloud server 620 or third party service 660 over at least one network, among other such options. In some cases, at least a portion of the processing, generation, synthesis, and / or determination of the content can be performed by one of these other devices, systems, or entities, then provided to the client device (or another such recipient) for presentation or other such use.

[0074] As an example, Figure 6An example network configuration 600 that can be used to provide, generate, modify, encode, process, and / or transmit image data or other such content is shown. In at least one embodiment, a client device 602 can generate or receive data for a session using components of a content application 604 on the client device 602 and data stored locally on the client device. In at least one embodiment, a content application 624 executing on a server 620 (e.g., a cloud server or edge server) can initiate a session associated with at least one client device 602 (as can utilize a session manager and user data stored in a user database 636), and can cause a content manager 626 to determine content from an asset repository 634, such as one or more digital assets (e.g., implicit and / or explicit object representations, which can include sparse voxel octree representations, meshes, and textures). The content manager 626 can work with a rendering module 628 to generate or select objects, digital assets, or other such content to be placed in a scene or other virtual environment. Views of these objects can be rendered by the rendering module 628 and provided to be presented via the client device 602. In at least one embodiment, the rendering module 628 can work with a content generator 630, which can determine image content (or other content representations) to be rendered by the rendering module 628 as part of a content provision, or generated by a sparse voxel hierarchy VAE discussed herein, among other such options. A training manager 632 can be used to train any or all generative models to be used. At least a portion of the rendered content (or representations to be used to render content) can be sent to the client device 602 using an appropriate transport manager 622, through a download, stream, or another such transmission channel. An encoder can be used to encode and / or compress at least some of this data before sending it to the client device 602. In at least one embodiment, a client device 602 receiving such content can provide the content to a corresponding content application 604, which can also or alternatively include a graphical user interface 610, content manager 612, and rendering module 614 for providing, synthesizing, rendering, combining, modifying, or using content for presentation on or by the client device 602 (or other purposes). A decoder can also be used to decode data received over one or more networks 640 for presentation via the client device 602, such as presenting image or video content through a display 606 and audio (such as sounds and music) through at least one audio playback device 608 (such as a speaker or headphones).In at least one embodiment, at least some of the content can already be stored on client device 602, rendered on client device 602, or accessible to client device 602, so at least that portion of the content does not need to be transmitted over network 640, such as can have been previously downloaded or stored locally on a hard drive or optical disc. In at least one embodiment, transmission mechanisms such as data streams can be used to transmit the content from server 620 or user database 636 to client device 602. In at least one embodiment, at least a portion of the content can be obtained, augmented, and / or streamed from another source, such as third party service 660 or other client device 650, which can also include content applications 662 used to generate, augment, or provide content. In at least one embodiment, portions of this functionality can be performed using multiple computing devices or multiple processors in one or more computing devices, such as can include a combination of CPUs and GPUs.

[0075] In this example, these client devices can include any appropriate computing device, such as can include a desktop computer, notebook computer, set-top box, streaming device, game console, smartphone, tablet computer, VR headset, AR eyewear, wearable computer, or smart television. Each client device can submit requests over at least one wired or wireless network, which can include the Internet, an Ethernet network, a local area network (LAN), or a cellular network, among other such options. In this example, these requests can be submitted to an address associated with a cloud provider, which can operate or control one or more electronic resources under a cloud provider environment, such as can include a data center or server farm. In at least one embodiment, the requests can be received or processed by at least one edge server located at the edge of a network and outside of at least one security layer associated with the cloud provider environment. In this way, latency can be reduced by enabling client devices to interact with servers that are closer in distance, while also improving security of resources in the cloud provider environment.

[0076] In at least one embodiment, such a system can be used to perform graphics rendering operations. In other embodiments, such a system can be used for other purposes, such as to provide image or video content to test or validate autonomous machine applications, or to perform deep learning operations. In at least one embodiment, such a system can be implemented using edge devices, or can incorporate one or more virtual machines (VMs). In at least one embodiment, such a system can be implemented at least partially in a data center or at least partially using cloud computing resources.

[0077] Inference and Training Logic

[0078] Figure 7AInference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are described in greater detail below. Figure 7A and / or Figure 7B Details regarding inference and / or training logic 715 are provided.

[0079] In at least one embodiment, inference and / or training logic 715 can include, without limitation, code and / or data storage 701 for storing forward and / or output weight and / or input / output data, and / or other parameters of neurons or layers of a neural network configured in aspects of one or more embodiments that are trained and / or used for inferencing. In at least one embodiment, training logic 715 can include or be coupled to code and / or data storage 701 for storing graph code or other software to control timing and / or order, where weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code, such as graph code, loads weight or other parameter information into processor ALUs based on an architecture of a neural network to which that code corresponds. In at least one embodiment, code and / or data storage 701 stores weight parameters and / or input / output data of each layer of a neural network trained or used in conjunction with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inferencing using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 701 can be included with other on-chip or off-chip data storage, including a processor’s Ll, L2, or L3 cache or system memory.

[0080] In at least one embodiment, any portion of code and / or data storage 701 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 701 can be cache memory, dynamic random access memory (“DRAM”), static random access memory (“SRAM”), non-volatile memory (such as flash memory), or other storage. In at least one embodiment, a choice of whether code and / or data storage 701 is internal or external to a processor, e.g., or comprised of DRAM, SRAM, flash or some other storage type, can depend on available storage space on-chip or off-chip, latency requirements of performing training and / or inferencing functions, batch size of data used in inferencing and / or training of a neural network, or some combination of these factors.

[0081] In at least one embodiment, inference and / or training logic 715 can include, without limitation, code and / or data storage 705 for storing backward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, code and / or data storage 705 stores weight parameters and / or input / output data of each layer of a neural network trained or used in conjunction with one or more embodiments during backward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, training logic 715 can include or be coupled to code and / or data storage 705 for storing graph code or other software to control timing and / or order, where weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code, such as graph code, loads weight or other parameter information into processor ALUs based on an architecture of a neural network to which that code corresponds. In at least one embodiment, any portion of code and / or data storage 705 can be included with other on-chip or off-chip data storage, including a processor’s Ll, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data storage 705 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 705 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., Flash memory), or other storage. In at least one embodiment, a choice of whether code and / or data storage 705 is internal or external to a processor, e.g., whether it is made up of DRAM, SRAM, Flash memory, or some other storage type, depends on whether available storage is on-chip or off-chip, latency requirements of training and / or inferencing functions being performed, batch size of data being used in inference and / or training of a neural network, or some combination of these factors.

[0082] In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 can be separate storage structures. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 can be the same storage structure. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 can be partially the same storage structure and partially separate storage structures. In at least one embodiment, any portion of code and / or data storage 701 and code and / or data storage 705 can be included with other on-chip or off-chip data storage, including a processor’s Ll, L2, or L3 cache or system memory.

[0083] In at least one embodiment, inference and / or training logic 715 can include, without limitation, one or more arithmetic logic units (“ALUs”) 710 (including integer and / or floating point units) for performing logical and / or mathematical operations based, at least in part, on training and / or inference code (e.g., graph code) or instructions therefrom. In at least one embodiment, results of such operations can produce activations (e.g., output values from layers or neurons within a neural network) stored in activation storage 720 that are functions of input / output and / or weight parameter data stored in code and / or data storage 701 and / or code and / or data storage 705. In at least one embodiment, activations stored in activation storage 720 are generated by ALUs 710 executing linear algebraic and / or matrix-based mathematics in response to executing instructions or other code, where weight values stored in code and / or data storage 705 and / or code and / or data storage 701 are used as operands along with other values such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which can be stored in code and / or data storage 705 or code and / or data storage 701 or other on-chip or off-chip storage.

[0084] In at least one embodiment, one or more ALUs 710 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment one or more ALUs 710 can be external to a processor or other hardware logic device or circuit using them (e.g., a co-processor). In at least one embodiment, one or more ALUs 710 can be included within execution units of a processor, or otherwise included in a group of ALUs accessible by execution units of a processor, which can be within a same processor or distributed between different types of processors (e.g., central processing units, graphics processing units, fixed function units, etc.). In at least one embodiment, code and / or data storage 701, code and / or data storage 705, and activation storage 720 can be on a same processor or other hardware logic device or circuit, while in another embodiment they can be in different processors or other hardware logic devices or circuits or some combination of same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 720 can be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory. Moreover, inference and / or training code can be stored with other code accessible to a processor or other hardware logic or circuitry, and can be fetched and / or processed using fetch, decode, schedule, execute, exit, and / or other logic circuitry of a processor.

[0085] In at least one embodiment, the active memory 720 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the active memory 720 may be wholly or partially located inside or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the active memory 720 is internal to or external to the processor may depend on the available on-chip or off-chip storage, the latency requirements for training and / or inference functions, the batch size of data used in inference and / or training the neural network, or some combination of these factors. For example, it may include DRAM, SRAM, flash memory, or other memory types. In at least one embodiment, Figure 7A The inference and / or training logic 715 shown can be used in conjunction with an application-specific integrated circuit (“ASIC”), such as those from Google. Processing unit, from Graphcore TM Inference processing unit (IPU) or from Intel Corp. (e.g., "LakeCrest") processor. In at least one embodiment, Figure 7A The inference and / or training logic 715 shown can be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as field programmable gate array (“FPGA”)

[0086] Figure 7B Inference and / or training logic 715 according to at least one or more embodiments is illustrated. In at least one embodiment, the inference and / or training logic 715 may include, but is not limited to, hardware logic, wherein computational resources are dedicated or otherwise uniquely used in conjunction with weight values ​​or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, Figure 7B The inference and / or training logic 715 shown can be used in conjunction with an application-specific integrated circuit (ASIC), such as those from Google. Processing unit, from Graphcore TM Inference processing unit (IPU) or from Intel Corp. (e.g., "LakeCrest") processor. In at least one embodiment, Figure 7BThe inference and / or training logic 715 shown in FIG. 11 can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware or other hardware, such as field-programmable gate arrays (FPGAs) for example. In at least one embodiment, the inference and / or training logic 715 includes, without limitation, code and / or data storage 701 and code and / or data storage 705, which can be used to store code (e.g., graph code), weight values and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In Figure 7B In at least one embodiment, each of code and / or data storage 701 and code and / or data storage 705 is associated with a dedicated computing resource, such as computing hardware 702 and computing hardware 706, respectively. In at least one embodiment, each of computing hardware 702 and computing hardware 706 includes one or more ALUs that perform mathematical functions (e.g., linear algebraic functions) only on information stored in code and / or data storage 701 and code and / or data storage 705, respectively, the results of which are stored in activation storage 720.

[0087] In at least one embodiment, each of code and / or data storage 701 and 705 and corresponding computing hardware 702 and 706, respectively, correspond to different layers of a neural network, such that activations resulting from one“storage / computing pair 701 / 702” of code and / or data storage 701 and computing hardware 702 are provided as input to the next“storage / computing pair 705 / 706” of code and / or data storage 705 and computing hardware 706 in order to reflect the conceptual organization of a neural network. In at least one embodiment, each storage / computing pair 701 / 702 and 705 / 706 can correspond to more than one neural network layer. In at least one embodiment, additional storage / computing pairs (not shown) can be included in inference and / or training logic 715 after or in parallel with storage / computing pairs 701 / 702 and 705 / 706.

[0088] Data Center

[0089] Figure 8 An example data center 800 that can use at least one embodiment is shown. In at least one embodiment, data center 800 includes a data center infrastructure layer 810, a framework layer 820, a software layer 830 and an application layer 840.

[0090] In at least one embodiment, as Figure 8As shown, the data center infrastructure layer 810 can include a resource orchestrator 812, grouped computing resources 814, and node computing resources (“node C.R.s”) 816(1)-816(N), where “N” represents any positive integer. In at least one embodiment, node C.R.s 816(1)-816(N) can include, but are not limited to, any number of central processing units (“CPUs” or “processors”), including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc., memory devices (e.g., dynamic random access memory), storage devices (e.g., solid state or disk drives), network input / output (“NWI / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more node C.R.s of node C.R.s 816(1)-816(N) can be a server having one or more of the above-described computing resources.

[0091] In at least one embodiment, grouped computing resources 814 can include individual groups of node C.R.s housed within one or more racks (not shown), or housed within a number of racks (also not shown) within various geographic locations. Individual groups of node C.R.s within grouped computing resources 814 can include groups of computing, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several node C.R.s including CPUs or processors can be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks can also include any number of power modules, cooling modules, and network switches, in any combination.

[0092] In at least one embodiment, resource orchestrator 812 can configure or otherwise control one or more node C.R.s 816(1)-816(N) and / or grouped computing resources 814. In at least one embodiment, resource orchestrator 812 can include a software design infrastructure (“SDI”) management entity for data center 800. In at least one embodiment, resource orchestrator 108 can comprise hardware, software, or some combination thereof.

[0093] In at least one embodiment, as Figure 8As shown, the framework layer 820 includes a job scheduler 822, a configuration manager 824, a resource manager 826, and a distributed file system 828. In at least one embodiment, the framework layer 820 can include a framework that supports software layer 830 software 832 and / or one or more applications 842 of application layer 840. In at least one embodiment, software 832 or applications 842 can include, respectively, web-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 820 can be, but is not limited to, a type of free and open-source software web application framework such as Apache Spark™ (hereinafter “Spark”) that can utilize the distributed file system 828 for large-scale data processing (e.g., “big data”). In at least one embodiment, the job scheduler 832 can include a Spark driver to facilitate scheduling workloads supported by various layers of the data center 800. In at least one embodiment, the configuration manager 824 can be capable of configuring different layers, such as the software layer 830 and the framework layer 820 including Spark and the distributed file system 828 for supporting large-scale data processing. In at least one embodiment, the resource manager 826 can be capable of managing clustered or grouped computing resources mapped to or allocated for supporting the distributed file system 828 and the job scheduler 822. In at least one embodiment, the clustered or grouped computing resources can include grouped computing resources 814 on the data center infrastructure layer 810. In at least one embodiment, the resource manager 826 can coordinate with the resource orchestrator 812 to manage these mapped or allocated computing resources.

[0094] In at least one embodiment, software 832 included in the software layer 830 can include software used by at least a portion of the node C.R.s 816(1)-816(N), the grouped computing resources 814, and / or the distributed file system 828 of the framework layer 820. One or more types of software can include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.

[0095] In at least one embodiment, one or more application programs 842 included in application layer 840 can include one or more types of application programs used by at least portions of node C.R.s 816(1)-816(N), grouped computing resources 814, and / or distributed file system 828 of framework layer 820. One or more types of application programs can include, but are not limited to, any number and / or type of genomics application programs, cognitive computing and machine learning application programs including training or inferencing software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning application programs used in conjunction with one or more embodiments.

[0096] In at least one embodiment, any of configuration manager 824, resource manager 826, and resource orchestrator 812 can implement any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible fashion. In at least one embodiment, self-modification actions can mitigate poor configuration decisions made by data center operators of data center 800 and can avoid underutilization and / or poorly performing portions of a data center.

[0097] In at least one embodiment, data center 800 can include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information in accordance with one or more embodiments described herein. For example, in at least one embodiment, a machine learning model can be trained according to a neural network architecture by computing weight parameters using software and computing resources described above with respect to data center 800. In at least one embodiment, using weight parameters computed through one or more training techniques described herein, a trained machine learning model corresponding to one or more neural networks can be used to infer or predict information using resources described above with respect to data center 800.

[0098] In at least one embodiment, a data center can use CPUs, application specific integrated circuits (ASICs), GPUs, FPGAs, or other hardware to perform training and / or inference using resources described above. Moreover, one or more software and / or hardware resources described above can be configured as a service to allow users to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.

[0099] Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. Inference and / or training logic 715 are used to process training data and / or inference data in accordance with a training process and / or an inference process. Inference and / or training logic 715 can be used in a variety of machine learning training and inference applications, including without limitation, speech recognition, image recognition, translation, and / or other applications. Figure 7A and / or Figure 7BDetails regarding inference and / or training logic 715 are provided. In at least one embodiment, inference and / or training logic 715 can be used in a system that uses neural network training operations, neural network functions and / or architectures, or neural network use cases described herein to infer or predict operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions, and / or architectures, or neural network use cases described herein. Figure 8

[0100] Such components can be used to generate a sparse voxel grid representation of 3D objects, such as for large-scale scenes.

[0101] Computer system

[0102] Figure 9 is a block diagram illustrating an example computer system, which can be a system with interconnected devices and components, a system on a chip (SOC), or some combination thereof formed with a processor that can include execution units to execute an instruction, according to at least one embodiment. In at least one embodiment, consistent with the present disclosure, for example, embodiments described herein, computer system 900 can include, without limitation, a component, such as processor 902, whose execution units include logic to perform an algorithm for process data. In at least one embodiment, computer system 900 can include a processor, such as a Pentium® processor family, Xeon™, XScale™, and / or StrongARM™, Core TM or Nervana TM microprocessor, although other systems (including PCs, workstations, set-top boxes, etc. with other microprocessors) can also be used. In at least one embodiment, computer system 900 can execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash., although other operating systems, embedded software, and / or graphical user interfaces can also be used.

[0103] ​Embodiments can be used in other devices such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, embedded applications can include a microcontroller, a digital signal processor ("DSP"), a system on a chip, a network computer ("NetPC"), a set-top box, a network hub, a wide area network ("WAN") switch, or any other system that can perform one or more instructions in accordance with at least one embodiment.

[0104] In at least one embodiment, computer system 900 can include, but not be limited to, processor 902, which can include, but not be limited to, one or more execution units 908 to perform machine learning model training and / or inferencing according to techniques described herein. In at least one embodiment, computer system 900 is a single processor desktop or server system, but in another embodiment, computer system 900 can be a multiprocessor system. In at least one embodiment, processor 902 can include, but not be limited to, a complex instruction set computer ("CISC") microprocessor, reduced instruction set computing ("RISC") microprocessor, very long instruction word ("VLIW") microprocessor, a processor implementing a combo of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 902 can be coupled to a processor bus 910 that can transmit data signals between processor 902 and other components in computer system 900.

[0105] In at least one embodiment, processor 902 can include, but not be limited to, level 1 ("Ll") internal cache memory ("cache") 904. In at least one embodiment, processor 902 can have a single -level internal cache or multi-level internal cache. In at least one embodiment, cache memory can reside in the processor 902's external. Other embodiments can include a combination of internal and external caches based on specific implementation and requirements. In at least one embodiment, register file 906 can store different types of data within various registers including, but not limited to, integer registers, floating point registers, status registers, and instruction pointer registers.

[0106] In at least one embodiment, execution unit 908 includes, without limitation, logic to perform integer and floating-point operations, including bit- wide operations. In at least one embodiment, processor 902 can also include a microcode (“ucode”) read only memory (“ROM”) that stores microcode for certain macro instructions. In at least one embodiment, execution unit 908 can also include logic to handle a packed instruction set 909. In at least one embodiment, by including packed instruction set 909 in a general-purpose processor, many multimedia applications can be accelerated by using full width of data bus of processor 902. In one or more embodiments, by using full width of data bus of processor for one or more operations on packed data, many multimedia applications can be executed more efficiently and speedier.

[0107] In at least one embodiment, execution unit 908 can also be used in a microcontroller, embedded processor, graphics device, DSP, and other types of logic circuits. In at least one embodiment, computer system 900 can include, without limitation, memory 920. In at least one embodiment, memory 920 can be implemented as a Dynamic Random Access Memory (“DRAM”) device, a Static Random Access Memory (“SRAM”) device, a flash memory device, or other memory device. In at least one embodiment, memory 920 can store instruction(s) 919 and / or data 921 represented by data signals that can be executed by processor 902.

[0108] In at least one embodiment, a system logic chip can be coupled to processor bus 910 and memory 920. In at least one embodiment, system logic chip can include, without limitation, a memory controller hub (“MCH”) 916 and processor 902 can communicate with MCH 916 via processor bus 910. In at least one embodiment, MCH 916 can provide a high bandwidth memory path 918 to memory 920 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, MCH 916 can also initiate data

[0109] In at least one embodiment, computer system 900 can use system I / O 922, which is a proprietary hub interface bus to couple MCH 916 to I / O controller hub (“ICH”) 930. In at least one embodiment, ICH 930 can provide a direct connection to some I / O devices and indirectly through the local I / O bus. In at least one embodiment, local I / O bus can include, without limitation, a high-speed I / O bus for connecting peripherals to memory 920, chipset, and processor 902. Examples can include, without limitation, audio controller 929, firmware hub (“Flash BIOS”) 928, wireless transceiver 926, data storage 924, legacy I / O controller 923 containing user input and keyboard interfaces, serial expansion port 927 (e.g., Universal Serial Bus (“USB”) port), and network controller 934. Data storage 924 can include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0110] In at least one embodiment, Figure 9 A system is shown that includes interconnected hardware devices or “chips,” while in other embodiments, Figure 9An exemplary system on a chip (SoC) can be shown. In at least one embodiment, devices can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 900 are interconnected using compute express link (CXL) interconnects.

[0111] Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7 A and / or 7B. In at least one embodiment, inference and / or training logic 715 can be used in Figure 7A and / or Figure 7B Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7 A and / or 7B. In at least one embodiment, inference and / or training logic 715 can be used in Figure 9 a system to infer or predict operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0112] Such components can be used to generate a sparse voxel grid representation of 3D objects, such as for large-scale scenes.

[0113] Figure 10 is a block diagram illustrating an electronic device 1000 for utilizing processor(s) 1010, in accordance with at least one embodiment. In at least one embodiment, electronic device 1000 can be, for example and without limitation, a laptop, a tower server, a rack server, a blade server, a laptop computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device.

[0114] In at least one embodiment, system 1000 can include, without limitation, processor(s) 1010 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor(s) 1010 are coupled using a bus or interface, such as a 1C bus, a System Management Bus (“SMBus”), a Low Pin Count (LPC) bus, a Serial Peripheral Interface (“SPI”), a High Definition Audio (“HDA”) bus, a Serial Advanced Technology Attachment (“SATA”) bus, a Universal Serial Bus (“USB”) (versions 1, 2, 3), or a Universal Asynchronous Receiver / Transmitter (“UART”) bus. In at least one embodiment, system 1000 includes, without limitation, a processor(s) 1010 that is a complex instruction set computer (CISC) or a reduced instruction set computer (RISC), or an extreme instruction set computer (EISC), or an Figure 10 A system is shown that includes interconnected hardware devices or “chips,” while in other embodiments, Figure 10 An exemplary system on a chip (SoC) can be shown. In at least one embodiment, Figure 10 devices shown in FIG. 9A can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, Figure 10 one or more components of computer system 900 are interconnected using compute express link (CXL) interconnects.

[0115] In at least one embodiment, Figure 10 It may include a display 1024, a touch screen 1025, a touchpad 1030, a near field communication unit (“NFC”) 1045, a sensor hub 1040, a thermal sensor 1046, a fast chipset (“EC”) 1035, a trusted platform module (“TPM”) 1038, a BIOS / firmware / flash (“BIOS, FWFlash”) 1022, a DSP 1060, a drive 1020 (e.g., a solid-state drive (“SSD”) or a hard disk drive (“HDD”)), a wireless local area network unit (“WLAN”) 1050, a Bluetooth unit 1052, a wireless wide area network unit (“WWAN”) 1056, a global positioning system (GPS) 1055, a camera (“USB 3.0 camera”) 1054 (e.g., a USB 3.0 camera) and / or a low-power double data rate (“LPDDR”) memory unit (“LPDDR3”) 1015 implemented in, for example, the LPDDR3 standard. These components can each be implemented in any suitable way.

[0116] In at least one embodiment, other components may be communicatively coupled to processor 1010 via the components described above. In at least one embodiment, accelerometer 1041, ambient light sensor (“ALS”) 1042, compass 1043, and gyroscope 1044 may be communicatively coupled to sensor hub 1040. In at least one embodiment, thermal sensor 1039, fan 1037, keyboard 1036, and touchpad 1030 may be communicatively coupled to EC 1035. In at least one embodiment, speaker 1063, earphone 1064, and microphone (“mic”) 1065 may be communicatively coupled to audio unit (“audio codec and Class D amplifier”) 1062, which in turn may be communicatively coupled to DSP 1060. In at least one embodiment, audio unit 1062 may include, for example, but not limited to, audio encoder / decoder (“codec”) and Class D amplifier. In at least one embodiment, SIM card (“SIM”) 1057 may be communicatively coupled to WWAN unit 1056. In at least one embodiment, components such as WLAN unit 1050, Bluetooth unit 1052, and WWAN unit 1056 can be implemented as next-generation form factor (NGFF).

[0117] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. The following is combined with... Figure 7A and / or Figure 7B Details are provided regarding the inference and / or training logic 715. In at least one embodiment, the inference and / or training logic 715 can be... Figure 10Inference and / or prediction operations in the system can be based, at least in part, on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0118] Such components can be used to generate sparse voxel grid representations of 3D objects, such as for large-scale scenes.

[0119] Figure 11 is a block diagram of a processing system in accordance with at least one embodiment. In at least one embodiment, system 1100 includes one or more processors 1102 and one or more graphics processors 1108, and can be a single processor desktop system, a multiprocessor workstation system, or a server system having many processors 1102 or processor cores 1107. In at least one embodiment, system 1100 is a processing platform incorporated within a system on a chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices.

[0120] In at least one embodiment, system 1100 can include or be incorporated within a server-based gaming platform, a game console, a media console, a mobile gaming console, a handheld game console, or an online game console that includes game and media processing consoles. In at least one embodiment, system 1100 is a mobile phone, a smart phone, a tablet device, or a mobile internet device. In at least one embodiment, processing system 1100 can also include or be

[0121] In at least one embodiment, one or more processors 1102 each include one or more processor cores 1107 to process instructions which, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 1107 is configured to process a specific instruction set 1109. In at least one embodiment, instruction set 1109 can facilitate Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or computing via a Very Long Instruction Word (VLIW). In at least one embodiment, processor cores 1107 can each process a different instruction set 1109, which can include instructions to facilitate emulation of other instruction sets. In at least one embodiment, processor core 1107 can include other processing devices, such a Digital Signal Processor (DSP).

[0122] In at least one embodiment, processor 1102 includes cache memory 1104. In at least one embodiment, processor 1102 can have single level cache or multiple levels of cache. In at least one embodiment, cache memory is shared among various components of processor 1102. In at least one embodiment, processor 1102 also uses an external cache (e.g., a level three (L3) cache or last level cache (LLC)) (not shown), which can be shared between processor cores 1107 using known cache coherency techniques. In at least one embodiment, processor 1102 additionally includes register file 1106, which can include different types of registers such as integer registers, floating point registers, status registers, and instruction pointer registers, to name a few.

[0123] In at least one embodiment, one or more processor(s) 1102 are coupled with one or more interface bus(es) 1110 for communicating data between processor 1102 and other components of system 1100, such as address, data, or control signals. In at least one embodiment, interface bus 1110 can be a version of a processor bus, such as a direct media interface (DMI) bus, in at least one embodiment. In at least one embodiment, interface bus 1110 is not limited to DMI bus, and can include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In at least one embodiment, processor 1102 includes integrated memory controller 1116 and platform controller hub 1130. In at least one embodiment, memory controller 1116 facilitates communication between memory devices and other components of processing system 1100, while platform controller hub 1130 provides connections to input / output devices via local I / O bus.

[0124] In at least one embodiment, memory device 1120 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase-change memory device, or a device with suitable performance for use as processor memory. In at least one embodiment, memory device 1120 may be used as system memory of processing system 1100 to store data 1122 and instructions 1121 for use when one or more processors 1102 execute an application or process. In at least one embodiment, memory controller 1116 is also coupled to an optional external graphics processor 1112, which may communicate with one or more graphics processors 1108 of processor 1102 to perform graphics and media operations. In at least one embodiment, display device 1111 may be connected to processor 1102. In at least one embodiment, display device 1111 may include one or more internal display devices, such as in mobile electronic devices or laptop devices, or external display devices connected via a display interface (e.g., DisplayPort). In at least one embodiment, the display device 1111 may include a head-mounted display (HMD), such as a stereoscopic display device for virtual reality (VR) or augmented reality (AR) applications.

[0125] In at least one embodiment, platform controller hub 1130 enables peripherals coupled to bridge 1122 to interact with a processor and / or each other over high-speed I / O buses 1120 and 1110. In at least one embodiment, I / O peripherals include, but are not limited to, an audio controller 1146, a network controller 1134, a firmware interface 1128, a wireless transceiver 1126, a touch sensor 1125, a data storage device 1124 (e.g., solid-state drive (SSD), floppy drive, optical drive, etc.). In at least one embodiment, data storage device 1124 can communicate via a storage interface (e.g., SATA) or via a peripheral bus, such as a Peripheral Component Interconnect bus (e.g., PCI, PCI Express). In at least one embodiment, touch sensor 1125 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, wireless transceiver 1126 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, firmware interface 1128 enables communication with system firmware, and can be, for example, a unified extensible firmware interface (UEFI) or the like. In at least one embodiment, network controller 1134 can enable a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) couples with interface bus 1110. In at least one embodiment, audio controller 1146 is a multi-channel high definition audio controller. In at least one embodiment, processing system 1100 includes an optional legacy I / O controller 1140 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to system 1100. In at least one embodiment, platform controller hub 1130 can also connect to one or more Universal Serial Bus (USB) controllers 1142 connect input devices, such as keyboard and mouse 1143 combinations, camera 1144, or other USB input devices.

[0126] In at least one embodiment, memory controller 1116 and instances of platform controller hub 1130 can be integrated into a discrete external graphics processor, such as external graphics processor 1112. In at least one embodiment, platform controller hub 1130 and / or memory controller 1116 can be external to one or more processor(s) 1102. For example, in at least one embodiment, system 1100 can include an external memory controller 1116 and platform controller hub 1130, which can be configured as a memory controller hub and a peripheral controller hub, respectively, in a system-on-a-chip (SoC) configuration.

[0127] Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7L and / or 7M. Figure 7A and / or Figure 7BDetails regarding the inference and / or training logic 715 are provided. In at least one embodiment, some or all of inference and / or training logic 715 can be incorporated with graphics processor 1100. For example, in at least one embodiment, the training and / or inference techniques described herein can use one or more ALUs embodied in a graphics processor. Further, in at least one embodiment, the inference and / or training operations described herein can be accomplished with logic other than that illustrated in Figure 7A and / or Figure 7B weight parameters can be stored in on-chip or off-chip memory and / or registers (shown or not) that configure the ALUs of the graphics processor to perform one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0128] Such components can be used to generate sparse voxel grid representations of 3D objects, such as for large-scale scenes.

[0129] Figure 12 is a block diagram of a processor 1200 having one or more processor cores 1202A-1202N, an integrated memory controller 1214, and an integrated graphics processor 1208, according to at least one embodiment. In at least one embodiment, processor 1200 can include additional cores, up to and including an additional core 1202N represented by a dashed lined in at least one embodiment. In at least one embodiment, each processor core 1202A-1202N includes one or more internal cache units 1204A-1204N. In at least one embodiment, each processor core can also include access to one or more shared cache units 1206.

[0130] In at least one embodiment, internal cache units 1204A-1204N and shared cache unit 1206 represent a cache memory hierarchy within processor 1200. In at least one embodiment, cache units 1204A-1204N can include at least one level of cache memory such as a level one (LI) instruction and data cache, a level two (L2) cache, a level three (L3) cache, a level four (L4) cache, or other levels of cache, in each processor core 1202A-1202N and a shared, multi-level cache unit 1206. In at least one embodiment, cache coherence logic maintains coherency between various cache units 1206 and 1204A-1204N.

[0131] In at least one embodiment, processor 1200 also includes a set of one or more bus controller units 1216 and a system agent core 1210. In at least one embodiment, one or more bus controller units 1216 manage a set of peripheral buses, such as one or more PCI or PCIe buses. In at least one embodiment, system agent core 1210 provides management functionality for various processor components. In at least one embodiment, system agent core 1210 includes one or more integrated memory controllers 1214 to manage access to various external memory devices (not shown), including support for data bus protocols such as DDR SDRAM.

[0132] In at least one embodiment, one or more processor cores 1202A-1202N include support to run in multiple threads simultaneously. In at least one embodiment, system agent core 1210 includes components for coordination and management of processor cores 1202A-1202N during multi-threaded processing. In at least one embodiment, system agent core 1210 can additionally include a power control unit (PCU), including logic and components to regulate one or more power states of processor cores 1202A-1202N and graphics processor 1208.

[0133] In at least one embodiment, processor 1200 also includes graphics processor 1208, which can be configured to perform a graphics processing operations. In at least one embodiment, graphics processor 1208 couples with shared cache unit 1206, and system agent core 1210, including one or more integrated memory controllers 1214. In at least one embodiment, system agent core 1210 also includes a display controller 1211 for driving one or more coupled displays to present graphics processor output. In at least one embodiment, display controller 1211 can also be a separate module coupled with graphics processor 1208 via at least one interconnect, or can be integrated within graphics processor 1208.

[0134] In at least one embodiment, ring based interconnect unit 1212 is used to couple the internal components of the processor 1200. In at least one embodiment, an alternative interconnect unit can be used, such as a point-to-point interconnect, a switched interconnect, or other technology. In at least one embodiment, graphics processor 1208 couples with ring based interconnect unit 1212 via I / O link 1213.

[0135] In at least one embodiment, I / O link 1213 represents at least one of a variety of I / O interconnects, including a package I / O interconnect that facilitates communication between various processor components and a high performance embedded memory module 1218 (e.g., an eDRAM module). In at least one embodiment, each of processor cores 1202A-1202N and graphics processor 1208 uses embedded memory module 1218 as a shared last level cache.

[0136] In at least one embodiment, processor cores 1202A-1202N are homogeneous cores executing a common instruction set architecture. In at least one embodiment, processor cores 1202A-1202N are heterogeneous with respect to instruction set architecture (ISA) in that one or more processor cores 1202A-1202N execute a common instruction set while one or more other processor cores 1202A-1202N execute a subset or a different instruction set. In at least one embodiment, processor cores 1202A-1202N are heterogeneous with respect to microarchitecture in that one or more cores have a relatively high power consumption while one or more power cores have a lower power consumption. In at least one embodiment, processor 1200 can be implemented on or as a SoC integrated circuit.

[0137] Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7L and / or 7M. Figure 7A and / or Figure 7B Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7L and / or 7M. In at least one embodiment, portions or all of inference and / or training logic 715 can be incorporated in processor 1200. For example, in at least one embodiment, training and / or inferencing techniques described herein can use one or more ALUs embodied in graphics processor 1208, graphics cores 1202A-1202N, or other components in Figure 12 In at least one embodiment, portions or all of inference and / or training logic 715 can be incorporated in processor 1200. For example, in at least one embodiment, training and / or inferencing techniques described herein can use one or more ALUs embodied in graphics processor 1208, graphics cores 1202A-1202N, or other components in Figure 7A and / or Figure 7B In at least one embodiment, weight parameters can be stored in on-chip or off-chip memory and / or registers (shown or not) that configure ALUs of graphics processor 1200 to perform one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0138] Such components can be used to generate sparse voxel grid representations of 3D objects, such as for large-scale scenes.

[0139] Virtualized computing platform

[0140] Figure 13 is an example data flow diagram of a process 1300 of generating and deploying image processing and inference pipelines, in accordance with at least one embodiment. In at least one embodiment, process 1300 can be deployed for use with imaging devices, processing devices, and / or other device types at one or more facilities 1302. Process 1300 can be executed within a training system 1304 and / or a deployment system 1306. In at least one embodiment, training system 1304 can be used to perform training, deployment, and implementation of machine learning models (e.g., neural networks, object detection algorithms, computer vision algorithms, etc.) for deployment system 1306. In at least one embodiment, deployment system 1306 can be configured to offload processing and computing resources in a distributed computing environment to reduce infrastructure requirements of facilities 1302. In at least one embodiment, one or more applications in a pipeline can use or call services (e.g., inference, visualization, computation, AI, etc.) of deployment system 1306 during application execution.

[0141] In at least one embodiment, some applications used in an advanced processing and inference pipeline can use machine learning models or other AI to perform one or more processing steps. In at least one embodiment, machine learning models can be trained at facilities 1302 using data 1308 (e.g., imaging data) generated at facilities 1302 (and stored on one or more picture archiving and communication systems (PACS) servers at facilities 1302), can be trained using imaging or sequencing data 1308 from another facility or facilities, or a combination thereof. In at least one embodiment, training system 1304 can be used to provide applications, services, and / or other resources to generate working, deployable machine learning models for deployment system 1306.

[0142] In at least one embodiment, model registry 1324 can be supported by object storage, which can support versioning and object metadata. In at least one embodiment, object storage can be accessed from within a cloud platform through, for example, a cloud storage compatible application programming interface (API). In at least one embodiment, machine learning models within model registry 1324 can be uploaded, listed, modified, or deleted by developers or partners of systems interacting with the API. In at least one embodiment, an API can provide access to methods that allow a user with appropriate credentials to associate a model with an application, such that the model can be executed as part of execution of containerized instantiation of the application.

[0143] In at least one embodiment, training system 1304 Figure 13) can include scenarios in which facility 1302 is training their own machine learning model, or has an existing machine learning model that needs to be optimized or updated. In at least one embodiment, imaging data 1308 generated by imaging devices, sequencing devices, and / or other types of devices can be received. In at least one embodiment, once imaging data 1308 is received, AI assisted annotation 1310 can be used to help generate annotations corresponding to imaging data 1308 to be used as ground truth data for a machine learning model. In at least one embodiment, AI assisted annotation 1310 can include one or more machine learning models (e.g., a convolutional neural network (CNN)) that can be trained to generate annotations corresponding to certain types of imaging data 1308 (e.g., from certain devices). In at least one embodiment, AI assisted annotation 1310 can then be used directly, or can be adjusted or fine-tuned using annotation tools to generate ground truth data. In at least one embodiment, AI assisted annotation 1310, labeled data 1312, or a combination thereof can be used as ground truth data to train a machine learning model. In at least one embodiment, a trained machine learning model can be referred to as output model 1316, and can be used by deployment system 1306, as described herein.

[0144] In at least one embodiment, a training pipeline can include scenarios in which facility 1302 requires a machine learning model for performing one or more processing tasks for one or more applications in deployment system 1306, but facility 1302 can not currently have such a machine learning model (or can not have a model that is optimized, efficient, or effective for this purpose). In at least one embodiment, an existing machine learning model can be selected from model registry 1324. In at least one embodiment, model registry 1324 can include machine learning models trained to perform a variety of different inferencing tasks on imaging data. In at least one embodiment, machine learning models in model registry 1324 can have been trained on imaging data from different facilities (e.g., facilities located remotely from facility 1302). In at least one embodiment, machine learning models can have been trained on imaging data from one location, two locations, or any number of locations. In at least one embodiment, when trained on imaging data from a particular location, training can occur at that location, or at least in a manner that protects confidentiality of the imaging data or limits transfer of the imaging data offsite. In at least one embodiment, once a model is trained, or partially trained, at a location, the machine learning model can be added to model registry 1324. In at least one embodiment, the machine learning model can then be retrained or updated at any number of other facilities, and the retrained or updated model can be used in model registry 1324. In at least one embodiment, a machine learning model can then be selected from model registry 1324 (and referred to as output model 1316), and can be used in deployment system 1306 to perform one or more processing tasks for one or more applications of deployment system.

[0145] In at least one embodiment, a scenario can include facility 1302 needing a machine learning model for performing one or more processing tasks for deploying one or more applications in deployment system 1306, but facility 1302 can not currently have such a machine learning model (or can not have an optimized, efficient, or effective model). In at least one embodiment, a machine learning model selected from model registry 1324 can not be fine-tuned or optimized for imaging data 1308 generated at facility 1302 due to population differences, robustness of training data used to train a machine learning model, diversity of training data anomalies, and / or other issues with training data. In at least one embodiment, AI assisted annotation 1310 can be used to help generate annotations corresponding to imaging data 1308 for use as ground truth data to train or update a machine learning model. In at least one embodiment, labeled data 1312 can be used as ground truth data to train a machine learning model. In at least one embodiment, retraining or updating a machine learning model can be referred to as model training 1314. In at least one embodiment, model training 1314 (e.g., AI assisted annotation 1310, labeled clinical data 1312, or a combination thereof) can be used as ground truth data to retrain or update a machine learning model. In at least one embodiment, a trained machine learning model can be referred to as output model 1316 and can be used by deployment system 1306, as described herein.

[0146] In at least one embodiment, deployment system 1306 can include software 1318, services 1320, hardware 1322, and / or other components, features, and functionality. In at least one embodiment, deployment system 1306 can include a software “stack” such that software 1318 can be built on top of services 1320, and can use services 1320 to perform some or all processing tasks, and services 1320 and software 1318 can be built on top of hardware 1322 and use hardware 1322 to perform processing, storage, and / or other computing tasks of deployment system. In at least one embodiment, software 1318 can include any number of different containers, where each container can execute an instantiation of an application. In at least one embodiment, each application can perform one or more processing tasks in a high-level processing and inference pipeline (e.g., inference, object detection, feature detection, segmentation, image enhancement, calibration, etc.). In at least one embodiment, a high-level processing and inference pipeline can be defined based on a selection of different containers desired or required to process imaging data 1308 (e.g., to convert output back to a usable data type, in addition to receiving and configuring containers for use by each container with imaging data for use and / or use by facility 1302 after processing through the pipeline. In at least one embodiment, a combination of containers within software 1318 (e.g., which make up a pipeline) can be referred to as a virtual instrument (as described in greater detail herein), and a virtual instrument can utilize services 1320 and hardware 1322 to perform some or all processing tasks of applications instantiated in containers.

[0147] In at least one embodiment, a data processing pipeline can receive input data (e.g., imaging data 1308) in a specific format in response to an inference request (e.g., a request from a user of deployment system 1306). In at least one embodiment, input data can represent one or more images, videos, and / or other data representations generated by one or more imaging devices. In at least one embodiment, data can be pre-processed as part of a data processing pipeline to prepare data for processing by one or more applications. In at least one embodiment, post-processing can be performed on output of one or more inference tasks or other processing tasks of a pipeline to prepare output data for a next application and / or to prepare output data for transmission and / or use by a user (e.g., in response to an inference request). In at least one embodiment, inference tasks can be performed by one or more machine learning models, such as trained or deployed neural networks, which can include output models 1316 of training system 1304.

[0148] In at least one embodiment, the tasks of the data processing pipeline can be encapsulated in containers, each container representing a discrete, fully functional instantiation of an application and a virtualized computing environment capable of referencing a machine learning model. In at least one embodiment, containers or applications can be published to a private (e.g., limited access) area of ​​a container registry (described in more detail herein), and trained or deployed models can be stored in a model registry 1324 and associated with one or more applications. In at least one embodiment, an image of an application (e.g., a container image) can be used in the container registry, and once a user selects an image from the container registry for deployment in the pipeline, that image can be used to generate containers for instantiation of the application for use by the user's system.

[0149] In at least one embodiment, a developer (e.g., a software developer, clinician, physician, etc.) can develop, publish, and store an application (e.g., as a container) for performing image processing and / or inference on provided data. In at least one embodiment, a software development kit (SDK) associated with the system can be used to perform development, publication, and / or storage (e.g., to ensure that the developed application and / or container conforms to or is compatible with the system). In at least one embodiment, the developed application can be tested locally using the SDK (e.g., at a first facility, testing data from a first facility), the SDK serving as a system (e.g., Figure 12 System 1200 may support at least some services 1320. In at least one embodiment, since DICOM objects may contain one to hundreds of images or other data types, and due to variations in the data, the developer may be responsible for managing (e.g., setting up constructs for preprocessing built into the application, etc.) the extraction and preparation of incoming data. In at least one embodiment, once verified by system 1300 (e.g., for accuracy), the application becomes available in the container registry for user selection and / or implementation to perform one or more processing tasks on data at the user's facility (e.g., a second facility).

[0150] In at least one embodiment, the developer can then share the application or container over a network for the system (e.g., Figure 13by users of the system 1300). In at least one embodiment, completed and validated applications or containers can be stored in a container registry, and related machine learning models can be stored in a model registry 1324. In at least one embodiment, a requesting entity (which provides an inference or image processing request) can browse the container registry and / or model registry 1324 to obtain applications, containers, datasets, machine learning models, etc., select a desired combination of elements to include in a data processing pipeline, and submit an image processing request. In at least one embodiment, a request can include input data necessary to perform the request (and, in some examples, data related to a patient), and / or can include a selection of applications and / or machine learning models to be executed in processing the request. In at least one embodiment, a request can then be passed to one or more components of the deployment system 1306 (e.g., a cloud) to perform processing of the data processing pipeline. In at least one embodiment, processing by the deployment system 1306 can include referencing elements (e.g., applications, containers, models, etc.) selected from the container registry and / or model registry 1324. In at least one embodiment, once results are generated through the pipeline, the results can be returned to a user for review (e.g., for review in a viewing application suite executed on a local, on-premises workstation or terminal).

[0151] In at least one embodiment, to help process or execute applications or containers in a pipeline, services 1320 can be utilized. In at least one embodiment, services 1320 can include computing services, artificial intelligence (Al) services, visualization services, and / or other service types. In at least one embodiment, services 1320 can provide functionality that is common to one or more applications in software 1318, and thus functionality can be abstracted as a service that can be called or utilized by applications. In at least one embodiment, functionality provided by services 1320 can run dynamically and more efficiently, while also allowing applications to process data in parallel (e.g., using Figure 12the parallel computing platform 1230) to scale well. In at least one embodiment, not every application that requires the same functionality provided by a shared service 1320 must have a corresponding instance of the service 1320, but rather the service 1320 can be shared among and between various applications. In at least one embodiment, as a non-limiting example, a service can include an inference server or engine that can be used to perform detection or segmentation tasks. In at least one embodiment, a model training service can be included that can provide machine learning model training and / or retraining capabilities. In at least one embodiment, a data augmentation service can be further included that can provide GPU-accelerated data (e.g., DICOM, RIS, CIS, REST-compliant, RPC, raw, etc.) extraction, resizing, scaling, and / or other augmentations. In at least one embodiment, a visualization service can be used that can add image rendering effects (e.g., ray tracing, rasterization, de-noising, sharpening, etc.) to add realism to two-dimensional (2D) and / or three-dimensional (3D) models. In at least one embodiment, a virtual instrument service can be included that provides beamforming, segmentation, inference, imaging, and / or support to other applications within a pipeline of a virtual instrument.

[0152] In at least one embodiment, where the services 1320 include an AI service (e.g., an inference service), as part of execution of an application, one or more machine learning models can be executed by invoking (e.g., as an API call) the inference service (e.g., inference server) to execute one or more machine learning models or processing thereof. In at least one embodiment, where another application includes one or more machine learning models for a segmentation task, the application can invoke the inference service to execute the machine learning model for performing one or more processing operations associated with the segmentation task. In at least one embodiment, software 1318 implementing an advanced processing and inference pipeline, including a segmentation application and an anomaly detection application, can be pipelined as each application can invoke the same inference service to perform one or more inference tasks.

[0153] In at least one embodiment, hardware 1322 may include a GPU, CPU, graphics card, AI / deep learning system (e.g., an AI supercomputer, such as NVIDIA's DGX), cloud platform, or a combination thereof. In at least one embodiment, different types of hardware 1322 may be used to provide efficient, specially built support for software 1318 and services 1320 in deployment system 1306. In at least one embodiment, GPU processing may be used to perform local processing (e.g., at facility 1302) within the AI / deep learning system, in the cloud system, and / or other processing components of deployment system 1306 to improve the efficiency, accuracy, and performance of image processing and generation. In at least one embodiment, as a non-limiting example, software 1318 and / or services 1320 may be optimized for GPU processing in relation to deep learning, machine learning, and / or high-performance computing. In at least one embodiment, at least some of the computing environment of deployment system 1306 and / or training system 1304 may be executed in a data center, one or more supercomputers, or high-performance computing systems with GPU-optimized software (e.g., a hardware and software combination of an NVIDIA DGX system). In at least one embodiment, as described herein, hardware 1322 may include any number of GPUs that can be invoked to perform data processing in parallel. In at least one embodiment, the cloud platform may also include GPU-optimized execution for deep learning tasks, GPU processing for machine learning tasks, or other computational tasks. In at least one embodiment, an AI / deep learning supercomputer and / or GPU-optimized software (e.g., as provided on NVIDIA's DGX systems) may be used as a hardware abstraction and scaling platform to execute the cloud platform (e.g., NVIDIA's NGC). In at least one embodiment, the cloud platform may integrate application container cluster systems or coordination systems (e.g., Kubernetes) across multiple GPUs to achieve seamless scaling and load balancing.

[0154] Figure 14 This is a system diagram of an example system 1400 for generating and deploying an imaging deployment pipeline according to at least one embodiment. In at least one embodiment, system 1400 can be used to implement Figure 13 The process 1300 and / or other processes include advanced processing and inference pipelines. In at least one embodiment, system 1400 may include training system 1304 and deployment system 1306. In at least one embodiment, training system 1304 and deployment system 1306 may be implemented using software 1318, service 1320 and / or hardware 1322, as described herein.

[0155] In at least one embodiment, system 1400 (e.g., training system 1304 and / or deployment system 1306) can be implemented in a cloud computing environment (e.g., using cloud 1426). In at least one embodiment, system 1400 can be implemented locally (with respect to a medical service facility), or as a combination of cloud computing resources and local computing resources. In at least one embodiment, access to APIs in cloud 1426 can be limited to authorized users by instituting security measures or protocols. In at least one embodiment, security protocols can include network tokens that can be signed by an authentication (e.g., AuthN, AuthZ, Gluecon, etc.) service and can carry appropriate authorization. In at least one embodiment, APIs (described herein) of a virtual instrument or other instances of system 1400 can be limited to a set of public IPs that have been vetted or authorized for interaction.

[0156] In at least one embodiment, various components of system 1400 can communicate information between each other using any of a plurality of different network types, including but not limited to local area networks (LANs) and / or wide area networks (WANs) via wired and / or wireless communication protocols. In at least one embodiment, communication between facilities and components of system 1400 (e.g., for sending inference requests, for receiving results of inference requests, etc.) can be communicated through one or more data buses, wireless data protocols (Wi-Fi), wired data protocols (e.g., Ethernet), etc.

[0157] In at least one embodiment, similar to training pipelines 1302 described herein with respect to Figure 13 In at least one embodiment, training system 1304 can execute training pipeline 1404. In at least one embodiment, where deployment system 1306 is to use one or more machine learning models in deployment pipeline 1410, training pipeline 1404 can be used to train or retrain one or more (e.g., pre-trained) models, and / or implement one or more pre-trained models 1406 (e.g., without retraining or updating). In at least one embodiment, as a result of training pipeline 1404, output model 1316 can be generated. In at least one embodiment, training pipeline 1404 can include any number of processing steps, such as but not limited to conversion or adaptation of imaging data (or other input data). In at least one embodiment, different training pipelines 1404 can be used for different machine learning models used by deployment system 1306. In at least one embodiment, similar to training pipeline 1302 described herein with respect to Figure 13 Training pipeline 1404 of a first example described herein with respect to Figure 13 Training pipeline 1404 of a second example described herein with respect to Figure 13The training pipeline 1404 of the third example described can be used for the third machine learning model. In at least one embodiment, any combination of tasks within the training system 1304 can be used depending on the requirements of each corresponding machine learning model. In at least one embodiment, one or more machine learning models can already be trained and ready for deployment, so the training system 1304 can not perform any processing on the machine learning model and the one or more machine learning models can be implemented by the deployment system 1306.

[0158] In at least one embodiment, the output model 1316 and / or the pre-trained model 1406 can include any type of machine learning model, depending on implementation or embodiment. In at least one embodiment and without limitation thereto, machine learning models used by the system 1400 can include using linear regression, logistic regression, decision trees, support vector machines (SVM), Naive Bayes, k- nearest neighbors (Knn), k-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutional, recurrent, perceptrons, long / short-term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolutional, generative adversarial, liquid state machines, etc.), and / or other types of machine learning models.

[0159] In at least one embodiment, the training pipeline 1404 can include AI-assisted annotation, as described herein with respect to at least Figure 14In at least one embodiment, labeled clinical data 1312 (e.g., traditional annotations) can be generated by any number of techniques. In at least one embodiment, labels or other annotations can be generated in a drawing program (e.g., annotation program), a computer aided design (CAD) program, a labeling program, another type of application suitable for generating annotations or labels for ground truth, and / or can be hand drawn, in at least one embodiment, ground truth data can be synthetically generated (e.g., from computer models or renderings), realistically generated (e.g., designed and generated from real world data), automatically generated by a machine (e.g., using feature analysis and learning to extract features from data and then generate labels), manually annotated (e.g., by a labeler or annotation specialist defining locations of labels), and / or combinations thereof. In at least one embodiment, for each instance of imaging data 1308 (or other data types used by a machine learning model), there can be corresponding ground truth data generated by training system 1304. In at least one embodiment, AI assisted annotation can be performed as part of deployment pipeline 1410; in addition to or instead of AI assisted annotation included in training pipeline 1404. In at least one embodiment, system 1400 can include a multi-tiered platform that can include a software tier of diagnostic applications (or other application types) (e.g., software 1318) that can perform one or more medical imaging and diagnostic functions. In at least one embodiment, system 1400 can be communicatively coupled to (e.g., via encrypted links) a PACS server network of one or more facilities. In at least one embodiment, system 1400 can be configured to access and reference data from a PACS server to perform operations such as training machine learning models, deploying machine learning models, image processing, inferencing, and / or other operations.

[0160] In at least one embodiment, software tier can be implemented as a secure, encrypted, and / or authenticated API through which applications or containers can be invoked (e.g., called) from an external environment (e.g., facility 1302). In at least one embodiment, applications can then call or execute one or more services 1320 to perform computing, AI, or visualization tasks associated with respective applications, and software 1318 and / or services 1320 can utilize hardware 1322 to perform processing tasks in an efficient and effective manner. In at least one embodiment, communications sent to or received by training system 1304 and deployment system 1306 can occur using a pair of DICOM adapters 1402A, 1402B.

[0161] In at least one embodiment, deployment system 1306 may execute deployment pipeline 1410. In at least one embodiment, deployment pipeline 1410 may include any number of applications, which may be sequential, non-sequential, or otherwise applied to imaging data (and / or other data types) – including AI-assisted annotation, the imaging data being generated by imaging devices, sequencing devices, genomics devices, etc., as described above. In at least one embodiment, as described herein, deployment pipeline 1410 for an individual device may be referred to as a virtual instrument for the device (e.g., a virtual ultrasound instrument, a virtual CT scanner, a virtual sequencing instrument, etc.). In at least one embodiment, for a single device, more than one deployment pipeline 1410 may exist, depending on the desired information from the data generated from the device. In at least one embodiment, a first deployment pipeline 1410 may exist if it is desired to detect an anomaly from an MRI machine, and a second deployment pipeline 1410 may exist if it is desired to perform image enhancement from the output of the MRI machine.

[0162] In at least one embodiment, the image generation application may include processing tasks that utilize machine learning models. In at least one embodiment, a user may wish to use their own machine learning model or select a machine learning model from the model registry 1324. In at least one embodiment, a user may implement their own machine learning model or select a machine learning model to be included in the application performing the processing tasks. In at least one embodiment, the application may be optional and customizable, and by defining the application's construction, the deployment and implementation of the application for a specific user is presented as a more seamless user experience. In at least one embodiment, by leveraging other features of system 1400 (e.g., service 1320 and hardware 1322), the deployment pipeline 1410 can be more user-friendly, provide easier integration, and produce more accurate, efficient, and timely results.

[0163] In at least one embodiment, deployment system 1306 may include a user interface (UI) 1414 (e.g., a graphical user interface, a web interface, etc.) that can be used to select applications to be included in deployment pipeline 1410, deploy applications, modify or change applications or their parameters or configurations, use and interact with deployment pipeline 1410 during setup and / or deployment, and / or otherwise interact with deployment system 1306. In at least one embodiment, although not shown with respect to training system 1304, UI 1414 (or different user interfaces) can be used to select models to be used in deployment system 1306, to select models to be trained or retrained in training system 1304, and / or to otherwise interact with training system 1304.

[0164] In at least one embodiment, in addition to the application coordination system 1428, a pipeline manager 1412 may be used to manage interactions between applications or containers deployed through the pipeline 1410 and services 1320 and / or hardware 1322. In at least one embodiment, the pipeline manager 1412 may be configured to facilitate interactions from application to application, from application to service 1320, and / or from application or service to hardware 1322. In at least one embodiment, although shown as included in software 1318, this is not intended to be limiting, and in some examples, the pipeline manager 1412 may be included in service 1320. In at least one embodiment, the application coordination system 1428 (e.g., Kubernetes, DOCKER, etc.) may include a container coordination system that can group applications into containers as logical units for coordination, management, scaling, and deployment. In at least one embodiment, by associating applications from the deployment pipeline 1410 (e.g., rebuilding applications, splitting applications, etc.) with individual containers, each application can execute in a self-contained environment (e.g., at the kernel level) to improve speed and efficiency.

[0165] In at least one embodiment, each application and / or container (or its image) can be developed, modified, and deployed independently (e.g., a first user or developer can develop, modify, and deploy a first application, and a second user or developer can develop, modify, and deploy a second application separate from the first user or developer). This allows focus on the tasks of a single application and / or container without being hindered by the tasks of another application or container. In at least one embodiment, the pipeline manager 1412 and the application coordination system 1428 can facilitate communication and collaboration between different containers or applications. In at least one embodiment, the application coordination system 1428 and / or the pipeline manager 1412 can facilitate communication and resource sharing between and within each application or container, provided that the expected inputs and / or outputs of each container or application are known to the system (e.g., based on the construction of the application or container). In at least one embodiment, since one or more applications or containers in the deployment pipeline 1410 can share the same services and resources, the application coordination system 1428 can coordinate, load balance, and determine the sharing of services or resources between and within the various applications or containers. In at least one embodiment, the scheduler can be used to track the resource requirements of applications or containers, the current or planned use of these resources, and resource availability. Therefore, in at least one embodiment, the scheduler can allocate resources to different applications and distribute resources between and among applications, taking into account the system's needs and availability. In some examples, the scheduler (and / or other components of the application coordination system 1428) can determine resource availability and distribution based on constraints imposed on the system (e.g., user constraints), such as Quality of Service (QoS), the urgency of data output (e.g., to determine whether to perform real-time processing or delayed processing), etc.

[0166] In at least one embodiment, service 1320, utilized and shared by applications or containers in deployment system 1306, may include computing service 1416, AI service 1418, visualization service 1420, and / or other service types. In at least one embodiment, an application may invoke (e.g., execute) one or more services 1320 to perform processing operations for the application. In at least one embodiment, an application may utilize computing service 1416 to perform supercomputing or other high-performance computing (HPC) tasks. In at least one embodiment, one or more computing services 1416 may be utilized to perform parallel processing (e.g., using parallel computing platform 1430) to process data substantially simultaneously through one or more applications and / or one or more tasks of a single application. In at least one embodiment, parallel computing platform 1430 (e.g., NVIDIA's CUDA) may implement general-purpose computing on a GPU (GPGPU) (e.g., GPU / graphics 1422). In at least one embodiment, the software layer of parallel computing platform 1430 may provide access to the GPU's virtual instruction set and parallel computing elements to execute computing kernels. In at least one embodiment, the parallel computing platform 1430 may include memory, and in some embodiments, memory may be shared between and within multiple containers, and / or between and within different processing tasks within a single container. In at least one embodiment, inter-process communication (IPC) calls may be generated for multiple containers and / or multiple processes within containers to enable the use of the same data (e.g., multiple different stages of one or more applications processing the same information) from a shared memory segment of the parallel computing platform 1430. In at least one embodiment, instead of copying data and moving it to different locations in memory (e.g., read / write operations), the same data in the same memory location can be used for any number of processing tasks (e.g., at the same time, at different times, etc.). In at least one embodiment, this information about the new location of the data can be stored and shared between applications because the resulting data from processing is used to generate new data. In at least one embodiment, the location of the data, and the location of the updated or modified data, may be part of the definition of how the payload in the container is understood.

[0167] In at least one embodiment, AI service 1418 may be used to perform an inference service for executing a machine learning model associated with the application (e.g., a task to perform one or more processing tasks of the application). In at least one embodiment, AI service 1418 may utilize AI system 1424 to execute a machine learning model (e.g., a neural network such as a CNN) for segmentation, reconstruction, object detection, feature detection, classification, and / or other inference tasks. In at least one embodiment, the application deploying pipeline 1410 may use one or more output models 1316 of self-training system 1304 and / or other models of the application to perform inference on imaging data. In at least one embodiment, two or more examples of inference using application coordination system 1428 (e.g., a scheduler) may be available. In at least one embodiment, a first category may include a high-priority / low-latency path that can implement a higher service level protocol, such as for performing inference on urgent requests in emergency situations or for radiologists during diagnostic procedures. In at least one embodiment, a second category may include a standard priority path that can be used for requests that may not be urgent or for situations where analysis can be performed at a later time. In at least one embodiment, the application coordination system 1428 may allocate resources (e.g., services 1320 and / or hardware 1322) based on priority paths for different inference tasks of the AI ​​service 1418.

[0168] In at least one embodiment, shared memory may be installed into AI service 1418 in system 1400. In at least one embodiment, shared memory may operate as a cache (or other storage device type) and may be used to process inference requests from applications. In at least one embodiment, when an inference request is submitted, a set of API instances of deployment system 1306 may receive the request and may select one or more instances (e.g., for best fit, for load balancing, etc.) to process the request. In at least one embodiment, to process the request, the request may be fed into a database, and if not already in the cache, a machine learning model may be located from model registry 1324. A verification step may ensure that an appropriate machine learning model is loaded into the cache (e.g., shared memory), and / or a copy of the model may be saved to the cache. In at least one embodiment, if the application is not already running or there are not enough instances of the application, a scheduler (e.g., the scheduler of pipeline manager 1412) may be used to start the application referenced in the request. In at least one embodiment, if an inference server has not yet been started to execute the model, an inference server may be started. Any number of inference servers may be started for each model. In at least one embodiment, in a pull model that clusters inference servers, the model can be cached whenever load balancing is favorable. In at least one embodiment, the inference servers can be statically loaded into the corresponding distributed servers.

[0169] In at least one embodiment, an inference server running in a container can be used to perform inference. In at least one embodiment, an instance of the inference server can be associated with a model (and optionally multiple versions of the model). In at least one embodiment, if an instance of the inference server does not exist when a request to perform inference on the model is received, a new instance can be loaded. In at least one embodiment, when the inference server is started, a model can be passed to the inference server, allowing the same container to be used to serve different models, as long as the inference server runs as different instances.

[0170] In at least one embodiment, during application execution, an inference request for a given application can be received, and a container (e.g., an instance of a hosted inference server) can be loaded (if not already loaded), and a launcher can be invoked. In at least one embodiment, preprocessing logic within the container can (e.g., using a CPU and / or GPU) load, decode, and / or perform any additional preprocessing on the incoming data. In at least one embodiment, once the data is ready for inference, the container can infer the data as needed. In at least one embodiment, this can include a single inference call for an image (e.g., a hand X-ray) or can request inference for hundreds of images (e.g., a chest CT scan). In at least one embodiment, the application can summarize the results before completion, which may include, but is not limited to, a single confidence score, pixel-level segmentation, voxel-level segmentation, generating visualizations, or generating text to summarize the results. In at least one embodiment, different priorities can be assigned to different models or applications. For example, some models may have a real-time (TAT less than 1 minute) priority, while other models may have a lower priority (e.g., TAT less than 10 minutes). In at least one embodiment, model execution time can be measured from the requesting agency or entity, and may include cooperative network traversal time and inference service execution time.

[0171] In at least one embodiment, the transfer of requests between service 1320 and the inference application can be hidden behind a software development kit (SDK) and robust transfer can be provided via queues. In at least one embodiment, requests are placed in queues via an API for individual application / tenant ID combinations, and the SDK pulls requests from the queues and provides them to the application. In at least one embodiment, the name of the queue can be provided in the environment where the SDK picks up the queue. In at least one embodiment, asynchronous communication via queues may be useful because it allows any instance of the application to pick up work when it becomes available. Results can be sent back via queues to ensure no data loss. In at least one embodiment, queues can also provide the ability to partition work, as the highest priority work can go into a queue connected to a majority of instances of the application, while the lowest priority work can go into a queue connected to a single instance that processes tasks in the order they are received. In at least one embodiment, the application can run on a GPU-accelerated instance generated in cloud 1426, and the inference service can perform inference on the GPU.

[0172] In at least one embodiment, visualization service 1420 can be used to generate visualizations for viewing the output of application and / or deployment pipeline 1410. In at least one embodiment, visualization service 1420 can utilize GPU / graphics 1422 to generate visualizations. In at least one embodiment, visualization service 1420 can implement rendering effects such as ray tracing to generate higher quality visualizations. In at least one embodiment, visualizations can include, but are not limited to, 2D image rendering, 3D volume rendering, 3D volume reconstruction, 2D tomographic slicing, virtual reality display, augmented reality display, etc. In at least one embodiment, a virtualized environment can be used to generate virtual interactive displays or environments (e.g., virtual environments) for system users (e.g., doctors, nurses, radiologists, etc.) to interact with. In at least one embodiment, visualization service 1420 can include an internal visualizer, cinematic and / or other rendering or image processing capabilities or functions (e.g., ray tracing, rasterization, internal optics, etc.).

[0173] In at least one embodiment, hardware 1322 may include GPU / graphics 1422, AI system 1424, cloud 1426, and / or any other hardware for performing training system 1304 and / or deployment system 1306. In at least one embodiment, GPU / graphics 1422 (e.g., NVIDIA's TESLA and / or QUADRO GPUs) may include any number of GPUs that can be used to perform processing tasks for any feature or function of computing service 1416, AI service 1418, visualization service 1420, other services, and / or software 1318. For example, for AI service 1418, GPU / graphics 1422 may be used to perform preprocessing on imaging data (or other data types used by machine learning models), postprocessing on the output of machine learning models, and / or inference (e.g., to execute machine learning models). In at least one embodiment, cloud 1426, AI system 1424, and / or other components of system 1400 may use GPU / graphics 1422. In at least one embodiment, cloud 1426 may include a GPU-optimized platform for deep learning tasks. In at least one embodiment, AI system 1424 may use a GPU, and one or more AI systems 1424 may be used to perform cloud 1426 (or at least part of a task for deep learning or inference). Similarly, although hardware 1322 is shown as a discrete component, this is not intended to be limiting, and any component of hardware 1322 may be combined with or utilized by any other component of hardware 1322.

[0174] In at least one embodiment, AI system 1424 may include a specially built computing system (e.g., a supercomputer or HPC) configured for inference, deep learning, machine learning, and / or other artificial intelligence tasks. In at least one embodiment, in addition to CPU, RAM, memory, and / or other components, features, or functions, AI system 1424 (e.g., NVIDIA's DGX) may also include software (e.g., a software stack) that can be used to perform GPU-optimized tasks using multiple GPUs / graphics 1422. In at least one embodiment, one or more AI systems 1424 may be implemented in a cloud 1426 (e.g., in a data center) to perform some or all of the AI-based processing tasks of system 1400.

[0175] In at least one embodiment, cloud 1426 may include GPU-accelerated infrastructure (e.g., NVIDIA's NGC) that can provide a GPU-optimized platform for performing processing tasks of system 1400. In at least one embodiment, cloud 1426 may include AI system 1424 for performing one or more AI-based tasks of system 1400 (e.g., as a hardware abstraction and scaling platform). In at least one embodiment, cloud 1426 may be integrated with application coordination system 1428 utilizing multiple GPUs to achieve seamless scaling and load balancing between and within applications and services 1320. In at least one embodiment, as described herein, cloud 1426 may be responsible for performing at least some of the services 1320 of system 1400, including computing service 1416, AI service 1418, and / or visualization service 1420. In at least one embodiment, cloud 1426 may perform large and small batch inference (e.g., perform NVIDIA's TENSORRT), provide accelerated parallel computing APIs and platform 1430 (e.g., NVIDIA's CUDA), perform application coordination system 1428 (e.g., KUBERNETES), provide graphics rendering APIs and platform (e.g., for ray tracing, 2D graphics, 3D graphics and / or other rendering techniques to produce higher quality cinematic effects), and / or provide other functionalities for system 1400.

[0176] Figure 15A A data flow diagram of a process 1500 for training, retraining, or updating a machine learning model according to at least one embodiment is shown. In at least one embodiment, a non-limiting example can be used. Figure 14 The system 1400 executes the process 1500. In at least one embodiment, the process 1500 may utilize services and / or hardware, as described herein. In at least one embodiment, the refined model 1512 generated by the process 1500 may be executed by a deployment system for one or more containerized applications in the deployment pipeline.

[0177] In at least one embodiment, model training 1514 may include retraining or updating the initial model 1504 (e.g., a pre-trained model) using new training data (e.g., new input data, such as customer dataset 1506, and / or new ground reality data associated with the input data). In at least one embodiment, to retrain or update the initial model 1504, the output or loss layer of the initial model 1504 may be reset or deleted, and / or replaced with an updated or new output or loss layer. In at least one embodiment, the initial model 1504 may have previously fine-tuned parameters (e.g., weights and / or biases) retained from previous training, so training or retraining 1514 may not require as much time or processing as training the model from scratch. In at least one embodiment, during model training 1514, when generating predictions on the new customer dataset 1506 by resetting or replacing the output or loss layer of the initial model 1504, the parameters of the new dataset may be updated and readjusted based on the loss calculation associated with the accuracy of the output or loss layer.

[0178] In at least one embodiment, the pre-trained model 1506 may be stored in a data store or registry. In at least one embodiment, the pre-trained model 1506 may have been trained at least partially at one or more facilities other than the facility executing process 1500. In at least one embodiment, to protect the privacy and rights of patients, subjects, or customers at different facilities, the pre-trained model 1506 may have been trained locally using locally generated customer or patient data. In at least one embodiment, the pre-trained model 1506 may be trained using cloud and / or other hardware, but confidential, privacy-protected patient data may not be transferred to, used by, or accessed by any component of the cloud (or other non-local hardware). In at least one embodiment, if the pre-trained model 1506 is trained using patient data from more than one facility, the pre-trained model 1506 may have been trained separately for each facility before training on patient or customer data from another facility. In at least one embodiment, such as when customer or patient data has been published for privacy reasons (e.g., by abandonment, for experimental purposes, etc.), or where customer or patient data is included in a public dataset, customer or patient data from any number of facilities can be used to train a pre-trained model 1506 locally and / or externally, such as in a data center or other cloud computing infrastructure.

[0179] In at least one embodiment, when selecting an application for use in the deployment pipeline, the user may also select a machine learning model for a specific application. In at least one embodiment, the user may not have a model available, so the user may select a pre-trained model to use with the application. In at least one embodiment, the pre-trained model may not be optimized to generate accurate results on the user facility's customer dataset 1506 (e.g., based on patient diversity, demographics, type of medical imaging equipment used, etc.). In at least one embodiment, the pre-trained model may be updated, retrained, and / or fine-tuned for use at various facilities before being deployed into the deployment pipeline for use with one or more applications.

[0180] In at least one embodiment, a user may select a pre-trained model to update, retrain, and / or fine-tune, and this pre-trained model may be referred to as the initial model 1504 of the training system in process 1500. In at least one embodiment, a client dataset 1506 (e.g., imaging data, genomic data, sequencing data, or other data types generated by equipment at the facility) may be used to perform model training (which may include, but is not limited to, transfer learning) on ​​the initial model 1504 to generate a refined model 1512. In at least one embodiment, ground-based data corresponding to the client dataset 1506 may be generated by the training system 1304. In at least one embodiment, the ground-based data may be generated at the facility, at least in part, by clinicians, scientists, physicians, or practitioners.

[0181] In at least one embodiment, AI-assisted annotation may be used to generate ground-based data in some examples. In at least one embodiment, AI-assisted annotation (e.g., implemented using an AI-assisted annotation SDK) may leverage machine learning models (e.g., neural networks) to generate suggested or predicted ground-based data for a customer dataset. In at least one embodiment, a user may use the annotation tool within a user interface (graphical user interface (GUI)) on a computing device.

[0182] In at least one embodiment, user 1510 can interact with the GUI via computing device 1508 to edit or fine-tune annotations or automatic annotations. In at least one embodiment, polygon editing features can be used to move the vertices of a polygon to more precise or fine-tuned positions.

[0183] In at least one embodiment, once the customer dataset 1506 has associated ground-based data, the ground-based data (e.g., from AI-assisted annotations, manual labeling, etc.) can be used to generate a refined model 1512 during model training. In at least one embodiment, the customer dataset 1506 can be applied to the initial model 1504 an arbitrary number of times, and the ground-based data can be used to update the parameters of the initial model 1504 until an acceptable level of accuracy is achieved for the refined model 1512. In at least one embodiment, once the refined model 1512 is generated, it can be deployed in one or more deployment pipelines at the facility to perform one or more processing tasks related to medical imaging data.

[0184] In at least one embodiment, the refined model 1512 can be uploaded to a pre-trained model registry for selection by another facility. In at least one embodiment, this process can be performed at any number of facilities, allowing the refined model 1512 to be further refined an arbitrary number of times on a new dataset to generate a more general model.

[0185] Figure 15B This is an example illustration of a client-server architecture 1532 for enhancing an annotation tool using a pre-trained annotation model, according to at least one embodiment. In at least one embodiment, an AI-assisted annotation tool 1536 may be instantiated based on the client-server architecture 1532. In at least one embodiment, the AI-assisted annotation tool 1536 in an imaging application can assist radiologists, for example, in identifying organs and abnormalities. In at least one embodiment, the imaging application may include software tools, as a non-limiting example, that help user 1510 identify several extreme points on a specific organ of interest in a raw image 1534 (e.g., in a 3D MRI or CT scan) and receive automatic annotation results for all 2D slices of that specific organ. In at least one embodiment, the results may be stored in a data store as training data 1538 and used as (e.g., but not limited to) ground-based data for training. In at least one embodiment, when computing device 1508 sends extreme points for AI-assisted annotation, for example, a deep learning model may receive this data as input and return inference results for segmenting organs or abnormalities. In at least one embodiment, a pre-instantiated annotation tool (e.g., Figure 15BThe AI-assisted annotation tool 1536 can be enhanced by making API calls (e.g., API call 1544) to a server (such as annotation assistant server 1540), which may include a set of pre-trained models 1542 stored, for example, in an annotation model registry. In at least one embodiment, the annotation model registry may store the pre-trained models 1542 (e.g., machine learning models, such as deep learning models) pre-trained to perform AI-assisted annotation on specific organs or anomalies. In at least one embodiment, these models can be further updated using a training pipeline. In at least one embodiment, the pre-installed annotation tool can be improved over time as new labeled data is added.

[0186] The various embodiments can be described by the following terms:

[0187] 1. A computer-implemented method, comprising:

[0188] Use a geometric mesh that approximates multiple volume particles to represent one or more objects in the scene;

[0189] Determine the intersection point of the light ray projected for the selected view of the scene with at least a portion of the geometric mesh corresponding to at least one of the volume particles;

[0190] Determine the response value of at least one volume particle corresponding to the intersection point of the light rays; and

[0191] The response value is used to determine the pixel values ​​of the image of the scene to be rendered from the selected view.

[0192] 2. The computer-implemented method as described in Clause 1, wherein the volume particle is a two-dimensional, three-dimensional, or more dimensional particle having anisotropy factors along different dimensions.

[0193] 3. The computer-implemented method as described in Clause 2, further comprising:

[0194] The multiple volume particles are generated in part based on multiple two-dimensional images obtained from multiple views of the scene.

[0195] 4. The computer-implemented method as described in Clause 3, wherein the selected view is different from any of the plurality of views for which the plurality of two-dimensional images are obtained.

[0196] 5. The computer-implemented method as described in Clause 1, wherein the volume particles represent different colors in different viewing directions.

[0197] 6. The computer-implemented method as described in Clause 1, wherein the volume particle corresponds to a local three-dimensional function, the local three-dimensional function comprising at least one of a linear function, a Lagrangian function, a Gaussian distribution function, a Gaussian kernel, or a Gabor kernel.

[0198] 7. The computer-implemented method as described in Clause 1, further comprising:

[0199] It is determined that the light ray intersects with multiple translucent particles; and

[0200] The pixel value corresponding to the light is determined in part based on the response value from one or more intersecting translucent particles, which are at least up to a transmission threshold.

[0201] 8. The computer-implemented method as described in Clause 1, wherein the view corresponds to a distorted or moving virtual camera with a rolling shutter.

[0202] 9. The computer-implemented method as described in Clause 1, wherein determining the intersection of the light rays is accelerated using hardware acceleration.

[0203] 10. The computer-implemented method as described in Clause 1, further comprising:

[0204] The image of the scene is generated to be provided to operations related to at least one of robotics, car navigation, realistic synthetic image generation, or synthetic image relighting.

[0205] 11. At least one processor, comprising:

[0206] Processing logic, the processing logic being used for:

[0207] Generate a geometric mesh approximating multiple volume particles for one or more objects in the scene; determine the intersection point of a ray cast to a selected view of the scene with the geometric mesh associated with at least one of the volume particles;

[0208] Determine the value of a local three-dimensional function represented by the at least one volume particle corresponding to the intersection point of the light rays; and

[0209] The determined values ​​are used to determine the pixel values ​​of the image of the scene to be rendered from the selected view.

[0210] 12. At least one processor as described in Clause 11, wherein the volume particle is a three-dimensional particle having anisotropy factors along different dimensions.

[0211] 13. At least one processor as described in Clause 11, wherein the volume particle corresponds to a local three-dimensional function, the local three-dimensional function comprising at least one of a linear function, a Lagrangian function, or a Gaussian distribution function.

[0212] 14. At least one processor as described in Clause 11, wherein the processing logic is further configured to:

[0213] It is determined that the light ray intersects with multiple translucent particles; and

[0214] The pixel value corresponding to the light is determined in part based on the response value from one or more intersecting translucent particles, which are at least up to a transmission threshold.

[0215] 15. At least one processor as described in Clause 11, wherein said at least one processor is included in at least one of the following:

[0216] A system used to perform simulation operations;

[0217] A system used to perform simulations to test or validate autonomous machine applications;

[0218] Systems used to perform digital twin operations;

[0219] A system for performing optical transmission simulation;

[0220] A system used for rendering graphics output;

[0221] A system used to perform deep learning operations;

[0222] Systems implemented using edge devices;

[0223] Systems used to generate or present virtual reality (VR) content;

[0224] Systems used to generate or present augmented reality (AR) content;

[0225] Systems used to generate or present mixed reality (MR) content;

[0226] A system containing one or more virtual machines (VMs);

[0227] A system that is at least partially implemented in a data center;

[0228] A system for performing hardware tests using simulation;

[0229] Systems for generating synthetic data;

[0230] Systems used to perform generative AI operations;

[0231] A system for performing one or more operations using a Large Language Model (LLM);

[0232] A system for performing one or more operations using a visual language model (VLM);

[0233] A collaborative content creation platform for 3D assets; or

[0234] A system that utilizes cloud computing resources at least in part.

[0235] 16. A system comprising:

[0236] One or more processors are configured to: determine pixel values ​​of an image of a scene, in part by projecting multiple rays corresponding to a specified view and determining the intersections of the multiple rays with a mesh representing volume particles of one or more objects in a scene to be rendered from the specified view, the pixel values ​​corresponding to a given ray calculated using the response values ​​of one or more individual particles that intersect with the rays.

[0237] 17. A system as described in Clause 16, wherein the specified view corresponds to a distorted virtual camera.

[0238] 18. The system as described in Clause 16, wherein the projection of the plurality of rays is accelerated using hardware acceleration.

[0239] 19. The system as described in Clause 16, wherein the volume particle is a three-dimensional particle having anisotropy factors along different dimensions.

[0240] 20. The system as described in Clause 16, wherein said system comprises at least one of the following:

[0241] A system used to perform simulation operations;

[0242] A system used to perform simulations to test or validate autonomous machine applications;

[0243] Systems used to perform digital twin operations;

[0244] A system for performing optical transmission simulation;

[0245] A system used for rendering graphics output;

[0246] A system used to perform deep learning operations;

[0247] Systems used to perform generative AI operations;

[0248] A system for performing one or more operations using a Large Language Model (LLM);

[0249] A system for performing one or more operations using a visual language model (VLM);

[0250] Systems implemented using edge devices;

[0251] Systems used to generate or present virtual reality (VR) content;

[0252] Systems used to generate or present augmented reality (AR) content;

[0253] Systems used to generate or present mixed reality (MR) content;

[0254] A system containing one or more virtual machines (VMs);

[0255] A system that is at least partially implemented in a data center;

[0256] A system for performing hardware tests using simulation;

[0257] Systems for generating synthetic data;

[0258] A collaborative content creation platform for 3D assets; or

[0259] A system that utilizes cloud computing resources at least in part.

[0260] Other variations are within the spirit of this disclosure. Therefore, while the disclosed technology is readily adaptable to various modifications and alternative constructions, certain embodiments thereof are illustrated in the accompanying drawings and have been described in detail above. However, it should be understood that the disclosure is not intended to be limited to one or more specific forms disclosed, but rather, it is intended to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of this disclosure as defined in the appended claims.

[0261] Unless otherwise stated or obviously contradicted by the context, the terms “a,” “an,” and “the,” and similar references, used in the context of describing the disclosed embodiments (particularly in the context of the appended claims), should be interpreted as encompassing both singular and plural forms, rather than as definitions of the terms. Unless otherwise stated, the terms “comprising,” “having,” “including,” and “containing” should be interpreted as open-ended terms (meaning “including, but not limited to”). The term “connection” (referring to a physical connection where not modified) should be interpreted as partially or wholly contained, attached to, or joined together, even with some intervention. Unless otherwise indicated herein, references to numerical ranges herein are intended only as a way of abbreviating each individual value falling within that range, and each individual value is incorporated into the specification as if it were separately described herein. Unless otherwise indicated or contradicted by the context, the use of the terms “set” (e.g., “item set”) or “subset” should be interpreted as a non-empty set comprising one or more members. Furthermore, unless otherwise indicated or contradicted by the context, the term "subset" of a corresponding set does not necessarily refer to an appropriate subset of the corresponding set, but rather the subset and the corresponding set can be equal.

[0262] Unless otherwise explicitly stated or clearly contradicted by the context, connective phrases such as “at least one of A, B, and C” or “at least one of A, B, and C” are understood in the context to generally refer to items, terms, etc., which can be A or B or C, or any non-empty subset of the set A, B, and C. For example, in an illustrative example of a set with three members, the connective phrases “at least one of A, B, and C” and “at least one of A, B, and C” refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Therefore, such connective language is generally not intended to imply that some embodiments require the presence of at least one of A, at least one of B, and at least one of C. Additionally, unless otherwise stated or contradicted by the context, the term “multiple” indicates a plural state (e.g., “multiple items” means multiple items). The number of items in a multiple item is at least two, but may be more if explicitly indicated or indicated by the context. Furthermore, unless otherwise stated or clearly understood from the context, the phrase “based on” means “at least partially based on” rather than “based on only”.

[0263] Unless otherwise indicated herein or clearly contradicted by the context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations thereof and / or combinations thereof) are executed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more application programs) that is executed jointly on one or more processors via hardware or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transient signals (e.g., propagating transient electrical or electromagnetic transmissions) but includes non-transitory data storage circuitry (e.g., buffers, caches, and queues). In at least one embodiment, code (e.g., executable code or source code) is stored on one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) on which executable instructions are stored, which, when executed by one or more processors of a computer system (i.e., as a result of execution), cause the computer system to perform the operations described herein. In at least one embodiment, the set of non-transitory computer-readable storage media comprises multiple non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media lack all the code, but the multiple non-transitory computer-readable storage media collectively store all the code. In at least one embodiment, the executable instructions are executed such that different instructions are executed by different processors; for example, the non-transitory computer-readable storage media store the instructions, and the main central processing unit (“CPU”) executes some instructions while the graphics processing unit (“GPU”) executes other instructions. In at least one embodiment, different components of the computer system have separate processors, and the different processors execute different subsets of the instructions.

[0264] Therefore, in at least one embodiment, the computer system is configured to implement one or more services that perform the operations of the processes described herein, either individually or collectively, and such a computer system is configured with suitable hardware and / or software to enable the implementation of the operations. Furthermore, the computer system implementing at least one embodiment of this disclosure is a single device, and in another embodiment it is a distributed computer system comprising multiple devices operating in different ways, such that the distributed computer system performs the operations described herein, and that a single device does not perform all the operations.

[0265] The use of any and all examples or exemplary language (e.g., “such as”) provided herein is intended only to better illustrate embodiments of this disclosure and does not constitute a limitation on the scope of the disclosure unless otherwise required. No language in the specification should be construed as indicating that any unclaimed element is essential to the practice of the disclosure.

[0266] All references cited in this article, including publications, patent applications and patents, are incorporated herein by reference as if each reference were individually and specifically indicated to be incorporated herein by reference and the entire contents of which are described herein.

[0267] The terms “coupled” and “connected”, and their derivatives, may be used in the specification and claims. It should be understood that these terms may not be intended to be synonyms with each other. Rather, in certain examples, “connected” or “coupled” may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. “Coupled” may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.

[0268] Unless otherwise expressly stated, it will be understood that throughout this specification, terms such as “processing,” “computing,” “determining,” etc., refer to the actions and / or processes of a computer or computing system or similar electronic computing device that process and / or convert data represented as physical quantities (e.g., electrons) in the registers and / or memory of the computing system into other data represented as physical quantities in the memory, registers, or other such information storage, transmission, or display devices of the computing system.

[0269] In a similar manner, the term "processor" can refer to any device or part of memory that processes electronic data from registers and / or memory and converts that electronic data into other electronic data that can be stored in registers and / or memory. As a non-limiting example, a "processor" can be a CPU or a GPU. A "computing platform" can include one or more processors. As used herein, a "software" process can include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Similarly, each process can refer to multiple processes that execute instructions sequentially or intermittently, sequentially, or in parallel. The terms "system" and "method" are used interchangeably herein, provided that a system can embody one or more methods, and a method can be considered a system.

[0270] This document refers to the process of acquiring, obtaining, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. Analog and digital data can be acquired, obtained, received, or input in various ways, such as by receiving data as a parameter to a function call or a call to an application programming interface (API). In some implementations, the process of acquiring, obtaining, receiving, or inputting analog or digital data can be accomplished by transmitting data via a serial or parallel interface. In another implementation, the process of acquiring, obtaining, receiving, or inputting analog or digital data can be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. Reference can also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data can be implemented by transmitting data as an input or output parameter to a function call, an API, or an inter-process communication mechanism.

[0271] While the discussion above illustrates example implementations of the described technologies, other architectures can be used to implement the described functionality and are intended to fall within the scope of this disclosure. Furthermore, although specific assignments of responsibilities have been defined above for discussion purposes, various functions and responsibilities can be assigned and divided in different ways depending on the circumstances.

[0272] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or actions described. Rather, specific features and actions are disclosed as exemplary forms for implementing the claims.

Claims

1. A computer-implemented method, comprising: Use a geometric mesh that approximates multiple volume particles to represent one or more objects in the scene; Determine the intersection point of the light ray projected for the selected view of the scene with at least a portion of the geometric mesh corresponding to at least one of the volume particles; Determine the response value of at least one volume particle corresponding to the intersection point of the light rays; as well as The response value is used to determine the pixel values ​​of the image of the scene to be rendered from the selected view.

2. The computer-implemented method of claim 1, wherein the volume particle is a two-dimensional, three-dimensional, or more dimensional particle having anisotropy factors along different dimensions.

3. The computer-implemented method as described in claim 2, further comprising: The multiple volume particles are generated in part based on multiple two-dimensional images obtained from multiple views of the scene.

4. The computer-implemented method of claim 3, wherein the selected view is different from any of the plurality of views for which the plurality of two-dimensional images are obtained.

5. The computer-implemented method of claim 1, wherein the volume particles represent different colors in different viewing directions.

6. The computer-implemented method of claim 1, wherein the volume particle corresponds to a local three-dimensional function, the local three-dimensional function including at least one of a linear function, a Lagrangian function, a Gaussian distribution function, a Gaussian kernel, or a Gabor kernel.

7. The computer-implemented method of claim 1, further comprising: It is determined that the light ray intersects with multiple translucent particles; as well as The pixel value corresponding to the light is determined in part based on the response value from one or more intersecting translucent particles, which are at least up to a transmission threshold.

8. The computer-implemented method of claim 1, wherein the view corresponds to a distorted or moving virtual camera with a rolling shutter.

9. The computer-implemented method of claim 1, wherein determining the intersection of the light rays is accelerated using hardware acceleration.

10. The computer-implemented method of claim 1, further comprising: The image of the scene is generated to be provided to operations related to at least one of robotics, car navigation, realistic synthetic image generation, or synthetic image relighting.

11. At least one processor, comprising: Processing logic, the processing logic being used for: Generates a geometric mesh that approximates multiple volume particles for one or more objects in the scene; Determine the intersection point of the light ray projected for the selected view of the scene with the geometric mesh associated with at least one of the volume particles; Determine the value of a local three-dimensional function represented by the at least one volume particle corresponding to the intersection point of the light rays; as well as The determined values ​​are used to determine the pixel values ​​of the image of the scene to be rendered from the selected view.

12. The at least one processor of claim 11, wherein the volume particle is a three-dimensional particle having anisotropy factors along different dimensions.

13. The at least one processor of claim 11, wherein the volume particle corresponds to a local three-dimensional function, the local three-dimensional function comprising at least one of a linear function, a Lagrangian function, or a Gaussian distribution function.

14. The at least one processor of claim 11, wherein the processing logic is further configured to: It is determined that the light ray intersects with multiple translucent particles; and The pixel value corresponding to the light is determined in part based on the response value from one or more intersecting translucent particles, which are at least up to a transmission threshold.

15. The at least one processor as claimed in claim 11, wherein the at least one processor is included in at least one of the following: A system used to perform simulation operations; A system used to perform simulations to test or validate autonomous machine applications; Systems used to perform digital twin operations; A system for performing optical transmission simulation; A system used for rendering graphics output; A system used to perform deep learning operations; Systems implemented using edge devices; Systems used to generate or present virtual reality (VR) content; A system for generating or presenting augmented reality (AR) content; A system for generating or presenting mixed reality (MR) content; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; A system for performing hardware tests using simulation; Systems for generating synthetic data; Systems used to perform generative AI operations; A system for performing one or more operations using a large language model (LLM); A system for performing one or more operations using a visual language model (VLM); A collaborative content creation platform for 3D assets; or A system that utilizes cloud computing resources at least in part.

16. A system comprising: One or more processors are configured to: determine pixel values ​​of an image of a scene, in part by projecting multiple rays corresponding to a specified view and determining the intersections of the multiple rays with a mesh representing volume particles of one or more objects in a scene to be rendered from the specified view, the pixel values ​​corresponding to a given ray calculated using the response values ​​of one or more individual particles that intersect with the rays.

17. The system of claim 16, wherein the specified view corresponds to a distorted virtual camera.

18. The system of claim 16, wherein the projection of the plurality of light rays is accelerated using hardware acceleration.

19. The system of claim 16, wherein the volume particle is a three-dimensional particle having anisotropy factors along different dimensions.

20. The system of claim 16, wherein the system comprises at least one of the following: A system used to perform simulation operations; A system used to perform simulations to test or validate autonomous machine applications; Systems used to perform digital twin operations; A system for performing optical transmission simulation; A system used for rendering graphics output; A system used to perform deep learning operations; Systems used to perform generative AI operations; Systems for performing one or more operations using a large language model (LLM); systems for performing one or more operations using a visual language model (VLM); systems implemented using edge devices; Systems used to generate or present virtual reality (VR) content; A system for generating or presenting augmented reality (AR) content; A system for generating or presenting mixed reality (MR) content; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; A system for performing hardware tests using simulation; Systems for generating synthetic data; A collaborative content creation platform for 3D assets; or A system that utilizes cloud computing resources at least in part.