HIERARCHICAL sparse voxel representation for synthetic scene generation

The 3D scene generation architecture addresses memory and computational constraints by using sparse voxel representations and depth prediction, enabling efficient one-shot 3D reconstruction for applications such as autonomous vehicle navigation.

DE102025112876A1Pending Publication Date: 2025-10-09NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102025112876
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-03
Filing Date
2025-04-02
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Conventional 3D scene generation methods require large voxel grids, leading to high computational effort and memory constraints, limiting the ability to capture fine scene details and necessitating iterative optimization schemes.

Method used

A 3D scene generation architecture that uses sparse voxel representations and a depth prediction network to construct a hierarchical volume representation, allowing one-shot prediction without iterative optimization, reducing memory usage and computational requirements.

Benefits of technology

Enables efficient and detailed 3D scene reconstruction in a single step, improving data processing for applications like autonomous vehicle navigation and reducing storage and processing demands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

In various examples, systems and methods are disclosed relating to generating each initial feature map from a plurality of feature maps based on a respective input image of an input data set, wherein each initial feature map incorporating depth data of the respective input image corresponds to a plurality of pixels of the respective input image, generating a sparse feature point cloud containing a plurality of features determined using the plurality of initial feature maps, transforming the sparse feature point cloud into multi-resolution sparse grids, wherein each of the multi-resolution sparse grids comprises a plurality of voxels, modeling the multi-resolution sparse grids using a plurality of neural networks according to a hierarchical architecture to construct a hierarchical volume representation,and providing constructed content based on the hierarchical volume representation.,
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Traditionally, methods for constructing three-dimensional scenes, including neural radiance fields and 3D Gaussian splats, require iterative optimization schemes to construct a 3D representation of the target scene. This limits their applicability for performing other tasks such as online environmental visualization and generative modeling. Traditional generative 3D scene models, such as 3D diffusion models, require an explicit data representation, such as 3D voxel grids.

[0002] The performance of diffusion-based 3D scene generation is influenced by the extent to which the data representation encodes scene details. Furthermore, due to memory constraints, diffusion-based 3D scene generation methods often only use smaller voxel grid representations (e.g., 128 × 128 × 32), limiting the models' ability to capture finer scene details. SUMMARY

[0003] The invention is defined by the claims. To illustrate the invention, aspects and embodiments are described herein, which may or may not be within the scope of the claims.

[0004] Various examples are disclosed that include systems and methods relating to generating an initial feature map from a plurality of initial feature maps based on a respective input image of an input data set, wherein each initial feature map incorporating depth data of the respective input image corresponds to a plurality of pixels of the respective input image, generating a sparse feature point cloud containing a plurality of features determined using the plurality of initial feature maps, transforming the sparse feature point cloud into multi-resolution sparse grids, wherein each of the multi-resolution sparse grids comprises a plurality of voxels, modeling the multi-resolution sparse grids using a plurality of neural networks according to a hierarchical architecture to construct a hierarchical volume representation,and providing constructed content based on the hierarchical volume representation.,

[0005] Approaches according to various embodiments relate to systems, methods, and non-transitory computer-readable media for improving efficiency and memory consumption in 3D scene generation, such as one-shot 3D scene generation from 2D images. In some embodiments, a pipeline is provided for constructing a hierarchical voxel representation of a 3D environment. The hierarchical voxel representation may be used, for example, to reconstruct an environmental view (e.g., for a first-person vehicle or a character). The improved 3D scene generation architecture enables one-shot prediction without iterative optimization, so that the prediction and construction of a 3D representation for any image can be performed in a single step. A hierarchical voxel representation can be constructed in a single shot from a set of given input images.In some examples, hierarchical voxel representation can be used in scene construction. The improved architecture for 3D scene representation described here therefore requires significantly less memory and processing overhead compared to conventional 3D scene generation models. Although currently available memory devices (e.g., the memory of a graphics processing unit (GPU)) are difficult to store the large number of voxels required for conventional 3D scene generation models, the improved architecture for 3D scene generation described here specifies voxels (e.g., volumetric pixels) in a hierarchical manner so that only occupied voxels are stored to reduce computational requirements, particularly during a volume rendering process where only occupied or filled voxels are queried.

[0006] At least one aspect relates to at least one processor. The processor may include one or more circuits for constructing each initial feature map from a plurality of initial feature maps based on a respective input image of an input dataset, wherein each initial feature map includes depth data of the respective input image and corresponds to a plurality of pixels of the respective input image. The one or more circuits of the processor, in one or more embodiments, may also construct a sparse feature point cloud including a plurality of features determined using the plurality of initial feature maps and transform the sparse feature point cloud into multi-resolution sparse grids, wherein each of the multi-resolution sparse grids comprises a plurality of voxels.In one or more embodiments, the one or more circuits of the processor may also model the multi-resolution sparse grids using a plurality of neural networks according to a hierarchical architecture to construct a hierarchical volume representation and provide constructed content based on the hierarchical volume representation.

[0007] At least one aspect relates to at least one processor. The processor may include one or more circuits for determining an initial feature map based on an input data set, wherein the initial feature map, including depth data, corresponds to a plurality of pixels of the input data set; determining a hierarchical volume representation based on multi-resolution sparse grids comprising a plurality of voxels corresponding to a transformed sparse feature point cloud; and providing constructed content based on the volume rendering of the hierarchical volume representation.

[0008] At least one aspect relates to at least one processor. The processor may include one or more circuits for constructing, using a model, a sparse feature point cloud comprising a plurality of features from the plurality of initial feature maps and a plurality of depth maps for a plurality of input images of an input dataset; constructing, using a model, a plurality of sparse grids with different resolutions using the sparse feature point cloud; combining, using a model, a plurality of features from the plurality of sparse grids to determine a hierarchical volume representation; constructing, using a model, an output image using the hierarchical volume representation, wherein the output image is constructed based on a pose of a first input image of the plurality of input images.determine a loss of the output image with respect to the first input image and update the model using the loss.

[0009] Disclosed embodiments may be included in a variety of different systems, such as automotive systems with control systems for an autonomous or semi-autonomous machine (e.g., an AI driver, an on-board infotainment system, and so on) and / or a perception system (e.g., sensor systems, and so on) for an autonomous or semi-autonomous machine, systems implemented using a robot, aviation systems, medical systems, boat systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for generating or presenting virtual reality (VR) content, augmented reality (AR) content, and / or mixed reality (MR) content, systems for performing digital twin operations, systems implemented using an edge device, systems including one or more virtual machines (VMs),Systems for performing operations to generate synthetic data, systems that are at least partially implemented in a data center, systems for performing operations using conversational AI, systems for performing operations using generative AI, systems that implement one or more language models - such as one or more large language models (LLMs) and / or one or more visual language models (VLMs), systems for hosting real-time streaming applications, systems for performing light transport simulations, systems for performing collaborative content creation for 3D assets, systems that are at least partially implemented using cloud computing resources, and / or other types of systems.

[0010] The disclosure extends to all novel aspects or features described and / or illustrated herein.

[0011] Further features of the disclosure are characterized by the independent and dependent claims.

[0012] Any feature of one aspect of the disclosure may be applied to other aspects of the disclosure, in any suitable combination. In particular, method aspects may be applied to device or system aspects, and vice versa.

[0013] Furthermore, functions implemented in hardware can also be implemented in software, and vice versa. Any reference to software and hardware features in this document should be interpreted accordingly.

[0014] Any system or device feature described herein may also be provided as a method feature, and vice versa. System and / or device aspects described functionally (including means plus functional features) may alternatively be expressed in terms of their corresponding structure, such as an appropriately programmed processor and associated memory.

[0015] It should also be noted that certain combinations of the various features described and defined in all aspects of the disclosure may be implemented and / or provided and / or used independently of one another.

[0016] The disclosure also provides computer programs and computer program products comprising software code that, when executed on a data processing device, is capable of performing any of the methods described herein and / or embodying any of the device and system features described herein, including any or all component steps of a method.

[0017] The disclosure also provides a computer or computer system (including networked or distributed systems) having an operating system that supports a computer program for performing any of the methods described herein and / or for embodying any of the device or system features described herein.

[0018] The disclosure also provides a computer-readable medium on which one or more of the aforementioned computer programs are stored.

[0019] The disclosure also provides a signal carrying one or more of the above-mentioned computer programs.

[0020] The disclosure extends to methods and / or devices and / or systems as described herein with reference to the accompanying drawings.

[0021] Aspects and embodiments of the disclosure will now be described, by way of example only, with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The present systems and methods for constructing a hierarchical voxel representation of a 3D environment are described in detail below with reference to the accompanying drawing figures, wherein: Fig. 1 illustrates an exemplary computing environment including a training system for training (e.g., updating) machine learning models and an application system for deploying machine learning models. Fig. 2 is a block diagram of an example model for determining an output image using a multi-view input dataset. Fig. 3 is a diagram representing a frustum constructed using a feature map. Fig. 4 is a block diagram of an example method for using a machine learning model to construct an output image. Fig. 5 is a block diagram of an example method for employing a machine learning model to construct an output image. Fig. 6 is a block diagram of an example method for training a machine learning model to construct an output image. Fig. 7 is a block diagram of an exemplary computing device. Fig. 8 shows an example data center. DETAILED DESCRIPTION

[0023] One of the challenges of conventional 3D scene generation is the large volume of voxel grids required for each scene in a dataset, often numbering in the billions. While conventional voxel grids can represent a scene with large dimensions (such as 1024 × 1024 × 128) in great detail, the resulting computational complexity is enormous—especially considering that computations grow cubically in voxel space. To address this problem, a scene construction model described here can generate 3D scenes using sparse voxel representations.In particular, unlike systems that iteratively update a 3D representation to generate the input image, the system described here uses the depth prediction network to obtain the initial depths and then sparsifies the input image into a sparse voxel grid based on the initial depth, which is then processed using a 3D neural network (e.g., convolutional neural network (CNN)), resulting in a more efficient and straightforward process.

[0024] Another challenge addressed by the 3D scene generation architecture described here is the efficient stitching of images from multiple cameras. The embodiments described here address challenges such as low-resolution images and high storage costs associated with voxel space by creating a sparse structure that eliminates the need to store all voxel entries while creating a hierarchical representation for information at different levels. Furthermore, the 3D scene generation architecture described here enables one-shot prediction without iterative optimization, allowing the prediction of a 3D representation for any arbitrary image to be performed in a single step.

[0025] The 3D scene generation architecture reduces the coarseness of 3D construction by combining different levels of granularity as defined in the hierarchy, resulting in smoother and more detailed output results. While many conventional systems have voxel size limitations, often limited by hardware memory capabilities, the 3D scene generation architecture described here introduces a sparse structure that allows for the growth of more detailed voxels within a scene. In some embodiments, features extracted from three or more levels of a hierarchy can be concatenated, with each level contributing to a component of a feature map.These components consist of vector values ​​and are combined during volume rendering to create a textured representation that results in the rendering of 3D features with improved detail and realism.

[0026] The 3D scene generation architecture described here is applicable to autonomous vehicle applications (e.g., training autonomous artificial intelligence (AI) drivers and calibrating sensors) that require highly accurate and detailed 3D representations of autonomous vehicles' environments to navigate safely. Conventional methods that use dense voxel grids (i.e., non-sparsified structures) often struggle to process and store the massive amounts of data required for high-resolution 3D imaging. By using the depth prediction network to determine initial depths and then converting the initial depths into a sparse voxel grid processed by a 3D neural network, the 3D scene generation architecture described here can improve the data processing process while enabling a one-shot prediction approach.For example, the entire 3D scene can be predicted and reconstructed in a single step, significantly increasing the efficiency and speed with which autonomous vehicles (e.g., their AI drivers) can interpret complex environments, including urban landscapes with multiple moving objects, varying topography, and varying lighting conditions. Accordingly, this one-shot capability ensures safer and more reliable navigation by enabling autonomous vehicles (e.g., their AI drivers) to quickly adapt to dynamic changes in the environment. In some examples, the one-shot 3D scene generation framework uses a single forward pass of neural networks from 2D input images. This contrasts with other scene construction methods, such as neural radiance fields (NeRF), which require an iterative optimization scheme.

[0027] A 3D scene generation model constructs neural fields of a 3D scene, from which a 2D image can be rendered from any viewpoint corresponding, for example, to a visual image sensor (e.g., a camera) in the 3D scene. In implementations related to autonomous vehicles, an autonomous vehicle may contain multiple cameras mounted within it in different poses (e.g., positions and orientations, thus different fields of view (FOVs). Each camera can capture a video or a sequence of images while the autonomous vehicle is moving. Based on the poses of the various cameras mounted in an autonomous vehicle, synthetic videos or a sequence of synthetic images can be constructed, which can be used to train an AI driver.For example, the AI ​​driver of the autonomous vehicle can use such synthetic videos or a sequence of synthetic images to construct instructions for various aspects of the autonomous vehicle (e.g., power supply, motor, steering, braking, suspension, etc.), and the instructions are evaluated to update the AI ​​driver. The synthetic videos or sequence of synthetic images are used instead of real videos / images to reduce the cost and improve the efficiency of training the AI ​​drivers.

[0028] With reference to Fig. 1 represents Fig. 1 illustrates an exemplary computing environment including a training system 100 for training (e.g., updating) machine learning models and an application system 150 for deploying machine learning models, in accordance with some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are provided as examples only. Other arrangements and elements (e.g., machines, interfaces, functions, orderings, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional units that may be implemented as individual or distributed components, or in conjunction with other components, and in any suitable combination and location.Various functions described herein as being performed by units may be performed by hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in memory.

[0029] The training system 100 may be a model 102 (e.g., the model 200 in Fig. 2) train or update. An example of the model 102 includes one or more encoders, neural networks, CNNs, one or more residual neural networks (ResNets), other network types, transformers, or various combinations thereof, and so on. The model 102 may include one or more neural networks. A neural network such as the CNN described here may include an input layer, an output layer, and / or one or more intermediate layers such as hidden layers, each of which may have corresponding nodes. Each component of the model 102 may include various neural network models, including models suitable for manipulating respective 2D data, 3D data, and so on. The model 102 and its components may include a scene construction model, which may include a statistical model that generates new data instances (e.g.,new, artificial, synthetic data (such as artificial, synthesized, or synthetic images or 3D representations and outputs described herein) using existing data (e.g., existing input images upon which a 3D scene is constructed). The new data instances are referred to as output data 106, such as output image 285.

[0030] The training system 100 may train or update the model 102 using the training data 104 as input. The training data 104 may include the input dataset 202, as described in detail herein. The model 102 is trained or updated using the training data 104 so that the model 102 can output the output data 106. The output data 106 may be used to evaluate whether the model 102 has been sufficiently trained / updated to meet a desired performance metric, such as a metric reflecting the accuracy of the model 102 in determining outputs. Such an evaluation may be performed based on various loss types, including reconstruction loss. An overall / aggregated loss may be calculated as the sum or combination of one or more loss types. In some embodiments, the loss function may be constructed using arbitrary target images.For example, for an arbitrary target image x and its camera pose p, a volume rendering can be performed on the hierarchical voxels (constructed based on an input image set) to obtain the corresponding output x'. The reconstruction loss between x' and x can be determined and used to update the model 102 in the manner described.

[0031] For example, the training system 100 may use a function—such as a loss function (e.g., the reconstruction loss or the overall loss)—to evaluate a condition to determine whether the model 102 is (sufficiently) configured to meet the target performance metric. The condition may be a convergence condition, such as a condition that is met when factors such as an output of the function reaching the target performance metric or threshold, a number of training iterations, the convergence of the training of the model 102, or various combinations thereof are taken into account. The function may, for example, be in the form of a mean error, a mean square error, or a mean absolute error function.

[0032] The training system 100 may iteratively apply the training data 104 to update the model 102, evaluate the loss in response to the application of the training data 104, and / or modify the model 102 (e.g., update one or more of its weights and biases). The training system 100 may modify the model 102 by modifying either a weight or a parameter of the model 102. The training system 100 may evaluate the function by comparing a result of the function to a threshold convergence condition, such as a minimum or minimized cost threshold, such that the model 102 is determined to be sufficiently trained (e.g., sufficiently accurate in determining results) when the output of the function is less than the threshold. The training system 100 may output the model 102 when the convergence condition is met.

[0033] The application system 150 may execute or deploy a model 180 to determine responses to the input data 154 (e.g., similar to the input data set 202). The application system 150 may be a system that provides output results (e.g., the output response 188) based on the input data 202, such as data with multiple views of a physical 3D scene. The application system 150 may be implemented by or communicatively coupled to the training system 100, or it may be separate from the training system 100.

[0034] Model 180 may be, or be received as, model 102, a portion thereof, or a representation thereof. For example, a data structure representing model 102 may be used by application system 150 as model 180. The data structure may represent parameters of trained model 102, such as weights or biases, used to configure model 180 based on the training of model 102.

[0035] Data processor 172 may be or include any function, operation, routine, logic, or instruction to perform functions such as processing input data 154 to determine or construct a structured output, such as the data structure of a structured image. Data processor 172 may provide the structured input to a data set generator 176.

[0036] The dataset generator 176 may be or include any function, operation, routine, logic, or instruction to perform functions such as determining input conforming to the model 180 based at least on the structured input. For example, the model 180 may be structured to receive input in a particular format, such as a particular 2D data format or file type that may be expected to contain certain value types. The particular format may include a format that is the same as or analogous to a format used to apply the training data 104 to the model 102 to train the model 102. The dataset generator 176 may identify the particular format of the model 180 and convert the structured input into the particular format.

[0037] The data processor 172 and the data set generator 176 can be implemented as discrete functions or in an integrated function. For example, a single functional processing unit can receive the images / videos and construct the input provided to the model 180 when the images / videos are received.

[0038] The model 180 may construct an output response 188 (e.g., the output image 285, and so on) when the input is received from the dataset generator 176. The output response 188 may represent a 2D image.

[0039] In some implementations, models 102, 180, and 200 may each construct neural fields of a 3D scene from which a 2D image can be represented from any viewpoint, corresponding, for example, to a visual image sensor (e.g., a camera) in the 3D scene. Synthetic videos or a sequence of synthetic images may be constructed or generated from the poses of the various cameras mounted in an autonomous vehicle, based on which an AI driver may operate or be trained. For example, the AI ​​driver of the autonomous vehicle may use such synthetic videos or a sequence of synthetic images to construct instructions for various aspects (e.g., power supply, motor, steering, braking, suspension, and so on) of the autonomous vehicle, and the instructions are evaluated to update the AI ​​driver.Such implementations are useful for constructing a 360-degree view of the autonomous vehicle's environment, such as stitching a 360-degree visualization to assist in automatic or manual parking of the vehicle. In some implementations, the 102, 180, and 200 models can facilitate task perception in autonomous driving research by allowing an AI driver to understand the surrounding 3D scene in a single shot. The 102, 180, and 200 models can be integrated into a 3D recognition model and a motion estimation model configured to obtain instantaneous information about the 3D scene, enabling immediate computation, object detection, and provision of the environmental view visualization, which is not possible using iterative procedures that require time-consuming calibration.

[0040] Fig. 2 is a block diagram of an example of the model 200 for determining an output image 285 using a multi-view input dataset 202 according to various embodiments. Each model described herein Fig. The block shown in Figure 2 may contain one or more data types or one or more types of computational processes that may be performed using any combination of hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. The model 200 includes one or more encoders 210a, 210b, ..., 210n, encoders 220a, 220b, ..., 220n, a voxelization function 250, a CNN 260, and a decoder 280. Each Fig. 2 may also be embodied as computer-usable instructions stored on computer storage media. Each block shown in Fig. The block shown in Figure 2 can be provided by a standalone application, a service, a hosted service (standalone or in combination with another hosted service), or a plug-in for another product, to name a few. Each Fig. The block shown in Figure 2 is also used as an example in relation to the system of Fig. 1. However, these blocks may additionally or alternatively be executed by any system or combination of systems, including but not limited to the systems described herein.

[0041] Fig. 2 illustrates a single forward pass of neural networks of model 200 (e.g., model 102 or 180) from 2D input dataset 202 (e.g., input data 154) to determine output image 285 (e.g., output response 188). This is in contrast to other scene construction methods such as NeRF, which require iterative optimization schemes where multiple iterations are needed to determine the output image. The inputs to model 200 include input dataset 202, which includes a plurality of input images 204a, 204b, ..., 204n (e.g., contents of a real 3D scene). In some embodiments, input dataset 202 includes multi-view inputs. For example, the input images 204a, 204b, ..., 204n are images (e.g., RGB images) that capture the same real, physical 3D scene using cameras arranged in different poses. That is, each of the input images 204a, 204b, ..., 204n is captured from a pose κ that differs from that of another input image. In some examples, the input images 204a, 204b, ..., 204n include multi-view images captured or otherwise obtained at each of a plurality of timestamps. In some examples, the input dataset 202 may include a plurality of multi-view input videos defined by a sequence of images captured in different poses and at multiple timestamps. In implementations related to autonomous vehicles, an autonomous vehicle may include multiple cameras, each mounted therein in different poses (e.g., positions and orientations, thus different fields of view (FOVs). Each camera may capture a video or image sequence (corresponding to a respective one of the input images 204a, 204b, ..., 204n) as the autonomous vehicle moves.

[0042] The input data set 202 is applied as input to a feature encoder (e.g., encoders 210a, 210b, ..., 210n). As shown, each of the input images 204a, 204b, ..., 204n is input to a respective one of the encoders 210a, 210b, ..., 210n to construct an output containing the initial feature maps 215a, 215b, ..., 215n, respectively. Although multiple encoders 210a, 210b, ..., 210n, as shown, process the input images 204a, 204b, ..., 204n in parallel, two or more of the input images 204a, 204b, ..., 204n may be processed sequentially using a same feature encoder, or all of the input images 204a, 204b, ..., 204n may be processed sequentially using one feature encoder. Each of the encoders 210a, 210b, ..., 210n may include a 2D CNN encoder or a scene autoencoder. For example, each of the encoders 210a, 210b, ..., 210n processes a respective one of the input images 204a, 204b, ..., 204n (e.g.,The input images 204a, 204b, ..., 204n are processed separately to construct a respective one of the initial feature maps 215a, 215b, ..., 215n. Each of the initial feature maps 215a, 215b, ..., 215n contains a 2D tensor with dimensions ℝ. H×W×(D+C) , where H and W are smaller than a size or dimension of a corresponding input image on the basis of which the initial feature map is constructed. In some examples, each of the initial feature maps 215a, 215b, ..., 215n includes at least one feature (e.g., a number vector) for each pixel of a corresponding one of the input images 204a, 204b, ..., 204n on the basis of which the initial feature map is constructed.

[0043] The input dataset 202 is applied as input to a depth prediction network (e.g., depth encoders 220a, 220b, ..., 220n). As shown, each of the input images 204a, 204b, ..., 204n is input to a respective one of the encoders 220a, 220b, ..., 220n to construct an output containing the depth maps 225a, 225b, ..., 225n (e.g., depth data, initial depth, and so on). Although multiple encoders 220a, 220b, ..., 220n, as shown, process the input images 204a, 204b, ..., 204n in parallel, two or more of the input images 204a, 204b, ..., 204n may be processed sequentially using the same depth encoder, or all of the input images 204a, 204b, ..., 204n may be processed sequentially using one depth encoder. Each of the encoders 220a, 220b, ..., 220n may include a depth prediction network that can predict a depth (e.g., a depth value) for each pixel of an image.For example, each of the encoders 220a, 220b, ..., 220n processes a respective one of the input images 204a, 204b, ..., 204n (e.g., the input images 204a, 204b, ..., 204n are processed separately) to construct a respective one of the depth maps 225a, 225b, ..., 225n. Each of the depth maps 225a, 225b, ..., 225n contains a depth value for each pixel of a corresponding one of the input images 204a, 204b, ..., 204n, based on which the depth map is constructed. In some examples, the encoders 220a, 220b, ..., 220n are pre-trained models (e.g., a MiDaS depth encoder) that output depth data based on the input images.

[0044] Each of the initial feature maps 215a, 215b, ..., 215n and a corresponding one of the depth maps 225a, 225b, ..., 225n constructed using the same input image 204a, 204b, ... or 204n are combined to form a respective one of the frusta 230a, 230b, ..., 230n. In other words, each of the initial feature maps 215a, 215b, ..., 215n is transformed (e.g., using Lift-Splat-Shoot (LSS)) into a corresponding one of the frusta 230a, 230b, ..., 230n using a corresponding one of the depth maps 225a, 225b, ..., 225n constructed using the same input image 204a, 204b, ... or 204n. For example, the depth map 225a is provided to the encoder 210a as a distortion, condition, or parameter to influence the result of the initial feature map 215a, so that the initial feature map 215a incorporates the depth map 225a. Likewise, the initial feature map 215b incorporates the depth map 225b, ..., and the initial feature map 215n includes the depth map 225n. Each of the frustums 230a, 230b, ..., 230n contains image features and density values ​​for each pixel of the input image, based on which the frustum is constructed along a predefined discrete set of D-depths. Each of the frustums 230a, 230b, ..., 230n is a discrete frustum (with discrete elements) with a size of H×W×D with the camera pose κ for a corresponding input image 204a, 204b, ..., 204n.

[0045] Fig. 3 is a diagram illustrating a frustum 320 constructed from a feature map 310 according to various embodiments. The frustum 320 is a simplified example of each of the frustums 230a, 230b, ..., 230n. The feature map 310 is a simplified example of each of the initial feature maps 215a, 215b, ..., 215n. Each block within the feature map 310 corresponds to a pixel in the input image and has a value corresponding to the image feature and a value corresponding to the density. The feature map 310 has a size of H × W, and the frustum 320 has a size of H × W × D, with the addition of the depth dimension D, which corresponds to the depth dimension along which the depth data of the depth maps 225a, 225b, ..., 225n are obtained. Conceptually, a ray is generated from each block (or pixel) of the feature map 310 (e.g.,Feature space) is projected into a 3D space of the frustum 320, wherein the directions of the rays are defined by the pose κ of the camera on the basis of which the corresponding input image 204a, 204b, ... or 204n is acquired. These rays define the FOV of the camera with which the input image is acquired or lie within this field of view. In other words, the pixel values ​​of the feature map 310 are voxelized based on the rays, or discretized into various entries 321, 322, 323, 324, 325, 326, 327, and 328 (or discrete elements or voxels) of the frustum 320. Each entry, discrete element, or voxel of the frustum 320 is identified using an index or identifier. In some examples, the value of each pixel in feature map 310 may be divided into multiple entries in frustum 320 along a direction of that ray.

[0046] As shown, the frustum 320 is not completely filled during this process. Some, but not all, entries of the frustum 320 are filled based on the feature map 310. The values ​​of the feature map 310 that correspond to depths within a depth range can be used to fill corresponding entries of the frustum 320, and values ​​of the feature map 310 that correspond to depths outside this range are omitted and not included in the frustum 320 and therefore are not stored or further processed. Accordingly, the depth maps 225a, 225b, ..., 225n are used to determine which pixel values ​​of the initial feature maps 215a, 215b, ..., 215n are included in the frustums 230a, 230b, ..., 230n. In some examples, the entries of the frustum 320 are filled with the depth corresponding to the predicted depths of the detected objects as shown in the depth maps 225a, 225b, ..., 225n are defined, is filled, and other entries of the frustum 320 remain unfilled. The depth range for a pixel may be set to contain the predicted depth of each detected object at that pixel. In the example where the predicted depth of a detected object at a pixel of the input image is 11 meters, the depth range (e.g., 10-12 meters) may contain a margin (e.g., 1 meter) that is greater or less than the predicted depth, or the depth range (e.g., 10-15 meters) may be one of a plurality of predefined depth ranges (e.g., 0-5 meters, 5-10 meters, 10-15 meters, and so on). The sparsity of the frustum 320 may be over 80%, 90%, 95%, or more.

[0047] The partially filled frusta 230a, 230b, ..., 230n are combined or merged to construct the sparse feature point cloud 240 (or sparse point cloud, a sparse voxel grid, and so on). The voxels of the frusta 230a, 230b, ..., 230n have physical meaning and are located in the same coordinate system as the 3D scene captured using the input images 204a, 204b, ..., 204n. Since the poses of the cameras capturing the input images 204a, 204b, ..., 204n are known, the voxels of the frusta 230a, 230b, ..., 230n can be merged using the poses of the respective cameras capturing the input images 204a, 204b, ..., 204n as reference points to construct the sparse feature point cloud 240 within a unified coordinate system. For example, the first terms of each of the frusta 230a, 230b, ..., 230n for the input images 204a, 204b, ..., 204n with different poses. The sparse feature point cloud 240 may also be referred to as a common voxel grid. For example, the sparse feature point cloud 240 may contain voxels whose feature is each obtained by combining or merging (e.g., adding) the features (e.g., the values) of the frusta 230a, 230b, ..., 230n at this position in the sparse feature point cloud 240. Each feature of the sparse feature point cloud 240 contains a number vector.

[0048] The resulting sparse feature point cloud 240 is also sparse because the output information of the frusta 230a, 230b, ..., 230n is sparse. Instead of storing all entries of the frusta 230a, 230b, ..., 230n and the sparse feature point cloud 240, only entries (e.g., voxels) that are occupied are stored, thereby significantly improving storage and computation efficiency. The sparse feature point cloud 240 may, for example, contain a large number of points corresponding to a 3D scene. The values ​​for a large number of these points remain unpopulated. The sparsity of the sparse feature point cloud 240 may be greater than 80%, 90%, or 95%. The sparse feature point cloud 240 and the frusta 230a, 230b, ..., 230n are referred to as sparse structures, which significantly reduce the computation and storage costs.

[0049] The feature point cloud 240 is voxelized at 250 into multi-resolution, sparse grids (e.g., sparse grids 255a, 255b, ..., 255n). The sparse grids 255a, 255b, ..., 255n have different resolutions and form a multi-resolution, hierarchical structure to provide different types and levels of detail of the 3D scene. For example, the sparse grid 255a has the highest resolution (e.g., 1024 3 Voxels for the 3D scene, smallest voxel size, highest granularity), the sparse grid 255b has the second highest resolution (e.g. 256 3 voxel, second smallest voxel size, second highest granularity), ..., and the sparse grid 225n has the lowest resolution (e.g. 64 3Voxel, largest voxel size, lowest granularity). The coarsest or lowest resolution sparse grid 225n can provide global properties of the 3D scene, such as the presence of a vehicle. The higher resolution sparse grid 225b can provide properties of group components of the 3D scene, such as a front portion of the vehicle. The highest resolution sparse grid 225c can provide detailed properties of the 3D scene, such as the handle of a door in the front portion of the vehicle.

[0050] The hierarchy of multi-resolution, sparse grids 255a, 255b, ..., 255n represents the same objects at different levels of granularity, which is useful for the scene construction model to understand the semantics of the 3D scene and improve the understanding of object placement and pixel coherence. This allows the scene construction model to construct an object-oriented output instead of a group of pixels without context or coherence.

[0051] Each of the sparse grids 255a, 255b, and 255n is independently processed using a respective one of the 3D CNNs 260a, 260b, ..., 260n to determine a hierarchical volume representation. In other words, the sparse grids 255a, 255b, ..., 255n are used as inputs to the respective CNNs 260a, 260b, ..., 260n for feature construction. Each feature constructed by the CNNs 260a, 260b, ..., 260n contains a number vector. For example, the CNN 260a can process the sparse grid 255a to construct at least one feature for each voxel of a 3D space corresponding to the highest resolution (e.g., 1024 3 voxels), the CNN 260b may process the sparse grid 255b to construct at least one feature for each voxel of a 3D space corresponding to the second resolution (e.g., 256 3voxels)..., and the CNN 260n may process the sparse grid 255n to construct at least one feature for each voxel of a 3D space corresponding to the lowest resolution (e.g., 64 3 Voxels). The hierarchical volume representation contains at least one feature for the 3D space that corresponds to the different resolutions of the hierarchy.

[0052] In some examples, the at least one output feature constructed by a CNN with a higher resolution is input as an additional input or constraint to the CNN, which is configured to construct at least one output feature at the resolution of the immediately lower rank in the hierarchy to provide contextual information. For example, the at least one output feature constructed by CNN 260a is provided to CNN 260b along with sparse grid 255b, the at least one output feature constructed by CNN 260n-1 is provided to CNN 260n along with sparse grid 255n, and so on.

[0053] In some examples, the sparse grids 255a, 255b, and 255n are passed through separate 3D CNN layers (e.g., CNNs 260a, 260b, ..., 260n corresponding to different resolutions, from a lower to a higher resolution) to determine the final processed sparse grids used for a volume rendering process. Each of the sparse grids 255a, 255b, and 255n is queried, and the retrieved features are concatenated together to form the volume-rendered feature map 275. In some examples, each of the CNNs 260a, 260b, ..., 260n includes a diffusion model that can construct an output based on inputs, including random noise. In some embodiments, random noise for a first resolution is applied as input to a first depth CNN to construct the at least one feature corresponding to the first resolution (e.g., 64 3voxels), and random noise for a second resolution is applied as input to a second deep CNN to construct, conditional on the at least one feature corresponding to the first resolution, the at least one feature corresponding to the second resolution (e.g., 256 3 voxels), and random noise for a third resolution is applied as input to a third depth CNN to construct, conditional on the at least one feature corresponding to the first resolution and the at least one feature corresponding to the second resolution, the at least one feature corresponding to the third resolution (e.g., 1024 3 voxels).

[0054] Volume rendering 270 of the hierarchical volume representation, including the combined output results of the CNNs 260a, 260b, ..., 260n, is performed with respect to a target pose (e.g., a target camera) to construct a volume-rendered feature map 275 (or a new feature map). The features output by the CNNs 260a, 260b, ..., 260n are combined (e.g., concatenated) to construct the hierarchical volume representation, which is applied as input to the volume rendering 270 to construct the volume-rendered feature map 275. The volume-rendered feature map 275 therefore has a component from each hierarchy level (e.g., from each of the sparse grids 255a, 255b, and 255n and each of the CNNs 260a, 260b, ..., 260n). For example, the vectors for the plurality of features constructed by the CNNs 260a, 260b, ..., 260n are combined or merged (e.g.,concatenated) to construct the hierarchical volume representation. Combining the vectors includes, for example, merging the features constructed by the CNNs 260a, 260b, ..., 260n. The volume-rendered feature map 275 is a 2D projection of the hierarchical volume representation with respect to a target capture device (e.g., in the target pose of the target camera).

[0055] The volume-rendered feature map 275 is input to a decoder 280 (which, in one or more embodiments, includes at least one neural network), which decodes the volume-rendered feature map 275 to output an output image 285 corresponding to the pose. Examples of the decoder 280 may be a CNN decoder. The combined vectors are volume-rendered at 270 and decoded with the decoder 280.

[0056] In some embodiments, the target camera pose on which volume rendering 270 is performed may be the same as the camera pose of one of the input images 204a, 204b, ..., 204n. In a training pipeline, a reconstruction loss for the output image 285 relative to the input image may be determined, wherein the output image 285 and the input image have the same camera pose. The model 200, e.g., one or more of the encoders 210a, 210b, ..., 210n, the encoders 220a, 220b, ..., 220n, the CNNs 260a, 260b, ..., 260n, and the decoder 280, may be updated using the reconstruction loss. For example, one or more of the encoders 210a, 210b, ..., 210n, the encoder 220a, 220b, ..., 220n, the CNNs 260a, 260b, ..., 260n and the decoder 280 may be modified (e.g., one or more weights and biases thereof may be updated) to minimize the reconstruction loss.

[0057] In some examples, an autoencoder model (including encoder 210a, 210b, ..., 210n) encodes multi-view input images of input dataset 202 into 3D density and feature voxel grids, such as sparse feature point cloud 240. Each voxel in the 3D voxel grid (e.g., sparse feature point cloud 240) has a feature vector (populated) or is empty (e.g., unpopulated). Volume rendering in voxel space may construct a 2D view (e.g., volume-rendered feature map 275) with respect to a camera pose passing through each voxel in the 3D voxel grid based on the feature vector, substantially flattening the 3D voxel grid into a 2D view based on the camera pose. A 2D CNN decoder 280 can render the 2D view into an output image 285. Accordingly, the model 200 constructs output images 285 (the reconstructions of input images 204a, 204b, ..., 204n, assuming the same camera poses) using a 3D space (e.g., the 3D voxel grid) constructed from the input images 204a, 204b, ..., 204n.

[0058] Fig. 4 is a block diagram of an example method 400 for employing a machine learning model (e.g., model 200) to construct output image 285. Each block of method 400 described herein may include one or more data types or one or more types of computational processes that may be performed using any combination of hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. Method 400 may also be embodied as computer-usable instructions stored on computer storage media. Method 400 may be provided by a standalone application, a service, or a hosted service (standalone or in combination with another hosted service), or a plug-in for another product, to name a few.Furthermore, the method 400 is exemplified with respect to the system of . Fig. 1 (e.g. model 102 and 180) and Fig. 2 (e.g., the Model 200). However, the method 400 may additionally or alternatively be performed by any system or combination of systems, including, but not limited to, the systems described herein.

[0059] At 410, the model 200 (e.g., a respective one of the encoders 210a, 210b, ..., and 210n) constructs at least one initial feature map from a plurality of initial feature maps 215a, 215b, ..., 215n based on a respective input image 204a, 204b, ..., or 204n of the input dataset 202. Each of the at least one initial feature map incorporates depth data (e.g., depth map 225a, 225b, ..., or 225n) of the respective input image and corresponds to a plurality of pixels of the respective input image. In some examples, each initial feature map includes at least one feature for each pixel of the respective input image. In some examples, the input dataset 202 includes the input image 204a, 204b, ..., or 204n of a 3D scene. In some examples, a depth encoder (e.g., an encoder 220a, 220b, ..., or 220n) constructs a depth map (e.g., a depth map 225a, 225b, ..., or 225n) of the respective input image with the respective input image as input.A feature encoder (e.g., encoder 210a, 210b, ..., or 210n) with the respective input image as input constructs each initial feature map. Each initial feature map is transformed into a frustum (e.g., frustum 230a, 230b, ..., or 230n) using the depth map.

[0060] At 420, the model 200 constructs a sparse feature point cloud 240 containing a plurality of features determined using the plurality of initial feature maps 215a, 215b, ..., 215n. Features of each of the plurality of initial feature maps 215a, 215b, ..., 215n that correspond to a depth within at least one depth range are used to populate entries in a respective one of the frusta 230a, 230b, ..., 230n, which are intermediate 3D structures. The depths of the features of the initial feature maps 215a, 215b, ..., 215n are specified by the integrated depth data.

[0061] At 430, the model 200 transforms (e.g., through the voxelization function 250) the at least one sparse feature point cloud 240 into a plurality of multi-resolution sparse grids, including the sparse grids 255a, 255b, ..., 255n. Each of the multi-resolution sparse grids includes a plurality of voxels. In some examples, a first multi-resolution sparse grid 255a includes a first voxel size (e.g., 1024 3 voxels for a 3D scene) corresponding to a first granularity. A second multi-resolution, sparse grid 255b contains a second voxel size (e.g., 256 3 Voxels for a 3D scene), which corresponds to a second granularity.

[0062] At 440, the model 200 models the multi-resolution, sparse grids 255a, 255b, ..., 255n using a plurality of neural networks (e.g., CNNs 260a, 260b, ..., 260n) according to a hierarchical architecture (e.g., the architecture with multiple resolution or granularity levels) to construct a hierarchical volume representation. In some examples, a first neural network (e.g., CNN 260a) of the plurality of neural networks processes the first multi-resolution, sparse grid 255a having the first voxel size. A second neural network (e.g., CNN 260b) of the plurality of neural networks processes the second multi-resolution, sparse grid 255b having the second voxel size.

[0063] At 450, the model 200 generates constructed content (e.g., the output image 285) based on the volume rendering 270 of the hierarchical volume representation. For example, the volume rendering 270 of the hierarchical volume representation constructs a new feature map (e.g., the volume-rendered feature map 275). The new feature map includes a 2D projection of the hierarchical volume representation with respect to a target capture device (e.g., the target pose of a target camera). Providing the constructed content based on the hierarchical volume representation includes decoding the new feature map using a decoder neural network (e.g., the decoder 280).

[0064] The volume-rendered feature map 275 includes a first component corresponding to a first level of the hierarchical architecture and a second component corresponding to a second level of the hierarchical architecture. The vectors for a plurality of features constructed by the plurality of neural networks (e.g., the CNNs 260a, 260b, ..., 260n) are combined to construct the hierarchical volume representation.

[0065] In some examples, method 400 further includes determining and updating a hierarchical encoder to reduce the dimensionality of each hierarchical voxel level of a hierarchical voxel representation and outputting the hierarchical voxel representation as compressed latent variables. In some examples, method 400 further includes determining and updating a multi-layer neural network by querying a subset of the plurality of voxels using coordinates. The plurality of features in the hierarchical volume representation are matched. A compressed representation of the hierarchical volume representation is output.In some examples, determining the hierarchical encoder and the multi-layer neural network includes a first stage corresponding to compressing each hierarchical voxel level and a second stage corresponding to compressing the hierarchical voxel representation into a final latent representation. In some examples, the plurality of neural networks includes a plurality of diffusion models. The plurality of diffusion models are used to model the plurality of voxels to construct the hierarchical volume representation.

[0066] Fig. 5 is a block diagram of an example method 500 for employing a machine learning model (e.g., model 200) to construct output image 285. Each block of method 500 described herein may include one or more data types or one or more types of computational processes that may be performed using any combination of hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. Method 500 may also be embodied as computer-usable instructions stored on computer storage media. Method 500 may be provided by a standalone application, a service, or a hosted service (standalone or in combination with another hosted service), or a plug-in for another product, to name a few.Furthermore, the method 500 is exemplified with respect to the system of . Fig. 1 (e.g. model 102 and 180) and Fig. 2 (e.g., the Model 200). However, the method 500 may additionally or alternatively be performed by any system or combination of systems, including, but not limited to, the systems described herein.

[0067] At 510, the model 200 determines an initial feature map 215a, 215b, ..., or 215n based on an input data set 202. The initial feature map, which includes depth data (e.g., depth map 225a, 225b, or 225n), corresponds to a plurality of pixels of the input data set 202. At 520, the model 220 determines a hierarchical volume representation based on multi-resolution, sparse grids 225a, 225b, ..., 225n and includes a plurality of voxels corresponding to a transformed sparse feature point cloud 240. At 530, the model 200 provides constructed content (e.g., output image 285) based on the volume rendering 270 of the hierarchical volume representation.

[0068] Fig. 6 is a block diagram of an example method 600 for training (e.g., updating) a machine learning model (e.g., model 200) to construct output image 285. Each block of method 600 described herein may include one or more data types or one or more types of computational processes that may be performed using any combination of hardware, firmware, and / or software. Various functions may be performed, for example, by a processor executing instructions stored in memory. Method 600 may also be embodied as computer-usable instructions stored on computer storage media. Method 600 may be provided by a standalone application, a service, or a hosted service (standalone or in combination with another hosted service), or a plug-in for another product, to name a few.Furthermore, the method 600 is exemplified with respect to the system of . Fig. 1 (e.g. model 102 and 180) and Fig. 2 (e.g., the Model 200). However, the method 600 may additionally or alternatively be performed by any system or combination of systems, including, but not limited to, the systems described herein.

[0069] At 610, the model 200 constructs a sparse feature point cloud 240 containing a plurality of features of the plurality of initial feature maps 215a, 215b, ..., 215n and a plurality of depth maps 225a, 225b, ..., 225n for a plurality of input images 204a, 204b, ..., 204n of an input data set 202.

[0070] At 620, the model 200 constructs a plurality of sparse grids 255a, 255b, ..., 255n with different resolutions using the sparse feature point cloud 240a. At 630, the model 200 combines a plurality of features from the plurality of sparse grids 255a, 255b, ..., 255n to determine a hierarchical volume representation.

[0071] At 640, the model 200 constructs an output image 285 using the hierarchical volume representation. The output image 285 is constructed based on a pose of a first input image (e.g., input image 204a) from the plurality of input images 204a, 204b, ..., 204n. The model 200 includes a decoder 280 for decoding the new feature map to construct the output image 285.

[0072] At 650, the training system determines a loss of the output image 285 relative to the first input image. At 660, the training system updates the model 200 using the loss. In some examples, the loss includes the reconstruction loss.

[0073] In some examples, features from the plurality of initial feature maps corresponding to depths within at least one depth range are used to populate entries in a respective one of a plurality of frusta, wherein the depths of the features of each of the plurality of initial features are indicated by a respective one of the plurality of depth maps.

[0074] In some examples, the plurality of sparse grids includes a first multi-resolution sparse grid 255a having a first voxel size and a second multi-resolution sparse grid 255b having a second voxel size. In some examples, the model 200 includes a first neural network (e.g., CNN 260a) of the plurality of neural networks to process the first multi-resolution sparse grid 255a having the first voxel size, and a second neural network (e.g., CNN 260b) of the plurality of neural networks to process the second multi-resolution sparse grid 255b having the second voxel size.

[0075] In some examples, the model 200 constructs a new feature map 275 by volume rendering 270 the hierarchical volume representation. The new feature map includes a 2D projection of the hierarchical volume representation corresponding to a target capture device in the pose. In some examples, the new feature map includes a first component corresponding to a first level of the hierarchical architecture and a second component corresponding to a second level of the hierarchical architecture. The method 600 further includes combining vectors for a plurality of features constructed by a plurality of neural networks 260a, 260b, ..., 260n to construct the hierarchical volume representation.

[0076] In some examples, the model 180 or 200 may be implemented in one or more systems, such as automotive systems with control systems for an autonomous or semi-autonomous machine (e.g., an AI driver, an on-board infotainment system, and so on) and / or a perception system (e.g., sensor systems, and so on) for an autonomous or semi-autonomous machine), systems implemented using a robot, aviation systems, medical systems, boat systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for generating or presenting virtual reality content, augmented reality content, and / or mixed reality content, systems for performing digital twin operations, systems implemented using an edge device, systems that include one or more VMs,Systems for performing operations to generate synthetic data, systems implemented at least partially in a data center, systems for performing operations using conversational AI, systems for performing operations using generative AI, systems implementing one or more language models—such as one or more large-scale language models (LLMs), one or more vision language models (VLMs), systems for hosting real-time streaming applications, systems for performing light transport simulations, systems for performing collaborative content creation for 3D assets, systems implemented at least partially using cloud computing resources, and / or other types of systems, or the application system 150 may include one or more such systems. EXAMPLE CALCULATION DEVICE

[0077] Fig. 7 is a block diagram of exemplary computing device(s) 700 suitable for use in implementing some embodiments of the present disclosure. The computing device(s) 700 are exemplary implementations of the training system 100 and / or the application system 150. The computing device 700 may include an interconnection system 702 that directly or indirectly couples the following devices: memory 704, one or more central processing units (CPUs) 706, one or more graphics processing units (GPUs) 708, a communications interface 710, input / output (I / O) ports 712, input / output components 714, a power supply 716, one or more presentation components 718 (e.g., display(s)), and one or more logic units 720.In at least one embodiment, the computing device(s) 700 may include one or more virtual machines (VMs), and / or each of the components thereof may include virtual components (e.g., virtual hardware components). As non-limiting examples, one or more of the GPUs 708 may include one or more vGPUs, one or more of the CPUs 706 may include one or more vCPUs, and / or one or more of the logic units 720 may include one or more virtual logic units. Thus, a computing device 700 may include discrete components (e.g., an entire GPU associated with the computing device 700), virtual components (e.g., a portion of a GPU associated with the computing device 700), or a combination thereof.

[0078] Although the different blocks of Fig. 7 as being connected to wires via the interconnect system 702, this is not to be construed as a limitation and is for clarity only. For example, in some embodiments, a presentation component 718, such as a display device, may be considered an I / O component 714 (e.g., if the display is a touchscreen). As another example, the CPUs 706 and / or GPUs 708 may include memory (e.g., the memory 704 may represent a storage device in addition to the memory of the GPUs 708, the CPUs 706, and / or other components). In other words, the computing device of Fig. 7 is merely exemplary. No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "game console", "electronic control unit (ECU)", "virtual reality system" and / or other device or system types, as they all fall within the scope of the computing device of Fig. 7 fall.

[0079] The interconnect system 702 may represent one or more connections or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 702 may include one or more types of buses or connections, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or connection. In some embodiments, there are direct connections between the components. For example, the CPU 706 may be directly connected to the memory 704. Further, the CPU 706 may be directly connected to the GPU 708. For a direct or point-to-point connection between components, the interconnect system 702 may include a PCIe connection to establish the connection.In these examples, a PCI bus need not be integrated into the computing device 700.

[0080] Memory 704 may include any of a variety of computer-readable media. The computer-readable media may be any available media accessible by the computing device 700. The computer-readable media may include both volatile and non-volatile media, as well as removable and non-removable media. By way of example and without limitation, the computer-readable media may include computer storage media and communication media.

[0081] The computer storage media may include both volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 704 may store computer-readable instructions (e.g., representing program(s) and / or program element(s)), such as an operating system.Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disks (DVDs) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing device 700. As used herein, the term "computer storage medium" does not per se include signals.

[0082] Computer storage media may embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave or other transport mechanism, and may include any information transmission media. The term "modulated data signal" may refer to a signal having one or more of its characteristics adjusted or altered to encode information in the signal. Computer storage media may include, for example, wired media, such as a wired network or a direct cable connection, and wireless media, such as acoustic, radio frequency (RF), infrared, and other wireless media, without limitation. Combinations of any of the foregoing should also be considered within the scope of computer-readable media.

[0083] The CPU(s) 706 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 700 to perform one or more of the methods and / or processes described herein. The CPU(s) 706 may each include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of concurrently processing a plurality of software threads. The CPU(s) 706 may include any type of processor and may include different types of processors depending on the type of computing device 700 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers).Depending on the type of computing device 700, the processor may be, for example, an Advanced RISC Machines (ARM) processor implemented with reduced instruction set computing (RISC) or an x86 processor implemented with complex instruction set computing (CISC). The computing device 700 may include one or more CPUs 706, in addition to one or more microprocessors or additional coprocessors, such as math coprocessors.

[0084] In addition to or alternatively to the CPU(s) 706, the GPU(s) 708 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 700 to perform one or more of the methods and / or processes described herein. One or more of the GPU(s) 708 may be an integrated GPU (e.g., with one or more of the CPU(s) 706 and / or one or more of the GPU(s) 708 may be a discrete GPU). In embodiments, one or more of the GPU(s) 708 may be a coprocessor of one or more of the CPU(s) 706. The GPU(s) 708 may be used by the computing device 700 to render graphics (e.g., 3D graphics) or to perform general-purpose computations.The GPU(s) 708 may be used, for example, for general-purpose computing on GPUs (GPGPU). The GPU(s) 708 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU(s) 708 may create pixel data for output images in response to rendering commands (e.g., rendering commands from the CPU(s) 706 received via a host interface). The GPU(s) 708 may include graphics memory, such as display memory, for storing pixel data or other suitable data, such as GPGPU data. The display memory may be included as part of the main memory 704. The GPU(s) 708 may include two or more GPUs operating in parallel (e.g., via an interconnect). The connection can connect the GPUs directly (e.g. with NVLINK) or via a switch (e.g. with NVSwitch).When combined, each GPU can generate 708 pixel data or GPGPU data for different parts of an output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can contain its own memory or share memory with other GPUs.

[0085] In addition to or alternatively to the CPU(s) 706 and / or the GPU(s) 708, the logic unit(s) 720 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 700 to perform one or more of the methods and / or processes described herein. In embodiments, the CPU(s) 706, the GPU(s) 708, and / or the logic unit(s) 720 may discretely or jointly execute any combination of the methods, processes, and / or portions thereof. One or more of the logic units 720 may be part of and / or integrated with one or more of the CPU(s) 706 and / or the GPU(s) 708, and / or one or more of the logic units 720 may be discrete components or otherwise external to the CPU(s) 706 and / or the GPU(s) 708.In embodiments, one or more of the logic units 720 may be a coprocessor of one or more of the CPU(s) 706 and / or one or more of the GPU(s) 708. Examples of the logic unit(s) 720 include the model 102, the training system 100, the data processor 172, the data set generator 176, the model 180, the application system 150, and so on.

[0086] Examples of the logic unit(s) 720 include one or more processing cores and / or components thereof, such as data processing units (DPUs), tensor cores (TCs), tensor processors (TPUs), pixel visual cores (PVCs), vision processing units (VPUs), graphics processing clusters (GPCs), texture processing clusters (TPCs), streaming multiprocessors (SMs), tree traversal units (TTUs), artificial intelligence accelerators (AIAs), deep learning accelerators (DLAs), arithmetic logic units (ALUs), application-specific integrated circuits (ASICs), floating point units (FPUs),Input / output (I / O) elements, peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) elements and / or the like.

[0087] The communication interface 710 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 700 to communicate with other computing devices over an electronic communication network, including wired and / or wireless communication. The communication interface 710 may include components and functions that enable communication over a variety of networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet or InfiniBand), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.In one or more embodiments, the one or more logic units 720 and / or the communication interface 710 may include one or more data processing units (DPUs) to communicate data received over a network and / or via the interconnect system 702 directly to one or more GPU(s) 708 (e.g., a memory).

[0088] Through the I / O ports 712, the computing device 700 can be logically coupled to other devices, including the I / O components 714, the presentation component(s) 718, and / or other components, some of which may be built into (e.g., integrated) the computing device 700. Example I / O components 714 include a microphone, a mouse, a keyboard, a joystick, a gamepad, a game controller, a satellite dish, a scanner, a printer, a wireless device, etc. The computing device 700 can include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof, for gesture capture and recognition. In addition, the computing device 700 can include accelerometers or gyroscopes (e.g., as part of an inertia measurement unit (IMU)) that enable the detection of motion.In some examples, the output of the accelerometers or gyroscopes from the computing device 700 may be used to render an immersive augmented reality or virtual reality.

[0089] Power supply 716 may include a hardwired power supply, a battery power supply, or a combination thereof. Power supply 716 may supply power to computing device 700 to enable operation of the components of computing device 700.

[0090] The presentation component(s) 718 may include a display (e.g., a monitor, a touchscreen, a television monitor, a heads-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component(s) 718 may receive data from other components (e.g., the GPU(s) 708, the CPU(s) 706, DPUs, etc.) and output the data (e.g., as an image, video, audio, etc.). EXEMPLARY DATA CENTER

[0091] Fig. 8 illustrates an example data center 800 that may be used in at least one embodiment of the present disclosure, such as to implement the training system 100 or the application system 150 in one or more examples of the data center 800. The data center 800 may include a data center infrastructure layer 810, a framework layer 820, a software layer 830, and / or an application layer 840.

[0092] As in Fig. 8, the data center infrastructure layer 810 may include a resource orchestrator 812, clustered compute resources 814, and node compute resources (“node CRs”) 816(1)-816(N), where “N” represents any positive integer. In at least one embodiment, the node CRs 916(1)-916(N) may include, but are not limited to, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power modules and / or cooling modules, etc.In some embodiments, one or more Node CRs among Node CRs 916(1)-916(N) may correspond to a server having one or more of the above-mentioned computing resources. Furthermore, in some embodiments, Node CRs 916(1)-916(N) may include one or more virtual components, such as vGPUs, vCPUs, and / or the like, and / or one or more of Node CRs 916(1)-916(N) may correspond to a virtual machine (VM).

[0093] In at least one embodiment, the grouped computing resources 814 may include separate groupings of node CRs 816 housed in one or more racks (not shown) or in many racks in data centers at different geographical locations (also not shown). Separate groupings of node CRs 816 within the grouped computing resources 814 may include grouped computing, networking, memory, or storage resources that may be configured or allocated to support one or more service units (workloads). In at least one embodiment, multiple node CRs 816, including CPUs, GPUs, DPUs, and / or other processors, may be grouped in one or more racks to provide computing resources to support one or more service units.The one or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.

[0094] The resource orchestrator 812 may configure or otherwise control one or more node CRs 816(1)-816(N) and / or grouped computing resources 814. In at least one embodiment, the resource orchestrator 812 may include a software design infrastructure (SDI) management entity for the data center 800. The resource orchestrator 812 may include hardware, software, or any combination thereof.

[0095] In at least one embodiment, as in Fig. 8, the framework layer 820 may include a job scheduler 828, a configuration manager 834, a resource manager 836, and / or a distributed file system 838. The framework layer 820 may include a framework to support the software 832 of the software layer 830 and / or one or more applications 842 of the application layer 840. The software 832 or the application(s) 842 may each include web-based service software or applications such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 820 may be, but is not limited to, a type of free and open source software for a web application framework such as Apache Spark™ (hereinafter referred to as "Spark"), which may utilize a distributed file system 838 for processing large amounts of data (e.g., "Big Data").In at least one embodiment, the job scheduler 828 may include a Spark driver to facilitate the scheduling of service units supported by various layers of the data center 800. The configuration manager 834 may be capable of configuring various layers such as the software layer 830 and the framework layer 820, including Spark and the distributed file system 838, to support the processing of large amounts of data. The resource manager 836 may be capable of managing clustered or grouped computing resources allocated or assigned to support the distributed file system 838 and the job scheduler 828. In at least one embodiment, clustered or grouped computing resources may include the clustered computing resource 814 on the data center infrastructure layer 810.The resource manager 836 may coordinate with the resource orchestrator 812 to manage these allocated or assigned computing resources.

[0096] In at least one embodiment, the software 832 included in software layer 830 may include software used by at least portions of node CRs 816(1)-816(N), clustered computing resources 814, and / or distributed file system 838 of framework layer 820. One or more types of software may include, but are not limited to, web page search software, email virus scanning software, database software, and streaming video content software.

[0097] In at least one embodiment, the application(s) 842 included in the application layer 840 may include one or more types of applications used by at least portions of the node CRs 816(1)-816(N), the clustered compute resources 814, and / or the distributed file system 838 of the framework layer 820. One or more types of applications may include any number of genomic applications, cognitive computation, and machine learning applications, including, but not limited to, training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in connection with one or more embodiments, such as to perform the training of the model 102 and / or the operation of the model 180.

[0098] In at least one embodiment, configuration manager 834, resource manager 836, and resource orchestrator 812 may implement any number and type of self-modifying actions based on at least any amount and type of data collected in any technically feasible manner. Self-modifying actions may relieve a data center operator of data center 800 from potentially making poor configuration decisions and potentially avoiding underutilized and / or malfunctioning portions of a data center.

[0099] Data center 800 may include tools, services, software, or other resources to train one or more machine learning models (e.g., train model 102) or to predict or infer information using one or more machine learning models (e.g., model 180) according to one or more embodiments described herein. For example, a machine learning model(s) may be trained by calculating weighting parameters according to a neural network architecture using software and / or computational resources described above with respect to data center 800.In at least one embodiment, trained or deployed machine learning models corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 800 by using weighting parameters calculated by one or more training techniques such as, but not limited to, those described herein.

[0100] In at least one embodiment, the data center 900 may utilize CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference (interferencing) using the resources described above. Furthermore, one or more of the software and / or hardware resources described above may be configured as a service to enable users to train or infer information such as image recognition, speech recognition, or other artificial intelligence services. EXAMPLE NETWORK ENVIRONMENTS

[0101] Network environments suitable for use in implementing embodiments of the disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be based on one or more instances of the computing device(s) 700 of Fig. 7 - e.g., each device may include similar components, features, and / or functions of the computing device(s) 700. Furthermore, if backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices may be included as part of a data center 800, an example of which is described herein with respect to Fig. 8 is described in more detail.

[0102] The disclosure of this application also contains the following numbered clauses:Clause 1: A system comprising at least one processor, the at least one processor comprising one or more circuits for: constructing at least one initial feature map from a plurality of initial feature maps based on a respective input image of an input data set, each of the at least one initial feature map incorporating depth data of the respective input image and corresponding to a plurality of pixels of the respective input image; constructing a sparse feature point cloud comprising a plurality of features determined using the plurality of initial feature maps; transforming the sparse feature point cloud into multi-resolution sparse grids, wherein at least one of the multi-resolution sparse grids comprises a plurality of voxels;to model the multi-resolution, sparse grids using a plurality of neural networks and according to a hierarchical architecture to construct a hierarchical volume representation; and to generate constructed content based on the hierarchical volume representation. Clause 2: The system of clause 1, wherein the input data set comprises a plurality of input images of a 3D scene, and wherein features of the plurality of initial feature maps corresponding to depths within at least one depth range are used to populate entries in a respective one of a plurality of frusta, the depths of the features being specified by incorporating the depth data. Clause 3: A system according to any one of clauses 1 or 2, wherein one multi-resolution sparse grid of the multi-resolution sparse grids comprises a first voxel size, and a second multi-resolution sparse grid of the multi-resolution sparse grids comprises a second voxel size. Clause 4: The system of clause 3, wherein the hierarchical architecture comprises: a first neural network of the plurality of neural networks processing the first multi-resolution sparse grid having the first voxel size; and a second neural network of the plurality of neural networks processing the second multi-resolution sparse grid having the second voxel size. Clause 5: The system of any preceding clause, wherein generating the constructed content further comprises: determining a new feature map by volume rendering the hierarchical volume representation, wherein the new feature map comprises a two-dimensional (2D) projection of the hierarchical volume representation corresponding to a target capture device. Clause 6: The system of Clause 5, wherein the new feature map comprises: a first component corresponding to a first level of the hierarchical architecture; a second component corresponding to a second level of the hierarchical architecture; and wherein the method further comprises combining vectors for a plurality of features constructed by the plurality of neural networks to construct the hierarchical volume representation. Clause 7: The system according to any one of clauses 5 or 6, wherein generating the constructed content based on the hierarchical volume representation further comprises decoding the new feature map using a decoder neural network. Clause 8: The system of any preceding clause, further comprising: determining, using a depth encoder with the respective input image as input, a depth map of the respective input image; determining, using a feature encoder with the respective input image as input, each initial feature map; and converting each initial feature map into a frustum using the depth map. Clause 9: The system of any preceding clause, wherein the at least one processor is further configured to: construct and update a hierarchical encoder to reduce the dimensionality of at least one hierarchical voxel level of a hierarchical voxel representation and output the hierarchical voxel representation as compressed latent variables; construct and update a multi-layer neural network by querying a subset of the plurality of voxels using coordinates, wherein the updating comprises matching the plurality of feature maps in the hierarchical volume representation and outputting a compressed representation of the hierarchical volume representation; and wherein determining the hierarchical encoder and the multi-layer neural network comprises: a first stage corresponding to compressing each hierarchical voxel level;and a second stage corresponding to the compression of the hierarchical voxel representation into a final latent representation.; Clause 10: The system of Clause 9, wherein the plurality of neural networks comprises a plurality of diffusion models, and wherein the modeling comprises using the plurality of diffusion models to model the plurality of voxels to construct the hierarchical volume representation. Clause 11: A system according to any preceding clause, wherein the at least one processor is included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system implemented using a robot; an aviation system; a medical system; a boat system; an intelligent area monitoring system; a system for performing deep learning operations; a system for performing simulation operations; a system for generating or presenting virtual reality (VR) content, augmented reality (AR) content, or mixed reality (MR) content; a system for performing digital twin operations; a system implemented using an edge device; a system including one or more virtual machines (VMs); a system for generating synthetic data;a system implemented at least in part in a data center; a system for performing conversational artificial intelligence (AI) operations; a system for performing generative AI operations; a system implementing language models; a system implementing large language models (LLMs); a system implementing vision language models (VLMs); a system for hosting one or more real-time streaming applications; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; or a system implemented at least in part using cloud computing resources. Clause 12: A system comprising at least one processor, the at least one processor comprising one or more circuits to: determine an initial feature map based on an input data set, the initial feature map incorporating depth data corresponding to a plurality of pixels of the input data set; determine a hierarchical volume representation based on multi-resolution sparse grids comprising a plurality of voxels corresponding to a transformed sparse feature point cloud; and provide constructed content based on the volume rendering of the hierarchical volume representation. Clause 13: A system comprising at least one processor, the at least one processor comprising one or more circuits for: constructing, using a model, a sparse feature point cloud comprising a plurality of features of the plurality of initial feature maps and a plurality of depth maps for a plurality of input images of an input dataset; constructing, using a model, a plurality of sparse grids with different resolutions using the sparse feature point cloud; combining, using a model, a plurality of features of the plurality of sparse grids to determine a hierarchical volume representation; constructing, using a model, an output image comprising the hierarchical volume representation, the output image being constructed based on a pose of a first input image of the plurality of input images;determine a loss of the output image with respect to the first input image; and update the model using the loss. Clause 14: System as defined in Clause 13, where the loss includes reconstruction loss. Clause 15: A system according to any one of clauses 13 or 14, wherein features from each of the plurality of initial feature maps corresponding to depths within at least one depth range are used to populate entries in a respective one of a plurality of frusta, the depths of the features of each of the plurality of initial features being indicated by a respective one of the plurality of depth maps. Clause 16: The system of any of clauses 13-15, wherein the plurality of sparse grids comprises: a first multi-resolution sparse grid having a first voxel size; and a second multi-resolution sparse grid having a second voxel size. Clause 17: The system of Clause 16, wherein the model comprises: a first neural network of the plurality of neural networks for processing the first multi-resolution, sparse grid having the first voxel size; and a second neural network of the plurality of neural networks for processing the second multi-resolution, sparse grid having the second voxel size. Clause 18: The system of any of clauses 13-17, further comprising determining a new feature map by volume rendering the hierarchical volume representation, wherein the new feature map comprises a two-dimensional (2D) projection of the hierarchical volume representation corresponding to a target capture device in the pose. Clause 19: The system of Clause 18, wherein the new feature map comprises: a first component corresponding to a first level of the hierarchical architecture; a second component corresponding to a second level of the hierarchical architecture; and wherein the method further comprises combining vectors for a plurality of features constructed by a plurality of neural networks to construct the hierarchical volume representation. Clause 20: A system according to either clause 18 or 19, further comprising decoding the new feature map using a decoder neural network.

[0103] The components of a network environment can communicate with each other over one or more networks, which can be wired, wireless, or both. The network can comprise multiple networks or a network of networks. For example, the network can comprise one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the Internet and / or a public switched telephone network (PSTN), and / or one or more private networks. If the network includes a wireless telecommunications network, components such as a base station, a communication tower, or even access points (as well as other components) can provide wireless connectivity.

[0104] Compatible network environments include one or more computer-to-computer (peer-to-peer) network environments—in which case a server may not be included in a network environment—and one or more client-server network environments—in which case one or more servers may be included in a network environment. In peer-to-peer network environments, the functionality described herein with respect to one or more servers may be implemented on any number of client devices.

[0105] In at least one embodiment, a network environment may include one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. A cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. A framework layer may include a framework for supporting software of a software layer and / or one or more applications of an application layer. The software or the application(s) may each include web-based service software or applications. In embodiments, one or more client devices may utilize the web-based service software or applications (e.g.,by accessing the service software and / or applications through one or more application programming interfaces (APIs). The framework layer may be a type of free and open source software for a web application framework that uses, for example, but is not limited to, a distributed file system for processing large amounts of data (e.g., "Big Data").

[0106] A cloud-based network environment may provide cloud computing and / or cloud storage performing any combination of the computing and / or data storage functions described herein (or one or more portions thereof). Each of these various functions may be performed across multiple locations of central or core servers (e.g., one or more data centers that may be distributed across a state, region, country, globe, etc.). Where a connection to a user (e.g., a client device) is relatively close to one or more edge servers, a core server(s) may allocate at least some functionality to the edge server(s). A cloud-based network environment may be private (e.g., restricted to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0107] The client device(s) may include at least some of the components, features, and functionalities of the devices described herein with respect to Fig.5. By way of example and without limitation, a client device may be implemented as a personal computer (PC), laptop, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, video camera, surveillance device or system, vehicle, boat, aircraft, virtual machine, drone, robot, handheld communication device, hospital device, gaming device or system, entertainment system, vehicle computing system, embedded system control device, remote control, appliance, consumer electronics device, workstation, edge device, any combination of these described devices, or any other suitable device.

[0108] The disclosure may be described in the general context of computer code or machine-usable instructions, including computer-executable instructions such as program modules, that are executed by a computer or other machine, such as a personal data assistant or other handheld device. In general, program modules, including routines, programs, objects, components, data structures, etc., refer to code that perform particular tasks or implement particular abstract data types. The disclosure may be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The disclosure may also be applied in distributed computing environments where tasks are performed by remote processing devices connected via a communications network.

[0109] As used herein, any reference to "and / or" in reference to two or more elements should be construed to mean only one element or a combination of elements. For example, "Element A, Element B, and / or Element C" may include only Element A, only Element B, only Element C, Element A and Element B, Element A and Element C, Element B and Element C, or Elements A, B, and C. Furthermore, "at least one of Element A or B" may include at least one of Element A, at least one of Element B, or at least one of Element A and at least one of Element B. Further, "at least one of Element A and Element B" may include at least one of Element A, at least one of Element B, or at least one of Element A and at least one of Element B.

[0110] The subject matter of the present disclosure is described in detail herein to satisfy legal requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors contemplated that the claimed subject matter could be implemented in other ways to include various steps or combinations of steps similar to those described herein, in conjunction with other present or future technologies. Although the terms "step" and / or "block" may be used herein to refer to various elements of the methods employed, the terms should not be construed to imply any particular order among or between the various steps disclosed herein, unless the order of each step is expressly described.

[0111] It is to be understood that aspects and embodiments described above are by way of example only and that changes in detail may be made within the scope of the claims.

[0112] Each device, method, and feature disclosed in the description and (where applicable) in the claims and drawings may be provided independently or in any suitable combination. Reference numerals included in the claims are for illustrative purposes only and are not intended to limit the scope of the claims.

Claims

[1] A system comprising at least one processor, wherein the at least one processor comprises one or more circuits to: constructing at least one initial feature map from a plurality of initial feature maps based on a respective input image of an input data set, wherein each of the at least one initial feature map includes depth data of the respective input image and corresponds to a plurality of pixels of the respective input image; construct a sparse feature point cloud comprising a plurality of features determined using the plurality of initial feature maps; transform the sparse feature point cloud into multi-resolution sparse grids, wherein at least one of the multi-resolution sparse grids comprises a plurality of voxels; to model the multi-resolution, sparse grids using a multitude of neural networks and according to a hierarchical architecture in order to construct a hierarchical volume representation; and to generate constructed content based on the hierarchical volume representation. [2] The system of claim 1, wherein the input data set comprises a plurality of input images of a 3D scene, and wherein features from the plurality of initial feature maps corresponding to depths within at least one depth range are used to populate entries in a respective one of a plurality of frusta, the depths of the features being indicated by incorporating the depth data. [3] The system of claim 1 or 2, wherein a multi-resolution sparse grid of the multi-resolution sparse grids comprises a first voxel size, and a second multi-resolution sparse grid of the multi-resolution sparse grids comprises a second voxel size. [4] The system of claim 3, wherein the hierarchical architecture comprises: a first neural network of the plurality of neural networks processes the first multi-resolution, sparse grid with the first voxel size; and a second neural network of the plurality of neural networks processes the second multi-resolution, sparse grid with the second voxel size. [5] The system of any preceding claim, wherein generating the constructed content further comprises: Determining a new feature map by volume rendering the hierarchical volume representation, wherein the new feature map comprises a two-dimensional (2D) projection of the hierarchical volume representation corresponding to a target capture device. [6] The system of claim 5, wherein the new feature map comprises: a first component corresponding to a first level of the hierarchical architecture; a second component corresponding to a second level of the hierarchical architecture; and wherein the method further comprises combining vectors for a plurality of features constructed by a plurality of neural networks to construct the hierarchical volume representation. [7] The system of claim 5 or 6, wherein generating the constructed content based on the hierarchical volume representation further comprises decoding the new feature map using a decoder neural network. [8] A system according to any preceding claim, further comprising: Determining, using a depth encoder with the respective input image as input, a depth map of the respective input image; Determining, using a feature encoder with the respective input image as input, each initial feature map; and Convert each initial feature map into a frustum using the depth map. [9] System according to one of the preceding claims, wherein the at least one processor is further provided for: Constructing and updating a hierarchical encoder to reduce the dimensionality of at least one hierarchical voxel level of a hierarchical voxel representation and output the hierarchical voxel representation into compressed latent variables; Constructing and updating a multi-layer neural network by querying a subset of the plurality of voxels using coordinates, wherein the updating comprises matching the plurality of feature maps in the hierarchical volume representation and outputting a compressed representation of the hierarchical volume representation; and wherein determining the hierarchical encoder and the multi-layer neural network comprises: a first stage corresponding to the compression of each hierarchical voxel level; and a second stage corresponding to the compression of the hierarchical voxel representation into a final latent representation. [10] The system of claim 9, wherein the plurality of neural networks comprises a plurality of diffusion models, and wherein the modeling comprises using the plurality of diffusion models to model the plurality of voxels to construct the hierarchical volume representation. [11] A system according to any one of the preceding claims, wherein the at least one processor is included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system implemented using a robot; an aviation system; a medical system; a boat system; an intelligent area monitoring system; a system for performing deep learning operations; a system for performing simulation operations; a system for generating or presenting augmented reality (AR) content, Virtual reality (VR) content or mixed reality (MR) content; a system for performing digital twin operations; a system implemented using an edge device; a system that contains one or more virtual machines (VMs); a system for generating synthetic data; a system that is at least partially implemented in a data center; a system for performing operations using conversational AI; a system for performing operations using generative AI; a system that implements language models; a system that implements large language models (LLMs) a system that implements visual language models (vision language models, VLMs); a system for hosting one or more real-time streaming applications; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; or a system implemented at least in part using cloud computing resources. [12] A system comprising at least one processor, wherein the at least one processor comprises one or more circuits to: determine an initial feature map based on an input data set, wherein the initial feature map incorporating depth data matches a plurality of pixels of the input data set; to determine a hierarchical volume representation based on multi-resolution sparse grids comprising a plurality of voxels corresponding to a transformed sparse feature point cloud; and to provide constructed content based on the volume rendering of the hierarchical volume representation. [13] A system comprising at least one processor, wherein the at least one processor comprises one or more circuits to: constructing, using a model, a sparse feature point cloud comprising a plurality of features of the plurality of initial feature maps and a plurality of depth maps for a plurality of input images of an input dataset; construct a variety of sparse grids with different resolutions using a model based on the sparse feature point cloud; to combine a multitude of features of the multitude of sparse grids using a model to determine a hierarchical volume representation; constructing an output image with the hierarchical volume representation using a model, the output image being constructed based on a pose of a first input image of the plurality of input images; to determine a loss of the output image with respect to the first input image; and to update the model using the loss. [14] The system of claim 13, wherein the loss comprises reconstruction loss. [15] A system according to any one of claims 13 or 14, wherein features from the plurality of initial feature maps corresponding to depths within at least one depth range are used to populate entries in a respective one of a plurality of frusta, the depths of the features of each of the plurality of initial features being indicated by a respective one of the plurality of depth maps. [16] The system of any of claims 13-15, wherein the plurality of sparse grids comprises: a first multi-resolution, sparse grid having a first voxel size; and a second multi-resolution, sparse grid with a second voxel size. [17] The system of claim 16, wherein the model comprises: a first neural network of the plurality of neural networks for processing the first multi-resolution, sparse grid having the first voxel size; and a second neural network of the plurality of neural networks for processing the second multi-resolution, sparse grid having the second voxel size. [18] The system of any of claims 13-17, further comprising determining a new feature map by volume rendering the hierarchical volume representation, wherein the new feature map comprises a two-dimensional (2D) projection of the hierarchical volume representation corresponding to a target capture device. [19] System according to claim 18, wherein The new feature map includes: a first component corresponding to a first level of the hierarchical architecture; a second component corresponding to a second level of the hierarchical architecture; and wherein the method further comprises combining vectors for a plurality of features constructed by a plurality of neural networks to construct the hierarchical volume representation. [20] The system of claim 18 or 19, further comprising decoding the new feature map using a decoder neural network.