INTERACTIVE THREE-DIMENSIONAL (3D) SCENE RECONSTRUCTION AND SIMULATION IN REAL TIME USING NEURAL REPRESENTATIONS

By using Gaussian splats and volumetric compression techniques, the method addresses the inefficiencies of existing 3D scene reconstruction methods, providing accurate and efficient simulations for real-time interactions in augmented and virtual reality environments.

DE102025138516A1Pending Publication Date: 2026-03-26NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing 3D scene reconstruction and interaction methods, such as neural radiance fields (NeRFs) and mesh-based approaches, face limitations in accuracy and efficiency, particularly in real-time applications, leading to geometric inaccuracies and challenges in simulating physical interactions in augmented and virtual reality environments.

Method used

The method involves generating 3D representations using neural representations like Gaussian splats, segmenting objects using depth maps and video data, refining these representations through inpainting and artifact removal, and performing volumetric compression to create accurate, efficient simulations by simulating interactions within voxelized volumes.

Benefits of technology

This approach enhances the realism and usability of 3D reconstructions in AR and VR by reducing computational bottlenecks and improving the quality of object representations and simulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Various examples, systems, and methods relating to the reconstruction, segmentation, and / or simulation of pipelines are disclosed. A first computing system can obtain at least one object segmented from video data. The first computing system can compress the at least one object. The first computing system can simulate one or more interactions of the voxelized volume of the at least one compressed object to update at least one physical attribute of the at least one object. The first computing system can generate at least one image that depicts at least one section of the at least one object with the at least one updated physical attribute.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Three-dimensional (3D) scene reconstruction and interaction often involve the use of neural representations such as neural radiance fields (NeRFs) or mesh-based methods to create 3D environments from image and video data. These existing methods are limited in terms of accuracy and efficiency, especially when applied to interactive or real-time applications. For example, mesh-based representations can introduce inaccuracies due to discretization errors when approximating continuous surfaces. This poses a challenge for accurately extracting surface data for simulating physical interactions. Neural reconstruction methods, such as neural radiance fields (NeRFs) or 3D Gaussian splats, lack explicit surface definitions, unlike traditional mesh models.Generating meshes from these neural representations is therefore computationally intensive and can lead to geometric inaccuracies that compromise simulation fidelity. Furthermore, simulating deformations or managing object interactions using these neural representations often requires algorithms capable of accurately processing volumetric data. These limitations reduce the realism and effectiveness of such simulations in augmented reality (AR) or virtual reality (VR) environments, where precise, real-time interaction models are crucial. While some methods can perform simulations on static meshes, neural representations such as NeRFs or 3D Gaussian splats can provide more detailed and realistic 3D reconstructions from images or multi-view videos.Simulating these neural representations can be challenging, however, as they often lack an explicit surface like traditional networks. At least one approach to overcoming this challenge is to extract a network from these representations and then sample within the network volume, but this conversion process can introduce errors and reduce fidelity. SUMMARY

[0002] The invention is defined in the claims. For the purpose of illustrating the invention, aspects and embodiments are described herein that may or may not fall within the scope of protection of the claims.

[0003] Various examples, systems, and methods relating to the reconstruction, segmentation, and / or simulation of a pipeline are disclosed. A first computing system can obtain at least one object segmented from video data. The first computing system can compress the at least one object. The first computing system can simulate one or more interactions of the voxelized volume of the at least one compressed object to update at least one physical attribute of the at least one object. The first computing system can generate at least one image that depicts at least one section of the at least one object with the at least one updated physical attribute.

[0004] Implementations of this disclosure relate to systems and methods for 3D scene reconstruction, segmentation, and / or simulation using neural representations combined with segmentation models and volumetric compression techniques. Systems and methods are disclosed that can use depth maps and video data to generate 3D representations depicting a scene. Segmentation models can be used to isolate objects within the 3D environment for manipulation and simulation. The implementations can further refine these 3D representations by performing operations such as inpainting or artifact removal to eliminate inconsistencies or inaccuracies and improve the quality of the reconstructed scenes.For example, systems and methods according to the present disclosure provide a pipeline for physical simulations by generating volumetric representations from 3D data, updating these volumes at least partially based on additional data inputs, and incorporating volumetric elements to perform realistic simulations of stiff and elastic objects.

[0005] Some implementations refer to a system containing one or more processors to perform one or more operations that involve receiving video data from a video source containing a depth map of a scene. The one or more processors perform one or more operations to reconstruct the scene into a three-dimensional (3D) representation using at least one or more Gaussian splat plots and the depth map. The one or more processors perform one or more operations to segment at least one object in the 3D representation. The one or more processors perform one or more operations to generate a two-dimensional (2D) segmentation mask of a reference view of the video data.One or more processors perform one or more operations to interpolate the 2D segmentation mask across a plurality of video frames. One or more processors perform one or more operations to map the 2D segmentation mask across the plurality of frames onto at least one corresponding region of a plurality of regions in the 3D representation, in order to segment the at least one object in the 3D representation from the scene. One or more processors perform one or more operations to update at least one of the plurality of regions in the 3D representation within a distance of the at least one object in the scene. One or more processors perform one or more operations to display the at least one object.

[0006] In some implementations, one or more processors are to perform one or more operations for the compaction of the at least one object. In some implementations, the compaction involves sampling a plurality of points on or approximately around the at least one object. In some implementations, the compaction involves generating a voxelized volume of the at least one object, based at least partially on the plurality of points. In some implementations, the compaction involves updating the voxelized volume at least partially based on the occupancy state of at least one voxel, a plurality of voxels of the voxelized volume, or at least partially based on at least one rendered depth map.

[0007] In some implementations, the one or more processors are to perform one or more operations to fill an interior space of the voxelized volume, based at least partially on injecting a plurality of volumetric elements into the interior space of the at least one object, including a plurality of interior regions. In some implementations, the one or more processors perform one or more operations to simulate one or more interactions of the voxelized volume of the at least one condensed object to update at least one physical attribute of the at least one object, where the at least one condensed object corresponds to a volumetric representation.

[0008] In some implementations, the simulation of one or more interactions involves performing a stiffness simulation. This simulation, using a first physics model, includes applying a first plurality of transformations to the at least one compressed object to obtain a plurality of stiff motions of the at least one compressed object. In some implementations, the application involves determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters. In some implementations, the application involves minimizing the energy function to determine a plurality of stiff states of the at least one compressed object. In some implementations, the application involves applying the plurality of stiff states to simulate the plurality of stiff motions of the at least one compressed object over time.

[0009] In some implementations, the simulation of one or more interactions involves performing an elasticity simulation. This simulation, using a second physics model, includes applying a second plurality of transformations to the at least one compressed object to obtain a plurality of deformed states of the at least one compressed object. In some implementations, the application involves determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters. In some implementations, the application involves minimizing the energy function to determine a plurality of updates for a plurality of control points.In some implementations, the application includes the calculation of one or more deformations of the at least one densified object, which are based at least partially on the plurality of updates of the plurality of control points and a plurality of corresponding skinning fields.

[0010] In some implementations, the reconstruction is further based on at least one of (i) at least one refined pose of the video source and a plurality of two-dimensional (2D) frames of the video data. In some implementations, the reconstruction further includes updating at least one initial pose of the video source to the at least one refined pose, at least partially based on aligning the 3D representation to the plurality of 2D frames of the video data.

[0011] In some implementations, one or more processors are to perform one or more operations that include generating an initial Gaussian sweep based at least partially on the depth map data and at least one initial pose from the video source. In some implementations, one or more processors are to perform one or more operations that include generating a 3D reconstruction based at least partially on the initial Gaussian sweep, which includes at least one refined pose from the video source and the plurality of 2D frames.

[0012] In some implementations, segmentation involves the use of a segmentation model. In some implementations, the reference view is based at least partially on user input that selects one or more objects. In some implementations, the reference view corresponds to a single frame from the majority of frames in the video data.

[0013] In some implementations, updating at least one or more regions of the 3D representation within the distance is based at least partially on filling at least one or more regions within the distance using sample data from one or more neighboring regions. In some implementations, updating at least one or more regions of the 3D representation within the distance is based at least partially on removing one or more elements from at least one or more regions within the distance. In some implementations, updating at least one or more regions of the 3D representation within the distance is based at least partially on updating at least one or more regions using sample data from one or more regions of the 3D representation.

[0014] Some implementations refer to one or more processors containing one or more circuits that receive video data, including a depth map of a scene. Using at least one or more Gaussian splat representations and the depth map, the one or more circuits reconstruct the scene into a three-dimensional (3D) representation. The one or more circuits segment at least one object in the 3D representation, at least partially, based on mapping a two-dimensional (2D) segmentation mask from a reference view of the video data across a plurality of frames onto at least one corresponding region or plurality of regions in the 3D representation. The one or more circuits update at least one or plurality of regions in the 3D representation within a certain distance of the at least one object in the scene.The one or more circuits should / should generate at least one image that depicts at least one section of the at least one object for display.

[0015] In some implementations, one or more circuits are intended to densify the at least one object. In some implementations, one or more circuits are intended to sample a plurality of points on or near the at least one object. In some implementations, one or more circuits are intended to generate a voxelized volume of the at least one object based at least partially on the plurality of points. In some implementations, one or more circuits are intended to update the voxelized volume at least partially based on the occupancy state of at least one voxel, a plurality of voxels of the voxelized volume, or at least partially based on at least one rendered depth map.

[0016] In some implementations, the one or more circuits are intended to fill an interior space of the voxelized volume, which is based at least partially on the injection of a plurality of volumetric elements into the interior of the at least one object, including a plurality of interior regions. In some implementations, the one or more circuits are intended to simulate one or more interactions of the voxelized volume of the at least one condensed object in order to update at least one physical attribute of the at least one object. In some implementations, the at least one condensed object corresponds to a volumetric representation.

[0017] In some implementations, the simulation of one or more interactions involves performing a stiffness simulation. This simulation, using a first physics model, includes applying a first plurality of transformations to the at least one compressed object to obtain a plurality of stiff motions of the at least one compressed object. In some implementations, the application involves determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters. In some implementations, the application involves minimizing the energy function to determine a plurality of stiff states of the at least one compressed object. In some implementations, the application involves applying the plurality of stiff states to simulate the plurality of stiff motions of the at least one compressed object over time.

[0018] In some implementations, the simulation of one or more interactions involves performing an elasticity simulation. This simulation, using a second physics model, includes applying a second plurality of transformations to the at least one compressed object to obtain a plurality of deformed states of the at least one compressed object. In some implementations, the application involves determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters. In some implementations, the application involves minimizing the energy function to determine a plurality of updates for a plurality of control points.In some implementations, the application includes the calculation of one or more deformations of the at least one densified object, which are based at least partially on the plurality of updates of the plurality of control points and a plurality of corresponding skinning fields.

[0019] In some implementations, the reconstruction is further based on at least one of: (i) at least one refined pose of the video source or (ii) a plurality of two-dimensional (2D) frames of the video data. In some implementations, the reconstruction further includes updating at least one initial pose of the video source to the at least one refined pose, at least partially based on aligning the 3D representation to the plurality of 2D frames of the video data.

[0020] Some implementations refer to a process. The process involves one or more processors receiving video data containing a depth map of a scene. The process involves the reconstruction of the scene into a three-dimensional (3D) representation by the one or more processors using at least Gaussian splatting and the depth map. The process involves the segmentation of at least one object in the 3D representation by the one or more processors. The process involves the updating of at least one of the plurality of regions of the 3D representation within a distance of the at least one object in the scene by the one or more processors.The method comprises the compression by one or more processors of the at least one object by generating a voxelized volume of the at least one object and updating the voxelized volume at least partially based on the occupancy state of at least one voxel or a plurality of voxels of the voxelized volume, or at least partially based on at least one rendered depth map. The method comprises the simulation by one or more processors of one or more interactions of the voxelized volume of the at least one compressed object. The method comprises the display by one or more processors, using a display device, of at least one rendered image depicting at least one section of the at least one object.

[0021] In some embodiments, the method further includes filling an interior space of the voxelized volume by one or more processors, based at least partially on injecting a plurality of volumetric elements into the interior space of the at least one object, including a plurality of interior regions. In some implementations, the simulation of the one or more interactions includes performing a stiffness simulation. In some implementations, the stiffness simulation includes applying a first plurality of transformations to the at least one densified object by one or more processors using a first physics model to obtain a plurality of stiff motions of the at least one densified object.In some implementations, the application includes determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters. In some implementations, the application includes minimizing the energy function to determine a plurality of stiff states of the at least one compressed object. In some implementations, the application includes applying the plurality of stiff states to simulate the plurality of stiff movements of the at least one compressed object over time.

[0022] In some implementations, the simulation of one or more interactions includes performing an elasticity simulation. In some implementations, the elasticity simulation includes the application by one or more processors of a second plurality of transformations to the at least one densified object using a second physics model to obtain a plurality of deformed states of the at least one densified object. In some implementations, the application includes determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters. In some implementations, the application includes minimizing the energy function to determine a plurality of updates for a plurality of control points.In some implementations, the application includes the calculation of one or more deformations of the at least one densified object, which are based at least partially on the plurality of updates of the plurality of control points and a plurality of corresponding skinning fields.

[0023] Some implementations refer to a system containing one or more processors. The one or more processes include at least one process for receiving and / or maintaining at least one object segmented from video data. The one or more processes include at least one process for compressing the at least one object. In some implementations, the compression involves sampling a plurality of points on or approximately around the at least one object. In some implementations, the compression involves generating a voxelized volume of the at least one object, based at least partially on the plurality of points.In some implementations, the compaction process includes updating the voxelized volume at least partially based on the occupancy state of at least one voxel, or of a plurality of voxels within the voxelized volume, at least partially based on at least one rendered depth map. The one or more operations include at least one operation for simulating interactions of the voxelized volume of the at least one compacted object to update at least one physical attribute of the at least one object. The one or more operations include at least one operation for generating an image that maps at least one section of the at least one object for display using a display device.

[0024] In some implementations, the one or more operations include at least one operation for filling an interior space of the voxelized volume, which is based at least partially on injecting a plurality of volumetric elements into the interior of the at least one object, including a plurality of interior regions. In some implementations, the at least one condensed object corresponds to a volumetric representation.

[0025] In some implementations, the simulation of one or more interactions involves performing a stiffness simulation. This simulation, using a first physics model, includes applying a first plurality of transformations to the at least one compressed object to obtain a plurality of stiff motions of the at least one compressed object. In some implementations, obtaining the plurality of stiff motions of the at least one compressed object involves determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters. In some implementations, obtaining the plurality of stiff motions of the at least one compressed object involves applying the minimization of the energy function to determine a plurality of stiff states of the at least one compressed object.In some implementations, preserving the plurality of stiff movements of the at least one compressed object involves applying the plurality of stiff states to simulate the plurality of stiff movements of the at least one compressed object over time.

[0026] In some implementations, simulating one or more interactions involves performing an elasticity simulation, which, using a second physics model, includes applying a second plurality of transformations to the at least one compressed object to obtain a plurality of deformed states of the at least one compressed object. In some implementations, obtaining the plurality of deformed states of the at least one compressed object involves determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters. In some implementations, obtaining the plurality of deformed states of the at least one compressed object involves minimizing the energy function to determine a plurality of updates for a plurality of control points.In some implementations, preserving the plurality of deformed states of the at least one densified object involves the application of calculating one or more deformations of the at least one densified object, based at least partially on the plurality of updates of the plurality of control points and a plurality of corresponding skinning fields.

[0027] Some implementations refer to one or more processors containing one or more circuits to receive and / or maintain at least one object segmented from video data. The one or more circuits are intended to compress the at least one object. In some implementations, the compression involves sampling a plurality of points on or approximately around the at least one object. In some implementations, the compression involves generating a voxelized volume of the at least one object, based at least partially on the plurality of points. In some implementations, the compression involves updating the voxelized volume at least partially based on the occupancy state of at least one voxel, a plurality of voxels of the voxelized volume, or at least partially based on at least one rendered depth map.The one or more circuits are intended to simulate one or more interactions of the voxelized volume of the at least one compressed object in order to update at least one physical attribute of the at least one object. The one or more circuits are intended to generate at least one image of the at least one object using the at least one updated physical attribute.

[0028] In some implementations, one or more circuits are intended to fill an interior space of the voxelized volume, which is at least partially based on the injection of a plurality of volumetric elements into the interior of the at least one object, including a plurality of interior regions. In some implementations, the at least one condensed object corresponds to a volumetric representation.

[0029] In some implementations, the simulation of one or more interactions involves performing a stiffness simulation. This simulation, using a first physics model, includes applying a first plurality of transformations to the at least one compressed object to obtain a plurality of stiff motions of the at least one compressed object. In some implementations, obtaining the plurality of stiff motions of the at least one compressed object involves determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters. In some implementations, obtaining the plurality of stiff motions of the at least one compressed object involves applying the minimization of the energy function to determine a plurality of stiff states of the at least one compressed object.In some implementations, preserving the plurality of stiff movements of the at least one compressed object involves applying the plurality of stiff states to simulate the plurality of stiff movements of the at least one compressed object over time.

[0030] In some implementations, simulating one or more interactions involves performing an elasticity simulation, which, using a second physics model, includes applying a second plurality of transformations to the at least one compressed object to obtain a plurality of deformed states of the at least one compressed object. In some implementations, obtaining the plurality of deformed states of the at least one compressed object involves determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters. In some implementations, obtaining the plurality of deformed states of the at least one compressed object involves minimizing the energy function to determine a plurality of updates for a plurality of control points.In some implementations, preserving the plurality of deformed states of the at least one densified object involves the application of calculating one or more deformations of the at least one densified object, based at least partially on the plurality of updates of the plurality of control points and a plurality of corresponding skinning fields.

[0031] Some implementations refer to a procedure. The procedure includes the reception by one or more processors of at least one object segmented from video data. The procedure includes the compression by the one or more processors of the at least one object. In some implementations, the compression includes the sampling of a plurality of points on or approximately around the at least one object. In some implementations, the compression includes the generation of a voxelized volume of the at least one object, based at least partially on the plurality of points. In some implementations, the compression includes the updating of the voxelized volume at least partially based on an occupancy state of at least one voxel, a plurality of voxels of the voxelized volume, or at least partially based on at least one rendered depth map.The method includes simulating, by one or more processors, one or more interactions of the voxelized volume of the at least one compressed object to update at least one physical attribute of the at least one object. The method includes displaying, by one or more processors and using a display device, at least one rendered image depicting at least one section of the at least one object.

[0032] In some embodiments, the method further includes filling an interior space of the voxelized volume by one or more processors, based at least partially on injecting a plurality of volumetric elements into the interior of the at least one object, including a plurality of interior regions. In some implementations, the at least one condensed object corresponds to a volumetric representation.

[0033] In some implementations, the simulation of one or more interactions includes performing a stiffness simulation. Using a first physics model, this simulation involves one or more processors applying a first plurality of transformations to the at least one compressed object to obtain a plurality of stiff movements of the at least one compressed object. In other implementations, the simulation of one or more interactions includes performing an elasticity simulation. Using a second physics model, this simulation involves one or more processors applying a second plurality of transformations to the at least one compressed object to obtain a plurality of deformed states of the at least one compressed object.

[0034] The processors, systems, and / or methods described herein may be implemented by or included in at least one system. The system may include a system for running games. The system may include a system for streaming content. The system may include a system for collaborative content creation. The system may include a system for running simulations. The system may include a system for collaborative content creation for 3D assets. The system may include a system for generating synthetic data. The system may include a system containing one or more Vision Language Models (VLMs). The system may include a system containing one or more Large Language Models (LLMs). The system may include a system for performing conversational AI operations.The system may include a system for performing light transport simulations. The system may include a system for performing deep learning operations. The system may include a system for performing digital twin operations. The system may include a control system for an autonomous or semi-autonomous machine. The system may include a perception system for an autonomous or semi-autonomous machine. The system may include a system that incorporates one or more virtual machines (VMs). The system may include a system that is implemented using a robot. The system may include a system that is implemented using an edge device. The system may include a system that is at least partially implemented in a data center. The system may include a system that is implemented at least partially using cloud computing resources.The system may include a system for generating interactive 3D visualizations. The system may include a system that is implemented, at least partially, using augmented reality (AR) or virtual reality (VR) platforms.

[0035] The revelation extends to all novel aspects or features described and / or illustrated herein.

[0036] Further features of the disclosure are characterized by the independent and dependent claims.

[0037] Any feature in one aspect of the disclosure can be applied in any suitable combination to other aspects of the disclosure. In particular, procedural aspects can be applied to apparatus or system aspects, and vice versa.

[0038] Furthermore, features implemented in hardware can be implemented in software and vice versa. Any reference to software and hardware features herein should be interpreted accordingly.

[0039] Each system or device feature described herein can also be provided as a process feature, and vice versa. System and / or device aspects that are functionally described (including means plus functional features) can alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and allocated working memory.

[0040] It is also understood that certain combinations of the various features described and defined in each aspect of the revelation can be implemented and / or provided and / or used independently of one another.

[0041] The disclosure also provides computer programs and computer program products comprising software code designed to perform one of the methods described herein when executed on a data processing device and / or to embody one of the device and system features described herein, including one or all component steps of a method.

[0042] The disclosure also provides a computer or computing system (including networked or distributed systems) with an operating system that supports a computer program for carrying out one of the methods described herein and / or for embodying one of the device or system features described herein.

[0043] The disclosure also provides a computer-readable medium on which one or more of the aforementioned computer programs are stored.

[0044] The revelation also provides a signal that carries one or more of the aforementioned computer programs.

[0045] The disclosure extends to methods and / or devices and / or systems as described herein with reference to the accompanying drawings.

[0046] Aspects and embodiments of the disclosure will now be described purely by way of example with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The systems and methods presented here for the reconstruction of and interaction with 3D environments are described in detail below with reference to the attached drawings, wherein: Fig. 1 is a block diagram of an example of a system according to some embodiments of the present disclosure; Fig. 2 is a block diagram of an exemplary reconstruction stage in an exemplary pipeline according to some implementations of the present disclosure; Fig. 3A is a block diagram of an exemplary segmentation stage in an exemplary pipeline according to some implementations of the present disclosure; Fig. 3B is a block diagram of another exemplary segmentation stage in an exemplary pipeline according to some implementations of the present disclosure; Fig. 4 is a block diagram of an exemplary preprocessing stage in an exemplary pipeline according to some implementations of the present disclosure; Fig. 5 is a block diagram of an exemplary compression stage in an exemplary pipeline according to some implementations of the present disclosure; Fig. 6A is a block diagram of an exemplary simulation stage in an exemplary pipeline according to some implementations of the present disclosure; Fig. 6B is a block diagram of an exemplary stiff simulation stage in an exemplary pipeline according to some implementations of the present disclosure; Fig. 6C is a block diagram of an exemplary elasticity simulation stage in an exemplary pipeline according to some implementations of the present disclosure; Fig. 7 is a block diagram of an exemplary simulation of an object in an exemplary pipeline according to some implementations of the present disclosure; Fig. 8A is a flowchart of an example of a procedure for scene reconstruction, segmentation, preprocessing, compression and / or simulation in an exemplary pipeline according to some implementations of the present disclosure; Fig. 8B is a flowchart of an example of a procedure for object compaction and / or simulation in an exemplary pipeline according to some implementations of the present disclosure; Fig. 9A is a block diagram of an exemplary generative language model system for use in implementing at least some implementations of the present disclosure; Fig. 9B is a block diagram of an exemplary generative language model that includes a transformer-encoder-decoder for use in implementing at least some implementations of the present disclosure; Fig. 9C is a block diagram of an exemplary generative language model that includes a decoder-transformer-only architecture for use in implementing at least some implementations of the present disclosure; Fig. 10 a block diagram of an exemplary computing device for use in the implementation of at least some implementations of the present disclosure; and Fig. 11 is a block diagram of an exemplary data center for use in the implementation of at least some implementations of the present disclosure. DETAILED DESCRIPTION

[0048] This disclosure relates to systems and methods for reconstructing, segmenting, and / or interacting with three-dimensional (3D) environments using volumetric representations, such as Gaussian splats, with enhanced implementations that segment, densify, and simulate objects within a scene. For example, systems and methods according to this disclosure include generating 3D representations from video data and depth information that can be used for object manipulation, simulation, and visualization in augmented reality (AR) and virtual reality (VR) platforms. That is to say, existing systems often fail to provide accurate real-time interaction and simulation capabilities due to limitations in object segmentation, volumetric representation, and physical simulations.Instead, the implementations described herein can use 3D representations, segmentation models, and volumetric densifications to create more accurate and efficient real-time 3D reconstructions that support object manipulation, realistic physical simulations, and interactive experiences in AR and VR.

[0049] Furthermore, generating mesh-based representations from neural data such as NeRFs or 3D Gaussian plots can be computationally intensive and lead to geometric inaccuracies, as these neural representations may lack an explicit surface definition. That is, while conventional mesh-based approaches can suffer from discretization errors, a challenge with neural-based reconstruction lies in obtaining a usable mesh representation. One approach, for example, might involve extracting a mesh from neural representations and using this mesh for simulations, which often employs multi-step conversion processes and can reduce the fidelity of the resulting model. In another example, neural representations can be simulated directly, avoiding surface extraction but employing techniques for simulating volumetric interactions within the data.The systems and methods address these challenges by simulating neural representations, managing the complexity of volumetric compression and interaction modeling, thereby improving the accuracy and efficiency of real-time 3D simulations in Augmented Reality (AR) and Virtual Reality (VR) environments.

[0050] Implementations of this disclosure provide systems and methods for simulating three-dimensional (3D) environments using neural representations, such as NeRFs and 3D Gaussian splats, which can generate high-quality 3D reconstructions from images or videos with multiple views. Unlike traditional mesh-based approaches, neural representations can pose a challenge for simulation because they often lack well-defined surfaces. The disclosed systems and methods utilize sampling within the volumetric data of these representations to facilitate accurate simulations of physical interactions without relying on mesh conversion. This technological solution reduces potential errors associated with conventional mesh extraction methods and supports efficient, realistic simulations for various applications in dynamic environments.

[0051] Some techniques for 3D scene reconstruction, segmentation, and / or interaction rely on neural radiation fields (NeRFs) or mesh representations, which often result in inaccurate or inefficient representations for object segmentation, interaction, and physical simulation. These techniques often fail to provide high-quality interactive 3D reconstructions because they are unable to adapt to the manipulation of real-time objects or accurately manage physical forces and deformations. The limitations include ineffective segmentation, inaccurate transformations, and inadequate volumetric representations. For example, mesh-based methods can lead to inaccuracies in representing the deformation of and interaction with objects under physical forces, reducing realism and usability.Furthermore, segmentation and compression approaches can prevent processing within the real-time limitations of AR and VR applications, leading to inefficiencies in rendering and interaction.

[0052] Systems and methods according to this disclosure can improve the accuracy and efficiency of 3D scene reconstruction, segmentation, and / or simulation by providing a framework using neural representations and volumetric compression. For example, multiple neural representations (e.g., Gaussian splats, collectively referred to herein as a "3D representation") can be generated to represent the 3D environment at least partially based on depth maps (e.g., low-resolution depth maps acquired by LiDAR sensors) and video data (e.g., RGB still frames with intrinsic parameters and camera poses). Additionally, one or more segmentation models (e.g., Segment Anything Model, SAM) can be used to isolate objects for manipulation and simulation. In some implementations, parameters such as depth maps, camera poses, and / or 2D segmentation masks (e.g.,Binary masks (generated for different object views) are used to represent the features of the 3D content with relevance and meaning. Implementations can further refine the 3D representation by updating regions of the neural representations within a certain distance threshold (e.g., proximity-based selection) to correct inconsistencies or inaccuracies, such as artifacts or missing data. Refining the 3D representation can include, for example, performing inpainting (e.g., filling gaps or holes using data from neighboring regions) and removing artifacts (e.g., discarding or replacing poorly reconstructed areas).

[0053] In some implementations, a densification process can be performed by generating a voxelized volume from sampled points (e.g., by converting neural representations to a voxel grid) and updating this volume, at least partially, based on rendered depth maps (e.g., by depth cutting to remove unoccupied regions) to provide an accurate volumetric mass for simulations. More generally, the densification process may involve voxelizing 3D Gaussian representations to create a voxelized shell (e.g., by populating only the voxels closest to the shape's surface). Additionally, depth maps can be used to cut out unoccupied regions around this voxelized shell, resulting in a dense volume representing the shape's interior.The dense volume can then be used to sample isotropic 3D Gaussian representations, which can be used to simulate physical interactions within the object. After densification, the implementation can fill the interior of the voxelized volume with additional volumetric elements (e.g., by injecting isotropic Gaussian representations) to enable realistic physical simulations such as stiffness simulations (e.g., simulating rigid objects) and elasticity simulations (e.g., modeling deformation under forces) to predict the behavior of objects under various forces. These enhancements provide improved accuracy and an interactive framework for 3D scene reconstruction, thereby increasing the realism and usability of AR and VR environments and other applications by reducing computational bottlenecks and improving the quality of object representations and simulations.

[0054] In some implementations, the video data captured by a device can contain RGB frames and camera information (e.g., intrinsic parameters and poses). For example, a low-resolution depth map can be converted into a point cloud, which can be used to generate neural representations (e.g., Gaussian splats that together form or create a 3D representation) capable of reconstructing the 3D scene. A segmentation model can be used to generate 2D segmentation masks that can be interpolated across multiple frames, and video tracking can be used to propagate the masks over time. That is, the segmentation process can be used to map 2D masks onto corresponding neural 3D representations, facilitating object segmentation in 3D space.In some implementations, the attributes of the 3D representation can be refined using densification. For example, points on object surfaces can be sampled to generate a voxelized volume. The voxelized volume (e.g., the voxelized shell) can be updated, at least partially, based on rendered depth maps. Furthermore, volumetric elements (e.g., isotropic Gaussian representations) can be injected into the interior of objects to generate realistic physical simulations based, at least partially, on the use of depth maps to intersect the space around the voxelized shell (e.g., with the remaining dense volume).

[0055] The systems and methods described herein can be used for a variety of purposes, including but not limited to the reconstruction of 3D environments, manipulation of objects in AR / VR, simulation-based training applications, digital twin creation, and the development of interactive content. These methods can improve efficiency in tasks involving 3D visualization, such as games, robotics, and automated driving simulations.

[0056] With reference to Fig. 1, is Fig. Figure 1 is an exemplary block diagram of a system 100 according to some embodiments of the present disclosure. It is noted that this and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as single or distributed components, or in conjunction with other components, in any combination and at any location. Various functions described herein, which are performed by entities, may be executed by hardware, firmware, and / or software.For example, various functions can be executed by a processor that carries out instructions stored in main memory. In some implementations, the systems, procedures, and processes described herein can be implemented using similar components, features, and / or functions to those of the exemplary generative language model system 900 by [source missing]. Fig. 9A, of the exemplary generative LM 930 from Fig. 9B-9C, the exemplary calculating device 1000 of Fig. 10, and / or the exemplary data center 1100 of Fig. 11 will be executed.

[0057] System 100 can implement at least one section of a 3D reconstruction, segmentation, and / or simulation (RCS) pipeline. For example, System 100 can process video data and depth maps to generate three-dimensional (3D) representations for object segmentation, manipulation, and simulation. System 100 can be used for real-time 3D reconstruction, object interaction, and simulation by various systems described herein, including but not limited to AR and VR systems, autonomous driving systems, robotics systems, gaming systems, and / or digital twin systems.

[0058] In general, the 3D RCS pipeline can contain operations performed by System 100. For example, the 3D RCS pipeline can include one or more stages: a video reception stage, a reconstruction stage, a segmentation stage, a preprocessing stage, a simulation stage, and / or a display stage.

[0059] The System 100 (e.g., for the implementation of the 3D RCS pipeline) can receive and / or process video data and depth information to reconstruct three-dimensional (3D) environments using neural representations and volumetric compression. Furthermore, the System 100 can process and segment objects in 3D space using generated 2D segmentation masks that can be interpolated across multiple frames. In some implementations, the System 100 can perform inpainting and artifact removal (e.g., prior to simulation) to refine specific regions of the 3D representation (e.g., within a certain distance of the segmented object). In this way, the 3D RCS pipeline can improve the quality of 3D environment reconstructions and facilitate accurate physical simulations by reducing inconsistencies in object representation and increasing the fidelity of object interactions.

[0060] In some implementations, the video reception stage can be the stage in the 3D RCS pipeline where the system 100 prepares the captured video data (e.g., RGB still images with intrinsic parameters and camera poses) and depth information for initial processing and / or alignment evaluation. For example, the video source 104 can provide data in formats such as raw RGB and / or depth maps, which the reconstructor 108 can process to extract pixel-level information for reconstructing the 3D environment. In some implementations, the video reception stage can perform operations that prepare depth maps by correcting any discrepancies in the camera poses that could affect the segmentation and / or simulation processes.

[0061] System 100 can contain or be coupled to at least one data source 104. Data source 104 can contain data such as video data, sensor data, and / or image data. Data source 104 can contain (or be implemented by) data from one or more sensors, such as one or more cameras (e.g., RGB-D cameras), LiDAR sensors, and / or depth sensors. For example, data source 104 can contain data structured as image frames and / or video frames, which may contain multiple pixels to represent information captured by the respective sensor(s) that output the data. Data source 104 can contain two-dimensional and / or three-dimensional image data and / or video data.

[0062] In some implementations, data source 104 contains training data (e.g., for training a segmentation model(s) and / or a simulation model(s)). For example, data source 104 might contain one or more sample frames, each with a tag assigned to it. The tag can specify at least one identifier of an object represented in the sample frame, such as a region of interest, a segmentation mask, or a classification (e.g., type, category). The tag can also include object data such as a 3D region, volumetric density, or metadata. In some implementations, the segmentation model and / or the simulation model can be configured, at least partially, based on data other than that contained in data source 104.System 100 can retrieve data from data source 104 as one or more data streams. The data can be retrieved, for example, according to a streaming protocol. The data from data source 104 can be encoded according to one or more encoding parameters.

[0063] In some implementations, the system 100 includes at least one reconstructor 108. During the reconstruction phase, the reconstructor 108 can apply various reconstruction operations to the data from the data source 104, such as performing reconstruction at least partially based on Gaussian splatting (e.g., one or more Gaussian splat representations) and the depth map. The reconstructor 108 can generate an initial set of one or more neural representations (e.g., Gaussian splats as 3D distributions) based at least partially on the depth data from the depth map and the at least one initial pose of the video source. The reconstructor 108 can further refine the 3D representation(s) by aligning the initial neural representations with the two-dimensional (2D) video frames using updated camera poses.The refined 3D representation(s) can be provided or used for further processing in subsequent stages.

[0064] In the reconstruction stage, Reconstructor 108 can generate a 3D representation (e.g., one or more neural representations) of a scene using video data and associated depth information. For example, Reconstructor 108 can convert low-resolution depth maps (e.g., 192 x 256 resolution) obtained from depth sensors (e.g., LiDAR sensors on mobile phones, tablets, and / or other smart devices) into at least one point cloud (e.g., representing the scene as a collection of 3D points). Reconstructor 108 can use the point cloud to generate volumetric neural representations (e.g., Gaussian splats, voxel grids, multi-resolution grids). Gaussian splats, for instance, can be a 3D Gaussian distribution that models the spatial properties of the scene. This means that the Reconstructor 108 can align the splats to 2D still images by adjusting the parameters (e.g.Mean, covariance, orientation and / or other shape information) are refined at least partially based on feedback from intrinsic and extrinsic parameters of the camera.

[0065] In some implementations, Reconstructor 108 can obtain (e.g., virtual) camera parameters, such as intrinsic parameters (e.g., focal length, optical center) and extrinsic parameters (e.g., position, orientation), using additional data sources, such as ARKit via NVIDIA iOS applications. For example, Reconstructor 108 can receive the parameters as initial estimates, which may be inaccurate, and perform optimization to refine them. Reconstructor 108 can, for instance, iteratively adjust the 3D point positions and camera parameters to reduce discrepancies between the projected 3D points and the observed 2D pixels. Reconstructor 108 can use bundle fitting to iteratively update the camera poses and 3D points to reduce reprojection errors between the observed 2D video frames and the projected 3D splats.In another example, the Reconstructor 108 can apply a non-linear least squares optimization to adjust both the Gaussian splats and the camera parameters simultaneously, ensuring more accurate alignment with the video frames.

[0066] Additionally, Reconstructor 108 can perform reconstruction into various types of 3D representations based at least partially on the specific use case. For example, Reconstructor 108 can generate neural radiation fields (NeRFs) if the application includes detailed volumetric renderings of the scene. In another example, Reconstructor 108 can generate a mesh representation by converting the point cloud into a polygonal surface model. In some implementations, Reconstructor 108 can determine the representation to be used, at least partially, based on various characteristics such as computational resources, desired fidelity, and the specific requirements of downstream processes (e.g., rendering, object manipulation).

[0067] In some implementations, Reconstructor 108 can partially optimize (also referred to herein as "reconstruct") a scene and provide the intermediate output to subsequent stages in the pipeline. That is, Segmentor 112 can begin processing the partially optimized scene at the segmentation stage while further optimizations (or reconstructions) continue to occur in the background. For example, Segmentor 112 can begin identifying and categorizing objects in the scene, at least partially, based on the initial reconstruction data. Furthermore, the simulation stage can run asynchronously, using the segmented data to simulate interactions and behaviors within the scene, while the visual quality is further improved as optimizations are applied to the reconstruction output. Reconstruction is described below with reference to Fig. 2 described in more detail.

[0068] In some implementations, the segmentation stage may refer to the stage in the 3D RCS pipeline where the system isolates 100 objects from the 3D representation. That is, Segmentor 112 can generate a two-dimensional (2D) segmentation mask (e.g., a binary mask that identifies specific regions corresponding to the object) of a reference view of the video data. The reference view may, for example, be based at least partially on a user selecting an object of interest in a specific frame of a video. In this example, the reference view may correspond to a single frame (e.g., a snapshot of the video) for object segmentation. The segmentation stage can interpolate the 2D segmentation mask across multiple frames of the video data (e.g., while maintaining consistency across frames).The segmentation stage can map the 2D segmentation mask across the majority of individual images to at least one corresponding region of the 3D representation in order to segment the at least one object in the 3D representation from the scene.

[0069] In some implementations, the segmentation stage may include a semi-interactive process where Segmentor 112 can generate 2D segmentation masks of objects within the video data. That is, Segmentor 112 can allow the user to select a reference view corresponding to a single frame of the video data and provide one or more selection options (e.g., mouse clicks, taps, etc.) to guide the segmentation model (e.g., an image segmentation model that can generate pixel-wise masks from input images, a model that can use user-supplied points to delineate objects, and / or any region-based model that can refine boundaries at least partially based on iterative user input) for identifying the foreground object. For example, the user can select different parts of an object (e.g., a rag doll) on a surface (e.g.,Click on a table to guide the algorithm in determining the regions that represent the object. In this example, the segmentation model can create a 2D mask that separates the object from surrounding objects or the background in the reference view.

[0070] During the segmentation phase, Segmentor 112 can perform additional operations by allowing the user to change the view and provide further selections for refining the segmentation mask. For example, the user can rotate the camera to view the back of the object, thus providing additional selections that allow the segmentation model to adjust the segmentation mask, at least partially, based on this new perspective. In another example, the user can select various features that distinguish the back of the object (e.g., a rag doll) from the rest of the scene, enabling Segmentor 112 to capture details not visible in the front view. Segmentor 112 can then use these multiple views to further refine the segmentation mask.

[0071] In the segmentation stage, the Segmentor 112 can generate a set of 2D segmentation masks for at least one (e.g., each) of the views for which user input has been provided. That is, the Segmentor 112 can compare these segmented views with the original video data to determine the points where the segmented masks align with the captured trajectory. For example, the Segmentor 112 can identify the frame in the video that corresponds to each segmented view and insert the segmentation mask into the video data at that point. In another example, the Segmentor 112 can facilitate the alignment of the segmentation mask with the spatial properties of the 3D representation to maintain consistency.

[0072] In some implementations, Segmentor 112 can use a video tracker model to interpolate the 2D segmentation masks over a sequence of frames in the video data (e.g., a temporal propagation model that can maintain object consistency across frames, a recurrent network-based model that can use memory to retrieve different frames, a feature-matching model that can align segmented regions over time, or any model that applies learned tracking algorithms to interpolate 2D segmentation masks over sequences of frames in video data). That is, Segmentor 112's video tracker can apply the segmentation mask (e.g.,(over time) propagate across the individual frames, using the reference view as a persistent memory input to maintain the identification of the segmented object throughout the video. For example, Segmentor 112 can input the segmented reference view and apply interpolation to project the segmentation onto the remaining video frames. In another example, the segmentation mask can be dynamically adjusted to adapt to changes in the object's appearance across multiple frames.

[0073] Furthermore, the Segmentor 112 can propagate segmentation masks to neural representations (e.g., 3D Gaussian splats) to facilitate the segmentation of the object of interest. That is, the Segmentor 112 can freeze the 3D representation (e.g., the Gaussian splat model) and then update it using the newly acquired segmentation masks to classify whether at least one (e.g., every) neural representation is part of the foreground object. The updated Gaussian splats can, for example, carry binary values ​​indicating the presence of the segmented object, enabling further processing in subsequent stages. In another example, the Segmentor 112 can repeat this process for multiple foreground objects in the scene, generating different segmentation outputs for at least one (e.g., every) object.

[0074] In some implementations, the segmentation stage may include a manual mode that uses the intersection of 2D bounding box queries to define regions of interest in the scene. That is, in a 2D view, the user can define a bounding box that selects all 3D Gaussian representations whose centers project within the defined region for the current camera view (e.g., regardless of their depth in 3D space). For example, a bounding box drawn around one or more objects (e.g., doll, planter, tree) in a view can select all Gaussian splats that represent parts of the objects, as well as, for example, splats of objects behind the doll. In this example, subsequent queries can be performed by Segmentor 112 by rotating the camera to new views and defining additional bounding boxes to refine the selection.The intersection of these bounding boxes across different views can be used by the Segmentor 112 to isolate the foreground object by removing background objects and retaining only the desired neural representations (e.g., Gaussian splats) that remain within the bounding box region across multiple views.

[0075] In some implementations, the segmentation stage can support iterative refinement by allowing the user to add, remove, or retain selections through multiple interaction steps. That is, at least one (e.g., every) interaction step can involve defining a new bounding box query or modifying an existing one, followed by the Segmentor 112 recalculating the intersection of selected 3D Gaussian representations across the views. For example, the user might first select a rough region containing the rag doll and surrounding objects, and then rotate the view to draw additional bounding boxes that exclude the unwanted objects.In another example, the Segmentor can retain 112 selected neural representations that remain consistently within the refined bounding boxes across all views, while removing those that lie outside an updated region. This means the iterative process can continue until the segmentation accurately isolates the object in question, at least partially, based on the user-defined queries and intersection terms.

[0076] In some implementations, the segmentation stage may include a semi-automatic mode that uses a combination of an image segmentation model and a video tracker model to provide more efficient and accurate segmentation with user guidance. That is, the semi-automatic mode can allow the user to interactively provide cues, such as clicks or selections, to guide the segmentor (e.g., in the implementation of the segmentation model) in distinguishing the object of interest from its background in a 2D view. The segmentation model can, for example, process the user input to generate an initial segmentation mask that identifies the desired object within the single frame.In another example, the user can change the view or perspective and provide additional input to further refine the segmentation mask, which can be used to account for variations in the object's appearance across different angles. Furthermore, the video tracker model can use the reference segmentation mask to propagate the segmentation to subsequent frames.

[0077] Segmentor 112 can incorporate one or more artificial intelligence models (e.g., machine learning models, supervised models, neural network models, deep neural network models), rules, heuristics, algorithms, functions, or various combinations thereof to perform operations involving the segmentation of one or more objects or features of one or more objects from data, such as from one or more frames of the data. In some implementations, Segmentor 112 can use the models to generate segmented masks or to delineate object boundaries, at least partially, based on input data. For example, Segmentor 112 can use different segmentation models to identify and isolate objects of interest in different frames or views, refining the segmentation boundaries across multiple perspectives as needed.The Segmentor 112 can use user input, such as clicks or bounding boxes, to guide the segmentation process.

[0078] In some implementations, the Segmentor 112 can maintain, execute, train, and / or update one or more segmentation-level machine learning models. In some implementations, the machine learning model(s) can include any type of image segmentation model configured to process single-frame data (e.g., single images) to identify and segment objects. For example, the machine learning models can be trained and / or updated to process single-frame inputs and account for variations in the appearance or perspective of objects. The machine learning model(s) can be or include a transformer-based model (e.g., encoder-decoder models) or other segmentation architectures for highly accurate object delimitation.Segmentor 112 can run the machine learning model to generate segmented outputs from the provided data. The segmentation is described below with reference to... Fig. 3A-3B described in more detail.

[0079] With further reference to Fig. System 100 can perform one of several preprocessing operations on the 3D representation output by Reconstructor 108 and segmented by Segmentor 112. For example, and without limitation, System 100 can perform inpainting, artifact removal, point sampling, or various combinations thereof on the 3D representation (e.g., neural representations such as Gaussian splats) during the preprocessing stage. That is, inpainting can include filling gaps or holes in the object model by sampling data from nearby regions. Preprocessor 116, for instance, can perform inpainting operations by sampling regions within a defined distance threshold from the segmented object. Furthermore, artifact removal can include replacing poorly reconstructed areas.For example, the preprocessor 116 can remove or replace artifacts to improve the visual and structural integrity of the 3D representation.

[0080] In some embodiments, the preprocessing stage can employ artifact processing, at least partially, based on artificial artifacts present in the 3D representation generated by the reconstructor 108 and segmented by the segmentor 112. That is, the preprocessing stage can include operations such as gap filling, correcting poorly reconstructed regions, and / or removing incorrect or unnecessary shadows or artifacts. For example, the preprocessor 116 can perform inpainting to fill gaps in the object model (e.g., in the form of Gaussian splats or other neural representations) by sampling from nearby, well-reconstructed regions within a defined distance threshold and / or by using Gaussian splats from adjacent surfaces that have a similar texture and lighting.In this example, the inpainting can include preprocessor 116 selecting Gaussian splats from a region (such as a clean area represented by a flat surface or uniform background) to cover regions that were not captured or were only inadequately captured in the original training views. Furthermore, preprocessor 116 can perform artifact removal in the preprocessing stage by replacing poorly reconstructed areas with sampled data from nearby regions to improve the visual and structural integrity of the 3D representation.

[0081] In some implementations, the preprocessing stage can use transformation techniques (e.g., affine transformations, non-stiff deformations, or arbitrary geometric transformations) to manipulate Gaussian splats, potentially revealing previously unseen regions that require correction. This means that transformations such as translation, rotation, or scaling (e.g., affine transformations) can be used by the preprocessor 116 to modify the position or covariance of Gaussian splats, possibly revealing regions that were not visible in the training images. For example, translating an object upwards might reveal a portion of the underlying surface that was poorly reconstructed due to lack of visibility during the training phase.In another example, the preprocessor can perform 116 rotations of an object to uncover hidden artifacts that must be addressed immediately during preprocessing to maintain visual consistency.

[0082] In some implementations, the training phase may refer to the process of providing the System 100 with a series of images or video frames of a scene from various viewpoints to create a 3D representation of the environment. During this phase, the System 100 can process the training images to generate neural representations, such as Gaussian splats, that capture the spatial and visual properties of the objects and surfaces in the scene. The training phase may include using the images to calculate parameters such as position, orientation, color, and depth information for the Gaussian splats that comprise the 3D scene. Following the training phase, the System 100 may perform stages in which the pre-built 3D object model is refined, prepared, and used in applications.

[0083] In some implementations, the 116 preprocessor can facilitate user-guided inpainting by allowing the user to mark or select a region to serve as a template for covering poorly reconstructed areas. That is, the user can interactively select well-reconstructed regions and instruct the 116 preprocessor to clone Gaussian splats from these regions and apply them to the exposed areas that need repair. For example, a user might identify a flat, well-textured area on a table surface near a region with visible artifacts and use it as the source for inpainting. In another example, the 116 preprocessor can automatically identify Gaussian splats within a specified distance threshold around the object and use these splats to fill gaps or replace faulty regions.

[0084] In some implementations, the preprocessing stage can perform shadow removal on the Gaussian splats that were captured along with the object during the initial training views. This means that shadows or color distortions that appear as artifacts in the 3D representation can be modified, updated, and / or removed. For example, if a segmented object, such as a doll, has shadows on the underlying surfaces of the Gaussian splats due to the lighting conditions during capture, the preprocessor can change the color of these Gaussian splats to remove the shadows. In another example, the preprocessing might involve sampling color data from nearby unshaded regions to provide uniform lighting throughout the 3D representation.

[0085] In some implementations, the 116 preprocessor can support multiple preprocessing actions and / or tasks to prepare the 3D representation for subsequent stages such as compression, simulation, and / or rendering. For example, the 116 preprocessor can first apply inpainting to repair poorly reconstructed regions, then perform shadow removal to ensure consistent lighting, and finally perform artifact removal to eliminate any remaining visual distortions. In another example, preprocessing can be prioritized, at least partially, based on the requirements of downstream processes, such as when a smooth and artifact-free surface is needed for accurate physical simulation.This means that by providing a 3D representation that is free (or nearly free) of artifacts and visually consistent, System 100 can enable more accurate interaction and simulation of objects within the scene. For example, a preprocessed 3D model can improve physical simulations where collisions and interactions are calculated, at least in part, based on accurate geometry. The preprocessing is described below with reference to [reference to relevant section]. Fig. 4 described in more detail.

[0086] In some implementations, the densification stage may refer to the phase in the 3D RCS pipeline where System 100 densifies the 3D representation to improve volumetric mass accuracy. That is, Simulator 120 can sample a plurality of points on or around the segmented object to generate a voxelized volume based at least partially on these points. The densification stage can update the voxelized volume, at least partially, based on rendered depth maps. For example, Simulator 120 can perform a depth slice to determine the occupancy state of voxels.

[0087] In some implementations, the densification stage may include converting the 3D Gaussian splats (3DGS) of the object representation into a dense voxel grid to simulate volumetric mass. That is, the densification stage may involve Simulator 120 voxelizing the space around the segmented object to determine the occupancy of each voxel, at least partially, based on the presence of Gaussian splats. For example, Simulator 120 may use a CUDA-based octree algorithm (which, for instance, accelerates space subdivision by using GPU processing power to create a hierarchical voxel grid of 3D Gaussian splats) to subdivide the space around the object into finer voxels, thus creating a hierarchical structure that efficiently represents the 3D occupancy.In this example, the axis-aligned bounding box of Gaussian splats can be enclosed in a root node of the octree, which can be recursively subdivided into smaller nodes, preserving a list of overlapping Gaussian splats for each subnode. In some implementations, Simulator 120 can employ a unified grid-based framework and / or a voxel hashing technique. For example, a unified grid-based approach can be used to subdivide space into voxels of fixed size. In another example, voxel hashing can be used to dynamically allocate voxels in sparse regions.

[0088] In some implementations, Simulator 120 can use the voxelization process to output a high-resolution representation of the object's interior. That is, the voxelization process can involve subdividing nodes containing Gaussian splats until the desired resolution is achieved, creating a grid representation (e.g., a sparse point cloud (SPC), a dense occupancy grid, or any hierarchical voxel grid) that represents the voxels occupied by the splats. For example, the nodes at the boundary of the octree can form the voxelized shell of the object (e.g., a voxel grid of voxels covering the approximate surface that are occupied), capturing the surface features as represented by the neural representations.In another example, the voxelized shell does not contain the object's inner voxels, which can be further processed to obtain a volumetric representation for accurate physical simulations.

[0089] In some implementations, Simulator 120 can perform a depth slice at the compaction stage to fill the voxelized shell with volumetric mass, thus approximating a solid interior. This depth slice can involve the use of rendered depth maps (e.g., rendered by an array of virtual cameras) from multiple viewpoints to determine the occupancy state of each voxel within the shell. Furthermore, the depth maps can be used to slice the space around the voxelized shell, preserving the object's dense volume. For example, Simulator 120 can perform ray tracing on the Sparse Point Cloud (SPC) from a collection of viewpoints to generate depth maps that capture the distance (e.g., the threshold distance) to the object's surface from different angles.In another example, the depth maps can be merged into a second thin SPC that can record the occupancy state for each voxel, such as empty, occupied, or unseen.

[0090] In some implementations, Simulator 120 can use the merged SPC to refine the voxelized volume by cutting out unoccupied spaces and preserving the solid regions. That is, Simulator 120 can update the occupancy state of at least one (e.g., every) voxel, at least partially, based on the depth maps to create a volumetrically dense representation of the object. For example, the cutting process can start from a fully occupied voxel grid and iteratively remove voxels that are determined to be empty, at least partially, based on their visibility in the depth maps. In this example, the cutting process can continue until only the occupied voxels remain, representing the solid shape of the object (e.g., filling the interior volume).

[0091] In some implementations, the compaction stage can be used to ensure that the 3D representation is suitable for physics-based simulations (e.g., when an accurate volumetric mass is important). That is, the compacted object representation can provide a realistic basis for simulating interactions, collisions, and physical behaviors. Once the compaction stage is complete, Simulator 120, for example, can accurately calculate forces, torques, and deformations, at least partially, based on the compacted voxelized volume. In another example, the provided volumetric approximation of mass can enable stable and realistic simulations. In some implementations, the compaction stage can be optimized for performance and integrated as a dedicated component into software frameworks, such as NVIDIA's Kaolin.This means the volume compaction block can be implemented as a CUDA kernel capable of performing voxelization and depth cutting. Compaction is described below with reference to... Fig. 5 described in more detail.

[0092] In some implementations, the simulation stage may refer to the phase in the 3D RCS pipeline where System 100 simulates the interactions of the voxelized volume. That is, Simulator 120 can inject a plurality of volumetric elements (e.g., isotropic Gaussian representations) into the interior of the voxelized volume to fill it. Simulator 120 can simulate one or more interactions of the voxelized volume of the compacted object to update at least one physical attribute, such as stiffness or elasticity, of the object. System 100 can contain at least one Simulator 120. Simulator 120 can contain one or more physics-based models, rules, heuristics, algorithms, functions, or various combinations thereof to perform operations that involve simulating one or more physical interactions (e.g., stiffness dynamics, elasticity) of the object.

[0093] To simulate one or more interactions of the voxelized volume, one or more processors are required to perform processing tasks. For example, Simulator 120 might use a primary physical model to obtain multiple stiff motions and a secondary physical model to obtain multiple deformed states of the object. In some implementations, Simulator 120 can maintain, execute, train, and / or update one or more simulation models during the simulation stage. In some implementations, the simulation model(s) might include any type of physics-based simulation model configured to process 3D representations to simulate physical behaviors. For example, the simulation model(s) might be trained and / or updated to process voxelized inputs.The simulation model(s) can be a physics engine model. The simulation model(s) can be configured to predict physical properties such as deformation under forces.

[0094] Simulator 120 can incorporate one or more physics-based models (e.g., mass-spring model, finite element model, neural network-based physics models), rules, heuristics, algorithms, functions, or various combinations thereof to perform operations involving the simulation of physical interactions (e.g., stiffness dynamics, elasticity, fluid dynamics) of objects within the 3D scene. In some implementations, Simulator 120 can simulate the behavior of objects, at least partially, based on various physical properties such as mass, density, stiffness, and elasticity. For example, Simulator 120 can apply physics-based rules to calculate interactions, deformations, and forces acting on the object.In another example, the Simulator 120 can use neural network models to predict the physical behavior of objects and generate simulations, at least partially, based on training data. In some implementations, the Simulator 120 can be trained independently of the models used by the Segmentor 112. In other implementations, the Simulator 120 can be trained together with the Segmentor 112. The Simulator 120 can be configured to perform both stiff and elastic simulations to model the physical behavior of objects in the scene.

[0095] The Simulator 120 can contain at least one physics model. The physics model can contain input parameters (e.g., object properties), transformation parameters, and / or one or more intermediate layers, such as skinning fields, at least one of which (e.g., each) can have respective control points. The System 100 can configure (e.g., train, update) the physics model by modifying or updating one or more parameters, such as weights and / or preloads of various nodes of the physics model, at least partially based on the evaluation of the estimated outputs. The Simulator 120 can consist of or contain various physics-based models that are effective for operation or data generation, including, but not limited to, object deformation, collision detection, or various combinations thereof.In some implementations, Simulator 120 can be configured (e.g., trained, updated, fine-tuned) at least partially based on training data derived from the 3D representation and segmentation results. For example, one or more sample scenes from the training data can be used as input for Simulator 120 to generate an estimated output. The estimated output can be evaluated and / or compared with one or more sample outputs (e.g., using cost functions, objective functions, or evaluation functions), and Simulator 120 can be updated at least partially based on this evaluation and / or comparison. For example, one or more parameters (e.g., weights) of Simulator 120 can be updated at least partially based on the output of an objective function.

[0096] With further reference to Fig. 1. The simulator 120 can receive and / or obtain one or more voxelized data volumes (e.g., by performing compression) and perform simulation operations (e.g., stiff or elastic simulation) on the voxelized volumes. For example, the simulator 120 can determine a representation (or a simulation result) of one or more interactions of the voxelized volume based on at least one given voxelized volume. The simulation representation can provide information regarding the physical properties and / or behavior of the segmented object.

[0097] In some implementations, the simulation stage can include the simulation of the physical interactions of voxelized volumes representing the interior of segmented objects. That is, Simulator 120 can perform physics-based simulations by injecting a plurality of volumetric elements (e.g., isotropic Gaussian elements, cubature points, particle-based elements, or any volumetric representations) into the interior of the voxelized volume (e.g., created during compaction) to fill the space with material properties for the simulation. For example, Simulator 120 can simulate the dynamics of a rigid body by treating the object as a rigid body with a control handle, allowing the entire object to move as a whole under external forces. In another example, Simulator 120 can simulate elastic deformations by applying a procedure (e.g.,deformation-based modeling, finite element analysis, or any physics-based simulation technique) to generate multiple control handles that guide the elastic properties and deformations of the object.

[0098] In some implementations, Simulator 120 can distinguish between (at least) two types of simulable objects: rigid and elastic. This means that, after segmentation, Simulator 120 can simulate rigid objects, treating the object as a single unit with a defined mass and inertia. For example, Simulator 120 can exert external forces, such as gravity, on the rigid object and calculate the resulting motion, at least partially, based on the object's mass and other properties. In another example, Simulator 120 can perform collision detection and reaction calculations to determine how the rigid object interacts with other objects in the scene.In some implementations, the Simulator 120 can apply techniques such as energy minimization, deformation field optimization and / or machine learning-based skinning methods to model the object with multiple control points and their associated weights in elastic simulations.

[0099] In some implementations, Simulator 120 can use object parameters and scene parameters as inputs to simulate physical interactions. That is, object parameters can include the initial state of the sampled cubic points, their rest positions, and physical material properties such as stiffness, density, and / or modulus of elasticity. The cubic points can, for example, represent the object's internal volume and provide the basis for simulating deformations and interactions. In another example, the scene parameters can include external forces such as gravity, wind forces, and / or contact forces, which can be modeled as constraints affecting the potential energy of each cubic point. Simulator 120 can minimize the system's potential energy by solving a Newton optimization problem (e.g.,...).iterative gradient descent) and determines how the object can deform under applied forces.

[0100] In some implementations, Simulator 120 can output a transformation that defines how the object can change over time. For example, Simulator 120 can compute a 12-degree-of-freedom (DoF) affine transformation that determines how a plurality of (e.g., all, some) Gaussian positions and covariances of an object can be transformed in each timeframe. For stiff objects, the transformation might be a combination of translation and rotation that can be applied uniformly to neural representations (e.g., Gaussian representations) within the object. In another example, for elastic objects, the transformation might vary in different parts of the object, providing non-uniform deformations such as bending, stretching, or twisting.In some implementations, Simulator 120 can perform elastic simulations to animate Gaussian splats according to the calculated transformations. This means that techniques such as Linear Blend Skinning (LBS), Dual Quaternion Skinning (DQS), skeleton-based deformation, and / or any mesh-based deformation technique can be used to perform elastic simulations. For example, Simulator 120 can use a deformation gradient obtained from LBS to transform both the Gaussian mean and the covariance. The skinning function can, for instance, induce weights for each Gaussian representation, determining how strongly it can respond to the movement of a control point. In another example, the transformations can be used to simulate large elastic deformations, facilitating movements of the objects such as squashing, stretching, and twisting.

[0101] In some implementations, Simulator 120 can operate interactively, allowing users to influence the simulation by providing input (e.g., clicking, dragging, selecting). For example, users can interact with the simulated objects in real time within an application (e.g., Application 124) by directly applying forces to the nearest neural representations (e.g., 3DGS) to simulate tensile and / or compressive forces (e.g., gravity, wind, pushing an object, etc.), rolling, bouncing, or other interactions. In another example, a user can click and drag a specific part of an elastic object to simulate a pulling motion, causing the object to deform accordingly. In yet another example, the interactive simulations can provide real-time (or near-real-time) visual feedback.In some implementations, Simulator 120 can improve simulation runtimes by using pre-computed data and efficient algorithms. That is, Simulator 120 can use techniques such as cached hessians (e.g., pre-computed second-order derivatives for energy minimization) to reduce the time required for complex physics calculations. For example, using NVIDIA Warp, Simulator 120 can reduce the training time for deformation modeling methods to 30–90 seconds per object, facilitating rapid setup for elastic simulations. In another example, optimization can enable the simulation to run online in an interactive mode, providing users with an enhanced experience for manipulating and testing various physics scenarios. The simulation is described below with reference to [reference missing]. Fig. 6A-6C described in more detail.

[0102] In some implementations, the display stage can refer to the stage in the 3D RCS pipeline where the data, including simulation outputs, is prepared for visualization or interaction. That is, System 100 can generate at least one image of the 3D representation that depicts at least one section of the at least one object for display. Generally, this at least one image can be a rendered single frame showing the geometry, deformations, and simulated behavior of the object, generated by applying rendering techniques such as rasterization, ray tracing, or neural rendering to the 3D representation. For example, System 100 can use ray tracing to generate photorealistic images of the object under different lighting conditions or camera angles.This means that System 100 can generate visual outputs of various physical attributes or interactions of the object, as simulated in the preceding stages. Furthermore, System 100 can process data from Simulator 120 to generate the image and apply shaders and materials to enhance visual fidelity. For example, System 100 can render the object's surface with detailed textures, reflections, and shadows, providing a realistic view of the simulated interactions and deformations. This rendered image can then be displayed on a graphical user interface, allowing users to interact with or analyze the simulated object in different states and from various perspectives.

[0103] System 100 can contain or be coupled to at least one application 124. This means that at least one application 124 can manage the generation, rendering, and display of the object output by the pipeline. For example, application 124 can facilitate the arrangement of the output data from the segmentation model 112 and / or the simulation model 120 into a structured format for display and presentation. In some implementations, application 124 can function as a simulation management and visualization system for rendering and displaying the results of the 3D RCS pipeline. That is, application 124 can generate (or render) and display simulation results (at least one image of a 3D representation) from the simulator 120 and convert the results into visual representations on a user interface (e.g., a graphic).3D visualization tool, interactive display panel, simulation dashboard, and / or any graphical user interface environment). For example, Application 124 can use GPU-accelerated rendering techniques to process the transformation matrices and deformation gradients generated during elastic or stiff-body simulations. In this example, Application 124 can maintain a real-time data stream between Simulator 120 and the rendering system or device, providing visualization updates when the simulation parameters are modified.

[0104] Furthermore, Application 124 can perform interpolation and blending of simulation frames. In some implementations, Application 124 can provide a programmable environment that supports user-defined inputs and real-time adjustments to simulation parameters. That is, Application 124 can expose an API or scripting layer that allows the user to programmatically control the simulation behavior, modify physical properties, and / or introduce new force fields or constraints. For example, Application 124 can use shader programming and parallel computing techniques to adapt the rendering of neural representations, at least in part, based on the deformation gradients calculated by Simulator 120. In another example, Application 124 can facilitate collision detection and reaction calculations in parallel with Simulator 120.The advertisement is listed below with reference to . Fig. 7 described in more detail.

[0105] With reference to Fig. Figure 2 is a block diagram of an exemplary reconstruction stage in an exemplary pipeline according to some implementations of the present disclosure. For example, the reconstructor 108 can reconstruct the scene into a three-dimensional (3D) representation using at least one or more Gaussian splat representations and a depth map. That is, the initialization stage 200 can refer to the beginning of the reconstruction stage of the 3D RCS pipeline, in which the reconstructor 108 can generate initial representations of the scene using imprecise camera poses and low-resolution depth maps in step 210. That is, the reconstructor 108 can receive and / or obtain imprecise camera poses (e.g., rough estimates of the camera's positions and orientations) and combine them with low-resolution depth maps to create initial Gaussian splat representations 220.For example, the low-resolution depth maps can provide sparse information about the depth of the scene, which facilitates the approximation of the spatial structure using the initial Gaussian splat representations 220 by the reconstructor 108. In another example, the inaccurate camera poses in step 210 can later be refined during the training phase 250 to improve the accuracy of the reconstruction.

[0106] In some implementations, the reconstruction may further be based, at least in part, on at least one of (i) at least one refined pose of the video source and a plurality of two-dimensional (2D) frames of the video data. For example, the reconstruction may further involve updating at least one initial pose (e.g., inaccurate camera poses) of the video source to the at least one refined pose, at least in part, based on aligning the 3D representation to the plurality of 2D frames of the video data. In this example, the alignment may involve determining correspondence points between the 3D representation and 2D frames to improve (or optimize) the camera positions. Additionally, Reconstructor 108 may generate a Gaussian splat representation (e.g.,Initialization phase inputs: imprecise camera poses, low-resolution depth maps; outputs: initial Gaussian splats), which is based at least partially on the depth data from the depth map and the at least one initial pose of the video source. That is, Reconstructor 108 can calculate the initial splat positions by projecting the depth values ​​from the depth map into 3D space using the initial camera poses. In some implementations, Reconstructor 108 can produce a 3D reconstruction (e.g., training phase inputs: initial Gaussian splats, RGB frames, and initial poses; output: refined 3D Gaussian splats) that aligns to the 2D video frames by using updated camera poses based at least partially on the initial Gaussian splat representation, the at least one refined pose of the video source, and the majority of 2D frames.This means that the Reconstructor 108 can iteratively improve (or optimize) the camera poses by minimizing the difference between the projected splat positions and the corresponding 2D features in the video frames.

[0107] In some implementations, training stage 250 may refer to the stage in the reconstruction phase of the 3D RCS pipeline where reconstructor 108 can refine the initial Gaussian splat representations 220 generated during initialization stage 200 using additional data such as RGB frames 260 and updated camera poses. That is, reconstructor 108 can use these RGB frames 260, which contain detailed color and texture information of the scene, along with refined camera poses, to update the positions and orientations of the Gaussian splat representations. For example, during training stage 250, reconstructor 108 can optimize the Gaussian splat representations to better align with the RGB frames 260, resulting in a more accurate reconstruction 270.In another example, the updated camera poses can provide improved geometric information for refining the placement of Gaussian plots, resulting in a 3D reconstruction that more accurately captures both the appearance and structure of the scene. The training stage improves the initial outputs by utilizing both the color data and the refined geometric information, ultimately producing a high-fidelity 3D reconstruction suitable for subsequent processing and simulation stages in the pipeline.

[0108] With reference to Fig. Figure 3A presents a block diagram of an exemplary segmentation stage in an exemplary pipeline according to some implementations of the present disclosure. For example, the Segmentor 112 can segment at least one object in the 3D representation (e.g., identify an object and isolate it from the 3D representation). The segmentation can include the generation of a two-dimensional (2D) segmentation mask (e.g., a binary mask that identifies specific regions corresponding to the object) of a reference view of the video data. Furthermore, the Segmentor 112 can interpolate the 2D segmentation mask over a plurality of frames of the video data (e.g., using a Tracker 322). In addition, the Segmentor 112 can map a 3D representation of the scene (e.g., by mapping 2D pixels in the mask to specific 3D regions of the scene).

[0109] In some implementations, Segmentor 112 can implement and / or use a segmentation model (e.g., a Segment Anything Model (SAM)). Furthermore, the reference view can be based, at least partially, on user input that selects the one or more objects. For example, the user can guide the segmentation process by selecting an object of interest in a specific frame to be used as the reference view. In this example, the reference view can correspond to a single frame (e.g., a snapshot of the video serving as the reference view) from the majority of frames in the video data.

[0110] In some implementations, the segmentation stage may refer to the phase in the 3D RCS pipeline where the Segmentor 112 generates segmented views from video frames by considering both user input and segmentation models. That is, the Segmentor 112 can perform a single-view process 310 (e.g., a single image or frame) in which a user selects an object within the frame. The Segmentation Model 312 can use these user selections to output a mask that identifies a part of the object (e.g., the head). After this initial segmentation, the user can make another selection (or the system can automatically refine the segmentation) to further refine the segmentation, for example, by selecting additional parts of the object (e.g., the body).This means that once the segmentation model 312 outputs this initial mask, another segmentation model 314 (or the same segmentation model 312) can be applied to segment the entire object. Furthermore, the single-view process 310 can segment a single image or frame, allowing the user to make a selection and apply segmentation to that specific image. In some implementations, a video process 320 can be performed on a plurality of frames, where a reference view can be selected by the user or automatically by the segmentor 112, and the tracker 322 can be applied to the entire video sequence. In some implementations, frame 330 maps the object before a user selection, and frame 332 maps the highlighted object after a user selection of the object to be segmented.

[0111] With reference to Fig. Section 3B presents a block diagram of another exemplary segmentation stage in an exemplary pipeline according to some implementations of this disclosure. In some implementations, the segmentation stage may refer to the stage in the 3D RCS pipeline where the Segmentor 112 isolates objects in different scenes using bounding box queries and 2D-to-3D projection techniques to maintain segmentation across frames. That is, the segmentation may receive a view 340 of an object, such as a pineapple, where the user defines a bounding box around the object to guide the segmentation model. The Segmentor 112 can use the initial selection of the bounding box to create a mask 342 that segments the object within the defined region.For example, the Segmentor 112 can use multiple bounding boxes from different views to facilitate occlusions or perspective changes. In another example, the Segmentor 112 can refine the segmentation by dynamically adjusting the bounding box 344 across different frames to maintain consistency even when the object appears in different contexts or angles.

[0112] With reference to Fig. Figure 4 shows a block diagram of an exemplary preprocessing stage in an exemplary pipeline according to some implementations of the present disclosure. In some implementations, the preprocessor 116 can update at least one of the plurality of regions of the 3D representation within a threshold distance to the at least one object in the scene. That is, the update of the at least one of the plurality of regions of the 3D representation within the threshold distance can be based at least partially on filling (e.g., inpainting) at least one of the plurality of regions within the threshold distance, at least partially, based on sample data from one or more neighboring regions.Furthermore, the updating of at least one of the plurality of regions of the 3D representation within the threshold distance can be based at least partially on the removal of one or more elements from at least one of the plurality of regions within the threshold distance and the updating of at least one of the plurality of regions at least partially on the basis of sample data from one or more regions of the plurality of regions of the 3D representation.

[0113] In general, the preprocessing stage can refer to the phase in the 3D RCS pipeline where the preprocessor 116 can modify and enhance the segmented object(s) for further compression, simulation, and / or display. Furthermore, the preprocessing stage can include identifying a segmented object in a quiescent view 400 (e.g., a doll) and applying transformation operations to adjust its position or orientation. For example, the preprocessor 116 can use a transformation tool to manipulate the object in the scene (e.g., or allow the user to manipulate it using a user interface), thereby revealing previously unseen or poorly reconstructed areas, as shown in the adjusted view 410.In another example, after repositioning or scaling the object, preprocessor 116 can perform inpainting or artifact removal to eliminate visual inconsistencies or artifacts that have appeared in the new position. The result can be a refined representation of the object in an updated view 420, where preprocessor 116 has corrected all visual errors and prepared the object for subsequent condensation or simulation stages in the 3D RCS pipeline.

[0114] With reference to Fig. Figure 5 is a block diagram of an exemplary compaction stage in an exemplary pipeline according to some implementations of this disclosure. In some implementations, Simulator 120 can compact an object by sampling a plurality of points on or approximately around the at least one object. In a first step, Simulator 120 can generate a voxelized volume of the at least one object (e.g., Gaussian representations of voxels) that is at least partially based on the plurality of points. In a second step, Simulator 120 can update the voxelized volume at least partially based on an occupancy state (e.g., occupied, half-occupied, or unoccupied) of at least one voxel or a plurality of voxels of the voxelized volume, based on at least one rendered depth map (e.g., depth section).Furthermore, the Simulator 120 can fill an interior space of the voxelized volume, which is based at least partially on the injection of a plurality of volumetric elements (e.g. isotropic Gaussian representations) into the interior space of the at least one object, including a plurality of interior regions.

[0115] In some implementations, the compaction stage in the 3D RCS pipeline may include Simulator 120 converting a set of 3D Gaussian splats into a dense voxel grid to simulate volumetric mass. That is, Simulator 120 can voxelize the Gaussian splats using a hierarchical algorithm (e.g., a CUDA-based octree, Kaolin's SPC) to partition the space around the object. For example, Simulator 120 can enclose the axis-aligned bounding box of the Gaussian splats in a cubic root node, which can be recursively subdivided (e.g., 12-fold, 8-fold, and / or 4-fold partitions to create smaller nodes, among other recursive processes). At least one (e.g., each) subnode can maintain a list of overlapping Gaussian plots for representing the object's surface. Subdivision can continue until Simulator 120 reaches the desired voxel resolution (e.g., 120).B. 64x64x64, 128x128x128 or any user-defined resolution), resulting in a voxelized shell that captures the outer surface of the object.

[0116] In some implementations, Simulator 120 can fill the interior of the voxelized shell by using depth maps generated from multiple viewpoints to create the filled object, which is then displayed in voxelized form. Fig. Figure 5 shows that Simulator 120 can perform ray tracing from different viewpoints (e.g., an icosahedral array) to create depth maps of the object. These depth maps can be merged to represent the object's internal structure. For example, Simulator 120 can merge these depth maps into a second sparse point cloud (SPC) and classify each voxel, at least partially, as empty, occupied, or unseen based on the depth information. Simulator 120 can then clip away unoccupied voxels to refine the volumetric representation and create the filled object, which is displayed in voxelized form in Figure 5. Fig. Figure 5 shows that Simulator 120 refines the interior by removing unoccupied spaces to create a volumetric representation, with the remaining voxels reflecting the object's mass and structure. As shown, the object's interior can be densely packed with voxels representing the object's volumetric properties. For example, the densely packed voxels can represent an object's internal structure, with each voxel corresponding to depth information from multiple viewpoints.

[0117] With reference to Fig. Figure 6A shows a block diagram of an exemplary simulation stage in an exemplary pipeline according to some implementations of this disclosure. In some implementations, the simulator 120 can simulate one or more interactions of the voxelized volume of the at least one densified object to update at least one physical attribute of the at least one object. The at least one densified object can, for example, correspond to a volumetric representation. The reconstruction, segmentation, preprocessing, densification, and / or simulation process 300 can include the simulator 120 performing the training and simulation stages of the 3D RCS pipeline after the 3D Gaussian splats (block 602) have been processed by the segmentation block 604 (and preprocessing), isolating the object for further operations.In some implementations, Simulator 120 can align the object for simulation in training block 606. That is, during training block 606, Simulator 120 can assign control points and neural weights to the object, based at least partially on its geometry and material properties. In some examples, Simulator 120 can determine the influence of each control point on the surrounding Gaussian plots. The manipulated object can then be output for simulation in simulation block 608.

[0118] In some implementations, Simulator 120 can apply physics-based simulations to the manipulated object and / or directly to the segmented object. That is, after segmentation block 604 (and / or after preprocessing), Simulator 120 can perform training on the object (e.g., if it is manipulated with control points for more complex simulations) and / or it can directly perform simulations on the object. For example, if the object is manipulated during training block 606, Simulator 120 can use control points and neural weights to provide elastic deformation or stiffness simulations (e.g., at least partially based on the object's material properties). In some implementations, elastic simulations may involve the use of a deformable model(s) that simulate the motion of objects (e.g., bending, stretching, compressing) under applied forces.In some implementations, stiff-body simulations can include maintaining the object's structural integrity and simulating its motion as a single, rigid unit. In both types of simulations, regardless of whether the object is manipulated, Simulator 120 can compute physics-based transformations and apply input forces (such as gravity or user interactions) to generate motion. For objects supplied by Segmentation Block 604 to Simulation Block 608, Simulator 120 can perform stiff-body simulations, treating the object as a single unit. For manipulated objects, the control points can allow for more extensive deformation and elastic simulations. The simulation results can provide realistic motion and interactions, preparing the object for subsequent rendering stages or further physical manipulation in the 3D RCS pipeline.

[0119] With reference to Fig. Figure 6B presents a block diagram of an exemplary stiffness simulation stage in an exemplary pipeline according to some implementations of this disclosure. In some implementations, the simulation of one or more interactions may include performing a stiffness simulation. The stiffness simulation may involve the simulator 120 applying a first plurality of transformations to the at least one compressed object using a first physics model (e.g., stiffness dynamics, mass-spring systems, collision detection algorithms) to obtain a plurality of stiff motions of the at least one compressed object. That is, the transformation may be based at least partially on determining an energy function using a plurality of scene parameters (e.g., external forces such as gravity and constraints such as boundary conditions) or a plurality of object parameters (e.g.,Physical properties of the object (i.e., the initial state of the sampled cubic points) are applied. Additionally, the transformation can include minimizing the energy function (e.g., minimizing the potential energy) to determine a plurality of stiff states of the at least one compressed object. In some implementations, the transformation can include applying the plurality of stiff states to simulate the plurality of stiff motions of the at least one compressed object over time.

[0120] In some implementations, preparation stage 610 may include the determination and / or definition of parameters for stiff simulation. This means that Simulator 120 can receive and / or obtain input parameters representing the object's physical properties, such as cubic points, material properties, and external forces. For example, cubic points can capture properties such as position, stiffness, and density (e.g., initial rest positions, density values) and be used to represent discrete points across the object. Scene forces can include multiple constraints and influences acting on the object, such as gravity and boundary conditions (e.g., gravitational fields, collision boundaries, static surfaces). Additionally, at least one (e.g., each) cubic point can be assigned a value representing the object's undisturbed configuration before forces are applied.

[0121] The Simulator 620 can model and / or determine transformations for the stiff object, at least partially, based on the prepared parameters. For example, the Simulator 620 can apply methods such as optimization algorithms (e.g., Newton optimization, gradient descent, boundary satisfaction) to minimize the potential energy in the system. This means the Simulator 620 can calculate the transformation for at least one (e.g., each) control handle representing movements such as translations, rotations, or scaling (e.g., 12 degrees of freedom (DoF), 6 DoF transformations, affine transformations). The Simulator 620 can, for example, apply one or more transformations to simulate the motion of mechanical components such as robot arms or articulated machines. In this example, the motion on one or more individual sections can be simulated independently while maintaining the overall stiffness of the object.In some implementations, the 620 simulator can use a per-point deformation formula (shown below) that combines neural weights trained in the preparation stage with affine transformations calculated in the simulation (per-point deformation formula):. ϕ_(X z)=X+∑j=1nWj_(X)Zj_[X1]

[0122] In some implementations, the simulator output 630 can contain a combined set of transformations for the object. This combined set can be animated using Linear Blend Skinning (LBS) techniques. That is, LBS can apply the transformations per grip Z. jInterpolate across cubic points. LBS can be used, for example, to animate rigid components in various examples, such as industrial machinery, articulated vehicles, and / or connected parts. The deformation-per-point formula integrates the influence of at least one (e.g., each) control handle by applying a weighted sum of affine transformations, facilitated by the neural weights acquired during training. That is, the deformation-per-point formula is used in such a way that the object can maintain its structural coherence (e.g., the points respond to the movements directed by the control handles).

[0123] In general, the training in the simulation process can include the calculation of neural weights for control handles assigned to cubic points to define the deformation and motion characteristics of the object. During training, the Simulator 120 can iteratively adjust the neural weights, at least partially, based on perturbations applied to control points, using optimization techniques (e.g., gradient descent, Newton's method) to minimize a predefined objective function, such as deformation error or potential energy. The influence of at least one (e.g., each) control handle on the surrounding cubic points can be quantified by the neural weights (e.g., by defining how motions or forces applied to the handle propagate to the geometry of the object).In rigid simulations, the training phase can be used, for example, to ensure that transformations applied to specific handles, such as moving or rotating a machine's arm, are reflected in all connected regions while maintaining structural integrity. That is, the training can result in a set of neural weights that can be applied during the simulation stage, allowing the simulator to animate the object, at least partially, with high fidelity based on the learned control points and their corresponding zones of influence.

[0124] The Simulator 120 can also provide interactive simulations, allowing users to influence the simulation in real time (or near real time) by manipulating control handles or adjusting parameters. This means users can update or modify inputs such as external forces or a time slider (e.g., applying directional forces, setting limits, modifying simulation speeds) to observe the object's response under different conditions. For example, a user can simulate the independent movement of specific parts of a rigid object (e.g., adjusting the arm of an excavator while the body remains stationary to observe the distribution of transformations across the object).This means that the Simulator 120 can apply the calculated transformations using the formula for deformation per point, incorporating neural weights and affine transformations to accurately animate the object.

[0125] With reference to Fig. Section 6C presents a block diagram of an exemplary elasticity simulation stage in an exemplary pipeline according to some implementations of the present disclosure. In some implementations, the simulation of one or more interactions of the object may include performing an elasticity simulation (e.g., elastic simulation: simulation of the motion and behavior of the compressed object). The elasticity simulation may involve the simulator 120 applying a plurality of transformations to the at least one compressed object using a second physics model (e.g., a finite element model, a mass-spring system, or any mesh-free method) to obtain a plurality of deformed states of the at least one compressed object.This means the transformation can include determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters. Additionally, the transformation can include minimizing the energy function (e.g., minimizing potential energy) to determine a plurality of updates for a plurality of control points. Furthermore, the transformation can include calculating one or more deformations (e.g., animating the Gaussian representations of the object) of the at least one densified object, based at least partially on the plurality of updates of the plurality of control points (e.g., points controlling the deformation) and a plurality of corresponding skinning fields (e.g., learned weights used to determine how the control points affect the object's deformation—the skinning fields can be cleared of deformation gradients).

[0126] In some implementations, the simulator can perform 120 elastic simulations by using a neural network (e.g., neural feedforward network, neural convolutional network, recurrent neural network) to implement a set of neural weights. W1i,W2i,…,Wni to model and / or calculate the influence of each control handle on the object's deformation. That is, at least one (e.g., each) control handle can correspond to a Gaussian point or a set of points, and the neural skinning functions can be used by Simulator 120 to determine how strongly (e.g., size, radius of influence, degree of deformation) a handle affects the movement of the surrounding Gaussian representations. For example, the neural weights can W1i They can be optimized to minimize an objective function that includes both an elastic loss term and an orthogonality loss term. In another example, the neural weights can be applied to various elastic models (e.g., linear blend skinning, dual quaternion skinning, spline-based deformation) to simulate different material properties (e.g., rubber-like elasticity or flexible material behavior).

[0127] The training phase for these neural weights can include the implementation of self-monitoring by the Simulator 120. For example, small perturbations can be applied to the control points, and the resulting deformations can be used to refine the weight values. Thus, the simulator can, for instance, perform small perturbations to determine the optimal weight field W* that minimizes the combined loss function. W*:=argmin(λelasticLeelastic+λorthoLortho) where Elastic refers to the energy required to elastically deform the object, to Lortho refers to ensuring that the deformations are orthogonal to each other (e.g., to prevent unnatural movements), λ elasctisch refers to a deformed target position of the object points under elastic forces and λ orthoThis refers to a limitation on maintaining orthogonality between deformation types (e.g., that one deformation does not interfere with other deformations). In some implementations, the training process can be accelerated using numerical gradient calculations or NVIDIA Warp (e.g., reducing training time to 30–90 seconds per object). This means that once neural skinning functions are trained, they can define a manipulation handle for each Gaussian representation, facilitating accurate and efficient simulation of elastic behavior.

[0128] With reference to Fig. Figure 7 shows a block diagram of an exemplary simulation of an object in an exemplary pipeline according to some implementations of the present disclosure. Fig. Figure 7 represents an interactive simulation mode 700 in which the application 124 and / or the simulator 120 can be used to simulate user interactions with an object represented by Gaussian splats. That is, the application 124 can process user inputs such as clicking and dragging to apply localized forces (e.g., tensile forces, torsional forces, compressive forces) to the nearest Gaussian representations, causing the object to deform or move in response to the applied forces. For example, as in Fig. As shown in Figure 7, the user can interactively manipulate the doll by clicking and dragging on different parts of its body, with the application 124 and / or the simulator 120 applying corresponding forces to generate realistic movements and deformations of the doll in the scene. The interactive simulation mode 700 can use the calculation configurations of the application 124 and / or the simulator 120 to provide real-time feedback.

[0129] With reference to Fig. Section 8A presents an exemplary flowchart illustrating a procedure for scene reconstruction, segmentation, preprocessing, compression, and / or simulation in an exemplary pipeline, according to some implementations of the present disclosure. It should be noted that this and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that can be implemented as single or distributed components, or in conjunction with other components, and in any combination and location.Various functions described herein, performed by entities, can be executed by hardware, firmware, and / or software. For example, various functions can be performed using one or more processors that execute instructions stored in one or more memory locations. For example, in some implementations, the systems and procedures described herein can be implemented using one or more generative language models (e.g., as in ). Fig. 9A-9C), one or more computing devices or components thereof (e.g., as described in Fig. 10) and / or one or more data centers or components thereof (e.g., as described in Fig. 11) will be implemented.

[0130] With reference to Fig. 8A, each block of Method 800 described herein contains a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed using one or more processors executing instructions stored in one or more memory locations. The method can also be embodied as computer-usable instructions stored on computer storage media. The method can be provided by a standalone application, service, or hosted service (alone or in combination with another hosted service), as microservices via an application programming interface (API), or as a plug-in for another product, to name just a few. Furthermore, Method 800 is exemplified with respect to the system of Fig. 1 described. However, this procedure can additionally or alternatively be performed by any system or any combination of systems, including, but not limited to, the systems described herein.

[0131] Fig. Figure 8A is a flowchart showing a method 800 for scene reconstruction, segmentation, preprocessing, compression, and / or simulation in an exemplary pipeline, according to some implementations of the present disclosure. Method 800 includes, in block 810, the reception (e.g., by the reconstructor 108) of video data from a video source (e.g., video source 104) containing a depth map of a scene. The video data can, for example, be captured RGB frames containing camera information such as intrinsic parameters and poses. Furthermore, the depth map can provide depth information for one or more (e.g., each) pixels.

[0132] Procedure 800 includes, in block 820, the reconstruction (e.g., by Reconstructor 108) of the scene into a three-dimensional (3D) representation (e.g., neural representations, 3D Gaussian splats). For example, one or more Gaussian splat representations can be used to reconstruct the scene as a series of Gaussian distributions (splats). In another example, the depth map can be converted into a point cloud to generate the 3D Gaussian splats. In some implementations, the processing circuitry can generate an initial Gaussian splat representation during an initialization phase of the reconstruction, based at least partially on the depth data from the depth map and the one or more initial poses of the video source. For example, inaccurate camera poses and low-resolution depth maps can be used to output the initial Gaussian splats.In some implementations, during a training phase (after initialization) of the reconstruction, the initial Gaussian splat representation, the RGB frames, and the optimized camera pose can be inputted to obtain (or receive) a 3DGS reconstruction. The 3DGS reconstruction can be based on at least one of (i) at least one refined pose of the video source and a plurality of two-dimensional (2D) frames of the video data. That is, the 3D reconstruction can be aligned to the 2D video frames by using updated camera poses that are at least partially based on the initial Gaussian splat representation (initial Gaussian), the at least one refined pose of the video source, and / or the plurality of 2D frames. Additionally, the processing circuitry can take at least one initial pose (e.g.,Update inaccurate camera poses) of the video source to at least one refined pose that is based at least partially on aligning the 3D representation with the majority of 2D frames of the video data.

[0133] Procedure 800, in block 830, includes the segmentation (e.g., by segmentor 112) of at least one object in the 3D representation and / or segmentation data corresponding to that at least one object. This means the processing circuitry can identify and isolate an object in the 3D representation. The segmentation can include generating a two-dimensional (2D) segmentation mask (e.g., a binary mask that identifies specific regions corresponding to the object) of a reference view of the video data. Additionally, the segmentation can include interpolating the 2D segmentation mask across multiple frames of the video data. For example, the processing circuitry can maintain consistency across frames by applying the 2D segmentation mask to multiple views. In this example, the mask can specify which pixels of the image are the object or the background.In some implementations, segmentation may involve mapping the 2D segmentation mask across multiple frames to at least one corresponding region of a plurality of regions in the 3D representation to segment the at least one object (e.g., in the 3D representation of the scene) and / or segmentation data (e.g., corresponding to the at least one object). For example, the processing circuitry may map 2D pixels in the mask to specific 3D regions of the scene (e.g., each pixel may correspond to a location in 3D space). In this example, the processing circuitry can identify which Gaussian splats correspond to the object in 3D space. In some implementations, segmentation may involve the use of a segmentation model (e.g., Segment Anything Model (SAM)). Furthermore, the reference view may be based, at least partially, on user input.For example, the user can control the segmentation process by selecting (or making multiple selections of) an object (or sections of the object) of interest in a specific frame. In this example, the reference view can correspond to a single frame (e.g., a snapshot of the video serving as the reference view) from the majority of frames in the video data.

[0134] Method 800 includes, in block 840, the updating (e.g., by preprocessor 116) of at least one of the plurality of regions of the 3D representation within a threshold distance of the at least one object in the scene. The 3D Gaussian splat regions can be preprocessed, for example, by performing inpainting and artifact removal. In some implementations, the processing circuitry can at least partially fill at least one of the plurality of regions within the threshold distance based on sample data from one or more neighboring regions. For example, the processing circuitry can fill gaps and / or holes in the object model by sampling data from nearby regions. In some implementations, the processing circuitry can remove one or more elements from at least one of the plurality of regions within the threshold distance.This means that poorly reconstructed areas can be discarded or replaced. Additionally, removal can involve updating at least one or more regions, at least partially, based on sample data from one or more regions of the 3D representation.

[0135] Method 800, in block 850, includes the densification (e.g., by Simulator 120) of the at least one object by sampling a plurality of points on or approximately around the at least one object. The processing circuits can, for example, convert sampled Gaussian splats into a voxel grid to represent the object's structure. Additionally, the densification can include the generation (e.g., Step 1: Gaussian Representations to Voxels) of a voxelized volume of the at least one object, based at least partially on the plurality of points. That is, the processing circuits can subdivide the space around the object into voxels, thus creating a structured representation of 3D space. In some implementations, the densification can include the updating (e.g.,Step 2: Depth slice) of the voxelized volume, at least partially based on the occupancy state of at least one voxel, or at least partially based on a rendered depth map. The processing circuits can, for example, remove unoccupied voxels by comparing the voxel positions with the depth map values ​​and refine the volume to match the object's geometry. The processing circuits can fill an interior of the voxelized volume, at least partially based on injecting a plurality of volumetric elements (e.g., isotropic Gaussian representations, particles, or any discrete sampling) into the interior of the at least one object, including a plurality of interior regions. That is, the volumetric elements can be injected into Block 860 for the purpose of performing the simulations.

[0136] Procedure 800 includes in Block 860 the simulation (e.g., by Simulator 120) of one or more interactions of the voxelized volume of the at least one compressed object to update at least one physical attribute of the at least one object. That is, the at least one compressed object may correspond to a volumetric representation (e.g., provided during compression in Block 850). In some implementations, the simulation of the one or more interactions may include performing a stiffness simulation (e.g., a stiffness simulation that simulates the motion and behavior of the compressed object). That is, the simulation may include the processing circuitry applying a first plurality of transformations to the at least one compressed object using a first physics model to obtain a plurality of stiff motions of the at least one compressed object.For example, the application of the transformation can be based, at least in part, on the processing circuitry determining an energy function using a plurality of scene parameters (e.g., external forces such as gravity and constraints such as boundary conditions) or a plurality of object parameters (e.g., physical properties of the object, such as the initial state of the sampled cubature points). In this example, the application of the transformation can be based, at least in part, on the processing circuitry minimizing the energy function (e.g., minimizing the potential energy) to determine a plurality of stiff states of the at least one compressed object.Furthermore, the application of the transformation can also be based, at least partially, on the processing circuits applying the majority of stiff states to simulate the majority of stiff movements of the at least one compressed object over time.

[0137] In some implementations, the simulation of one or more interactions may include performing an elasticity simulation (e.g., elastic simulation, simulation of the motion and behavior of the compressed object). That is, one or more operations for simulating the one or more interactions include at least one operation for performing an elasticity simulation, which, using a second physics model, involves applying a second plurality of transformations to the at least one compressed object to obtain a plurality of deformed states of the at least one compressed object.The simulation may, for example, involve the processing circuitry applying a second plurality of transformations to the at least one densified object using a second physics model to obtain a plurality of deformed states of the at least one densified object. For example, the application of the transformation may be based, at least in part, on the processing circuitry determining an energy function, using at least one parameter from a plurality of scene parameters or a plurality of object parameters. In this example, the application of the transformation may be based, at least in part, on the processing circuitry minimizing the energy function (e.g., minimizing potential energy) to determine a plurality of updates for a plurality of control points.Furthermore, the application of the transformation can also be based, at least partially, on the processing circuits calculating one or more deformations (e.g., animating the Gaussian representations of the object) of the at least one densified object, which are based, at least partially, on the plurality of updates of the plurality of control points (e.g., points that can control the deformation) and a plurality of corresponding skinning fields (e.g., learned weights that are used to determine how the control points affect the deformation of the object).

[0138] Procedure 800 includes, in block 870, the generation of at least one image of the 3D representation that depicts at least one section of the at least one object for display. In some implementations, the processing circuitry can generate the at least one object for display (e.g., on application 124). The generation might, for example, involve rendering the voxelized volume or the 3D Gaussian splat representation of the object into a visual output from multiple viewpoints to capture different angles and aspects of the object's structure and behavior. In this example, the at least one image might be a sequence of still images showing the object's deformations and interactions over time, and the 3D representation might be a model that includes physical properties and simulated effects.This means that the processing circuitry can generate (e.g., render) the simulated object for visualization or interaction on a user interface. For example, the application can display the object in various states, based at least partially on the simulation results, allowing the user to observe the object's behavior under different conditions. In some implementations, generation may also include the use of rendering techniques such as shadow mapping or global illumination to enhance the realism of the visual output. Furthermore, display can facilitate updates or modifications by the user. For instance, the user can modify simulation parameters or apply new forces to the object to see how it reacts.This means that the processing circuits can update the visual display, at least partially, in real time based on user input.

[0139] The disclosed implementations can be found in a variety of different systems, such as automotive systems (e.g., a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine), systems implemented using a robot or robotic platform, antenna systems, media systems, boat systems, intelligent area surveillance systems, systems for performing deep learning operations, systems for performing simulation operations (e.g., in a driving or vehicle simulation, in a robot simulation, in a smart city or surveillance simulation, etc.), systems for performing digital twin operations (e.g.,in conjunction with a platform or system for collaborative content creation, such as, without limitation, NVIDIA OMNIVERSE and / or any other platform, system, or service that uses USD or OpenUSD data types; systems implemented using an edge device; systems that include one or more virtual machines (VMs); systems for performing operations to generate synthetic data (e.g., using one or more neural rendering fields (NERFs), neural representation techniques, diffusion models, transformer models, etc.); systems that are at least partially implemented in a data center; systems for performing conversational AI operations; systems that implement one or more language models—such as one or more large language models (LLMs), one or more vision language models (VLMs), one or more multimodal language models, etc., systems for performing light transport simulations, systems for performing collaborative content creation for 3D assets (e.g., using Universal Scene Describer (USD) data such as OpenUSD, computer-aided design (CAD) data, 2D and / or 3D graphics or design data and / or other data types), systems that are implemented at least partially using cloud computing resources, and / or other types of systems.

[0140] With reference to Fig. Section 8B presents an exemplary flowchart illustrating a procedure for object compaction and / or simulation in an exemplary pipeline, according to some implementations of the present disclosure. It is noted that this and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as single or distributed components, or in conjunction with other components, in any combination and location. Various functions described herein, performed by entities, may be executed by hardware, firmware, and / or software.For example, various functions can be executed using one or more processors that execute instructions stored in one or more memory locations. For example, in some implementations, the systems and procedures described herein can be implemented using one or more generative language models (e.g., as in ). Fig. 9A-9C), one or more computing devices or components thereof (e.g., as described in Fig. 10) and / or one or more data centers or components thereof (e.g., as described in Fig. 11) will be implemented.

[0141] With reference to Fig. 8B, each block of Method 880 described herein contains a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed using one or more processors executing instructions stored in one or more memory locations. The method can also be embodied as computer-usable instructions stored on computer storage media. The method can be provided by a standalone application, service, or hosted service (alone or in combination with another hosted service), as microservices via an application programming interface (API), or as a plug-in for another product, to name just a few. Furthermore, Method 880 is exemplified with respect to the system of Fig. 1 described. However, this procedure can additionally or alternatively be performed by any system or any combination of systems, including, but not limited to, the systems described herein.

[0142] Fig. Figure 8B is a flowchart showing Method 880 for object compression and / or simulation in an exemplary pipeline, according to some implementations of this disclosure. Method 880 includes, in block 882, the reception of at least one object (e.g., segmentation data corresponding to the at least one object) that has been segmented from video data. That is, the process circuits can receive and / or obtain (or identify) a 3D representation of the segmented object (e.g., a Gaussian splat representation and / or neural representation). The segmented object can, for example, be represented by a set of Gaussian splats corresponding to its surface structure. In some implementations, the segmented object can contain metadata, such as object boundaries and pose information.The object metadata can include, for example, camera poses and depth information, which are used during reconstruction and segmentation.

[0143] Method 880, in block 884, includes the densification of the at least one object (e.g., of segmentation data corresponding to the at least one object). In general, the densification can involve converting a set of 3D Gaussian splats into a dense voxel grid to simulate volumetric mass. That is, the processing circuitry can inject a plurality of volumetric elements (e.g., isotropic Gaussian representations) into the interior of the voxelized volume, filling regions that were previously hollow to create a solid representation. The processing circuitry can, for example, use a hierarchical voxelization method (e.g., CUDA-based Octree, Kaolin's SPC) to subdivide the space around the object into a structured voxel grid that captures both the surface and the interior of the object.In some implementations, the processing circuitry can initialize voxelization by enclosing the axis-aligned bounding box of the Gaussian splats in a cubic root node (which can, for example, be recursively subdivided into finer nodes). Furthermore, the voxel grid can be updated, at least partially, based on occupancy states determined from depth maps rendered from different viewpoints. For example, the processing circuitry can perform a depth slice to refine the voxel grid and remove unoccupied voxels, at least partially, based on the rendered depth map information. In some implementations, the output voxelized volume can provide a high-resolution representation of the object (e.g., for physics simulations).

[0144] Method 880, in block 886, involves sampling a plurality of points on or near the at least one object. That is, the processing circuitry can sample points on or around the object to generate a grid of voxels representing the object's structure. For example, the processing circuitry can distribute the sampled points evenly across the object's surface and within its interior, creating a voxelized representation that captures the object's geometric and volumetric properties. In some implementations, the processing circuitry can use Gaussian splats as sampling points to generate the voxel grid, converting the splats to voxels, at least partially, based on their positions and covariances. Furthermore, the sampling process can be refined to achieve a desired resolution for the voxelized volume.For example, the processing circuitry can adjust the sampling density to ensure that the voxel grid accurately represents the object's surface and internal features. In some implementations, the voxelized volume can be further refined by aligning the sampled points with depth maps rendered from different viewpoints.

[0145] Method 880, in block 888, contains the generation of a voxelized volume of the at least one object, based at least partially on the plurality of points. That is, the processing circuits can generate a structured voxel grid that represents the geometry of the object, at least partially, based on the sampled points. For example, the processing circuits can assign at least one (e.g., each) point to a voxel, at least partially, based on its position and covariance, thereby creating a dense grid that captures both surface and internal features of the object. In some implementations, the processing circuits can generate a voxelized shell of the object by defining the boundaries of each voxel, at least partially, based on the positions of the Gaussian splats.Furthermore, the processing circuits can adjust the voxel grid resolution to achieve a desired level of detail. For example, the processing circuits can subdivide the voxel grid into smaller nodes to refine the rendering of certain regions (e.g., complex, highly pixelated ones). In some implementations, the processing circuits can use a hierarchical voxel grid to store the generated volume.

[0146] Method 880, in block 890, includes updating the voxelized volume at least partially based on the occupancy state of at least one voxel and at least partially based on a rendered depth map of multiple voxels within the voxelized volume. This means the processing circuits can determine the occupancy state of each voxel, at least partially, based on depth information obtained from multiple viewpoints. For example, the processing circuits can perform ray tracing of an icosahedral array of viewpoints to generate depth maps that capture the threshold distance to the object's surface from different angles.In some implementations, the processing circuitry can merge one or more depth maps to classify each voxel as occupied, empty, or unseen, and refine the voxelized volume to match the object's geometry. Furthermore, the processing circuitry can update the voxelized volume by at least partially removing unoccupied voxels based on the depth map values. For example, the processing circuitry can use a sparse point cloud (SPC) representation to store the voxel occupancy states. In this example, the SPC can enable efficient querying and manipulation of the voxelized volume.

[0147] Method 880, in block 892, includes the simulation of one or more interactions of the voxelized volume of the at least one compressed object to update at least one physical attribute of the at least one object. That is, the processing circuits can simulate interactions such as deformation, collision, or motion, at least partially, based on the object's material properties and external forces. For an elastic object, for example, the processing circuits can simulate deformations using control points and / or neural weights to animate the object, at least partially, based on input forces such as gravity or user interactions. In another example, for a rigid object, the processing circuits can simulate the dynamics of the rigid body by applying a set of affine transformations to the voxelized representation of the object (e.g.,a rigid object). In some implementations, the processing circuitry can use different physics models for rigid and elastic simulations and adjust parameters such as stiffness, density, or external forces. Furthermore, the simulation results can include updates to the object's position, orientation, and shape over time. For example, the processing circuitry can animate the object using techniques such as linear blend skinning or dual quaternion skinning to visualize the simulated interactions. In another example, the processing circuitry can generate visual outputs that depict the simulated behavior of the object under various conditions.

[0148] Method 880, in block 894, includes the generation of at least one image of the 3D representation that depicts at least one section of the at least one object for display. This means that the at least one object can be displayed. The generation can, for example, involve rendering the voxelized volume or the 3D Gaussian splat representation of the object into one or more images from different viewpoints. In this example, the at least one image can be a rendered single frame depicting the geometry, material properties, and all simulated interactions of the object, and the 3D representation can be a visual model of the object with textures and lighting effects. This means that the processing circuitry can generate (or render) and display the simulated object within a graphical user interface or visualization platform.In some implementations, the generation process may incorporate photorealistic rendering techniques such as ray tracing or path tracing. For example, the processing circuitry can visualize the object's state over one or more timeframes, depicting deformations, movements, or other physical changes. In some implementations, the display may include interactive features that allow the user to manipulate the object, modify simulation parameters, or view the simulation from different perspectives. Furthermore, the processing circuitry can update the visualization in real time while one or more simulations are running. The display may, for instance, provide comparative views of different simulation results, highlighting changes in object properties or behavior under varying conditions. EXEMPLARY LANGUAGE MODELS

[0149] In at least some implementations, language models such as large language models (LLMs), visual language models (VLMs), multimodal language models (MMLMs), and / or other types of generative artificial intelligence (AI) can be implemented. Generally, the language models can be used to process, analyze, and generate multimodal content (e.g., text, images, video, 3D models) in various applications, as in the 3D RCS pipeline described above. That is, the models can interpret and produce outputs that meet the specific requirements of the reconstruction, segmentation, condensation, and / or simulation stages. These models can be capable of processing text (e.g., natural language text, code, etc.), images, videos, computer-aided design (CAD) assets, OMNIVERSE and / or METAVERSE file information (e.g.,in USD format, such as OpenUSD) and / or the like, to understand, summarize, translate, and / or otherwise generate at least partially based on the context provided in prompts or queries. These language models can be considered "large" in implementations, at least partially, based on the fact that the models are trained on massive datasets and have architectures with a large number of machine learning parameters (weights and biases)—such as millions or billions of parameters. The LLMs / VLMs / MMLMs / etc. can be implemented to summarize textual data, analyze data and derive insights from it (e.g., text, image, video, etc.), and generate new text / image / video / etc. in user-specified styles, tones, and / or formats. The LLMs / VLMs / MMLMs / etc.The present disclosure may be used in implementations solely for text processing, while in other implementations multimodal LLMs may be implemented to generate text and / or other types of content such as images, audio, 2D and / or 3D data (e.g., in USD formats), and / or video. For example, visual language models (VLMs) or, more generally, multimodal language models (MMLMs) may be implemented to accept input data types such as image, video, audio, text, 3D design (e.g., CAD), and / or others, and / or to generate or output data types such as image, video, audio, text, 3D design, and / or others.

[0150] Different types of architectures for LLMs / VLMs / MMLMs, etc., can be implemented in various implementations. For example, different architectures can be implemented that use different techniques for understanding and generating outputs—such as text, audio, video, still images, 2D and / or 3D design, or asset data, etc. In some implementations, architectures for LLMs / VLMs / MMLMs, etc., can be used, such as recurrent neural networks (RNNs) or long-short-term memory networks (LSTMs), while in other implementations, transformer architectures—such as those based on self-attention and / or cross-attention mechanisms (e.g., between context data and text data)—are used to understand and recognize relationships between words or tokens and / or context data (e.g., other text, video, image, design data, USD, etc.).One or more generative processing pipelines containing LLMs / VLMs / MMLMs / etc. may also contain one or more diffusion blocks (e.g., denoisers). The LLMs / VLMs / MMLMs / etc. of this disclosure may contain one or more encoder and / or decoder blocks. For example, discriminative or encoder-only models such as BERT (Bidirectional Encoder Representations from Transformers) may be implemented for tasks involving language comprehension, such as classification, sentiment analysis, question answering, and named entity recognition. As another example, generative or decoder-only models such as GPT (Generative Pretrained Transformer) may be implemented for tasks involving speech and content generation, such as text completion, action generation, and dialogue generation., which contain both encoder and decoder components, such as T5 (Text-to-Text Transformer), can be implemented to understand and generate content, such as for translations and summaries. These examples are not intended to be restrictive, and any type of architecture—including, but not limited to, the one described herein—can be implemented depending on the specific implementation and the task(s) performed using the LLMs / VLMs / MMLMs / etc.

[0151] In various implementations, LLMs / VLMs / MMLMs / etc. can be trained using unsupervised learning, where an LLM / VLM / MMLM / etc. learns patterns from large amounts of unlabeled text / audio / video / image / design / USD / etc. data. Due to the extensive training, the models in some implementations may not require task-specific or domain-specific training. LLMs / VLMs / MMLMs / etc. that have undergone extensive pretraining with enormous amounts of unlabeled data can be considered foundational models and can be suitable for a variety of tasks, such as answering questions, summarizing, filling in missing information, translating, and generating images / videos / designs / USD / data. Some LLMs / VLMs / MMLMs / etc.They can be tailored for a specific use case using techniques such as prompt tuning, fine-tuning, retrieval augmented generation (RAG), adding adapters (e.g., custom-tailored neural networks and / or layers of neural networks that tune or adapt prompts or tokens to align the language model with a particular task or domain), and / or using other fine-tuning or tailoring techniques that optimize the models for use in specific tasks and / or domains.

[0152] In some implementations, the LLMs / VLMs / MMLMs / etc. of this disclosure can be implemented using various model alignment techniques. For example, guardrails can be implemented in some implementations to identify impermissible or unwanted inputs (e.g., requests) and / or outputs of the models. The system can use the guardrails and / or other model alignment techniques to either prevent a specific unwanted input from being processed using the LLMs / VLMs / MMLMs / etc., and / or to prevent the output or presentation (e.g., display, audio output, etc.) of information generated using the LLMs / VLMs / MMLMs / etc. In some implementations, one or more additional models—or layers thereof—can be implemented to identify problems with the inputs and / or outputs of the models.For example, these “security models” can be trained to identify inputs and / or outputs that are “safe” or otherwise acceptable or desirable, and / or that are “unsafe” or otherwise undesirable for the particular application / implementation. As a result, the LLMs / VLMs / MMLMs / etc. of this disclosure are less likely to output speech / text / audio / video / design data / USD data / etc. that is offensive, vulgar, inappropriate, unsafe, non-technical, and / or otherwise undesirable for the particular application / implementation.

[0153] In some implementations, the LLMs / VLMs / etc. can be configured to access or use one or more plugins, application programming interfaces (APIs), databases, data stores, repositories, etc. For example, for certain tasks or operations for which it is not ideally suited, the model may have instructions (e.g., as a result of training and / or at least partially based on instructions in a particular request) to access one or more plugins (e.g., third-party plugins) that assist in processing the current input. In such an example, where at least part of a request is related to restaurants or the weather, the model may access one or more restaurant or weather plugins (e.g., via one or more APIs) to retrieve the relevant information.As another example where at least part of a response requires a mathematical calculation, the model can access one or more mathematical plugins or APIs to assist in solving the one or more problems and then use the plugin's and / or API's response in the model's output. This process can be repeated—e.g., recursively—for any number of iterations and using any number of plugins and / or APIs until a response can be generated that addresses each question / request / requirement / process / operation / etc. Therefore, the one or more models can rely not only on their own knowledge gained from training on one or more large datasets but also on the expertise or optimized nature of one or more external resources—such as APIs, plugins, and / or the like.

[0154] In some implementations, multiple language models (e.g., LLMs / VLMs / MMLMs / etc.), multiple instances of the same language model, and / or multiple prompts served to the same language model or instance of the same language model can be implemented, executed, or accessed (e.g., using one or more plugins, user interfaces, APIs, databases, data stores, repositories, etc.) to provide output in response to the same query or in response to separate parts of a query. In at least one implementation, multiple language models, such as language models with different architectures or language models trained on different (e.g., updated) datasets, can be served with the same input query and the same prompt (e.g., a set of constraints, conditioners, etc.).In one or more implementations, the language models can be different versions of the same base model. In one or more implementations, at least one language model can be instantiated as multiple agents—for example, more than one prompt can be provided to constrain, direct, or otherwise influence the style, content, or character of the output provided. In one or more non-constraint implementations, the same language model can be instructed to provide output corresponding to a different role, perspective, character, or with a different knowledge base, etc., as defined by a provided prompt.

[0155] In each of these implementations, the output of two or more (e.g., each) language models, two or more versions of at least one language model, two or more instantiated agents of at least one language model, and / or two further requests provided to at least one language model can be further processed, e.g., aggregated, compared, or filtered, or used to determine (and provide) a consensus response. In one or more implementations, the output of a language model—or version, instance, or agent—can be provided as input to another language model for further processing and / or validation. In one or more implementations, a language model can be requested to produce or otherwise receive output with respect to input source material, with the output being associated with the input source material.Such an assignment might involve, for example, generating a label or a portion of text that is embedded (e.g., as metadata) in input source text or image. In one or more implementations, an output from a language model can be used to determine the validity of input source material for further processing or insertion into a dataset. For example, a language model can be used to assess the presence (or absence) of a target word in a portion of text or an object in an image, annotating the text or image to indicate such presence (or absence). Alternatively, the determination from the language model can be used to determine whether the source material should be included in a maintained dataset, for example, with or without restrictions.

[0156] Fig. Figure 9A is a block diagram of an exemplary generative language modeling system 900, suitable for use in implementing at least some implementations of the present disclosure. In general, the generative language modeling system 900 can be used with various stages of the 3D RCS pipeline. That is, the system can generate parameters, refine segmentation outputs, and simulate the behavior of objects during the reconstruction, segmentation, condensation, and / or simulation stages. In the Fig. 9A illustrated example contains the generative language model system 900 a Retrieval Extended Generation (RAG) component 992, an Input Processor 905, a Tokenizer 910, an Embed Component 920, Plug-ins / APIs 995 and a Generative Language Model (LM) 930 (which may contain an LLM, a VLM, a Multimodal LM, etc.).

[0157] Generally speaking, the Input 905 can receive an Input 901 containing text and / or other types of input data (e.g., audio, video, image, sensor data (e.g., LiDAR, RADAR, ultrasound, etc.), 3D design data, CAD data, Universal Scene Describer (USD) data—such as OpenUSD, etc.), depending on the architecture of the generative LM 930 (e.g., LLM / VLM / MMLM / etc.). In some implementations, the Input 901 contains plain text in the form of one or more sentences, paragraphs, and / or documents. Additionally or alternatively, the Input 901 can contain numeric sequences, pre-computed embeddings (e.g., word or sentence embeddings), and / or structured data (e.g., in tabular formats, JSON, or XML).In some implementations where the generative LM 930 is capable of processing multimodal inputs, the input 901 can combine text with image data, audio data, video data, design data, USD data, and / or other types of input data, such as, but not limited to, those described herein (or it may contain no text at all). Using raw input text as an example, the input processor 905 can prepare raw input text in various ways. For instance, the input processor 905 can perform various types of text filtering to remove noise (e.g., special characters, punctuation, HTML markup, stop words, parts of one or more images, parts of audio, etc.) from relevant text content.In an example that includes stop words (frequent words that tend to convey little semantic meaning), the Input Processor 905 can remove stop words to reduce noise and allow the generative LM 930 to focus on more meaningful content. The Input Processor 905 can apply text normalization, such as converting all characters to lowercase, removing accents, and / or handling special cases like contractions or abbreviations to ensure consistency. These are just a few examples, and other types of input processing can be applied as well.

[0158] In some implementations, a RAG component 992 (which may contain one or more RAG models and / or be performed using the generative LM 930 itself) can be used to retrieve additional information to be used as part of the input 901 or the prompt. RAG can be used to enhance the input to the LLM / VLM / MMLM / etc. with external knowledge, making the answers to specific questions, requests, or requirements more relevant—as in a case where specific knowledge is needed. The RAG component 992 can retrieve this additional information (e.g., grounding information such as grounding text / image / video / audio / USD / CAD / etc.) from one or more external sources, which can then be fed into the LLM / VLM / MMLM / etc. along with the prompt to improve the accuracy of the model's responses or outputs.

[0159] For example, in some implementations, input 901 can be generated using the query or input into the model (e.g., a question, a request, etc.) in addition to the data retrieved using RAG component 992. In some implementations, input processor 905 can parse input 901 and communicate with RAG component 992 (or RAG component 992 can be part of input processor 905 in some implementations) to identify relevant text and / or other data to provide to generative LM 930 as additional context or information from which to identify the reaction, response, or output 990 in general.For example, if the input indicates that the user is interested in a desired tire pressure for a specific make and model of vehicle, the RAG component 992—for example, using a RAG model that performs a vector search in an embedding space—can retrieve the tire pressure information or the relevant text from a digital (embedded) version of the user manual for that specific vehicle make and model. Similarly, if a user revisits a chatbot in connection with a specific product offering or service, the RAG component 992 can retrieve a previously stored conversation history—or at least a summary thereof—and include the previous conversation history, along with the current question / request, as part of the input 901 in the generative LM 930.

[0160] The RAG component 992 can employ various RAG techniques. For example, naive RAG can be used when documents are indexed, split into pieces, and applied to an embedding model to generate embeddings that correspond to the pieces. A user request can also be applied to the embedding model and / or another embedding model of the RAG component 992, and the piece embeddings can be compared with the request's embeddings to identify the most similar embeddings that can be fed to the generative LM 930 to produce output.

[0161] In some implementations, more advanced RAG techniques can be used. For example, the pieces can undergo pre-fetching processes (e.g., redirection, rewriting, metadata analysis, extension, etc.) before being passed to the embedding model. Furthermore, post-fetching processes (e.g., re-ranking, request compression, etc.) can be performed on the outputs of the embedding model before the final embeddings are generated and used as a comparison to an input query.

[0162] Another example is modular RAG techniques, such as those similar to naive and / or extended RAG, but which may also include features like hybrid search, recursive retrieval and query engines, step-back approaches, subqueries, and hypothetical document embedding.

[0163] As another example, Graph-RAG can use knowledge graphs as a source of contextual or factual information. Graph-RAG can be implemented using a graph database as a source of contextual information, which is then sent to the LLM / VLM / MMLM / etc. Instead of providing the model with (or in addition to) snippets of data extracted from larger documents—which can result in a lack of context, factual accuracy, linguistic precision, etc.—Graph-RAG can also provide structured entity information to the LLM / VLM / MMLM / etc. by combining the structured text description of the entity with its many properties and relationships, allowing the model deeper insights. In implementing Graph-RAG, the systems and procedures described herein use a graph as a content store, extract relevant snippets from documents, and request them from the LLM / VLM / MMLM / etc.to respond using this. In such implementations, the knowledge graph can contain relevant text content and metadata about the knowledge graph and be integrated into a vector database. In some implementations, the graph RAG can use a graph as a domain expert, extracting descriptions of concepts and entities relevant to a query / request and passing them to the model as semantic context. These descriptions can include relationships between the concepts. In other examples, the graph can be used as a database, where part of a query / request can be allocated to a graph query, the graph query can be executed, and the LLM / VLM / MMLM / etc. can summarize the results.In such an example, the graph can store relevant factual information, and a query (natural language query) to a graph query tool (NL-to-Graph-query tool) and an entity join can be used. In some implementations, graph RAG (e.g., using a graph database) can be combined with standard RAG (e.g., vector database) and / or other RAG types to benefit from multiple approaches.

[0164] In all implementations, the RAG component 992 can implement a plug-in, an API, a user interface, and / or other functionality to perform RAG. For example, a graph RAG plug-in can be used by the LLM / VLM / MMLM / etc. to query the knowledge graph to extract relevant information for feeding into the model, and a standard or vector RAG plug-in can be used to query a vector database. For example, the graph database can interact with a plug-in's REST interface, thus decoupling the graph database from the vector database and / or the embedding models.

[0165] The Tokenizer 910 can segment (e.g., processed) text data into smaller units (tokens) for subsequent analysis and processing. Depending on the implementation, the tokens can represent individual words, partial words, characters, parts of audio / video / images, etc. Word-based tokenization divides the text into individual words, with each word treated as a separate token. Partial word tokenization breaks words down into smaller meaning-bearing units (e.g., prefixes, suffixes, stems), enabling the generative LM 930 to understand morphological variations and more effectively handle words not in the vocabulary. Character-based tokenization represents each character as a separate token, allowing the generative LM 930 to process text at a fine-grained level.The choice of tokenization strategy can depend on factors such as the language being processed, the task at hand, and / or the characteristics of the training dataset. Therefore, Tokenizer 910 can convert the (e.g., processed) text into a structured format according to the tokenization scheme implemented in the respective implementation.

[0166] The Embedding Component 920 can use any known embedding technique to convert discrete tokens into (e.g., dense, continuous vector) representations of semantic meaning. For example, the Embedding Component 920 can use pre-trained word embeddings (e.g., Word2Vec, GloVe, or FastText), one-hot coding, term frequency-inverse document frequency (TF-IDF) coding, one or more neural network embedding layers, and / or other techniques.

[0167] In some implementations where the input 901 contains image data / video data / etc., the input processor 901 can resize the data to a standard size compatible with the format of a corresponding input channel and / or normalize pixel values ​​to a common range (e.g., 0 to 1) to ensure uniform representation. The embedding component 920 can encode the image data using any known technique (e.g., using one or more convolutional neural networks (CNNs) for visual feature extraction). In some implementations where the input 901 contains audio data, the input processor 901 can resample an audio file to a uniform sampling rate for consistent processing. The embedding component 920 can use any known technique to extract and encode audio features, for example, in the form of a spectrogram (e.g., a spectrogram).(of a Mel spectrogram). In some implementations where the input contains video data, the input processor can extract frames or apply resizing to extracted frames, and the embedding component can extract features such as optical flow or video embeddings and / or encode temporal information or sequences of frames. In some implementations where the input contains multimodal data, the embedding component can fuse representations of the different data types (e.g., text, image, audio, USD, video, design, etc.) using techniques such as early fusion (concatenation), late fusion (sequential processing), attention-based fusion (e.g., self-attention, cross-attention), etc.

[0168] The generative LM 930 and / or other components of the generative LM System 900 can use different types of neural network architectures, depending on the implementation. For example, transformer-based architectures, such as those used in models like GPT, can be implemented and include self-attention mechanisms that weigh the importance of different words or tokens in the input sequence, and / or feedforward networks that process the output of the self-attention layers, applying nonlinear transformations to the input representations and extracting higher-level features. Some non-restrictive example architectures include transformers (e.g.,Encoding-decoding (encoder-only, decoder-only, multimodal) RNNs, LSTMs, fusion models, diffusion models, crossmodal embedding models that learn shared embedding spaces, graph neural networks (GNNs), hybrid architectures that combine different types of architectures, adversarial networks such as generative adversarial networks (GANs) or adversarial autoencoders (AAEs) for joint distributional learning, and others. Therefore, depending on the implementation and architecture, the embedding component 920 can apply a coded representation of the input 901 to the generative LM 930, and the generative LM 930 can process the coded representation of the input 901 to produce an output 990 that may contain response text and / or other types of data.

[0169] As described herein, in some implementations, the generative LM 930 may be configured to access or use plug-ins / APIs 995—or have the capability to access or use them (which may include one or more plug-ins, application programming interfaces (APIs), databases, data stores, repositories, etc.). For example, for certain tasks or operations for which the generative LM 930 is not ideally suited, the model may have instructions (e.g., as a result of training and / or at least partially based on instructions in a given request, such as those retrieved using the RAG component 992) to access one or more plug-ins / APIs 995 (e.g., third-party plug-ins) to obtain assistance in processing the current input.In such an example, where at least part of a request is related to restaurants or weather, the model can access one or more restaurant or weather plugins (e.g., via one or more APIs), send at least part of the request related to the specific API 995 to the plugin / API 995, the plugin / API 995 can process the information and return a response to the generative LM 930, and the generative LM 930 can use the response to generate the output 990. This process can be repeated for any number of iterations and using any number of plugins / APIs 995—e.g., recursively—until an output 990 can be generated that addresses every question / request / request / process / operation / etc. from the input 901.Therefore, one or more models can rely not only on their own knowledge from training with a large dataset(s) and / or from data retrieved using the RAG component 992, but also on the expertise or optimized nature of one or more external resources - such as the plug-ins / APIs 995.

[0170] Fig. Figure 9B is a block diagram of an example implementation where the generative LM 930 includes a transformer-encoder-decoder. In general, the generative LM 930 can generate model parameters and processing rules for the stages of the 3D RCS pipeline. That is, the generative LM 930 can generate segmentation masks, update camera poses, and adjust simulation parameters, at least partially, based on input data. For example, suppose an input text such as "Who discovered gravity" is tokenized (e.g., by the Tokenizer 910). Fig. 9A) into tokens such as words, and each token is encoded into a corresponding embedding (e.g., of size 512) (e.g., by the embedding component 920 of Fig. 9A). Since these token embeddings do not typically represent the token's position in the input sequence, any known technique can be used to add positional encoding to each token embedding to encode the sequential relationships and context of the tokens in the input sequence. As such, the (e.g., resulting) embeddings can be applied to one or more encoders 935 of the generative LM 930.

[0171] In an exemplary implementation, the one or more encoders form an encoder stack, with each encoder containing a self-attention layer and a feedforward net. In an exemplary transformer architecture, each token (e.g., word) flows through a separate path. Therefore, each encoder can accept a sequence of vectors, with each vector passing through the self-attention layer, then through the feedforward net, and then up to the next encoder in the stack. Any known self-attention technique can be used.For example, to calculate a self-attention score for each token (word), a query vector, a key vector, and a value vector can be created for each token. A self-attention score for pairs of tokens can be calculated by taking the dot product of the query vector with the corresponding key vectors, normalizing the resulting numerical values, multiplying them by the corresponding value vectors, and summing the weighted value vectors. The encoder can employ multi-headed attention, where the attentional mechanism is applied multiple times in parallel with different learned weight matrices. Any number of encoders can be cascaded to generate a context vector that encodes the input. An attentional projection layer 940 can convert the context vector into attentional vectors (keys and values) for the decoder(s) 945.

[0172] In an exemplary implementation, the one or more decoders 945 form a decoder stack, with each decoder containing a self-attention layer, an encoder-decoder self-attention layer that uses the encoder's attention vectors (keys and values) to focus on relevant parts of the input sequence, and a feedforward network. As with the encoder(s) 935, in an exemplary transformer architecture, each token (e.g., word) flows through a separate path in the decoder(s) 945. During a first pass, the decoder(s), a classifier 950, and a generation mechanism 955 can generate an initial token, and the generation mechanism 955 can apply the generated token as input during a second pass. The process can repeat in a loop, successively processing tokens (e.g., words ...words) are generated and added to the output of the previous pass, and the token embeddings of the compound sequence with positional encodings are applied as an input to decoder(s) 945 during a subsequent pass, generating one token at a time (known as autoregression) until a symbol or token is predicted that represents the end of the response. Within each decoder, the self-attention layer is typically restricted to focusing only on previous positions in the output sequence by applying a masking technique (e.g., setting future positions to negative infinity) before the softmax operation. In an exemplary implementation, the encoder-decoder attention layer works similarly to the (e.g.,multi-headed) self-awareness in the encoder(s) 935, except that it creates its queries from the layer below and takes the keys and values ​​(e.g. matrix) from the output of encoder(s) 935.

[0173] Therefore, the decoder(s) 945 can output a decoded (e.g., vector) representation of the input applied during a given iteration. The classifier 950 can include a multi-class classifier containing one or more layers of a neural network that project the decoded (e.g., vector) representation into an appropriate dimensionality (e.g., one dimension for each supported word or token in the output vocabulary) and a softmax operation that converts logits into probabilities. Therefore, the generation mechanism 955 can select or sample a word or token, at least partially, based on an appropriate predicted probability (e.g., selecting the word with the highest predicted probability) and append it to the output of a previous iteration, generating each word or token sequentially.The generation mechanism 955 can repeat the process, triggering successive decoder inputs and corresponding predictions until a symbol or token is selected or sampled that represents the end of the reaction, at which point the generation mechanism 955 can output the generated reaction.

[0174] Fig. Figure 9C is a block diagram of an exemplary implementation where the generative LM 930 incorporates a decoder-transformer-only architecture. For example, the decoder(s) 960 of Fig. 9C similar to the decoder 945 from Fig. 9B work, with the exception that each of the decoder(s) 960 of Fig. 9C omits the encoder-decoder self-attention layer (since there is no encoder in this implementation). Therefore, the decoder(s) 960 can form a decoder stack, with each decoder containing a self-attention layer and a feedforward network. Furthermore, instead of encoding the input sequence, a symbol or token representing the end of the input sequence (or the beginning of the output sequence) can be appended to the input sequence, and the resulting sequence (e.g., corresponding embeddings with positional encodings) can be applied to the decoder(s). As with the decoder(s) 945 of Fig. 9B allows each token (e.g., word) to flow through a separate path in the decoder(s) 960, and the decoder(s), a classifier 965, and a generation mechanism 970 can use autoregression to sequentially generate one token after another until a symbol or token is predicted that represents the end of the response. The classifier 965 and the generation mechanism 970 can function similarly to the classifier 950 and the generation mechanism 955 of Fig. 9B, wherein the generation mechanism 970 selects or samples each successive output token at least partially based on a corresponding predicted probability and appends it to the output from a previous pass, each token being generated sequentially until a symbol or token is selected or sampled that represents the end of the response. This and other architectures described herein are intended only as examples, and other suitable architectures may be implemented within the scope of protection of this disclosure. EXAMPLE CALCULATION DEVICE

[0175] Fig. Figure 10 is a block diagram of an exemplary computing device(s) 1000 suitable for use in the implementation of at least some implementations of the present disclosure. In general, the exemplary computing devices 1000 can perform various stages of the 3D RCS pipeline, such as reconstruction, segmentation, preprocessing, compaction, and / or simulation. That is, the computing device(s) 1000 can perform calculations to generate 3D representations, segment objects, process volumetric data, compact objects, and / or simulate object interactions within the pipeline.The computing device 1000 can include a connection system 1002 that directly or indirectly couples the following devices: main memory 1004, one or more central processing units (CPUs) 1006, one or more graphics processing units (GPUs) 1008, a communication interface 1010, input / output (I / O) ports 1012, input / output components 1014, a power supply 1016, one or more presentation components 1018 (e.g., display(s)), and one or more logic units 1020. In at least one implementation, the computing device(s) 1000 can include one or more virtual machines (VMs), and / or each of its components can include virtual components (e.g., virtual hardware components).As non-restrictive examples, one or more of the GPUs 1008 can contain one or more vGPUs, one or more of the CPUs 1006 can contain one or more vCPUs, and / or one or more of the logic units 1020 can contain one or more virtual logic units. Thus, a computing device (or devices) 1000 can contain discrete components (e.g., a complete GPU allocated to computing device 1000), virtual components (e.g., a portion of a GPU allocated to computing device 1000), or a combination thereof.

[0176] Although the various blocks of Fig. Where components 10 are shown as being connected via lines through the connection system 1002, this is not intended as a limitation and is only for clarity. In some implementations, for example, a presentation component 1018, such as a display device, may be considered an I / O component 1014 (e.g., if the display is a touchscreen). As another example, the CPUs 1006 and / or GPUs 1008 may contain memory (e.g., the memory 1004 may represent a storage device in addition to the memory of the GPUs 1008, the CPUs 1006, and / or other components). Thus, the computing device of Fig. 10. This is for illustrative purposes only. No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "game console", "electronic control unit (ECU)", "virtual reality system" and / or other device or system types, since all within the scope of the computing device of Fig. 10.

[0177] The 1002 connection system can represent one or more connections or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The 1002 connection system can include one or more bus or connection types, such as an Industry Standard Architecture (ISA) bus, an Extended ISA bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI Express (PCIe) bus, and / or another type of bus or connection. In some implementations, there are direct connections between components. For example, the CPU 1006 can be directly connected to the RAM 1004. Furthermore, the CPU 1006 can be directly connected to the GPU 1008.In a direct or point-to-point connection between components, the connection system 1002 can include a PCIe link to establish the connection. In these examples, a PCI bus does not need to be included in the computing device 1000.

[0178] The main memory 1004 can contain any of a variety of computer-readable media. Computer-readable media can be any available media that the computing device 1000 can access. Computer-readable media can include both volatile and non-volatile media, and removable and non-removable media. For example, and without limitation, computer-readable media can include computer storage media and communication media.

[0179] Computer storage media can include both volatile and non-volatile media, and / or removable and non-removable media, implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, and / or other types of data. For example, main memory can store 1004 computer-readable instructions (e.g., representing a program and / or program element, such as an operating system).Computer storage media may, but are not limited to, include RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, Digital Versatile Discs (DVDs) or other optical disk storage, magnetic cartridges, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that the computing device 1000 can access. As used herein, computer storage media do not per se contain signals.

[0180] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other types of data in a modulated data signal, such as a carrier wave or other transport mechanism, and may include any media for transmitting information. The term "modulated data signal" can refer to a signal in which one or more of its properties are set or modified to encode information within the signal. Computer storage media may include, but are not limited to, wired media, such as a wired network or a direct-wired connection, and wireless media, such as acoustic, RF, infrared, and other wireless media. Combinations of the foregoing should also be included in the scope of protection of the computer-readable media.

[0181] The CPU(s) 1006 can be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1000 to perform one or more of the procedures and / or processes described herein. The CPU(s) 1006 can each contain one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing a plurality of software threads simultaneously. The CPU(s) 1006 can contain any type of processor and can contain different types of processors depending on the type of computing device 1000 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers).Depending on the type of computing device 1000, the processor can be, for example, an Advanced RISC Machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC), or an x86 processor implemented with Complex Instruction Set Computing (CISC). The computing device 1000 can contain one or more CPUs 1006, in addition to one or more microprocessors or additional coprocessors, such as mathematical coprocessors.

[0182] In addition to or as an alternative to the CPU(s) 1006, the CPU(s) 1008 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1000 to perform one or more of the procedures and / or processes described herein. One or more of the GPUs 1008 may be an integrated GPU (e.g., with one or more of the CPU(s) 1006) and / or one or more of the GPUs 1008 may be a discrete GPU. In implementations, one or more of the GPUs 1008 may be a coprocessor of one or more of the CPU(s) 1006. The GPU(s) 1008 may be used by the computing device 1000 to render graphics (e.g., 3D graphics) or to perform general-purpose calculations. The GPU(s) 1008 can be used, for example, for general-purpose computing on GPUs (GPGPU).The GPU(s) 1008 can contain hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU(s) 1008 can generate pixel data for output images in response to rendering commands (e.g., rendering commands from the CPU(s) 1006 received via a host interface). The GPU(s) 1008 can include graphics memory, such as display memory, for storing pixel data or other suitable data, such as GPGPU data. The display memory can be included as part of the main memory 1004. The GPU(s) 1008 can contain two or more GPUs operating in parallel (e.g., via a link). The link can connect the GPUs directly (e.g., using NVLINK) or connect them via a switch (e.g., using NVSwitch).When combined, each GPU can generate 1008 pixel data or GPGPU data for different sections of an output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can have its own dedicated memory or share memory with other GPUs.

[0183] In addition to or as an alternative to the CPU(s) 1006 and / or the GPU(s) 1008, the logic unit(s) 1020 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1000 to perform one or more of the procedures and / or processes described herein. In implementations, the CPU(s) 1006, the GPU(s) 1008, and / or the logic unit(s) 1020 may discretely or jointly perform any combination of the procedures, processes, and / or sections thereof. One or more of the logic units 1020 may be part of and / or integrated within one or more of the CPU(s) 1006 and / or the GPU(s) 1008, and / or one or more of the logic units 1020 may be discrete components or otherwise separate from the CPU(s) 1006 and / or the GPU(s) 1008.In implementations, one or more of the logic units 1020 can be a co-processor of one or more of the CPUs 1006 and / or one or more of the GPUs 1008.

[0184] Examples of Logic Unit(s) 1020 include one or more processing cores and / or components thereof, such as Data Processing Units (DPUs), Tensor Cores (TCs), Tensor Processing Units (TPUs), Pixel Visual Cores (PVCs), Vision Processing Units (VPUs), Graphics Processing Clusters (GPCs), Texture Processing Clusters (TPCs), Streaming Multiprocessors (SMs), Tree Traversal Units (TTUs), Artificial Intelligence Accelerators (AIAs), Deep Learning Accelerators (DLAs), and Programmable Vision Accelerators (PVAs) – which contain one or more Direct Memory Access (DMA) systems,one or more vision or vector processing units (VPUs), one or more pixel processing engines (PPEs) - e.g. B. including a 2D array of processing elements, each communicating north, south, east and west with one or more other processing elements in the array, one or more decoupled accelerators or units (e.g., decoupled lookup table (DLUT) accelerators or units), etc., vision processing units (VPUs), optical flow accelerators (OFAs), field programmable gate arrays (FPGAs), neuromorphic chips, quantum processing units (QPUs), associative process units (APUs), arithmetic logic units (ALUs),Application-specific integrated circuits (ASICs), floating-point units (FPUs), input / output (I / O) elements, peripheral component interconnect (PCI) or PCI Express (PCIe) elements, and / or the like.

[0185] The Communication Interface 1010 can include one or more receivers, transmitters, and / or transceivers that enable the Computing Device 1000 to communicate with other Computing Devices over an electronic network, including wired and / or wireless communication. The Communication Interface 1010 can include components and functions that enable communication over a variety of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., Ethernet or InfiniBand communication), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.In one or more implementations, the logic unit(s) 1020 and / or the communication interface 1010 may contain one or more data processing units (DPUs) to transfer data received via a network and / or the connection system 1002 directly to one or more GPUs 1008 (e.g., one of their memory units).

[0186] The I / O ports 1012 enable the computing device 1000 to be logically coupled with other devices, including the I / O components 1014, the presentation component(s) 1018, and / or other components, some of which may be built into (e.g., integrated with) the computing device 1000. Illustrative I / O components 1014 include a microphone, mouse, keyboard, joystick, gamepad, game controller, satellite dish, scanner, printer, wireless device, etc. The I / O components 1014 can provide a natural user interface (NUI) that processes air gestures, speech, or other physiological inputs generated by a user. In some examples, the inputs can be passed to a suitable network element for further processing.A NUI can implement any combination of speech capture, stylus capture, face capture, biometric capture, gesture capture (both on-screen and off-screen), air gestures, head and eye tracking, and touch capture (as described in more detail below) associated with a display of the Computing Device 1000. The Computing Device 1000 can include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof, for gesture capture and recognition. Additionally, the Computing Device 1000 can include accelerometers or gyroscopes (e.g., as part of an inertial measurement unit (IMU)) that enable motion detection. In some examples, the output from the accelerometers or gyroscopes can be used by the Computing Device 1000 to render immersive augmented reality or virtual reality.

[0187] The power supply 1016 can include a hardwired power supply, a battery power supply, or a combination of both. The power supply 1016 can power the computing device 1000 to enable the operation of the components of the computing device 1000.

[0188] The presentation component(s) 1018 can include a display (e.g., a monitor, a touchscreen, a television screen, a head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component(s) 1018 can receive data from other components (e.g., the GPU(s) 1008, the CPU(s) 1006, DPUs, etc.) and output the data (e.g., as an image, video, sound, etc.). EXEMPLARY DATA CENTER

[0189] Fig. Figure 11 illustrates an exemplary data center 1100 that can be used in at least one implementation of the present disclosure. In general, the exemplary data center 1100 can support the execution of computations and storage for the 3D RCS pipeline. That is, the data center 1100 can process and store data for steps such as reconstruction, segmentation, compression, and / or simulation. The data center 1100 can include a data center infrastructure layer 1110, a framework layer 1120, a software layer 1130, and / or an application layer 1140.

[0190] As in Fig. As shown in Figure 11, the infrastructure layer 1110 of the data center can include a resource orchestrator 1112, grouped compute resources 1114 and node compute resources (“node RRs”) 1116(1)-1116(N), where “N” is any positive integer. In at least one implementation, the Node RRs 1116(1)-1116(N) may include, but are not limited to, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic solid-state memory), storage devices (e.g., solid-state or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power supply modules and / or cooling modules, etc.In some implementations, one or more node RRs among node RRs 1116(1)-1116(N) may correspond to a server that has one or more of the aforementioned compute resources. Furthermore, in some implementations, node RRs 1116(1)-1116(N) may include one or more virtual components, such as vGPUs, vCPUs, and / or the like, and / or one or more of node RRs 1116(1)-1116(N) may correspond to a virtual machine (VM).

[0191] In at least one implementation, the grouped compute resources 1114 can contain separate groupings of node RRs 1116, which are housed in one or more racks (not shown) or in many racks in data centers at different geographic locations (also not shown). Separate groupings of node RRs 1116 within grouped compute resources 1114 can include grouped compute, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one implementation, multiple node RRs 1116, including CPUs, GPUs, DPUs, and / or other processors, can be grouped in one or more racks to provide compute resources to support one or more workloads.The one or more racks can also contain any number of power supply modules, cooling modules and / or network switches in any combination.

[0192] The Resource Orchestrator 1112 can configure or otherwise control one or more Node RRs 1116(1)-1116(N) and / or grouped compute resources 1114. In at least one implementation, the Resource Orchestrator 1112 can include a Software Design Infrastructure (SDI) management entity for the data center 1100. The Resource Orchestrator 1112 can include hardware, software, or a combination thereof.

[0193] In at least one implementation, as in Fig. As shown in Figure 11, the framework layer 1120 can contain a job scheduler 1128, a configuration manager 1134, a resource manager 1136, and / or a distributed file system 1138. The framework layer 1120 can contain a framework that supports the software 1132 of the software layer 1130 and / or an application 1142 of the application layer 1140. The software 1132 or the application 1142 can each contain web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 1120 can be a type of free and open-source software web application framework, such as Apache Spark™ (hereinafter "Spark"), which can use a distributed file system 1138 for processing large amounts of data (e.g., "Big Data"), but is not limited to it.In at least one implementation, the job scheduler 1128 can include a Spark driver to facilitate the scheduling of workloads supported by different layers of the data center 1100. The configuration manager 1134 can be able to configure different layers, such as the software layer 1130 and the framework layer 1120, which contains Spark and the distributed file system 1138, to support the processing of large amounts of data. The resource manager 1136 can be able to manage clustered or grouped compute resources allocated or assigned to support the distributed file system 1138 and the job scheduler 1128. In at least one implementation, the clustered or grouped compute resources can include the grouped compute resource 1114 on the infrastructure layer 1110 of the data center.The resource manager 1136 can coordinate with the resource orchestrator 1112 to manage these allocated or assigned computing resources.

[0194] In at least one implementation, the software contained in software layer 1130 may include software 1132 that is used by at least sections of the node RRs 1116(1)-1116(N), the grouped compute resources 1114, and / or the distributed file system 1138 of framework layer 1120. One or more types of software may include, but are not limited to, web page search software, email virus scanning software, database software, and streaming video content software.

[0195] In at least one implementation, the application(s) contained in application layer 1140 may include one or more types of applications used by at least sections of the node RRs 1116(1)-1116(N), the clustered compute resources 1114, and / or the distributed file system 1138 of framework layer 1120. One or more types of applications may include, but are not limited to, any number of genome applications, cognitive computations, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more implementations.

[0196] In at least one implementation, a configuration manager 1134, resource manager 1136, and resource orchestrator 1112 can implement any number and type of self-modifying actions based on any amount and type of data collected in any technically feasible way. Self-modifying actions can relieve a data center operator of data center 1100 of potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly functioning sections of a data center.

[0197] Data Center 1100 may contain tools, services, software, or other resources to train one or more machine learning models or to predict or infer information using one or more machine learning models according to one or more implementations described herein. For example, a machine learning model(s) may be trained by calculating weighting parameters according to a neural network architecture, using software and / or computing resources described above with reference to Data Center 1100.In at least one implementation, trained or deployed machine learning models corresponding to one or more neural networks can be used to infer or predict information using the resources described above with reference to Computing Center 1100, by using weighting parameters calculated by one or more training techniques such as, but not limited to, those described herein.

[0198] In at least one implementation, the data center can utilize 1100 CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or equivalent virtual computing resources) to perform training and / or inference using the resources described above. Additionally, one or more of the software and / or hardware resources described above can be configured as a service to allow users to train or infer information, such as image capture, speech capture, or other artificial intelligence services. EXEMPLARY NETWORK ENVIRONMENTS

[0199] Network environments suitable for implementing the disclosure may include one or more client devices, servers, a network-attached storage (NAS) device, other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may run on one or more instances of the computing device(s). Fig. 10. Implemented - e.g., each device may contain similar components, features, and / or functionality to the computing device(s) 1000. If backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices may also be included as part of a data center 1100, an example of which is given herein with reference to Fig. 11 is described in more detail.

[0200] The components of a network environment can communicate with each other over a network, which can be wired, wireless, or both. The network can contain multiple networks or a network of networks. For example, the network can contain one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the internet and / or a public switched telephone network (PSTN), and / or one or more private networks. If the network includes a wireless telecommunications network, components such as a base station, a communications tower, or even access points (as well as other components) can provide wireless connectivity.

[0201] Compatible network environments can contain one or more peer-to-peer network environments—in which case a server cannot be included in a network environment—and one or more client-server network environments—in which case one or more servers can be included in a network environment. In peer-to-peer network environments, the functionality described herein can be implemented on any number of client devices with reference to a server.

[0202] In at least one implementation, a network environment can include one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. A framework layer can include a framework for supporting software of a software layer and / or one or more applications of an application layer. The software or application(s) can each include web-based service software or applications. In implementations, one or more of the client devices can use the web-based service software or applications (e.g.,by accessing the service software and / or applications via one or more application programming interfaces (APIs). The framework layer can be a type of free and open-source software web application framework, such as one that uses a distributed file system for processing large amounts of data (e.g., "Big Data"), but is not limited to that.

[0203] A cloud-based network environment can provide cloud computing and / or cloud storage, performing any combination (or parts thereof) of the computing and / or data storage functions described herein. Each of these different functions can be distributed across multiple locations of central or core servers (e.g., one or more data centers, which may be distributed across a state, region, country, the globe, etc.). If a connection to a user (e.g., a client device) is relatively close to an edge server, the core server(s) may offload at least some functionality to the edge server(s). A cloud-based network environment can be private (e.g., restricted to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0204] The client device(s) may include at least some of the components, features, and functions described herein with respect to Fig.The 10 exemplary computing device(s) described may contain 1000. By way of example, and not as a limitation, a client device may be a personal computer (PC), a laptop, a mobile device, a smartphone, a tablet computer, a smartwatch, a portable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or global positioning device, a video player, a video camera, a surveillance device or surveillance system, a vehicle, a boat, a hydrofoil, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or gaming system, an entertainment system, a vehicle computer system, an embedded system controller, a remote control, a device, a consumer electronics device, a workstation, an edge device,any combination of these described devices or any other suitable device may be embodied.

[0205] Disclosure can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions such as program modules that are executed by a computer or other machine, such as a personal data assistant or other handheld device. Generally, program modules contain routines, programs, objects, components, data structures, etc., and refer to code that performs specific tasks or implements certain abstract types of data. Disclosure can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. Disclosure can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected to each other via a network for communication.

[0206] As used herein, any mention of "and / or" in relation to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, "element A, element B and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Furthermore, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Additionally, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0207] The subject matter of this disclosure is specifically described herein to satisfy legal requirements. However, the description itself is not intended to limit the scope of protection afforded by this disclosure. Rather, the inventors have considered that the claimed subject matter may also be embodied in other ways to include various steps or combinations of steps similar to those described in this document, in conjunction with other present or future technologies. Although the terms "step" and / or "block" may be used herein to denote various elements of the methods employed, these terms should not be interpreted as implying any particular sequence among or between the various steps disclosed herein, except where the sequence of each step is expressly described.

[0208] The disclosure of this application also contains the following numbered clauses: Clause 1. System comprising the following: one or more processors to perform one or more operations for the following: Receive at least one object segmented from the video data; Compacting the at least one object by the following: Scanning a plurality of points on or approximately around the at least one object; Generation of a voxelized volume of the at least one object, which is based at least partially on the plurality of points; Updating the voxelized volume at least partially based on the occupancy state of at least one voxel of a plurality of voxels of the voxelized volume at least partially based on at least one rendered depth map; Simulation of one or more interactions of the voxelized volume of the at least one compressed object to update at least one physical attribute of the at least one object; and Generate at least one image that depicts at least one section of the at least one object with at least one updated physical attribute. Clause 2. System according to Clause 1, wherein the one or more operations include at least one operation for the following: Filling an interior space of the voxelized volume, which is based at least partially on the injection of a plurality of volumetric elements into the interior space of the at least one object comprising a plurality of interior regions. Clause 3. System according to Clause 1 or 2, wherein at least one condensed object corresponds to a volumetric representation. Clause 4. System according to any of the preceding clauses, wherein the one or more operations for simulating the one or more interactions includes at least one operation for performing a stiffness simulation by: Applying a first plurality of transformations to the at least one compressed object using a first physics model to obtain a plurality of stiff motions of the at least one compressed object. Clause 5. System according to Clause 4, wherein the plurality of stiff movements of the at least one compressed object are obtained by the following: Determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters; Minimizing the energy function to determine a plurality of stiff states of the at least one compressed object; and Application of the plurality of stiff states to simulate the plurality of stiff movements of the at least one compressed object over time. Clause 6. System according to any of the preceding clauses, wherein the one or more operations for simulating the one or more interactions includes at least one operation for performing an elasticity simulation, which includes the following: Applying a second plurality of transformations to the at least one condensed object using a second physics model to obtain a plurality of deformed states of the at least one condensed object. Clause 7. System according to Clause 6, wherein the plurality of deformed states of the at least one compacted object are obtained by the following: Determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters; Minimizing the energy function to determine a plurality of updates for a plurality of control points; and Calculation of one or more deformations of the at least one densified object, which are based at least partially on the plurality of updates of the plurality of control points and a plurality of corresponding skinning fields. Clause 8. System according to any of the preceding clauses, wherein the one or more processors are contained in at least one of the following: a system for conducting games; a system for conducting content streaming; a system for conducting collaborative content creation; a system for carrying out simulation processes; a system for conducting collaborative content creation for 3D assets; a system for generating synthetic data; a system that includes one or more large visual language models (VLMs); a system that includes one or more large language models (LLMs); a system for performing conversational AI operations; a system for performing light transport simulations; a system for performing deep learning processes; a system for performing digital twin operations; a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system that embodies one or more virtual machines (VMs); a system that is implemented using a robot; a system implemented using an edge device; a system that is at least partially implemented in a data center; a system that is implemented at least partially using cloud computing resources. a system for generating interactive 3D visualizations; or a system that is implemented at least partially using Augmented Reality (AR) or Virtual Reality (VR) platforms. Clause 9. One or more processors comprising the following: one or more circuits for the following: Receive at least one object segmented from the video data; Compacting the at least one object by the following: Scanning a plurality of points on or approximately around the at least one object; Generation of a voxelized volume of the at least one object, which is based at least partially on the plurality of points; Updating the voxelized volume at least partially based on the occupancy state of at least one voxel of a plurality of voxels of the voxelized volume at least partially based on at least one rendered depth map; Simulation of one or more interactions of the voxelized volume of the at least one compressed object to update at least one physical attribute of the at least one object; and Displaying at least one item. Clause 10. One or more processors according to Clause 9, wherein the one or more circuits are used for the following: Filling an interior space of the voxelized volume, which is based at least partially on the injection of a plurality of volumetric elements into the interior space of the at least one object comprising a plurality of interior regions. Clause 11. One or more processors according to Clause 9 or 10, wherein the at least one compressed object corresponds to a volumetric representation. Clause 12. One or more processors according to any one of Clauses 9 to 11, wherein for the simulation of the one or more interactions of the one or more processors, a stiffness simulation shall be performed by: Applying a first plurality of transformations to the at least one compressed object using a first physics model to obtain a plurality of stiff motions of the at least one compressed object. Clause 13. One or more processors according to Clause 12, wherein the plurality of stiff movements of the at least one compressed object are obtained by the following: Determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters; Minimizing the energy function to determine a plurality of stiff states of the at least one compressed object; and Application of the plurality of stiff states to simulate the plurality of stiff movements of the at least one compressed object over time. Clause 14. One or more processors according to any one of Clauses 9 to 13, wherein for the simulation of the one or more interactions of the one or more processors, an elasticity simulation shall be performed by: Applying a second plurality of transformations to the at least one condensed object using a second physics model to obtain a plurality of deformed states of the at least one condensed object. Clause 15. One or more processors according to Clause 14, wherein the plurality of deformed states of the at least one compacted object are obtained by the following: Determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters; Minimizing the energy function to determine a plurality of updates for a plurality of control points; and Calculation of one or more deformations of the at least one densified object, which are based at least partially on the plurality of updates of the plurality of control points and a plurality of corresponding skinning fields. Clause 16. Procedure, which includes the following: Reception by one or more processors of segmentation data corresponding to at least one object segmented from the video data; Compression by the one or more processors of the at least one object by the following: Scanning a plurality of points on or approximately around the at least one object; Generation of a voxelized volume of the at least one object, which is based at least partially on the plurality of points; Updating the voxelized volume at least partially based on the occupancy state of at least one voxel of a plurality of voxels of the voxelized volume at least partially based on at least one rendered depth map; Simulation by one or more processors of one or more interactions of the voxelized volume of the at least one compressed object to update at least one physical attribute of the at least one object; and Display by one or more processors of the at least one object. Clause 17. Procedure according to Clause 16, which further includes the following: Filling by one or more processors of an interior space of the voxelized volume, which is based at least partially on the injection of a plurality of volumetric elements into the interior space of the at least one object comprising a plurality of interior regions. Clause 18. Procedure according to Clause 16 or 17, wherein the at least one condensed object corresponds to a volumetric representation. Clause 19. Procedure according to any one of Clauses 16 to 18, wherein the simulation of one or more interactions includes performing a stiffness simulation which includes the following: Application by one or more processors of a first plurality of transformations to the at least one compressed object using a first physics model to obtain a plurality of stiff movements of the at least one compressed object. Clause 20. Procedure according to any one of Clauses 16 to 19, wherein the simulation of one or more interactions includes performing an elasticity simulation comprising the following: Application by one or more processors of a second plurality of transformations to the at least one condensed object using a second physics model to obtain a plurality of deformed states of the at least one condensed object.

[0209] It is understood that the aspects and embodiments described above are only exemplary and that changes to details may be made within the scope of protection of the claims.

[0210] Each device, method and feature disclosed in the description and (where applicable) in the claims and drawings may be provided independently or in any suitable combination.

[0211] The reference numerals appearing in the claims are for illustrative purposes only and are not intended to have any limiting effect on the scope of protection of the claims.

Claims

[1] System comprising the following: one or more processors to perform one or more operations for the following: Receive at least one object segmented from the video data; Compacting the at least one object by the following: Scanning a plurality of points on or approximately around the at least one object; Generation of a voxelized volume of the at least one object, which is based at least partially on the plurality of points; Updating the voxelized volume at least partially based on the occupancy state of at least one voxel of a plurality of voxels of the voxelized volume at least partially based on at least one rendered depth map; Simulation of one or more interactions of the voxelized volume of the at least one compressed object to update at least one physical attribute of the at least one object; and Generate at least one image that depicts at least one section of the at least one object with at least one updated physical attribute. [2] System according to claim 1, wherein one or more operations comprise at least one operation for the following: Filling an interior space of the voxelized volume, which is based at least partially on the injection of a plurality of volumetric elements into the interior space of the at least one object comprising a plurality of interior regions. [3] System according to claim 1 or 2, wherein the at least one compressed object corresponds to a volumetric representation. [4] System according to any of the preceding claims, wherein the one or more processes for simulating the one or more interactions comprise at least one process for performing a stiffness simulation by: Applying a first plurality of transformations to the at least one compressed object using a first physics model to obtain a plurality of stiff motions of the at least one compressed object. [5] System according to claim 4, wherein the plurality of stiff movements of the at least one compressed object are obtained by the following: Determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters; Minimizing the energy function to determine a plurality of stiff states of the at least one compressed object; and Application of the plurality of stiff states to simulate the plurality of stiff movements of the at least one compressed object over time. [6] System according to any of the preceding claims, wherein the one or more processes for simulating the one or more interactions comprise at least one process for performing an elasticity simulation comprising: Applying a second plurality of transformations to the at least one condensed object using a second physics model to obtain a plurality of deformed states of the at least one condensed object. [7] System according to claim 6, wherein the plurality of deformed states of the at least one compacted object are obtained by the following: Determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters; Minimizing the energy function to determine a plurality of updates for a plurality of control points; and Calculation of one or more deformations of the at least one densified object, which are based at least partially on the plurality of updates of the plurality of control points and a plurality of corresponding skinning fields. [8] System according to any of the preceding claims, wherein the one or more processors are contained in at least one of the following: a system for conducting games; a system for conducting content streaming; a system for conducting collaborative content creation; a system for carrying out simulation processes; a system for conducting collaborative content creation for 3D assets; a system for generating synthetic data; a system that includes one or more large Visual Language Models (VLMs); a system that includes one or more large language models (LLMs); a system for performing conversational AI operations; a system for performing light transport simulations; a system for performing deep learning processes; a system for performing digital twin operations; a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system that embodies one or more virtual machines (VMs); a system that is implemented using a robot; a system implemented using an edge device; a system that is at least partially implemented in a data center; a system that is implemented at least partially using cloud computing resources. a system for generating interactive 3D visualizations; or a system that is implemented at least partially using Augmented Reality (AR) or Virtual Reality (VR) platforms. [9] One or more processors comprising the following: one or more circuits for the following: Receive at least one object segmented from the video data; Compacting the at least one object by the following: Scanning a plurality of points on or approximately around the at least one object; Generation of a voxelized volume of the at least one object, which is based at least partially on the plurality of points; Updating the voxelized volume at least partially based on the occupancy state of at least one voxel of a plurality of voxels of the voxelized volume at least partially based on at least one rendered depth map; Simulation of one or more interactions of the voxelized volume of the at least one compressed object to update at least one physical attribute of the at least one object; and Displaying at least one item. [10] One or more processors according to claim 9, wherein the one or more circuits serve to: Filling an interior space of the voxelized volume, which is based at least partially on the injection of a plurality of volumetric elements into the interior space of the at least one object comprising a plurality of interior regions. [11] One or more processors according to claim 9 or 10, wherein the at least one compressed object corresponds to a volumetric representation. [12] One or more processors according to any one of claims 9 to 11, wherein for the simulation of the one or more interactions of the one or more processors a stiffness simulation shall be performed by the following: Applying a first plurality of transformations to the at least one compressed object using a first physics model to obtain a plurality of stiff motions of the at least one compressed object. [13] One or more processors according to claim 12, wherein the plurality of stiff movements of the at least one compressed object are obtained by the following: Determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters; Minimizing the energy function to determine a plurality of stiff states of the at least one compressed object; and Application of the plurality of stiff states to simulate the plurality of stiff movements of the at least one compressed object over time. [14] One or more processors according to any one of claims 9 to 13, wherein for the simulation of the one or more interactions of the one or more processors, an elasticity simulation is to be performed by: Applying a second plurality of transformations to the at least one condensed object using a second physics model to obtain a plurality of deformed states of the at least one condensed object. [15] One or more processors according to claim 14, wherein the plurality of deformed states of the at least one compressed object are obtained by the following: Determining an energy function using at least one of a plurality of scene parameters or a plurality of object parameters; Minimizing the energy function to determine a plurality of updates for a plurality of control points; and Calculation of one or more deformations of the at least one densified object, which are based at least partially on the plurality of updates of the plurality of control points and a plurality of corresponding skinning fields. [16] Method comprising the following: Reception by one or more processors of segmentation data corresponding to at least one object segmented from the video data; Compression by the one or more processors of the at least one object by the following: Scanning a plurality of points on or approximately around the at least one object; Generation of a voxelized volume of the at least one object, which is based at least partially on the plurality of points; Updating the voxelized volume at least partially based on the occupancy state of at least one voxel of a plurality of voxels of the voxelized volume at least partially based on at least one rendered depth map; Simulation by one or more processors of one or more interactions of the voxelized volume of the at least one compressed object to update at least one physical attribute of the at least one object; and Display by one or more processors of the at least one object. [17] The method of claim 16, further comprising: Filling by one or more processors of an interior space of the voxelized volume, which is based at least partially on the injection of a plurality of volumetric elements into the interior space of the at least one object comprising a plurality of interior regions. [18] Method according to claim 16 or 17, wherein the at least one compressed object corresponds to a volumetric representation. [19] Method according to any one of claims 16 to 18, wherein the simulation of one or more interactions comprises performing a stiffness simulation which includes: Application by one or more processors of a first plurality of transformations to the at least one compressed object using a first physics model to obtain a plurality of stiff movements of the at least one compressed object. [20] Method according to any one of claims 16 to 19, wherein the simulation of one or more interactions comprises performing an elasticity simulation which includes: Application by one or more processors of a second plurality of transformations to the at least one condensed object using a second physics model to obtain a plurality of deformed states of the at least one condensed object.