Automated asset generation for robotic assembly tasks
By generating assembly asset pairs through a three-stage pipeline, the problems of diverse part shapes and limited datasets in robot assembly are solved, achieving efficient and generalizable assembly strategy generation and improved fault tolerance.
Patent Information
- Application Number
- CN202511331087.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-07-21
- Filing Date
- 2025-09-17
- Publication Date
- 2026-03-20
AI Technical Summary
Existing technologies struggle to train robots to reliably assemble parts with a wide variety of geometries, especially both familiar and unfamiliar ones. Furthermore, the limited methods for generating existing assembly asset datasets result in insufficient generalization capabilities of assembly strategies.
A three-stage pipeline is used to generate paired parts, including contact surface extraction, shape completion, and gap specification stages. Assembly asset pairs are automatically generated using a visual language model and a 3D generative model, and the assembly strategy is trained and tested in a simulation environment.
It enables the rapid and efficient generation of diverse assembly asset pairs, improves the generalization ability and fault tolerance of robot assembly tasks, and enhances assembly performance in real-world environments.
Smart Images

Figure CN121706931A_ABST
Abstract
Description
[0001] Cross Reference to Related Applications
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 696,112, filed on September 18, 2024, the entire contents of which are incorporated herein by reference. Background Technology
[0003] Neural motion control refers to the use of neural networks (or another type of machine learning model) to move and / or animate physical and / or virtual entities. For example, a deep learning model can be trained to perform motion control by generating sequences of poses (i.e., positions and orientations) corresponding to movements in humans, animals, robots, and / or other types of jointed objects. The output poses can then be incorporated into robots or other autonomous machines, games, animations, simulations, and / or other applications involving jointed objects.
[0004] Neuromotor control can also be extended to robotic assembly, where the movement of a robot (or another type of articulated object) is used to physically assemble discrete parts (also referred to herein as assets) into a functional product. For example, the robot might use a machine learning model trained with neuromotor control techniques to perform tasks such as (but not limited to) connecting electronic devices to charging cables, assembling furniture, installing nuts and bolts, inserting bearings and / or fastening components.
[0005] While assembly tasks typically involve specific objectives (e.g., moving two or more parts into a fixed relative pose), teaching autonomous robots to perform robotic assembly robustly presents numerous challenges. More specifically, robots must perceive, grasp, and insert parts precisely and accurately under varying environmental conditions and uncertainty regarding the position and orientation of parts. To address these challenges, recent approaches have been developed including: rapid simulations for contact-rich scenarios; techniques for transferring contact-rich assembly strategies trained in simulations to the real world; and assembly strategy learning techniques capable of handling individual part pairs.
[0006] However, existing approaches are unable to train a general strategy that can reliably assemble a range of previously seen and / or unseen parts having a variety of geometrical shapes. Further, while an increase in the quantity and diversity of training data can improve the ability of a robotic assembly strategy to generalize across scenarios, tasks, and learning techniques, existing datasets of assembly pairs (e.g., different plug and socket combinations) are typically generated and / or curated via manual processes, which limits the number and diversity of assembly problems that can be used for strategy learning. Additionally, datasets of assembly parts generated via traditional techniques can include pairs of assets that penetrate and / or have insufficient clearance from one another, which interferes with the use of the pairs of assets in high-accuracy simulation environments and / or real-world environments having non-penetration constraints (e.g., because the pairs of assets cannot be physically assembled after manufacture).
[0007] As explained previously, there is a need in the art for more effective techniques for training robots to perform assembly tasks. BRIEF DESCRIPTION OF DRAWINGS
[0008] The present systems and methods for automated assembly asset generation for robotic technology systems and applications are described below in more detail with reference to the following figures, in which:
[0009] Figure 1 illustrates a block diagram of a computing system configured to implement one or more aspects of at least one embodiment;
[0010] Figure 2 a more detailed illustration of the data generation engine, training engine, and execution engine of Figure 1
[0011] Figure 3A illustrates how the data generation engine of Figure 1 determines attributes associated with a part, in accordance with at least one embodiment;
[0012] Figure 3B illustrates how the data generation engine of Figure 1 determines contact surfaces associated with a part, in accordance with at least one embodiment;
[0013] Figure 4A illustrates an example set of parts and a corresponding set of paired parts generated by the data generation engine of Figure 1
[0014] Figure 4B illustrates how the data generation engine of Figure 1 performs clearance specification for a part and a corresponding paired part, in accordance with at least one embodiment;
[0015] Figure 4C illustrates how the data generation engine 122 performs gap specifications for parts and corresponding mating parts, in accordance with at least one embodiment; Figure 1
[0016] Figure 5 illustrates a flowchart of a method for generating robotic assembly data, in accordance with at least one embodiment;
[0017] Figure 6A is an illustration of example sensor locations with respective fields of view or sensing fields of, for example, autonomous or semi-autonomous machines, in accordance with at least some embodiments of the present disclosure;
[0018] Figure 6B is an illustration of example component and sensor locations on an autonomous or semi-autonomous vehicle, in accordance with at least some embodiments of the present disclosure;
[0019] Figure 6C is a block diagram of an example system architecture of an autonomous or semi-autonomous vehicle, robot, and / or other machine type, in accordance with at least some embodiments of the present disclosure;
[0020] Figure 6D is a block diagram of an example architecture of a computing system (e.g., a system on a chip (SoC)), in accordance with at least some embodiments of the present disclosure;
[0021] Figure 6E is a system diagram of communications between a cloud-based server and an example autonomous or semi-autonomous vehicle, robot, and / or other machine type, in accordance with at least some embodiments of the present disclosure;
[0022] Figure 7 is a system diagram showing three computer ecosystems for generating or creating artificial intelligence (AI) (e.g., AI training and validation data), training artificial intelligence, and deploying AI at the edge, in accordance with at least some embodiments of the present disclosure;
[0023] Figure 8 is a block diagram of an example computing system for generating generative artificial intelligence (AI), in accordance with at least some embodiments of the present disclosure; and
[0024] Figure 9 is a block diagram of an example computing device, in accordance with at least some embodiments of the present disclosure. DETAILED DESCRIPTION
[0025] Systems and methods related to automated assembly asset generation for robotic technology systems and applications are disclosed. While the present disclosure can be with respect to example autonomous or semi-autonomous vehicles, robots, and / or other machine types 600 (alternatively referred to herein as examples thereof Figures 6A-6E “vehicles 600,” “ego vehicles 600,” “machines 600,” “ego machines,” “robots,” and / or “ego robots 600” are described, but this is not intended to be limiting. For example, the systems and methods described herein can be implemented by, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms (e.g., autonomous mobile robots (AMRs), humanoid robots, robotic arms, and / or end effectors), warehouse vehicles, off-road vehicles, vehicles coupled with one or more trailers, aircraft, watercraft, shuttles (e.g., autonomous taxis), emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, engineering vehicles, underwater vehicles (e.g., manned or unmanned submarines), drones, and / or other vehicle, robot, or vehicle types. Additionally, although the present disclosure can be described with respect to the generation of robotic assembly assets, this is not intended to be limiting, and the systems and methods described herein can be used for augmented reality (AR), virtual reality (VR), mixed reality (MR), robotics, safety and surveillance (e.g., smart cities), autonomous or semi-autonomous machine applications, industrial manufacturing, simulation, and / or any other technical space in which assembly assets can be used. In some embodiments, the systems, methods, and processes described herein can be implemented using components, features, and / or functionality similar to those of the example machine 600 of Figures 6A-6E the example computing ecosystem 700 of Figure 7 the example generative language model system 800 of Figure 8 and / or the example computing device 900 of Figure 9
[0026] As discussed herein, it can be difficult to train a general-purpose strategy that can reliably assemble a range of seen and / or unseen parts having a wide variety of geometries. Further, while an increase in the quantity and diversity of training data can improve the ability of a robotic assembly strategy to generalize across scenarios, tasks, part types, and / or learning techniques, existing training datasets for robotic assembly tasks are limited in the number and / or types of assembly asset pairs that can be used to simulate and / or real-world environments.
[0027] To address the aforementioned limitations, the disclosed technology includes a three-stage pipeline for generating mating parts in an automated assembly. The pipeline includes a first contact surface extraction stage in which a set of contact surfaces are extracted from a first part based on a visual representation and / or another representation of the first part by a visual language model (VLM) and / or another type of machine learning model. The pipeline also includes a second shape completion stage in which the contact surfaces are used to regulate the operation of a diffusion model and / or another type of three-dimensional (3D) generative model in generating a shape of a second part that is complementary to the first part. The pipeline further includes a third clearance specification stage in which the shape of a given part is updated to satisfy minimum clearance distances with other parts. The mating parts generated via the pipeline can then be used to train policies with respect to assembly tasks in simulated environments, evaluate the performance of deployed policies in real-world environments, for real-world assembly tasks, and / or for other tasks related to assembly.
[0028] One advantage of the disclosed technology over existing approaches is the ability to automatically generate a large and diverse set of mating parts that can be used for assembly tasks. Thus, the disclosed technology can be used to generate pairs of assembly assets more quickly and efficiently than conventional approaches that involve hand-crafting and / or curating datasets for robotic assembly. The generated pairs of assets can also be used to train, test, and / or evaluate robots on assembly tasks in a more comprehensive manner than robots trained using much smaller and / or less diverse sets of assembly components. Further, robots trained using the generated pairs of assets can be more fault-tolerant and / or able to generalize to different scenarios than robots trained using more limited sets of paired components. Additionally, the disclosed technology can be used to “repair” interpenetrating assets generated via other techniques, thereby further increasing the number and types of assets that can be used for assembly tasks.
[0029] In some embodiments, the systems and methods described herein can be performed using simulated data (e.g., simulated environment data and simulated sensor data of simulated sensors of virtual or simulated vehicles, robots, or machines within a simulated environment) within a simulated environment (e.g., NVIDIA’s Drive SIM, ISAAC Sim, ISAAC Gym, ISAAC Lab, etc.). For example, simulated input data (e.g., map data, perception data, ego-motion data, haptics data, and / or any other data described herein) can be used to determine assembly assets and / or environments associated with a robotic assembly task, and this information can be used to perform operations associated with a virtual machine within a simulated environment. These simulated operations can be used to test the performance of underlying algorithms, systems, and / or processes before they are deployed in the real world. In some instances, simulation can be used to generate synthetic training data, e.g., from pairs of assembly assets within a simulation. The synthetic training data can then be used or processed (in addition to or as an alternative to real-world data) to train and / or deploy a policy for performing a robotic assembly task.
[0030] In any example (such as where the simulation environment is used for testing, validation, training, etc.), the simulation environment and / or associated training data may be rendered or otherwise generated using one or more optical transport simulation algorithms (such as one or more ray tracing and / or path tracing algorithms). When using optical transport simulation, the simulation system may employ one or more dedicated ray tracing hardware accelerators and / or processors (e.g., NVIDIA's RTX or another real-time ray tracing GPU, such as a GPU including one or more ray tracing (RT) cores) optimized for performing real-time or near-real-time optical transport simulation operations in conjunction with one or more other processors of the system (e.g., GPUs, CPUs, accelerators, etc.). In some embodiments, the simulation environment and / or one or more of its objects, features, or components may be generated or managed in a 3D content collaboration platform (e.g., NVIDIA's OMNIVERSE) that can be optimized or suitable for industrial digitization, generative physics, artificial intelligence, and / or other use cases, applications, and / or services. For example, a content collaboration platform or system may include systems for managing objects, features, scenes, etc., within simulated environments, digital environments, etc., using or developing generic scene descriptors (USD) (e.g., OpenUSD). The platform may include realistic physical simulations (e.g., using NVIDIA's PhysX software development kit (SDK)) to simulate real physical phenomena and physical interactions with simulations hosted by the platform. The platform may integrate OpenUSD with ray tracing / path tracing / light transport simulations (e.g., NVIDIA's RTX rendering technology) into software tools and simulation workflows for building, training, deploying, and / or testing AI systems, such as systems for testing, validating, training (e.g., machine learning models, neural networks, etc.) and / or other tasks related to automobiles, robots, other machine types, and / or other systems and applications. In some examples, the simulated environment may include digital twins of real-world environments, such as specific road segments, warehouses, data centers, airports, geographic areas, ocean areas, and / or any other real-world environments that autonomous or semi-autonomous vehicles or machines might operate in.
[0031] In some embodiments, a remote control or teleoperation system may be used to perform teleoperation or remote control of vehicles, robots, and / or other machines. For example, the systems and methods described herein may be used to generate and / or place pairs of component assets included in a visualization or mapping of an environment to assist a remote operator in controlling an autonomous or semi-autonomous machine through the environment (or to provide waypoints or other indications for control or navigation). Thus, a remote operator may use visual, auditory, textual, and / or other cues or indicators generated by the systems and methods described herein to assist in navigating vehicles, robots, machines, etc., through a real-world environment using a teleoperation system.
[0032] In some embodiments, the systems and methods described herein can be deployed in robotic applications. For example, a robot or robotic system can include one or more on-board processors (e.g., CPUs, GPUs, hardware-based deep learning accelerators (DLAs), clusters of deep learning accelerators (XNNs), neural processing units (NPUs), neural network accelerators (NNAs), hardware-based programmable vision accelerators (PVAs) that can include one or more vector processing units (VPUs), direct memory access (DMA) systems, and / or pixel processing engines (PPEs), hardware-based optical flow accelerators (OFAs), SoCs, etc.) as well as memory and / or storage (e.g., for storing control algorithms, sensor data, and one or more machine learning models). The robotic system can use these processors to implement one or more machine learning models (e.g., language models, visual language models (VLMs), large language models (LLMs), visual language action (VLA) models, multi-modal language models (MMLMs), etc.) that allow the robotic system to autonomously or semi-autonomously perform complex tasks, such as interacting with static and / or dynamic objects for assembly and / or manipulation, or navigating an environment using sensors such as cameras, LiDARs, RADARs, ultrasonic sensors, etc. The system can use sensor fusion techniques to combine data from multiple sensors (e.g., cameras, infrared, LiDARs, RADARs, accelerometers) together to create a comprehensive model of the environment around the robot. This data can be processed locally on the robot, or sent to a remote server for more computationally intensive tasks such as 3D mapping or SLAM (simultaneous localization and mapping). In one or more embodiments, data from various robots (e.g., sensor data, task status, or environmental conditions) can be uploaded to the cloud, where a centralized AI model can analyze optimized commands and distribute them to the entire fleet. In some embodiments, one or more machine learning models described herein (e.g., language models, VLMs, VLAs, LLMs, MMLMs, diffusion models, NeRF models, DNNs, etc.) can be used to allow a robot to perceive and reason about an environment and / or to communicate with one or more other robots and / or humans in the environment. In some embodiments, a robot can communicate with one or more locally-hosted servers / computing devices and / or with one or more remotely-located servers / computing devices (e.g., in one or more data centers), for example, using one or more network interface cards (NICs) and / or data processing units (DPUs).
[0033] In some embodiments, the systems and methods described herein can be deployed in an in-vehicle infotainment (IVI) system or in-cabin experience (IX) application. For example, an infotainment system within a vehicle (e.g., an automobile, a truck, a drone, an engineering device, a robot, a semi-autonomous vehicle, or an autonomous vehicle) can include one or more on-board processors (e.g., CPUs, GPUs, hardware-based deep learning accelerators (DLAs), clusters of deep learning accelerators (XNNs), neural processing units (NPUs), neural network accelerators (NNAs), hardware-based programmable vision accelerators (PVAs) that can include one or more vector processing units (VPUs), direct memory access (DMA) systems, and / or pixel processing engines (PPEs), hardware-based optical flow accelerators (OFAs), SoCs, etc.), as well as memory and / or storage (e.g., for storing control algorithms, sensor data, and one or more machine learning models), and memory and / or storage (e.g., for storing entertainment content, navigation data, and user preferences). The system can use these processors to execute one or more machine learning models (e.g., language models) to enable features such as voice control, personalized media recommendations, dynamic navigation, and real-time communication with other services over a network connection. The in-vehicle infotainment system can also use natural language processing (NLP) models to enable voice-based interactions. The one or more machine learning models can be stored locally or accessed through one or more APIs connected to a cloud service, enabling the system to process requests in real-time or near real-time.
[0034] In some examples, one or more machine learning models described herein (e.g., deep neural networks, language models, LLMs, VLMs, multi-modal language models, visual-language-action (VLA) models, perception models, tracking models, fusion models, transformer models, diffusion models, encoder-only models, decoder-only models, encoder-decoder models, neural rendering field (NERF) models, etc.) can be packaged as microservices, such as inference microservices (e.g., NVIDIA NIM), which can include containers (e.g., operating system (OS) level virtualization packages), which can include an application programming interface (API) layer, a server layer, a runtime layer, and / or a model “engine.” For example, an inference microservice can include the container itself and one or more models (e.g., weights and biases). In some instances, such as where one or more machine learning models are small enough (e.g., have a small enough number of parameters), the one or more models can be included within the container itself. In other examples, such as where one or more models are large, the one or more models can be hosted / stored in the cloud (e.g., in a data center) and / or can be hosted on-premises and / or at the edge (e.g., on a local server or computing device, but outside of the container). In these embodiments, the one or more models can be accessible via one or more APIs, such as REST APIs. As such and in some embodiments, one or more machine learning models described herein can be deployed as inference microservices to accelerate deployment of the one or more models on any cloud, data center, or edge computing system, while ensuring security of data. For example, an inference microservice can include one or more APIs, pre-configured containers for ease of deployment, an optimized inference engine (e.g., built with standardized AI model deployment to implement software, such as NVIDIA’s Triton Inference Server, and / or one or more APIs for high-performance deep learning inference, which can include an inference runtime and model optimization that provides low latency and high throughput for production applications, such as NVIDIA’s TensorRT), and / or enterprise management data for telemetry (e.g., including identity, metrics, health checks, and / or monitoring). One or more machine learning models described herein can be included as part of a microservice with an acceleration infrastructure that has the ability to be deployed with a single command and / or orchestrated and auto-scaled on the acceleration infrastructure (e.g., up to data center scale on a single device) using a container orchestration system.As such, an inference microservice can include one or more machine learning models (e.g., one or more models that have been optimized for high performance inference), inference runtime software that executes the one or more machine learning models and provides an output / response to an input (e.g., a user query, a prompt, etc.), and enterprise management software that provides health checks, authentication, and / or other monitoring. In some embodiments, an inference microservice can include software that performs on-the-fly replacement and / or updates to the one or more machine learning models. When replaced or updated, the software that performs the replacement / update can maintain user configurations of the inference runtime software and the enterprise management software.
[0035] Although examples can be described herein with respect to using machine learning models, such as neural networks, this is not intended to be limiting. For example, and without limitation, any of the various machine learning models and / or neural networks described herein can include any type of machine learning model, such as using linear regression, logistic regression, decision trees, support vector machines (SVM), Naive Bayes, k-nearest neighbors (Knn) (K meaning clustering), random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoder neural networks, artificial neural networks (ANN), convolutional neural networks (CNN), recurrent neural networks (RNN), perceptron, long / short-term memory (LSTM) networks, multilayer perceptron (MLP) networks, deep stacking networks (DSN), generative pre-trained (GPT) models or networks, feedforward networks, radial basis function ANNs, self-organizing maps (SOM), Kohonen maps, Hopfield networks, Boltzmann machines, deep belief neural networks, deconvolutional neural networks, generative adversarial networks (GAN), liquid state machines, modular neural networks, liquid state machines, sequence-to-sequence models, networks using transformer architectures, state space models (SSM) (e.g., networks using Mamba architectures (e.g., Mamba-1, Mamba 2, etc.), networks using selective state space models, networks using structured state space sequence models, etc.), diffusion models (e.g., diffusion probability models, score-based generative models, etc.), neural radiance fields (NeRF) models, Gaussian Splatting models, Kolmogorov-Arnold networks (KAN), models with encoder-only architectures, models with decoder-only architectures, models with encoder-decoder architectures, generative machine learning models, language models, large language models (LLM), visual language models (VLM), multi-modal language models (MMLM), large action models (LAM), visual language action (VLA) models, etc.), and / or other types of machine learning models.
[0036] The systems and methods described herein can be used with, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled with one or more trailers, aircraft, watercraft, shuttles (e.g., autonomous taxis), emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, engineering vehicles, underwater vehicles (e.g., manned or unmanned submarines), drones, and / or other vehicle types. Further, the systems and methods described herein can be used for various purposes, by way of example but not limitation, for machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, safety and supervision, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environment simulation, object or actor simulation and / or digital twinning, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets (e.g., NVIDIA’s Omniverse), cloud computing, and / or any other suitable application.
[0037] The disclosed embodiments can be included in various different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines, etc.), systems implemented using robots, aviation systems, medical systems, marine systems, smart district monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twinning operations, systems implemented using edge devices, systems implementing language models (such as large language models (LLMs), visual language models (VLMs), visual language action (VLA) models, and / or multi-modal language models), systems using or deploying one or more inference microservices, systems including one or more machine learning models deployed in services or microservices and OS-level virtualization packages (e.g., containers), systems including one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in data centers, systems for performing conversational AI operations, systems for performing light transport simulation, systems for performing collaborative content creation for 3D assets, systems for performing generative AI operations, systems implemented at least partially using cloud computing resources, and / or other types of systems.
[0038] System Overview
[0039] Figure 1A block diagram illustrating a computing system 100 configured to implement one or more aspects of at least one embodiment is shown. In at least one embodiment, computing system 100 can include any type of computing device, including, but not limited to, a server machine, a server platform, a desktop machine, a laptop machine, a handheld / mobile device, a digital kiosk, an in-vehicle infotainment system, a smart speaker or display, a television, and / or a wearable device. In at least one embodiment, computing system 100 is a server machine operating in a data center or cloud computing environment that provides scalable computing resources as a service over a network.
[0040] In various embodiments, computing system 100 includes, without limitation, one or more processors 102 and one or more memories 104 coupled to parallel processing subsystem 112 via memory bridge 105 and communication path 113. Memory bridge 105 is also coupled to I / O (input / output) bridge 107 via communication path 106, which in turn is coupled to translator 116.
[0041] In one embodiment, I / O bridge 107 is configured to receive user input information from optional input device 108, such as, but not limited to, a keyboard, a mouse, a touchscreen, sensor data analysis (e.g., evaluating gestures, speech, or other information about one or more uses of a field of view or sensing field of one or more sensors), a VR / MR / AR headset, a gesture recognition system, a steering wheel, a mechanical, digital, or touch sensitive button or input component, and / or a microphone, and to forward the input information to one or more processors 102 for processing. In at least one embodiment, computing system 100 can be a server machine in a cloud computing environment. In such an embodiment, computing system 100 can omit input device 108 and receive input information as commands (e.g., in response to one or more inputs from a remote computing device) and / or messages transmitted over a network and received via network adapter 118. In at least one embodiment, translator 116 is configured to provide connectivity between I / O bridge 107 and other components of computing system 100, such as network adapter 118 and various additional cards 120 and 121.
[0042] In at least one embodiment, I / O bridge 107 is coupled to system disk 114, which can be configured to store content and applications and data for use by one or more processors 102 and parallel processing subsystem 112. In one embodiment, system disk 114 provides non-volatile storage for applications and data and can include fixed or removable hard disks, flash memory devices, and CD-ROM (compact-disc read-only memory), DVD-ROM (digital versatile-disc ROM), Blu-ray, HD-DVD (high definition DVD) or other magnetic, optical, or solid state storage devices. In various embodiments, other components such as a universal serial bus or other port connections, compact flash storage devices, digital versatile disc drives, movie recording devices, etc. can also be connected to I / O bridge 107.
[0043] In various embodiments, memory bridge 105 can be a northbridge chip, while I / O bridge 107 can be a southbridge chip. Additionally, communication paths 106 and 113, as well as other communication paths within computing system 100, can be implemented using any technically suitable protocol including, but not limited to, AGP (accelerated graphics port), HyperTransport, or any other bus or point-to-point communication protocol known in the art.
[0044] In at least one embodiment, parallel processing subsystem 112 includes a graphics subsystem that will deliver pixels to optional display device 110, which can be any conventional cathode ray tube, liquid crystal display, light emitting diode display, or the like. In such embodiment, parallel processing subsystem 112 can contain circuitry optimized for graphics and video processing, including, for example, video output circuitry. Such circuitry can be incorporated into a
[0045] In at least one embodiment, parallel processing subsystem 112 includes circuitry (e.g., optimized circuitry) optimized for general and / or compute processing. Further, such circuitry can be incorporated within one or more PPUs included within parallel processing subsystem 112 that are configured to perform such general and / or compute operations. In other embodiments, one or more PPUs included within parallel processing subsystem 112 can be configured to perform graphics processing, general purpose processing, and / or compute processing operations. One or more memories 104 include at least one device driver configured to manage processing operations of one or more PPUs within parallel processing subsystem 112. Additionally, one or more memories 104 include instructions implementing data generation engine 122, training engine 124, and execution engine 126 that can be executed by one or more processors and / or parallel processing subsystem 112.
[0046] In various embodiments, parallel processing subsystem 112 can be integrated with one or more of other elements of system 100 to form a single system. For example, parallel processing subsystem 112 can be integrated with one or more processors 102 and other connectivity circuitry on a single chip to form a system on a chip (SoC). Figure 1
[0047] One or more processors 102 can include any suitable processor, implemented as a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), an artificial intelligence (AI) accelerator, a deep learning accelerator (DLA), a parallel processing unit (PPU), a data processing unit (DPU), a vector or visual processing unit (VPU), a programmable visual accelerator (PVA) that can include one or more VPUs, a pixel processing engine (PPE), and / or a direct memory access (DMA) system, any other suitable type of processing unit, or a combination of different processing units such as one or more CPUs configured to operate with one or more GPUs. In general, one or more processors 102 can include any technically feasible hardware unit capable of processing data and / or executing software applications. Further, in the context of the present disclosure, computing elements shown in computing system 100 can correspond to physical computing systems (e.g., systems in a data center or machine) and / or can correspond to virtual computing instances executing in a computing cloud.
[0048] In at least one embodiment, one or more processors 102 issue commands that control operation of PPUs. In at least one embodiment, communication path 113 is a Peripheral Component Interconnect Express (PCIe) link in which dedicated lanes are allocated to each PPU. Other communication paths can also be used. PPUs advantageously implement a highly parallel processing architecture, and any number of local parallel processing memories (PP memories) can be provided to a PPU.
[0049] It will be appreciated that systems shown herein are illustrative and that variations and modifications are possible. The connection topology, including the number and arrangement of bridges, the number of processors 102, and the number of parallel processing subsystems 112, can be modified as desired. For example, in at least one embodiment, one or more memories 104 can be connected directly to one or more processors 102 rather than through memory bridge 105, and other devices can communicate with the one or more memories 104 via memory bridge 105 and processor 102. In other embodiments, parallel processing subsystems 112 can be connected to I / O bridge 107 or directly to one or more processors 102 rather than to memory bridge 105. In still other embodiments, I / O bridge 107 and memory bridge 105 can be integrated into a single chip rather than existing as one or more discrete devices. In certain embodiments, Figure 1 One or more components shown in FIG. 1 can not be present. For example, transducers 116 can be eliminated, and network adapters 118 and additional cards 120, 121 can be connected directly to I / O bridge 107. Further, in certain embodiments, Figure 1 One or more components shown in FIG. 1 can be implemented as virtualized resources in a virtual computing environment, such as a cloud computing environment. Specifically, in at least one embodiment, parallel processing subsystem 112 can be implemented as a virtualized parallel processing subsystem. For example, parallel processing subsystem 112 can be implemented as one or more virtual graphics processing units (vGPUs) that render graphics on one or more virtual machines (VMs) that execute on one or more server machines whose one or more GPUs and other physical resources are shared across the one or more VMs.
[0050] Automated assembly asset generation
[0051] Figure 2 is in accordance with at least one embodiment Figure 1A more detailed illustration of the data generation engine 122, training engine 124, and execution engine 126 is provided. As discussed herein, the data generation engine 122, training engine 124, and execution engine 126 are configured to perform automated data generation, training, and / or inference for robot assembly tasks. Each of these components is described in further detail below.
[0052] In one or more embodiments, robotic assembly refers to the physical combination of discrete parts (also referred to herein as assets or components) into a functional product using a robot (or another type of articulated object). For example, robotic assembly may involve using a dual-arm robot, a human robot, and / or another type of robot to perform tasks such as (but not limited to) connecting electronic devices to charging cables, assembling furniture, installing nuts and bolts, inserting bearings, and / or fastening components.
[0053] Additionally, robot assembly may be performed within a neuromotor control framework, which includes a strategy 220 of one or more neural networks (or other types of machine learning models) used to generate different actions 226 for each frame or time step in the robot's motion. For example, each action 226 generated by strategy 220 may be used to update the configuration of the robot's joints and / or other parts at the corresponding time step, thereby generating the corresponding motion in the robot. The sequence of actions generated by strategy 220 for the corresponding time step sequence may be used to manipulate multiple arms and / or end effectors on each arm to perform one or more tasks involving manipulating and / or assembling parts.
[0054] like Figure 2 As shown, each action 226 is generated by strategy 220 based on the corresponding perception 224. In one or more embodiments, perception 224 includes data representing the robot's current and / or past states. For example, perception 224 may include (but is not limited to) current position, angle, velocity, angular velocity, linear acceleration, contact force, motion and / or joints, end effectors, and / or other physical properties of other parts of the robot. This data may be derived from inertial data, torque data, force data, and / or other sensor inputs with jointed objects available in real-world and / or simulated environments.
[0055] Perception 224 may also include, or alternatively include, observations of the real-world and / or simulated environment available to the robot. For example, perception 224 may include (but is not limited to) one or more camera views of the environment surrounding the jointed object, one or more visualizations generated by combining multiple camera views of the environment (e.g., bird's-eye view, perspective, 360-degree, etc.), semantic labels associated with the camera views and / or visualizations (e.g., segmented map, detected objects, boundary shapes, etc.), a three-dimensional (3D) representation of the environment (e.g., point cloud, mesh, generic scene description (USD), etc.), and / or other representations of the environment.
[0056] In one or more embodiments, strategy 220 generates motion distribution 218 based on input including sensing 224. Corresponding actions 226 may be sampled from motion distribution 218 as position, orientation, motor commands, and / or other types of output related to motion in the robot's joints and / or end effectors. Actions 226 may then be used (e.g., via the robot's controller) to actuate degrees of freedom in the robot, thereby generating updated sensing 224 for the next time step. This process may be repeated until an overall target objective (e.g., completing one or more assembly tasks) is achieved and / or another predetermined termination condition is met.
[0057] It should be understood that strategy 220 may generate motion distribution 218 based on additional inputs associated with the manipulation of the articulated object. For example, this additional input may include a goal to be achieved by performing motion 226 during the current time step. This goal may include the target position, orientation, angle, linear velocity, angular velocity, and / or other attributes of the joints, end effectors, and / or other parts of the articulated object at different time steps. This goal may also, or alternatively, include higher-level task objectives, such as (but not limited to) achieving a specific configuration of two or more parts and / or completing a given task by manipulating two or more parts using the arm and / or end effector of the articulated object.
[0058] Data generation engine 122 generates asset pairs that can be physically assembled in simulated and / or real environments. Each asset pair includes first parts 232(1) to 232(X) (each of which is individually referred to herein as part 232) and second mating parts 238(1) to 238(X) complementary to the first part 232 (each of which is individually referred to herein as mating part 238). For example, a given part 232 and a corresponding mating part 238 may include screwdrivers and screws, nuts and bolts, containers and caps, and / or another type of plug and socket that can mate with each other to form corresponding components.
[0059] In some embodiments, the data generation engine 122 performs assembly asset generation through a pipeline comprising three stages. For example...Figure 2 As shown, the pipeline includes a shape characterization stage 212 in which the data generation engine 122 determines a different set of attributes 234(1) through 234(X) (each of which is individually referred to herein as an attribute 234) and a set of contact surfaces 236(1) through 236(X) (each of which is individually referred to herein as a contact surface 236) for each part 232.
[0060] In one or more embodiments, the shape characterization stage 212 involves interactions between the data generation engine 122 and a VLM, LLM, MMLM, and / or another type of machine learning model capable of universal understanding and / or generation of natural language and visual content. The input to the machine learning model includes a visual representation and / or another representation of a given part 232 to be characterized. For example, the part 232 can include a computer-aided design (CAD) model and / or another three-dimensional (3D) representation of an independent asset sampled from a dataset and / or generated (e.g., by a user, another machine learning model, a procedural generation technique, etc.). The data generation engine 122 can use a pre-specified and / or user-defined set of parameters to generate a rendering of the part 232 (e.g., normalize the part 232 to fit into a bounding box [-1, 1] 3 from a top-front view at coordinates (5, 5, 0) towards (0, 0, 0). The data generation engine 122 can then input the rendered part 232 into the machine learning model and obtain a prediction of one or more attributes 234 associated with the rendered part 232 as a corresponding output of the machine learning model.
[0061] The input to the machine learning model also or instead includes additional information that can be used to determine the attributes 234 of a given part 232. Continuing the above example, the data generation engine 122 can also input the rendered part 232 into the instruct the machine learning model to output one or more hints of the attributes 234 based on the rendered part 232. The data generation engine 122 can also or instead input a textual description of the rendered part 232 and / or additional context associated with the rendered part 232 to help the machine learning model determine the attributes 234.
[0062] Figure 3A FIGURE 1 illustrates how the data generation engine 122 determines attributes 234 associated with a part in accordance with at least one embodiment. Figure 1 Figure 3A As shown, the data generation engine 122 generates a series of inputs 302(1) to 302(4) (each of which is individually referred to herein as input 302) into the VLM 300 and obtains a corresponding series of outputs 304(1) to 340(3) (each of which is individually referred to herein as output 304) that include attributes 234 associated with a given part 232.
[0063] More particularly, the data generation engine 122 uses a chain-of-thought prompt to guide the VLM 300 to generate outputs 304 that specify different attributes 234 of the part 232. During this chain-of-thought prompt, the data generation engine 122 generates a first input 302(1) that includes a rendering of the selected part 232 and a second input 302(2) that includes prompts to describe the geometry and functionality of the part 232. Based on the inputs 302(1) to 302(2), the VLM 300 generates a first output 304(1) that identifies the part 232 as a Phillips head screw with certain geometric and functional characteristics.
[0064] Next, the data generation engine 122 generates a third input 302(3) that asks the VLM 300 to identify the part 232 as either a plug or a socket and includes definitions of “plug” and “socket.” Based on this third input 304(3), the VLM 300 generates a second output 304(2) that identifies the part 232 as a socket.
[0065] The data generation engine 122 then generates a fourth input 302(4) that asks the VLM 300 to select an insertion axis and direction from a list. Based on this fourth input 302(4), the VLM 300 generates a third output 304(3) that identifies the insertion axis and direction as “top down.”
[0066] By iteratively prompting the VLM 300 to discern different attributes 234 of the part 232, the data generation engine 122 can improve the quality and / or accuracy of the output attributes 234. For example, the data generation engine 122 can prompt the VLM 300 to predict increasingly specific attributes of the part 232 such that each predicted attribute can be included in a context that informs the prediction of subsequent attributes 234 of the same part 232.
[0067] Returning to the discussion of Figure 2 After the attributes 234 of a given part 232 have been determined, the data generation engine 122 uses some or all of the attributes 234 to determine a set of contact surfaces 236 of the same part 232. In some embodiments, the contact surfaces 236 include surfaces on the part 232 that are predicted to come into contact with a corresponding mating part 238 during assembly. As discussed below with respect to Figure 3BAs described in further detail, the data generation engine 122 may use an analysis program that utilizes the insertion axis and direction specified in attribute 234 to identify the contact surface 236 of part 232.
[0068] Figure 3B The illustration shows an embodiment according to at least one of the embodiments. Figure 1 How does the data generation engine 122 determine the relationship with parts (e.g., Figure 3A The contact surface 236 associated with the Phillips head screwdriver. Figure 3B As shown, the data generation engine 122 performs a first step 312 of generating a voxel mesh above the part. For example, the data generation engine 122 may align the assembly direction of the part with the z-axis in a "top-down" orientation. The data generation engine 122 may also align the part's assembly direction with the z-axis in a aligned manner above the part (e.g., 512). 3 Initialize a voxel mesh in 1000 dimensional dimensions.
[0069] Next, the data generation engine 122 performs a second step 314: moving the voxel mesh downwards until the top layer of the voxel mesh is aligned with the top of the part. As the voxel mesh is moved downwards, the data generation engine 122 also performs a step 316: removing mesh cells that are in contact with the part. The data generation engine 122 also performs a step 318: removing “outlier” mesh cells in the voxel mesh that are not directly above the part.
[0070] The data generation engine 122 then performs two steps, 320 and 322, to locate the contact surface 236 on the part using the remaining mesh cells. In step 320, the data generation engine 122 removes additional mesh cells located outside the convex hull of the part (e.g., when the part is identified as a socket). In step 322, the data generation engine 122 traverses the faces in the part to extract the contact surface 236 as the contact surface that is in full contact with the remainder of the voxel mesh.
[0071] Back Figure 2discussed above, after the contact surface 236 of a given part 232 has been identified, the data generation engine 122 performs a second shape completion stage 214 to generate a counterpart part 238. In one or more embodiments, the shape completion stage 214 uses the contact surface 236 to regulate the operation of a three-dimensional (3D) generative model in generating a shape of the counterpart part 238 that is complementary to the part 232. For example, the data generation engine 122 can perform the shape completion stage 214 by using a transformer-based diffusion model to generate a 3D computer-aided design (CAD) model in a unit cube and in a boundary representation (B-rep) format that includes a graph in which geometric primitives (e.g., faces and edges) are represented by graph nodes and topological relationships between the geometric primitives are represented by graph edges. The data generation engine 122 can initialize a denoising process performed by the diffusion model by moving the contact surface 236 to the bottom of the unit cube. During at least a portion of the time steps in the denoising process, the data generation engine 122 can replace a subset of face tokens processed by the diffusion model with the contact surface 236. As a result, the denoising process can be used to complete a shape that includes the contact surface 236.
[0072] Figure 4A FIG. 1 illustrates an example of a data generation engine 122 generating a set of parts 232(1) through 232(10) and a set of counterpart parts 238(1) through 238(10) in accordance with at least one embodiment. Figure 1 Figure 4A As shown, the top of each part 232(1) through 232(10) and the bottom of the counterpart part 238(1) through 238(10) are complementary to each other and can fit together in a particular configuration (e.g., via inserting the bottom of a given counterpart part 238(1) through 238(10) into the top of the corresponding part 232(2) through 232(10)). The geometry of the top of the counterpart part 238(1) through 238(10) can vary (e.g., based on the operation and / or output of one or more machine learning models used for the shape completion stage 214), thereby allowing for training, testing, and / or evaluating a robot that performs assembly using the parts 232 and the counterpart parts 238 in an integrated manner on an assembly task.
[0073] Returning to FIG. 1, the data generation engine 122 can generate the set of parts 232(1) through 232(10) and the set of counterpart parts 238(1) through 238(10) by performing the shape completion stage 214 for each part 232(1) through 232(10) in the set of parts 232(1) through 232(10) and generating a corresponding counterpart part 238(1) through 238(10) for each part 232(1) through 232(10). Figure 2 discussed, the data generation engine 122 also performs a third gap specification stage 216 in which a minimum gap distance 244(1) through 244(X) (each of which is referred to herein as a gap distance 244) is enforced between a given part 232 and a corresponding mating part 238 to ensure that each part 232 and corresponding mating part 238 can successfully mate in a real-world and / or simulated environment. Each gap distance 244 can be specified by a user, set to a default value, determined by a machine learning model, set to a value based on standards and / or rules associated with the type and / or application of the part 232 and / or mating part 238, and / or determined via another technique.
[0074] During the gap specification stage 216, the data generation engine 122 can convert the given part 232 and / or the corresponding mating part 238 to an occupancy grid representation and initialize the part 232 and the corresponding mating part 238 in an assembled state. The data generation engine 122 can also remove grid cells in the occupancy grid that are in contact with or fall within the gap distance 244 of another part until all grid cells in the occupancy grid satisfy the minimum gap distance 244. The data generation engine 112 can then convert the resulting “pruned” occupancy grid to a mesh (e.g., via a marching cubes technique) or another 3D format for an updated part 242(1) through 242(X) (each of which is referred to herein individually as an updated part 242) corresponding to the part 232 and / or an updated mating part 240(1) through 240(X) (each of which is referred to herein individually as an updated mating part 240) corresponding to the mating part 238.
[0075] Figure 4B FIG. 13 illustrates how the data generation engine 122 performs gap specification on the part 232 and the corresponding mating part 238, according to at least one embodiment. Figure 1 As shown, the region 402 between the part 232 and the mating part 238 includes portions of the part 232 and the mating part 238 that interpenetrate each other, which can cause the part 232 and the mating part 238 to explode when loaded into a simulated environment and / or prevent the part 232 and the mating part 238 from mating in a real-world environment. Figure 4B
[0076] To address the interpenetration, the data generation engine 122 uses the gap specification stage 216 to remove portions of the part 232 and / or the mating part 238 that are within one millimeter of each other. The resulting contact surfaces of the updated part 242 and the updated mating part 240 are at least one millimeter apart from each other.
[0077] Figure 4C The illustration shows an embodiment according to at least one of the embodiments. Figure 1 How does the data generation engine 122 execute the gap specification for part 232 and its corresponding paired part 238? Figure 4C Part 232 and mating part 238 with Figure 4B Those are the same, and they interpenetrate within region 402.
[0078] To resolve interpenetration, the data generation engine 122 uses a gap specification stage 216 to remove portions of part 232 and / or mating part 238 that are within a four-millimeter gap distance 244 from each other. The resulting updated part 242 and updated mating part 240 have contact surfaces separated from each other by at least four millimeters.
[0079] continue Figure 2 While the operation of data generation engine 122 has been described in relation to generating paired component assets (such as plugs and sockets), it should be understood that data generation engine 112 can be used to generate parts that can be included in other types of components. For example, data generation engine 122 may be configured to generate components with more than two parts by including a first part having more than one plug, more than one socket, and / or at least one plug and one socket in a given component and generating more than one additional part that mates with one or more plugs and / or one or more sockets in the first part. Data generation engine 122 may also, or alternatively, use more than one part to form plugs and / or sockets and generate one or more additional parts that mate with plugs and / or sockets. In another example, data generation engine 122 may include the ability to generate other types of parts, such as (but not limited to) interlocking components, puzzle pieces, gears, shafts, bearings, races, frames, panels, brackets, tracks, hinges, clamps, couplings, hoses, and / or rotating components.
[0080] After the shape representation stage 212, shape completion stage 214, and gap specification stage 216 are used to generate a given updated part 242 and a corresponding updated mating part 240, the data generation engine 122 adds the updated part 242 and the updated mating part 240 to the training component set 204 included in the training data 200 for one or more machine learning models. For example, after verifying that a given updated part 242 and a corresponding updated mating part 240 can successfully mate in a simulated and / or real-world environment, the data generation engine 122 may add these two parts to the training component 204.
[0081] The training engine 124 uses training data 200 and one or more training objectives 208 to update the model parameters 206 of one or more machine learning models that implement the robot assembly strategy 220. For example... Figure 2As shown, the training data 200 includes training components 204 generated by the data generation engine 122, and training trajectories 202 associated with the training components 204. The training trajectories 202 can include sequences of actions performed to assemble mating parts (e.g., updated part 242 and updated mating part 240 generated by the data generation engine 122) in the training components 204. For example, a given training trajectory can include a sequence of training actions for a given end effector of a robot over a number of time steps. Each action can include a set of positions, orientations, motor commands, and / or other outputs that describe the configuration of a joint object at a particular “frame” or time step within a joint, end effector, and / or corresponding motion. A given training trajectory can also include, and / or be associated with, a set of training perceptions (e.g., states and / or observations) for each “frame” or time step. These training perceptions can include, but are not limited to, positions, angles, velocities, angular velocities, linear accelerations, contact forces, actions, and / or other state representations of the joint, end effector, and / or other portions of the joint object. These training perceptions can also or instead include, but are not limited to, camera views of the environment surrounding the joint object, visualizations generated by combining multiple camera views of the environment, semantic labels associated with the camera views and / or visualizations, 3D representations of the environment, and / or other information related to observations of the environment. The training actions and / or training perceptions in a given training trajectory can be generated by the data generation engine 122, the training engine 124, and / or another component via assembly-by-disassembly, via human demonstration of a teleoperation system, dataset aggregation, and / or other techniques.
[0082] Additionally, the training engine 124 can use various training techniques and / or corresponding training objectives 208 to update the model parameters 206 using the training data 200. For example, the training engine 124 can use behavior cloning, dataset aggregation, trajectory matching, and / or other types of reinforcement learning (RL) and / or imitation learning techniques and corresponding training objectives 208 to update the model parameters 206 of a multilayer perceptron (MLP), a long short-term memory (LSTM) neural network, a recurrent neural network (RNN), an RNN-Gaussian Mixture Model (RNN-GMM), a policy 220, and / or another type of machine learning model corresponding to the policy 220 based on the training data 200 and the one or more training objectives 208. During training of the policy 220, the training engine 124 initializes the simulated environment with the updated parts 242 and corresponding updated counterpart parts 240 from the training component 204 in a pre-defined and / or randomized pose. The training engine also inputs the training perception associated with various time steps in the training trajectory 202 into the policy 220. The training engine 124 uses the model parameters 206 of the policy 220 to generate the training output 210 corresponding to the action for the same time step. The training engine 124 uses the generated training output 210 to compute the one or more training objectives 208 and updates the parameters of the policy 220 in a manner that optimizes the training objectives 208.
[0083] After training of the policy 220 is complete, the execution engine 126 uses the trained policy 220 to perform various robotic assembly tasks. During the tasks, the execution engine 126 inputs the perception 224 for a given time step into the policy 220 and uses the policy 220 to generate a corresponding action distribution 218. The execution engine 126 samples a corresponding action 226 from the action distribution 218 output by the policy 220 and converts the action 226 into commands and / or corresponding motions in the articulated object. The execution engine 126 repeats the process for a next time step with a new perception 224 until a certain number of time steps have elapsed, an overall target goal (e.g., completion of one or more assembly tasks) is achieved, and / or another pre-defined termination condition is satisfied.
[0084] The execution engine 126 can additionally incorporate each generated action 226 into various applications. For example, the execution engine 126 can simulate the articulated object performing the robotic assembly within a game, a video, a virtual world, a visualization, and / or another setting. The execution engine 126 can also or instead generate commands that cause a robot corresponding to the articulated object to perform each action 226 in a real-world environment.
[0085] While the operations of the data generation engine 122, the training engine 124, and the execution engine 126 have been described above with respect to generating paired assembly assets for training machine learning models, it should be appreciated that the data generation engine 122, the training engine 124, and / or the execution engine 126 can be used to perform other types of tasks. For example, the data generation engine 122 can use the gap specification stage 216 to repair assets and / or components that include interpenetrating parts and / or that otherwise cannot be used in a simulation and / or real-world environment. In another example, one or more updated parts 242 and corresponding updated counterpart parts 240 can be manufactured for assembly and / or use in a real-world environment by a robot, a human, and / or another type of articulating object.
[0086] It should be understood that the arrangements and other arrangements described herein are stated by way of example only. Other arrangements, components, features, and elements (e.g., machines, interfaces, functions, orders, groups of functions, etc.) can be used in Figures 6A-6E The example machine 600, Figure 7 The example computing ecosystem 700, Figure 8 The example generative language model system 800, and / or Figure 9The example computing device 900 uses components, features, and / or functions similar to those of other computing devices, features, and / or functions to implement this.
[0087] Now, for reference Figure 5 Each block of the method 500 described herein includes a computational process that may be performed using any combination of hardware, firmware, and / or software. For example, various functions may be performed using one or more processors (such as, but not limited to, the processors described herein) that execute instructions stored in one or more memories or storage systems. In some embodiments, the computer process may also be embodied as computer-usable instructions stored on a computer storage medium. The method may be provided by a standalone application, a service (standalone or in combination with another managed service), a managed service, an application programming interface (API), and / or a plug-in to another product, etc. Additionally, examples are provided regarding... Figures 1-2 Method 500 is described herein. However, these methods may be implemented additionally or alternatively by any system or any combination of systems, including but not limited to the systems described herein.
[0088] Figure 5 A flowchart illustrating a method 500 for generating robot assembly data according to at least one embodiment is shown. Figure 5 As shown, method 500 begins with operation 502, where data generation engine 122 determines one or more attributes of a first part in a component by executing a first machine learning model. For example, data generation engine 122 may select a first part from a dataset, receive a first part from a user, and / or generate the first part. Data generation engine 122 may also input a rendering and / or another visual representation of the first part, along with one or more instructions describing the geometry, function, type (e.g., plug, socket, etc.), insertion axis, insertion direction, and / or other attributes of the first part, into a VLM and / or other type of machine learning model. Upon receiving a given instruction, the machine learning model may generate an output including an answer to the instruction based on the visual representation of the first part and / or the context including previous interactions between data generation engine 122 and the machine learning model.
[0089] In operation 504, the data generation engine 122 determines a set of contact surfaces on the first part based on the one or more properties and the alignment of the mesh to the top of the 3D geometry of the first part. For example, the data generation engine 122 can rotate the 3D geometry of the first part to assign the assembly direction identified in operation 502 with the z-axis in a top-down orientation. The data generation engine 122 can also initialize a voxel mesh above the rotated first part and lower the voxel mesh until the top layer of the mesh is aligned with the top surface of the first part. Thus, the data generation engine 122 can “project” or “superimpose” the voxel mesh over the first part. As the voxel mesh is lowered, the data generation engine 122 can remove mesh cells that contact the first part. The data generation engine 122 can also or instead remove mesh cells that are outside the convex hull of the first part (e.g., when the first part is identified as a socket). The data generation engine 122 can then identify the contact surfaces as the set of faces in the first part that are in full contact with the remaining mesh cells in the voxel mesh.
[0090] In operation 506, the data generation engine 122 generates a second part that mates with the first part by executing a second machine learning model based on the contact surfaces. For example, the data generation engine 122 can use a diffusion model and / or another type of generative model to generate a 3D representation of the first part. The data generation engine 122 can also adjust the operation and / or output of the generative model on the contact surfaces (e.g., by using the contact surfaces to replace portions of the output shape generated by a denoising step performed by the generative model).
[0091] In operation 508, the data generation engine 122 updates the contact surfaces on the first part and / or corresponding contact surfaces on the second part based on a clearance distance between the parts. For example, the data generation engine 122 can convert a given part to an occupancy grid representation and remove mesh cells of the occupancy grid along the assembly axis of the occupancy grid until all mesh cells in the occupancy grid are at least a minimum “clearance distance” away from other parts. The data generation engine 122 can then convert the occupancy grid to a triangular mesh and / or another “final” representation of the part.
[0092] In operation 510, the data generation engine 122 adds the first part and the second part to the robotic assembly dataset. For example, after verifying that the parts can mate in a simulated and / or real environment, the data generation engine 122 can add the parts to the robotic assembly dataset.
[0093] In operation 512, the data generation engine 122 determines whether to continue generating assembly assets. For example, the data generation engine 122 can determine to continue generating assembly assets until a certain number of components and / or assembly assets have been generated, a certain number of components and / or assembly assets of a certain type have been generated, and / or another condition is met. When the data generation engine 122 determines to continue generating assembly assets, the data generation engine 112 repeats operations 502, 504, 506, 508, and 510 to generate additional pairs of assembly assets that fit together.
[0094] After the data generation engine 122 determines to no longer continue generating assembly assets (or while the data generation engine 122 continues generating assembly assets), the training engine 124 and / or the execution engine 126 perform operation 514 in which the training engine 124 and / or the execution engine 126 train, evaluate, and / or execute machine learning models using the robotic assembly dataset. For example, the training engine 124 can train generalist and / or specialist strategies for robotic assembly using the paired parts in the robotic assembly dataset. In another example, the execution engine 126 can evaluate the performance of robotic assembly strategies and / or robots on various assembly tasks involving the parts in the robotic assembly dataset. In a third example, the execution engine 126 can use the robotic assembly strategies and / or robots to assemble assets in the robotic assembly dataset in simulated and / or real-world environments.
[0095] In summary, the disclosed technology automates the generation of paired parts (also referred to herein as components) in assembly components via a three-stage pipeline. The pipeline includes a first contact surface extraction stage in which a set of contact surfaces are extracted from a first part based on a visual representation and / or another representation of the first part by a visual language model (VLM) and / or another type of machine learning model. The pipeline also includes a second shape completion stage in which the contact surfaces are used to regulate the operation of a diffusion model and / or another type of three-dimensional (3D) generative model in generating a shape of a second part that is complementary to the first part. The pipeline also includes a third clearance specification stage in which the shape of a given part is updated to satisfy a minimum clearance distance from other parts. The paired parts generated via the pipeline can then be used to train strategies for assembly tasks in simulated environments, evaluate the performance of deployed strategies in real-world environments, for real-world assembly tasks, and / or for other tasks related to assembly.
[0096] One advantage of the disclosed technology over existing approaches is the ability to automatically generate large, diverse sets of paired parts that can be used for assembly tasks. Thus, the disclosed technology can be used to generate pairs of assembly assets more quickly and efficiently than traditional approaches that involve manually generating and / or curating sets of robot assembly data. The generated pairs of assets can also be used to train, test, and / or evaluate robots on assembly tasks in a more comprehensive manner than robots trained using much smaller and / or less diverse sets of components. Further, robots trained using the generated pairs of assets can be more fault-tolerant and / or able to generalize to different scenarios than robots trained using more limited sets of paired components. Additionally, the disclosed technology can be used to "repair" interpenetrating assets generated via other techniques, thereby further increasing the number and types of assets that can be used for assembly tasks.
[0097] The systems and methods described herein can be used by, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled with one or more trailers, aircraft, watercraft, shuttles (e.g., autonomous taxis), emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, engineering vehicles, underwater vehicles (e.g., manned or unmanned submarines), drones, and / or other vehicle types. Further, the systems and methods described herein can be used for a variety of purposes, by way of example but not limitation, for machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, safety and supervision, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environment simulation, object or actor simulation and / or digital twinning, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets (e.g., NVIDIA’s Omniverse), cloud computing, and / or any other suitable application.
[0098] The disclosed embodiments can be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines, etc.), systems implemented using robots, aviation systems, medical systems, marine systems, smart zone monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems implementing language models (such as large language models (LLM), visual language models (VLM), and / or multimodal language models), systems using or deploying one or more inference microservices, systems containing one or more machine learning models deployed in services or microservices and OS-level virtualization packages (e.g., containers), systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems for performing optical transmission simulations, systems for performing collaborative content creation for 3D assets, systems for performing generative AI operations, systems implemented at least partially using cloud computing resources, and / or other types of systems.
[0099] Example autonomous or semi-autonomous machines
[0100] Figure 6A Examples of sensor positions with corresponding fields of view or sensing fields of view of autonomous or semi-autonomous vehicles 600a, autonomous mobile robots (AMRs) 600b, and humanoid robots 600c according to some embodiments of this disclosure. While three types of machines 600 are shown, this is not intended to be limiting, and the machines 600 described herein may include vehicles, cars, trucks, buses, first-response vehicles, shuttle buses, electric or motorized bicycles, motorcycles, fire trucks, police or emergency vehicles, ambulances, boats, construction vehicles, underwater vehicles, robots (e.g., AMRs, humanoid robots, robotic arms, end effectors, forklifts, etc.), drones, aircraft, vehicles coupled to trailers (e.g., semi-trailer tractors for hauling goods), and / or another type of vehicle or machine (e.g., driverless and / or vehicles or machines accommodating one or more passengers). In some cases, vehicles 600a, AMRs 600b, humanoid robots 600c, and / or other machine types may be collectively referred to herein as machines 600.
[0101] With respect to vehicle 600A, autonomous and semi-autonomous vehicles are often described in terms of an automation level defined by the National Highway Traffic Safety Administration (NHTSA), a department of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (Standard J3016-201806 published June 15, 2018, Standard J3016-201609 published September 30, 2016, and prior and future versions of the standard). Machine 600 can have functionality that complies with one or more of automation levels 3-5. Machine 600 can have functionality that complies with one or more of automation levels 1-5. For example, machine 600 can be capable of providing driver assistance (level 1), partial automation (level 2, 2+, 2++), conditional automation (level 3), high automation (level 4), and / or full automation (level 5), depending on the embodiment. The term “autonomous” as used herein can include any and / or all types of autonomy of machine 600 or other machines, such as fully autonomous, highly autonomous, conditionally autonomous, partially autonomous, providing assistance autonomous, semi-autonomous, primarily autonomous, or other designations.
[0102] With respect to vehicle 600A, autonomous and semi-autonomous vehicles are often described in terms of an automation level defined by the National Highway Traffic Safety Administration (NHTSA), a department of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (Standard J3016-201806 published June 15, 2018, Standard J3016-201609 published September 30, 2016, and prior and future versions of the standard). Machine 600 can have functionality that complies with one or more of automation levels 3-5. Machine 600 can have functionality that complies with one or more of automation levels 1-5. For example, machine 600 can be capable of providing driver assistance (level 1), partial automation (level 2, 2+, 2++), conditional automation (level 3), high automation (level 4), and / or full automation (level 5), depending on the embodiment. The term “autonomous” as used herein can include any and / or all types of autonomy of machine 600 or other machines, such as fully autonomous, highly autonomous, conditionally autonomous, partially autonomous, providing assistance autonomous, semi-autonomous, primarily autonomous, or other designations. Figure 6A , the sensors and their respective fields of view (not shown for clarity) or sensing fields (not shown for clarity) are one example embodiment and are not intended to be limiting. Although not shown, each sensor can have a corresponding field of view (e.g., 360 degrees for a surround camera 668D, 180 degrees for a wide-angle camera 668B, 360 degrees for a LiDAR sensor 664, etc.). For example, only a subset of the illustrated sensors can be included, additional sensors can be included, alternative sensors can be included, the number of each sensor modality can be different, the sensor modalities can be different (e.g., LiDAR or RADAR can not be included, SONAR, thermal sensors, etc. can be included), the sensor locations can be different from those illustrated on vehicle 600a, AMR 600b, and / or humanoid robot 600c, etc. For example, for vehicle 600a, the locations, number, modalities, and / or other sensor information can be different depending on the type (e.g., SUV, truck, sedan, robot, motorcycle, etc.), size (e.g., 18-wheeler, forklift, compact car, etc.), and related functionality (e.g., L2 vs. L5). Similarly, for AMR 600b and / or humanoid robot 600c, the shape, size, purpose, implementation, model, etc. can dictate the number and type of sensors used.
[0103] As Figure 6AAs shown, autonomous or semi-autonomous vehicles 600A, AMR 600B, and humanoid robots 600C may include different sensor types, numbers, and locations. As a non-limiting example, vehicle 600A may include twelve cameras 668, such as a front wide-angle camera (e.g., 120-degree field of view (FOV)), a front telephoto camera (e.g., 30-degree FOV), a side-rear left camera (e.g., 70-degree FOV), a side-rear right camera (e.g., 70-degree FOV), a front fisheye camera (e.g., 200-degree FOV), a rear fisheye camera (e.g., 200-degree FOV), a left fisheye camera (e.g., 200-degree FOV), a right fisheye camera (e.g., 200-degree FOV), a front telephoto satellite camera (e.g., 30-degree FOV), a rear telephoto camera (e.g., 30-degree FOV), a cross-left camera (e.g., 120-degree FOV), and a cross-right camera (e.g., 120-degree FOV). In this embodiment, the camera 668 may use a Gigabit Multimedia Serial Link (GMSL) interface (e.g., GMSL2) as input / output (I / O).
[0104] In some embodiments, although Figure 6A As not shown, vehicle 600A may include an in-cabin occupant and / or driver monitoring system, which may include various sensors. For example, in-cabin sensors may include various cameras 668, such as a driver monitoring camera (e.g., located in front of the driver's seat and facing the driver's seat at a 55-degree FOV), a front occupant monitoring camera (e.g., located in front of the front occupant seat and facing the front occupant seat at a 190-degree FOV), and a rear occupant monitoring camera (e.g., located in front of the rear occupant seat and facing the rear occupant seat at a 190-degree FOV). Similar to external cameras 668, in embodiments, internal cameras 668 may use a GMSL (e.g., GMSL2) interface for I / O.
[0105] As another non-limiting example, vehicle 600A may also include nine RADAR sensors 660. For example, vehicle 600A may include a front center imaging RADAR sensor (e.g., 120-degree FOV or sensing field), a left front corner RADAR sensor (e.g., 160-degree FOV or sensing field), a right front corner RADAR sensor (e.g., 160-degree FOV or sensing field), a right rear corner RADAR sensor (e.g., 160-degree FOV or sensing field), a left RADAR sensor (e.g., 160-degree FOV or sensing field), a right RADAR sensor (e.g., 160-degree FOV or sensing field), a left rear RADAR sensor (e.g., 50-degree FOV or sensing field), and a right rear RADAR sensor (e.g., 50-degree FOV or sensing field). In embodiments, the RADAR sensors 660 may use an Ethernet interface as I / O.
[0106] As a non-limiting example, vehicle 600A can also include twelve ultrasonic sensors 662. As shown, the ultrasonic sensors can be placed along the front and rear bumpers of vehicle 600A, as well as along the sides of vehicle 600A, and can be used to detect objects (static and dynamic) in close proximity to vehicle 600A. In some embodiments, ultrasonic sensors 662 can use a DS13 interface as I / O. Figure 6A
[0107] As a non-limiting example, vehicle 600A can also include LiDAR sensors 664, such as a front center LiDAR sensor (e.g., 120 degree horizontal FOV or sensing field and 30 degree vertical FOV or sensing field). In some embodiments, such as where additional or alternative LiDAR sensors are used, the LiDAR sensors can have different horizontal and vertical fields of view or sensing fields. For example, LiDAR sensors 664 can include a 360 degree horizontal FOV or sensing field (e.g., in a rotating LiDAR sensor) and a 90 degree vertical FOV or sensing field. In some embodiments, LiDAR sensors 664 can use an Ethernet interface as I / O.
[0108] As a non-limiting example, autonomous mobile robot (AMR) 600B can include three LiDAR sensors 664. For example, the topmost shown LiDAR sensor 664 can include a beam or 3D LiDAR sensor (e.g., 360 degree horizontal and 90 degree vertical FOV or sensing field), and the front and rear LiDAR sensors can include planar or 2D LiDAR sensors (e.g., 180 degree horizontal FOV or sensing field).
[0109] As a non-limiting example, AMR 600B can further include eight cameras 668, such as a front stereo camera (e.g., 120 degree FOV), a rear stereo camera (e.g., 120 degree FOV), a left stereo camera (e.g., 120 degree FOV), a right stereo camera (e.g., 120 degree FOV), a front fisheye camera (e.g., 202 degree +-3 degree FOV), a rear fisheye camera (e.g., 202 degree +-3 degree FOV), a left fisheye camera (e.g., 202 degree +-3 degree FOV), and a right fisheye camera (e.g., 202 degree +-3 degree FOV).
[0110] The AMR 600B can also include a charging port, a charging port contact, a status indicator light, one or more (e.g., four) RGB LEDs, one or more IMU sensors 666, a magnetometer, and a barometer. The AMR 600B is capable of high-precision time synchronization between sensors using hardware timestamps and PTP over Ethernet with less than 10 microseconds of sensor acquisition time. In embodiments, the AMR 600B provides simultaneous camera capture within 100 microseconds of a single hardware trigger on all cameras 668, and can write sensor captures to disk at 4 GB / second to write packets (e.g., to ROSbags for a Robot Operating System (ROS)). Thus, the AMR 600B is capable of running ROS (e.g., Isaac ROS by NVIDIA), can be teleoperated (as described herein), can map an environment, and can navigate in the environment using vision cameras 668, LiDAR 664, and / or other sensor types or modalities.
[0111] The humanoid robot 600C can include (as non-limiting examples) one LiDAR sensor 664. For example, the LiDAR sensor 664 can include a light beam or 3D LiDAR sensor (e.g., 360 degree horizontal and 90 degree vertical FOV or sensing field), or can include a planar or 2D LiDAR sensor (e.g., 180 degree horizontal FOV or sensing field).
[0112] As non-limiting examples, the humanoid robot 600C can also include four cameras 668, such as a front-facing stereo camera (e.g., 120 degree FOV), a rear-facing stereo camera (e.g., 120 degree FOV), a front-facing fisheye camera (e.g., 202 degree +-3 degree FOV), and a rear-facing fisheye camera (e.g., 202 degree +-3 degree FOV).
[0113] As non-limiting examples, the humanoid robot 600C can also include four ultrasonic sensors 662, such as a left arm ultrasonic sensor, a right arm ultrasonic sensor, a left leg ultrasonic sensor, and a right leg ultrasonic sensor.
[0114] The humanoid robot 600C can also include any number of actuators, such as allow for control and manipulation of joints. For example, the humanoid robot 600C can include actuators that allow for various degrees of freedom (DoF) to be implemented depending on the design. In non-limiting embodiments, the humanoid robot 600C can have a total of 40 degrees of freedom (DoF) (e.g., arm 6 DoF x2, hand 6 DoF x2, leg 6 DoF x2, torso 2 DoF, and neck 2 DoF). The actuators can convert energy into physical motion, allowing for actions such as joint movement, locomotion, and grasping / manipulation. For example, joint movement can be performed using motors and servos to control the rotation of joints in the arms or manipulators and allow for reaching, grasping, and manipulating objects. Locomotion can be achieved using wheels, tracks, or other mobility devices (robotic legs) to move in the environment. Grasping and manipulation can be performed using end effectors or hands / fingers, which can be equipped with actuators to grasp objects, apply forces, and perform specific tasks. In some examples, the humanoid robot 600C can include position and orientation sensors, such as encoders, gyroscopes, etc., to determine the position of the robot 600C in space, enabling position determination and motion tracking. In embodiments, the humanoid robot 600C can include force and pressure sensors to detect environmental interactions, enabling the robot 600C to grasp objects with appropriate force and avoid obstacles along the way. Sensory sensors (e.g., cameras, LiDAR, RADAR, ultrasound, SONAR, etc.) can be used with tactile sensors to enable the robot 600C to perceive objects, shapes, and textures and understand when to start and stop touching (along with force sensors to regulate the force used during touching). As non-limiting examples, the humanoid robot 600C can have a height of about 1-2 meters (e.g., 1.7 meters or 5'6" inches), a weight of 50-70 kilograms, be able to move at 8 kilometers / hour or higher, and be able to carry a payload of 20-100 kilograms, depending on the design and requirements of the system.
[0115] In embodiments, the humanoid robot 600C can include a dialog system, such as driven by a language model (e.g., LLM, VLM, MMLM, VLA, etc.), to help understand the environment, reason, and communicate with humans, animals, devices, and / or other robots, and / or make planning, control, and navigation decisions. Thus, in addition to performing various tasks, the humanoid robot 600C can also use onboard sensors, microphones, and speakers to understand speech, audio, and visual cues, among others, while also being able to communicate with the environment.
[0116] Referring to the cameras 668 of the machine 600, the camera type of the cameras 668 can include, but is not limited to, a digital camera applicable to components and / or systems of the machine 600. For the vehicle 600a implementation, the cameras 668 can operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. The cameras can be capable of any image capture rate, such as 30 frames / second (fps), 60 fps, 120 fps, 240 fps, etc., depending on the embodiment. The cameras can use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array can include a Red Colorless Colorless Color (RCCC) color filter array, a Red Colorless Colorless Blue (RCCB) color filter array, a Red Blue Green Colorless (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, achromatic pixel cameras (e.g., cameras with a
[0117] The field of view includes cameras (e.g., front-facing cameras) that can be used for surround view to help identify the path and obstacles ahead, as well as provide information critical to generating an occupancy grid and / or determining preferred machine motion, trajectories, and / or paths with the help of one or more controllers 636 and / or control SoCs. The front-facing cameras can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. The front-facing cameras can also be used for ADAS functions and systems, including Lane Departure Warning (“LDW”), Automatic Cruise Control (“ACC”), and / or other functions, such as traffic sign recognition.
[0118] Various cameras can be used in a front-facing configuration, including, for example, a monocular camera platform that includes a Complementary Metal-Oxide Semiconductor (“CMOS”) color imager. Another example can be a wide-angle camera 668B, which can be used to perceive objects (e.g., pedestrians, warehouse vehicles, other robots, cross-traffic, or bicycles) entering the field of view from the periphery. In addition, any number of long-range cameras 668E (e.g., long-range stereo camera pairs) can be used for depth-based object detection, especially for objects for which a neural network has not been trained. The long-range cameras 668E can also be used for object detection and classification, as well as basic object tracking.
[0119] Any number of stereo cameras 668A can also be included in front and / or other (e.g., rear) configurations. In at least one embodiment, one or more stereo cameras 668A can include an integrated control unit that includes a scalable processing unit that can provide programmable logic (“FPGA”) and a multi-core microprocessor with integrated controller area network (“CAN”) or Ethernet interfaces on a single chip. Such a unit can be used to generate 3D maps of the machine 600 environment, including distance estimates for points in the image (e.g., disparity or depth images). Alternative stereo cameras 668A can include compact stereo vision sensors that can include two camera lenses (one on the left and one on the right) and an image processing chip that can measure distances from the vehicle to target objects and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 668A can be used in addition to or as an alternative to the stereo cameras described herein. For example, in some embodiments, stereo depth estimation can be performed using other cameras than stereo cameras (e.g., two monocular cameras with at least partially overlapping fields of view).
[0120] Cameras with fields of view that include portions of the environment to the sides of the machine 600 (e.g., side-view cameras) can be used, e.g., for surround view, to provide information for creating and updating the occupancy grid, and to generate side collision warnings and / or to indicate, e.g., side- present objects, features, and / or people to the AMR 600B or humanoid robot 600C. For example, surround cameras 668D can be provided on the machine 600. The surround cameras 668D can include wide-view cameras 668B, fisheye cameras, 360-degree cameras, etc. For example, four fisheye cameras can be provided on the front, back, and sides of the machine 600. In alternative arrangements, the machine 600 can use three surround cameras 668D (e.g., left, right, and back), and can utilize one or more other cameras (e.g., front-facing cameras) as a fourth surround view camera.
[0121] Cameras 668 with fields of view that include portions of the environment to the rear of the machine 600 (e.g., rear-view cameras) can be used to understand objects, features, personnel, and / or other information to the rear of the machine 600, e.g., for parking assistance, surround view, rear collision warnings, planning, control, and navigation determinations, and / or to create and update occupancy grids, BEV images representing the environment, height maps, etc. A wide variety of cameras 668 can be used, including but not limited to cameras that are also suitable for use as front-facing cameras (e.g., long-range and / or mid-range cameras 668E, stereo cameras 668A, infrared cameras 668C, etc.), rear-facing cameras, side-facing cameras, downward-facing cameras, upward-facing cameras, and / or similar cameras 668, as described herein.
[0122] Similarly, for LiDAR sensors 664, RADAR sensors 660, ultrasonic sensors 662, and / or other sensor modalities or types, the location and placement of the sensors and their corresponding fields of view or sensing fields can be determined based on the use case, implementation, or design of the particular machine 600.
[0123] For example, the machine 600 includes RADAR sensors 660, which can be used by the machine 600 for long-range object detection, even in darkness and / or adverse weather conditions. In embodiments, the RADAR functional safety level can be ASIL B. The RADAR sensors 660 can use CAN and / or the bus 602 (e.g., to transmit data generated by the RADAR sensors 660) for control and access to object tracking data, in some examples, raw data can be accessed using Ethernet. A variety of types of RADAR sensors can be used. For example, but not limited to, the RADAR sensors 660 can be suitable for front, rear, and side radar use. In some examples, pulsed Doppler RADAR sensors are used.
[0124] The RADAR sensors 660 can include different configurations, such as with narrow field of view long range, with wide field of view short range, short range side coverage, etc. In some examples, long range radar can be used for adaptive cruise control (ACC) functionality. Long range RADAR systems can provide a broad field of view implemented by two or more independent scans, such as over a 250 m range. The RADAR sensors 660 can help distinguish between static and moving objects, which can be used by ADAS systems for emergency brake assist and forward collision warning, and by robots for detecting dynamic objects in a variety of environments (e.g., in environments with low or no light). Long range RADAR sensors can include a single static multi-modal RADAR with multiple (e.g., six or more) fixed RADAR antennas and high speed CAN and FlexRay interfaces. In examples with six antennas, the central four antennas can create a focused beam pattern aimed at recording the machine’s 600 surroundings at higher speeds while minimizing peripheral interference (e.g., from traffic in adjacent lanes). The other two antennas can expand the field of view, enabling quick detection of objects entering or leaving the machine’s direct path (e.g., a lane).
[0125] A mid-range RADAR system can include, for example, a range of up to 660 m (front) or 80 m (rear), and a field of view of up to 42 degrees (front) or 150 degrees (rear). A short-range RADAR system can include, but is not limited to, RADAR sensors designed to be mounted on the side surfaces at both ends (e.g., rear bumper) such that two beams can be used to constantly monitor the blind spots behind and to the side of the machine 600 (e.g., vehicle, robot, etc.). Thus, the short-range RADAR system can be used in an ADAS system for blind spot detection and / or lane change assist.
[0126] The machine 600 can also include ultrasonic sensors 662. The ultrasonic sensors 662 can be located on the front, rear, and / or sides of the machine 600 and can be used to assist in near-field perception, such as for parking assist, collision avoidance (e.g., for robotic components), and / or to create and update occupancy grids, evidence grid maps (EGMs), height maps, BEV images, and / or other representations of objects and features in the machine 600 environment. A variety of ultrasonic sensors 662 can be used, and different ultrasonic sensors 662 can be used for different detection ranges (e.g., 2.5 m, 4 m). For example, the ultrasonic sensors 662 can operate at an ASIL B functional safety level.
[0127] The machine 600 can include LiDAR sensors 664. The LiDAR sensors 664 can be used for object and feature detection, pedestrian and other robot detection, emergency braking, collision avoidance, simultaneous localization and mapping (SLAM), free space detection, and / or other functions. In embodiments, the LiDAR sensors 664 can be at a functional safety level of ASIL B. In some examples, the machine 600 can include multiple LiDAR sensors 664 (e.g., two, four, six, etc.) which can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).
[0128] In some examples, the LiDAR sensors 664 can be capable of providing a list of objects and their distances for a 360-degree field of view. For example, a commercially available LiDAR sensor 664 can have a advertised range of about 600 m, a precision of 2 cm - 3 cm, and support for a 600 Mbps Ethernet connection. In some examples, one or more flush LiDAR sensors 664 can be used. In such examples, the LiDAR sensors 664 can be implemented as small devices that can be embedded into the front, rear, sides, top, and / or corners of the machine 600. In such examples, the LiDAR sensors 664 can provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, and can reach a range of 200 m even for low reflectivity objects. A front-facing LiDAR sensor 664 can be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0129] In some examples, LiDAR technology can also be used, such as 3D Flash LiDAR. 3D Flash LiDAR uses a laser flash as a transmission source, illuminating the vehicle’s surroundings up to about 200 m away. The Flash LiDAR unit includes a receiver that records the laser pulse transmission time and reflected light on each pixel, which in turn corresponds to the distance from the vehicle to the object. Flash LiDAR can allow for high-accuracy and distortion-free images of the surroundings to be generated every laser flash. In some examples, four Flash LiDAR sensors can be deployed, one on each side of the machine 600. Available 3D Flash LiDAR systems include solid-state 3D staring array LiDAR cameras with no moving parts other than a fan (e.g., non-scanning LiDAR devices). The Flash LiDAR device can use 5 nanosecond Class I (eye-safe) laser pulses per frame and can capture reflected laser light in the form of 3D range point clouds and co-registered intensity data. By using Flash LiDAR, and because Flash LiDAR is a solid-state device with no moving parts, the LiDAR sensor 664 can be less susceptible to motion blur, vibration, and / or shock.
[0130] Figure 6B FIG. 1 is an illustration of sensor and component locations of an example autonomous or semi-autonomous vehicle 100A (also referred to herein as “vehicle 100,” “my vehicle 100,” “my machine 100,” or “machine 100”) according to some embodiments of the present disclosure. While vehicle 100A is illustrated in the figure, this is not meant to be limiting, as similar components and / or sensors can be included on any other machine type without departing from the scope of the present disclosure. For example, similar sensors and / or components can be used on a vehicle, car, truck, bus, first responder vehicle, shuttle, electric or motorized bicycle, motorcycle, fire truck, police car, ambulance, watercraft, construction vehicle, underwater vehicle, robot (e.g., AMR, humanoid robot, robotic arm, end effector, forklift, etc.), drone, airplane, vehicle coupled to a trailer (e.g., a semi-trailer for towing cargo), and / or another type of vehicle or machine (e.g., driverless and / or capable of accommodating one or more passengers).
[0131] Figure 6Cis a block diagram of an example system architecture of a machine 600 (e.g., an autonomous or semi-autonomous vehicle 600A, an autonomous mobile robot (AMR) 600B, a humanoid robot 600C, and / or other types of machines) in accordance with some embodiments of the present disclosure. It should be understood that this and other arrangements described herein are set forth only as examples. Other arrangements, components, elements, features, and elements (e.g., machines, interfaces, functions, order, groupings of functions, etc.) can be used in addition to or instead of the arrangements illustrated, and some elements can be wholly omitted from some embodiments. Further, many of the arrangements, components, elements, features, etc. described herein are functional entities that can be implemented as discrete or distributed components or with other components, and can be implemented in any suitable combination and location (e.g., on a local device, vehicle, or edge machine, on-premise (e.g., locally-hosted server), remote location (e.g., in one or more computing or server devices in one or more data centers in the cloud), and / or other locations). The various functions described herein as being performed by an entity can be performed by hardware, firmware, and / or software. For example, various functions can be implemented using one or more processors (e.g., central processing units (CPUs), graphics processing units (GPUs), microprocessors, microcontrollers, embedded processors, digital signal processors (DSPs), image signal processors (ISPs), physical processing units (PPUs), field programmable gate arrays (FPGAs), accelerators (e.g., deep learning accelerators (DLAs), clusters of deep learning accelerators (XNNs), neural network accelerators (NNAs), and / or neural processing units (NPUs), programmable visual accelerators (PVAs), optical flow accelerators (OFAs), etc.), application specific integrated circuits (ASICs), data processing units (DPUs), quantum processors, etc.) executing instructions stored in memory. In some embodiments, the systems, methods, and processes described herein can be implemented using components, features, and / or functionality similar to those of the example machine 600 of Figures 6A-6E the example computing ecosystem 700 of Figure 7 the example generative language model system 800 of Figure 8 and / or the example computing device 900 of Figure 9 .
[0132] Figure 6CEach component, feature, and system of the machine 600 in FIG. 6 is shown connected by a bus 602 (or referred to as a “machine communication network 602” or simply “communication network 602”). The bus 602 can include a controller area network (CAN) data interface (or referred to herein as a “CAN bus”). The CAN can be a network within the machine 600 that is used to help control various features and functions of the machine 600, such as actuation of brakes, acceleration, braking, steering, windshield wipers, etc. The CAN bus can be configured to have tens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus can be read to find steering wheel angle, ground speed, engine revolutions per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus can comply with ASIL B standards. In some embodiments, in addition to or as an alternative to the CAN bus, the bus 602 can include FlexRay, embedded buses (e.g., SPI, I2C), Local Interconnect Link (LIN), NVLink by NVIDIA, USB (2.0, 3.0 and above), radio frequency (RF), Ethernet (e.g., 10BASE / 100BASE, 1000BASE, 10G, etc.), and / or other communication protocols or functions. Moreover, while a single line is used to represent the bus 602, this is not limiting. For example, there can be any number of buses 602, which can include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 602 can be used to perform different functions, and / or can be used for redundancy. For example, a first bus 602 can be used for collision avoidance functions, while a second bus 602 can be used for actuation control. In any example, each bus 602 can be in communication with any component of the machine 600, and two or more buses 602 can be in communication with the same component. In some examples, each SoC 604, each controller 636, and / or each computer or computing engine within the machine 600 can have access to the same input data (e.g., input from sensors of the machine 600) and can be connected to a common bus, such as a CAN bus.
[0133] The machine 600 can include components such as a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, a battery, side mirrors, and / or other components of a vehicle or machine. The machine 600 can include a propulsion system 650, such as an internal combustion engine, a hybrid electric power plant, an all-electric motor, a hydrogen fuel engine, and / or another propulsion system type. The propulsion system 650 can be connected to a transmission system of the machine 600, which can include a transmission to effectuate propulsion of the machine 600. The propulsion system 650 can be controlled in response to receiving a signal from a throttle / accelerator 652.
[0134] The steering system 654, which can include a steering wheel and / or other steering mechanism (e.g., remote steering and / or local steering), can be used to steer the machine 600 (e.g., along a desired path or route) while the propulsion system 650 is operating (e.g., when the vehicle is in motion). The steering system 654 can receive a signal from a steering actuator 656. In some embodiments, a steering wheel or other steering mechanism can not be included, for example, for machines 600 capable of fully automated (e.g., level 5) functionality.
[0135] The brake sensor system 646 can be used to operate vehicle brakes in response to receiving a signal from a brake actuator 648 and / or brake sensors.
[0136] The machine 600 can include one or more controllers 636, such as the controller 100 described herein with respect to FIG. 1, for example. The controller 636 can be configured to receive signals from the various sensors and / or actuators of the machine 600, and to control the operation of the machine 600 in response to the received signals. Figure 6AThe described controllers. The controllers 636 can be used for various functions and can be coupled to any of the various other components and systems of the machine 600. For example, the controllers 636 can be used to control the machine 600, artificial intelligence executing on the machine 600, infotainment of the machine 600, etc. For example, one controller 636 can be used for some or all of the functions, or different controllers 636 can be used for different functions, e.g., to ensure separation of availability and security between various controllers for different tasks. For example, the controllers 636 can use system-computed plans (e.g., paths or trajectories for the vehicle 600A or AMR 600B, or motions, component trajectories, motion positions or displacements for the joints or components (e.g., manipulators, end effectors, limbs, hands, fingers, legs, feet, etc.) of the humanoid robot 600C, etc.) to control the machine 600 in an environment. In some cases, the controllers 636 can include proportional-integral-derivative (PID) controllers, fuzzy logic controllers, neural controllers (e.g., controllers embodied as one or more neural networks), force control controllers, programmable logic controllers (PLCs), and / or other types of controllers. For example, in the humanoid robot 600C, the controllers 636 can act as the brain, responsible for analyzing sensor data, making decisions, and sending commands to actuators. The controllers 636 can include low-level controllers that handle basic motor control, ensuring accurate and precise motion of individual joints and actuators. The controllers 636 can include high-level controllers to coordinate multiple actuators and sensors, plan complex motions, and adapt to changing environments.
[0137] In embodiments, the controllers 636 can include artificial intelligence controllers that can use AI algorithms (e.g., DNNs, MLMs, etc.) to learn, make decisions, and autonomously perform tasks of the machine 600. In some embodiments, the controllers 636 can use fixed open-loop control algorithms and not adjust actions based on the environment. In other embodiments, closed-loop control can be used that incorporates feedback mechanisms to monitor the performance of the robot and make necessary adjustments. In examples, the controllers 636 can implement reactive control to directly respond to sensory inputs, enabling fast reflexive actions and real-time changes. Additionally, in some examples, deliberative control can be implemented that uses internal models and planning algorithms to generate high-level actions that can be suitable for complex tasks that require reasoning, decision making, and long-term planning.
[0138] The controllers 636 can include one or more system-on-chips (SoCs) 604 Figure 6C and Figure 6D), CPUs, GPUs, accelerators, etc., the controller 636 can provide signals (e.g., representative of commands or messages) to one or more components and / or systems of the machine 600. Although the controller 636 is listed separately from the SoC 604, this is not meant to be limiting, and in some embodiments one or more components of the SoC 604 can perform the operations of the controller 636. For example, the controller can send signals to operate machine brakes by one or more brake actuators 648, to operate a steering system 654 by one or more steering actuators 656, to operate a propulsion system 650 by one or more throttle / accelerator 652, etc. The controller 636 can include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operational commands (e.g., signals representative of commands) to enable autonomous or semi-autonomous navigation and movement and / or to assist a human operator using the machine 600. The controller 636 can include a first controller 636 for autonomous control and navigation functions, a second controller 636 for functional safety functions, a third controller 636 for artificial intelligence functions (e.g., computer vision), a fourth controller 636 for infotainment functions, a fifth controller 636 for redundancy in emergency situations, and / or other controllers. For example, hardware for safety monitoring and other safety functions (e.g., functional safety island) can be discrete or partitioned (physically or by processing separation) relative to hardware for processing sensor data for perception and making vehicle control decisions. Similarly, hardware (e.g., controllers, SOCs, etc.) for controlling in-vehicle infotainment and / or in-cabin monitoring can be separate or partitioned from hardware for vehicle perception and control. In some examples, a single controller 636 can handle two or more of the functions described above, two or more controllers 636 can handle a single function, and / or any combination thereof.
[0139] The controller(s) 636 can provide signals to control one or more components and / or systems of the machine 600 in response to sensor data (e.g., sensor inputs) received from one or more sensors. The sensor data can be received from, for example and without limitation, a global navigation satellite system (“GNSS”) sensor 658 (e.g., a global positioning system sensor), a RADAR sensor 660, an ultrasonic sensor 662, a LiDAR sensor 664, an inertial measurement unit (IMU) sensor 666 (e.g., an accelerometer, a gyroscope, a magnetic compass, a magnetometer, etc.), a microphone 696, a camera 668 (e.g., a stereo camera 668A, a wide-angle camera 668B (e.g., a fisheye camera), an infrared camera 668C, a surround camera 668D (e.g., a 360-degree camera), a long-range and / or mid-range camera 668E, and / or other types of cameras), a speed sensor 644 (e.g., to measure a speed of the machine 600), a vibration sensor 642, a steering sensor 640, a brake sensor (e.g., as part of a brake sensor system 646), an actuator, and / or other types of sensors.
[0140] The controller(s) 636 can receive inputs (e.g., represented by input data) from an instrument cluster 632 of the machine 600 and provide outputs (e.g., represented by output data, display data, etc.) through a human-machine interface (HMI) display 634 (e.g., a screen, a heads-up display, a mirror display, a face display, a robotic display, etc.), a sound alarm, a loudspeaker, a speaker, and / or through other components of the machine 600. The outputs can include information such as a machine speed, a velocity, a time, a map 622 (e.g., map data corresponding to a navigation map, a standard definition (SD) map, a high definition (“HD”) map, etc.), a location of the machine 600 (e.g., a location on the map 622), a direction, a location of other vehicles (e.g., an occupancy map, a height map, a bird’s eye view (BEV) image, a grid, etc.), information about objects and object states perceived by the system, system state information, etc. Figure 6C The HMI display 634 can display information about the presence of one or more objects (e.g., a street sign, a warning sign, a traffic signal change, etc.), and / or information about a driving maneuver that the vehicle has made, is making, or will make (e.g., now changing lanes, taking exit 34B in two miles, etc.).
[0141] The machine 600 can include one or more system-on-chips (SoCs) 604 (in Figure 6DSoC 604 can include CPU 606, GPU 608, processor 610, cache 612, accelerator 614, data store 616, and / or other components and features. SoC 604 can be used to process and provide data for various operations of the machine 600, such as navigation, planning, reasoning, inference, perception, control, and / or actuation operations in various platforms and systems. For example, SoC 604 can process live perception data (e.g., from cameras, LiDAR, RADAR, ultrasound, etc.) as well as map data corresponding to one or more maps 622 (e.g., HD map, SD map, navigation map, occupancy map, etc.) in order to conduct or assist in performing various operations of the machine 600. When using maps and / or AI (e.g., model parameter updates, fine-tuning, etc.), the maps and / or AI are refreshed and / or updated from one or more servers (e.g., servers 678) (e.g., one or more servers of a cloud-based data center) via network interface 624. Figure 6E
[0142] Although SoC 604 is shown in Figures 6A-6E , additional or alternative components and / or architectures can be used, such as a multi-chip module (MCM), an application-specific integrated circuit (ASIC), a system-in-a-package (SiP), a field-programmable gate array (FPGA), a heterogeneous integration (HI), a single-board computer (SBC), without departing from the scope of the present disclosure. For example, depending on the type of machine 600, the use of machine 600, the model of machine 600, and the capabilities required of machine 600, one or more SoCs 604 and / or alternative architectures and / or components can be used to satisfy a particular implementation.
[0143] Machine 600 can include CPU 618 (e.g., a discrete CPU or dCPU) that can be coupled to SoC 604 via a high-speed interconnect (e.g., PCIe). CPU 618 can include, for example, an X86 processor. CPU 618 can be used to perform any of a variety of functions, including arbitrating potentially inconsistent results between ADAS sensors and SoC 604, and / or monitoring the status and health of controller 636 and / or infotainment SoC 630.
[0144] Machine 600 can include GPU 620 (e.g., a discrete GPU or dGPU) that can be coupled to SoC 604 via a high-speed interconnect (e.g., NVIDIA’s NVLink). GPU 620 can provide additional artificial intelligence functionality, such as by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based on input (e.g., sensor data) from sensors of machine 600.
[0145] The machine 600 can also include a network interface 624, which can include one or more wireless antennas 626 and / or modems (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). The network interface 624 can be used to implement wireless connections over the Internet with the cloud (e.g., with the server 678 and / or other network devices), with other vehicles and / or computing devices (e.g., a client device of a passenger). For communication with other vehicles, a direct link can be established between two vehicles and / or an indirect link established (e.g., across a network and over the Internet). The direct link can be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide the machine 600 with information about vehicles in the vicinity of the machine 600 (e.g., vehicles in front of, to the side of, and / or behind the machine 600). This functionality can be part of a cooperative adaptive cruise control functionality of the machine 600.
[0146] The network interface 624 can include a SoC that provides modulation and demodulation functionality and enables the controller 636 to communicate over wireless networks. The network interface 624 can include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. The frequency conversion can be performed through well-known processes and / or can be performed using a superheterodyne process. In some examples, the radio frequency front end functionality can be provided by a separate chip. For example, the network interface 624 can be capable of communicating over Long-Term Evolution (“LTE”), Wideband Code-Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile communications (“GSM”), IMT-CDMA Multi-Carrier (“CDMA2000”), fifth generation mobile communication technology (5G), sixth generation mobile communication technology (6G), and / or other cellular and / or wireless communication standards. The wireless antennas 626 can also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area networks (e.g., Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, etc.) and / or low power wide area networks (“LPWAN”) (e.g., LoRaWAN, SigFox, etc.).
[0147] The machine 600 can also include a data storage 628, which can include storage off-chip (e.g., off of the SoC 604). The data storage 628 can include one or more storage elements, including RAM, SRAM, DRAM, VRAM, Flash, hard disks, and / or other components and / or devices that can store at least one bit of data.
[0148] The machine 600 can also include a GNSS sensor 658. The GNSS sensor 658 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) is used to assist mapping, perception, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 658 can be used, such as but not limited to using a GPS with an Ethernet-to-serial (RS-232) bridge using a USB connector.
[0149] The machine 600 can also include an IMU sensor 666. In some examples, the IMU sensor 666 can be located at the center of the rear axle of the machine 600. The IMU sensor 666 can include, for example and without limitation, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, such as in six-axis applications, the IMU sensor 666 can include an accelerometer and a gyroscope, while in nine-axis applications, the IMU sensor 666 can include an accelerometer, a gyroscope, and a magnetometer.
[0150] In some embodiments, the IMU sensor 666 can be implemented as a micro-electro-mechanical systems (MEMS) based, high-performance GPS-aided inertial navigation system (GPS / INS) that combines MEMS inertial sensors, high-sensitivity GPS receivers, and advanced Kalman filtering algorithms to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 666 can enable the machine 600 to estimate heading by directly observing and correlating changes in velocity from the GPS to the IMU sensor 666 without input from a magnetic sensor. In some examples, the IMU sensor 666 and the GNSS sensor 658 can be combined in a single integrated unit.
[0151] The vehicle can include one or more microphones 696 placed within and / or around the machine 600. The microphones 696 can be used for emergency vehicle detection and identification, among others.
[0152] The machine 600 can also include a vibration sensor 642. The vibration sensor 642 can measure vibrations of a machine component, such as an arm or leg of a humanoid robot 600C, or an axle of a vehicle 600A or AMR 600B. For example, changes in vibration can indicate changes in road, walking, or traversable surfaces. In another example, when two or more vibration sensors 642 are used, differences between the vibrations can be used to determine the friction or slip of a surface (e.g., when the vibration difference is between an electrically driven axle and a freely rotating axle).
[0153] The machine 600 can include an ADAS system 638, for example when the machine 600 is a vehicle 600A. In some examples, the ADAS system 638 can include a dedicated SoC. The ADAS system 638 can include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision or collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), blind spot monitoring (BSM), rear cross traffic warning (RCTW), pedestrian detection, driver monitoring, collision warning system (CWS), traffic sign recognition, speed limit detection, automatic parking, lane centering (LC), high beam safety system, and / or other features and functionality.
[0154] The machine 600 can also include an infotainment SoC 630 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as a SoC, the infotainment system can not be a SoC and can include one or more discrete components, such as a multi-chip module (MCM), an application-specific integrated circuit (ASIC), a system-in-a-package (SiP), a heterogeneous integrated (HI), a single-board computer (SBC), and / or the like. The infotainment SoC 630 can include a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, and / or the like), video (e.g., television, movies, streaming media, and / or the like), telephony (e.g., hands-free calling), network connectivity (e.g., wireless, Wi-Fi, and / or the like), and / or information services (e.g., navigation systems, rear park assist, radio data system, vehicle related information (e.g., fuel level, total distance traveled, brake fluid level, oil level, doors open / close, air filter information, and / or the like)) to the machine 600. For example, the infotainment SoC 630 can be a radio, a disc player, a navigation system, a video player, USB and Bluetooth connectivity, an in-car computer, in-car entertainment, Wi-Fi, steering wheel audio controls, hands-free voice controls, a heads-up display (HUD), the HMI display 634, a telematics device, a control panel (e.g., to control and / or interact with various components, features, and / or systems), and / or other components. The infotainment SoC 630 can also be used to provide information (e.g., visually and / or audibly) to a user of the vehicle, such as information from the ADAS system 638, autonomous driving information (e.g., planned vehicle maneuvers, trajectories), surrounding environment information (e.g., intersection information, vehicle information, road information, and / or the like), and / or other information.
[0155] The infotainment SoC 630 can include GPU functionality. The infotainment SoC 630 can communicate with other devices, systems, and / or components of the machine 600 over the bus 602 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 630 can be coupled to a supervisory MCU such that the GPU of the infotainment system can perform some autonomous driving functionality in the event of a failure of the main controller 636 (e.g., the main computer and / or backup computer of the machine 600). In such examples, the infotainment SoC 630 can place the machine 600 into a driver safe stop mode, as described herein.
[0156] In some embodiments, the infotainment system can provide a digital or virtual assistant, which can be voice-only, or can have a visual component (e.g., in the form of a digital human or avatar). The assistant can provide basic functionality, such as texting, adjusting vehicle settings, music or video control, navigation functions, etc., and / or can provide more advanced functionality, such as supported by one or more language models (e.g., large language models (LLMs), visual language models (VLMs), multi-modal language models (MMLMs), etc.). For example, the driver and / or occupants can interact with the assistant in a manner similar to how a user interacts with a language model, such as asking general questions, specific questions, requesting restaurants, gas stations, and / or other recommendations and / or locations, understanding vehicle functionality or troubleshooting (e.g., asking for tire pressure information, oil change information, battery swap information, etc.). Thus, the machine 600 (whether a vehicle 600A, AMR 600B, humanoid robot 600C, or other type of machine) can include a locally stored language model and / or communicate with a remotely hosted language model (e.g., via one or more APIs) to provide more detailed and in-depth communication functionality to users of the machine 600.
[0157] In some examples, the infotainment SoC 630, SoC 604, and / or another SoC or computing / processing system can perform driver and / or occupant monitoring within the vehicle cabin. For example, the computing system can perform facial recognition, and vehicle owner identification can use data from cameras and / or other sensors to identify the presence of an authorized driver and / or owner of the machine 600. An always-on sensor processing engine can be used to unlock the vehicle and turn on the headlights when the owner approaches the driver’s door, and disable the vehicle in a safe mode when the owner leaves the vehicle. In this way, the SoC 604 can provide security against theft and / or carjacking.
[0158] In some embodiments, the in-cabin monitoring camera sensors can be monitored using one or more neural networks running on another or a dedicated SoC (e.g., an in-vehicle infotainment or in-vehicle monitoring SoC) configured to identify in-cabin events and respond accordingly. The in-cabin system can activate cellular service and place a call by reading lips, dictate emails, change a vehicle’s destination, activate or change the vehicle’s infotainment system and settings, or provide voice-activated web browsing. The in-cabin system can also include one or more in-cabin AI agents or assistants that can interact with one or more LLMs, VLMs, MMLMs, etc. of the cloud using one or more APIs or plugins. For example, the in-cabin AI agents or assistants can provide directions, vehicle or machine feedback information, answer general questions, handle music / video and / or other requests, activate windows, doors, and / or other vehicle components, etc. Thus, one or more dedicated SoCs and / or processor groups can be used to perform in-cabin infotainment and / or in-cabin monitoring (e.g., as an occupant monitoring system (OMS)) of the machine 600.
[0159] The machine 600 can also include an instrument cluster 632 (e.g., a digital instrument cluster, an electronic instrument cluster, a digital dashboard, etc.). The instrument cluster 632 can include a controller and / or supercomputer (e.g., a discrete controller or supercomputer). The instrument cluster 632 can include a set of instruments, such as a speedometer, fuel level, oil pressure, tachometer, odometer, steering indicator, gearshift position indicator, seatbelt warning light, parking brake warning light, engine malfunction light, supplemental restraint system (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information can be displayed and / or shared between the infotainment SoC 630 and the instrument cluster 632. In other words, the instrument cluster 632 can be included as part of the infotainment SoC 630, and vice versa.
[0160] Figure 6D is a block diagram of an example architecture of a computing system (with respect to Figure 6C a subset of the systems described). While illustrated as an SoC 604, this is not meant to be limiting, and the computing system can additionally or instead include a multi-chip module (MCM), an application-specific integrated circuit (ASIC), a system-in-a-package (SiP), a heterogeneous integrated (HI), a single-board computer (SBC), and / or other components and / or architectures without departing from the scope of the present disclosure.
[0161] SoC 604 can be an end-to-end platform with flexible architecture that spans automation levels 2-5, or SoC 604 can be specifically designed for a particular automation level (e.g., a first SoC 604 for level 2 to level 2++, a second SoC 604 for level 3, a third SoC 604 for level 4, etc.), thereby providing a comprehensive functional safety architecture that leverages and efficiently uses computer vision, neural network inference, robotic planning, control and navigation, ADAS technology, etc., with diversity and redundancy to provide a flexible, reliable platform for driving or robotic control software stacks, as well as deep learning tools. SoC 604 can be faster, more reliable, and even more energy efficient and space saving than traditional systems. For example, when accelerator 614 is used in conjunction with CPU 606, GPU 608, and data store 616, it can provide a fast, efficient platform for level 2-5 autonomous vehicles, as well as for AMR 600B, humanoid robot 600C, and / or other robots or machines of the type for safe planning, navigation, and control.
[0162] In some embodiments, for example, SoC 604 includes: a GPU 608 with 2000 or more cores (e.g., 2048 cores), 60 or more tensor cores (e.g., 64 tensor cores), and a GPU maximum frequency over 1 GHz (e.g., 1.3 GHz); a CPU 606 including 10 or more cores (e.g., 12 cores) with 64-bit, 3 MB L2, and 6 MB L3 cache, and a maximum frequency of 2 GHz or more (e.g., 2.2 GHz); one or more deep learning accelerators (DLAs); a deep learning accelerator cluster (XNN); a neural network accelerator (NNA) or neural processing unit (NPU) 609 (e.g., 2 DLAs / XNN / NNA / NPUs 609); and a vision accelerator (e.g., programmable vision accelerator (PVA) 607), a single SoC 604 can be capable of achieving 275 trillion operations per second (TOPS) of AI performance. For example, NVIDIA’s Jetson AGX Orin 64GB SoC meets these criteria and achieves such performance.
[0163] Similarly, in embodiments, the SoC 604 includes: a GPU 608 with 1700 or more cores (e.g., 1792 cores), 50 or more tensor cores (e.g., 56 tensor cores), and a GPU maximum frequency over 900 MHz (e.g., 930 MHz); a CPU 606 including 8 or more cores (e.g., 8 cores), with 64-bit, 2 MB L2, and 4 MB L3 cache memory, and a maximum frequency of 2 GHz or more (e.g., 2.2 GHz); one or more deep learning accelerators (DLAs); a deep learning accelerator cluster (XNN); a neural network accelerator (NNA), or a neural processing unit (NPU) 609 (e.g., 2 DLAs / XNN / NNA / NPUs 609), and a visual accelerator (e.g., a programmable visual accelerator (PVA) 607), a single SoC 604 can be capable of achieving an AI performance of 200 trillion operations per second (TOPS). For example, NVIDIA’s Jetson AGX Orin 32GB SoC meets these criteria and achieves such performance.
[0164] In some embodiments, for example, the SoC 604 includes: a GPU 608 with 1000 or more cores (e.g., 1024 cores), 28 or more tensor cores (e.g., 32 tensor cores), and a GPU maximum frequency over 900 MHz (e.g., 1173 MHz); a CPU 606 including 8 or more cores (e.g., 8 cores), with 64-bit, 2 MB L2, and 4 MB L3 cache memory, and a maximum frequency of 2 GHz or more (e.g., 2 GHz); one or more deep learning accelerators (DLAs); a deep learning accelerator cluster (XNN); a neural network accelerator (NNA), or a neural processing unit (NPU) 609 (e.g., 1 DLA / XNN / NNA / NPU 609), and a visual accelerator (e.g., a programmable visual accelerator (PVA) 607), a single SoC 604 can be capable of achieving an AI performance of 157 trillion operations per second (TOPS). For example, NVIDIA’s Jetson AGX Orin NX 16GB SoC meets these criteria and achieves such performance.
[0165] In various embodiments, for example, the SoC 604 includes a GPU 608 with 1000 or more cores (e.g., 1024 cores), 28 or more tensor cores (e.g., 32 tensor cores), and a GPU maximum frequency exceeding 900 MHz (e.g., 1020 MHz); a CPU 606 including 6 or more cores (e.g., 6 cores), with 64-bit, 1.5 MB L2, and 4 MB L3 cache memory, and a maximum frequency of 1.5 GHz or more (e.g., 1.7 GHz), a single SoC 604 can achieve an AI performance of 67 trillion operations per second (TOPS). For example, NVIDIA’s Jetson Orin Nano 8GB SoC meets these criteria and achieves such performance.
[0166] The SoC 604 can include one or more CPUs 606. In embodiments, the CPU 606 can include a CPU cluster or CPU complex (also referred to herein as a “CCPLEX”). The CPU 606 can include multiple cores and / or caches (e.g., L2, L3). For example, in some embodiments, the CPU 606 can include twelve cores in a coherent multi-processor configuration. In some embodiments, the CPU 606 can include four dual-core clusters, with each cluster having a dedicated L2 cache (e.g., 3 MB L2 cache). The CPU 606 (e.g., CCPLEX) can be configured to support simultaneous cluster operation, such that any combination of clusters of the CPU 606 are active at any given time.
[0167] The SoC 604 can include any type and number of GPUs 608. For example, an integrated GPU (also referred to herein as an “iGPU”) can be used in some embodiments. The GPU 608 can be programmable and can be efficiently used for parallel workloads. In some examples, the GPU 608 can use an enhanced tensor instruction set. The GPU 608 can include one or more streaming microprocessors, where each streaming microprocessor can include a cache (e.g., an L1 cache with at least 96 KB of storage capacity), and two or more streaming microprocessors can share an L2 cache (e.g., an L2 cache with 512 KB of storage capacity). In some embodiments, the GPU 608 can include at least eight streaming microprocessors. The GPU 608 can use a compute application programming interface (API). Further, the GPU 608 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA’s CUDA).
[0168] The GPU 608 can be power-optimized for best performance in cars, robots, and / or other embedded use cases. For example, the GPU 608 can be fabricated on a fin- field effect transistor (FinFET). However, this is not intended to be limiting, and the GPU 608 can be fabricated using other semiconductor fabrication or manufacturing processes. Each streaming microprocessor can contain multiple mixed-precision processing cores divided into multiple blocks. For example, but not by way of limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA TENSOR COREs for deep learning matrix algorithms, one (e.g., L0) instruction cache, one thread warp scheduler, one dispatch unit, and / or one (e.g., 64 KB) register file. Additionally, the streaming microprocessor can include independent parallel integer and floating-point data paths for efficient execution of mixed compute and addressing workloads. The streaming microprocessor can include independent thread scheduling functionality to enable finer-grain synchronization and cooperation between parallel threads. The streaming microprocessor can include a combined LI data cache and shared memory unit to improve performance while simplifying programming.
[0169] The GPU 608 can include a high bandwidth memory (HBM) and / or (e.g., 16 GB) HBM2 memory subsystem to provide, in some examples, a peak memory bandwidth of approximately 900 GB / sec. In some examples, in addition to or instead of HBM memory, a synchronous graphics random access memory (SGRAM) such as a graphics double data rate fifth generation synchronous random-access memory (GDDR5) can be used.
[0170] The GPU 608 can include a unified memory technology, including an access counter, to allow more accurate migration of memory pages to the processor that accesses them most frequently, improving efficiency of memory ranges shared between processors. In some examples, address translation services (ATS) support can be used to allow the GPU 608 to directly access CPU 606 page tables. In such examples, when a GPU 608 memory management unit (MMU) miss occurs, an address translation request can be transmitted to the CPU 606. In response, the CPU 606 can look up a virtual-to-physical mapping for the address in its page tables and transmit the translation back to the GPU 608. Thus, the unified memory technology can allow the memory of the CPU 606 and the GPU 608 to use a single unified virtual address space, simplifying GPU 608 programming and porting of applications to the GPU 608.
[0171] SoC 604 can include any number of caches 612, including the caches described herein. For example, caches 612 can include L0 caches, LI caches, L2 caches, L3 caches (e.g., usable by CPU 606 and GPU 608 (e.g., connected to CPU 606 and GPU 608)), etc. Caches 612 can include write-back caches that can track the state of lines, such as by using one or more cache coherency protocols (e.g., MEI, MESI, MSI, etc.). According to embodiments, (e.g., L3) caches can include 4MB or more, although smaller or larger cache sizes can be used.
[0172] SoC 604 can include one or more arithmetic logic units (ALUs) 665 that can be used to perform processing related to various tasks or operations of machine 600 (e.g., computer vision, machine learning or deep learning processing, world model management, etc.). In addition, SoC 604 can include floating point units (FPUs) 667 or other mathematical co-processor or digital co-processor types for performing mathematical operations within the system. For example, SoC 604 can include one or more FPUs 667 integrated as execution units within CPU 606 and / or GPU 608.
[0173] SoC 604 can include one or more accelerators 614 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, SoC 604 can include a hardware acceleration cluster that can include optimized hardware accelerators and / or large on-chip memory. Large on-chip memory 615 (e.g., 4MB SRAM, 32GB and / or 64GB 256-bit LPDDR5 (204.8GB / s), 8GB and / or 16GB 128-bit LPDDR5 (102.4GB / s), and / or other memory types and sizes) can enable the hardware acceleration cluster to accelerate neural network processing, transformer processing, optical flow processing, vision processing, and / or other computations or processing. The hardware acceleration cluster can be used to supplement GPU 608 and offload some tasks of GPU 608 (e.g., freeing up more cycles of GPU 608 to perform other tasks). As an example, accelerators 614 can be used for target workloads (e.g., perception, convolutional neural networks (CNNs), deep neural networks (DNNs), language models (LLMs, VLMs, MMLMs, VLAs, etc.), transformer models, diffusion models, encoder-only models, encoder-decoder models, etc.) that are stable enough to be suitable for acceleration.
[0174] The accelerator(s) 614 (e.g., hardware acceleration cluster) can include a deep learning accelerator (DLA) 609 (also referred to herein as “deep learning accelerator cluster (XNN) 609,” “neural network accelerator (NNA) 609,” or “neural processing unit (NPU) 609”). The DLA 609 can include one or more tensor processing units (TPUs) 641 that can be configured to provide additional, e.g., deep learning application and inference tera operations per second. The TPU 641 can be an accelerator configured to perform and optimized for image processing functions (e.g., for CNNs, RCNNs, DNNs, etc.). The DLA 609 can be further optimized for a specific set of neural network types and floating point operations and inference. The design of the DLA can provide higher performance per mm than general purpose GPUs and significantly outperform CPUs. The TPU 641 can perform several functions including single instance convolution functions, support for INT8, INT16, and FP16 data types for features and weights, e.g., and post-processor functions. Although the TPU 641 is described as being included as part of the DLA 609, this is not intended to be limiting and the TPU 641 can be included in additional or alternative accelerators 614 and / or other components, and / or can be included as a discrete processing component.
[0175] The DLA 609 can quickly and efficiently execute neural networks on processed or unprocessed data to implement various functions including, but not limited to: object and feature recognition and detection using data from one or more sensor modalities (e.g., vehicles, pedestrians, other robots, lane lines, road boundary lines, debris, potholes, boxes, warehouse items, etc.); distance estimation using data from one or more sensor modalities; emergency vehicle detection and identification using data from microphones and / or vision-based sensors; facial recognition; pick and place operations; manipulation operations; occupant monitoring; vehicle owner identification; and / or other in-vehicle operations using data from in-vehicle cameras and / or other sensor types; and / or safety and / or safety-related events, to name a few.
[0176] The DLA 609 can perform any of the functions of the GPU 608 and by using an inference accelerator, e.g., the designer can anchor the DLA 609 or GPU 608 for any function. For example, the designer can concentrate the processing of DNNs and floating point operations on the DLA 609, while leaving other functions to the GPU 608 and / or other accelerators 614. The DLA 609 can be used to run any type of network to enhance control and safety, including, e.g., a neural network that outputs a confidence metric for each object detection.
[0177] The accelerator 614 (e.g., hardware acceleration cluster) can include a programmable vision accelerator (PVA) 607, which can alternatively be referred to herein as a computer vision accelerator or generally as a vision accelerator. The PVA 607 can be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), semi-autonomous driving, autonomous driving, robotics applications, security and surveillance applications, augmented reality (AR), virtual reality (VR), and / or mixed reality (MR) applications, etc. The PVA 607 can provide a balance between performance and flexibility. For example, each PVA 607 can include, for example and without limitation, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA) systems, pixel processing engines (PPE), vector processors or vector processing units (VPU), and / or other components. The PVA engine can include an advanced very large instruction word (VLIW), single instruction multiple data (SIMD) digital signal processor. The PVA 607 can be optimized for image processing and computer vision algorithm acceleration tasks. For example, the PVA 607 provides excellent performance with very low power consumption and can be used asynchronously and concurrently with the CPU 606, GPU 608, and / or other accelerators in the system (e.g., vehicle, robot, etc.) as part of a heterogeneous computing pipeline.
[0178] The PVA 607 can include one or more (e.g., two) vector processing subsystems (VPSs), where each VPS can include one or more vector processing unit (VPU) cores, one or more decoupled lookup units (DLUTs), one or more shared or vector memories (VMEMs), and one or more instruction caches (I-caches). The VPU cores can be the main processing units and can include vector SIMD VLIW DSPs 643 optimized for computer vision. The VPU cores can fetch instructions through the I-cache and can access data through the VMEM. The DLUTs can include specialized hardware components that enhance the efficiency of parallel lookup operations. For example, the DLUTs allow parallel lookups using a single copy of a lookup table by performing the lookups in a decoupled pipeline that is independent of the main processor pipeline. By doing so, the DLUTs can minimize or reduce memory usage and increase throughput while avoiding data-dependent memory bank conflicts, ultimately leading to improved overall system performance. The VPU VMEMs can provide local data storage for the VPU, allowing for efficient implementation of various image processing and computer vision algorithms. The VPU VMEMs can support access from external VPS hosts, such as direct memory access (DMA) and the CPU 606 (e.g., an ARM Cortex-R5 processor), facilitating data exchange with the CPU 606 and other system-level components. The VPU I-caches can provide instruction data to the VPU on request, can request missing instruction data from system memory, and / or can maintain temporary instruction storage for the VPU. For each VPU task, the CPU 606 can configure the DMA system, optionally prefetch VPU programs into the VPU I-cache, and / or initiate each VPU-DMA pair to process the task. The PVA 607 can also include an L2 SRAM memory that will be shared between one or more (e.g., two) sets of VPSs and DMA. In some embodiments, one or more (e.g., two) DMA devices are used to move data between external memory, PVA L2 memory, VMEM (e.g., one in each VPS), CPU tightly coupled memory (TCM), DMA descriptor memory, and / or PVA-level configuration registers. In light-load systems, two parallel DMA accesses to DRAM can achieve read / write bandwidths of up to 15 GB / s, while in heavy-load systems, this bandwidth can reach up to 10 GB / s. In terms of compute capacity, INT8 gigamultiply-accumulate operations (GMACs) per second can be 2048 or greater, not including the DLUTs. FP32 GMACs can include 32 per PVA instance.
[0179] The RISC core can interact with image sensors (e.g., image sensors of any of the cameras described herein), image signal processors, and the like. Each RISC core can include any number of memories. The RISC core can use any of a variety of protocols, depending on the embodiment. In some examples, the RISC core can execute a real-time operating system (RTOS). The RISC core can be implemented using one or more integrated circuit devices, application specific integrated circuits (ASICs), and / or memory devices. For example, the RISC core can include an instruction cache and / or a tightly coupled RAM.
[0180] The DMA system can enable components of the PVA 607 to access system memory independently of the CPU 606. The DMA can support any number of functions for providing optimization for the PVA 607, including but not limited to supporting multi-dimensional addressing and / or circular addressing. In some examples, the DMA can support addressing in up to six or more dimensions, which can include block width, block height, block depth, horizontal block stride, vertical block stride, and / or depth stride.
[0181] A vector processor or VPU can be a programmable processor that can be designed to efficiently and flexibly execute programming of computer vision algorithms and provide signal processing capabilities. In some examples, the PVA 607 can include a PVA core and two vector processing subsystem partitions. The PVA core can include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripherals. The vector processing subsystems can operate as the main processing engines of the PVA 607 and can include one or more vector processing units (VPUs), one or more pixel processing engines (PPEs) (which can include a 2D layout of interconnected (e.g., north, south, east, west intercommunicating) processing elements), one or more instruction caches, and / or one or more shared or vector memories (e.g., VMEM). The VPU core can include a digital signal processor, such as a single instruction, multiple data (SIMD), very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can improve throughput and speed.
[0182] In some embodiments, each vector processor can include an instruction cache and can be coupled to a dedicated memory. Thus, in some examples, each vector processor can be configured to execute independently of other vector processors. In other examples, the vector processors contained in a particular PVA 607 can be configured to employ data parallelism. For example, in some embodiments, multiple vector processors contained in a single PVA 607 can execute the same computer vision algorithm but for different regions of an image. In other examples, the vector processors contained in a particular PVA 607 can simultaneously execute different computer vision algorithms on the same image, or even different algorithms on successive images or portions of images. Among other things, any number of PVAs 607 can be included in a hardware acceleration cluster, and any number of vector processors can be included in each PVA. Moreover, the PVAs 607 can include additional error-correcting code (ECC) memory to enhance overall system security.
[0183] The accelerator 614 (e.g., hardware acceleration cluster) has broad use for autonomous and semi-autonomous machine control. The PVA 607 can be a programmable vision accelerator that can be used for key processing stages in perception, robotic understanding and reasoning, ADAS, semi-autonomous and autonomous vehicles, etc. The functionality of the PVA 607 is well suited for algorithm domains that require predictable processing, with low power consumption and low latency. In other words, the PVA 607 performs well on semi-dense or dense rule computations, even on small data sets that require predictable run-time, low latency, and low power consumption. Thus, in the context of autonomous vehicle and robotic platforms, the PVA 607 is designed to run classical computer vision algorithms as they are very effective at object detection and integer math operations.
[0184] For example, according to one embodiment of the technology, the PVA 607 is used to perform computer stereo vision. A semi-global matching based algorithm can be used in some examples, although this is not intended to be limiting. Many level 3-5 autonomous driving applications require motion estimation / stereo matching to be performed on the fly (e.g., motion structure, pedestrian recognition, lane detection, etc.). The PVA 607 can perform computer stereo vision functions on inputs from two monocular cameras.
[0185] In some examples, the PVA 607 can be used to perform dense optical flow. Processed RADAR is provided according to processing raw RADAR data (e.g., using a 4D fast Fourier transform). In other examples, the PVA 607 is used for time-of-flight depth processing, e.g., by processing raw time-of-flight data to provide processed time-of-flight data.
[0186] While the VPU, DMA, RISC Core, VMEM, and decoupled co-processor (e.g., DLUT) are described as included within the PVA 607, this is not meant to be limiting. In some embodiments, these components can be included in alternative or additional processing components and / or accelerators 614, and / or can be included as discrete components of the SoC 604 and / or other computing system architectures.
[0187] In some examples, the SoC 604 can include a real-time ray tracing hardware accelerator (RTA) 651, which can be used to quickly and efficiently determine the location and extent of objects (e.g., within a world model) to generate real-time or near real-time visualizations simulations for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for simulation of SONAR, RADAR, LiDAR, camera, and / or other sensor modalities in simulations, for general wave propagation simulations, for comparison with LiDAR data for localization, to generate ground truth training data for training neural networks, and / or other functions and uses. In some embodiments, one or more tree traversal units (TTUs) can be used to perform one or more ray tracing related operations. For example, the machine 600 (or another machine or device) can perform a simulation in a simulated environment, and can use one or more light transport simulation algorithms (e.g., ray tracing, path tracing, etc.) to generate the simulated environment. Thus, the ray tracing accelerator 651 and / or a ray tracing optimized GPU 608 (e.g., NVIDIA’s RTX GPU) can be used to accelerate these ray tracing algorithms.
[0188] The accelerators 614 (e.g., in a hardware acceleration cluster) can include one or more optical flow accelerators (OFAs) 611. For example, the OFAs 611 can be used to compute optical flow and stereo disparity between frames of sensor data (e.g., images). The optical flow can be accelerated on the OFAs 611 for uses such as object detection and tracking, and / or for stereo depth estimation, where stereo disparity between stereo image frames (e.g., two or more frames captured using two or more image sensors with at least partially overlapping fields of view) is computed.
[0189] The SoC 604 can include one or more camera serial interfaces (CSIs) 623. For example, the CSI 623 can include a Mobile Industry Processor Interface (MIPI) camera serial interface (CSI) for receiving video and input from a camera, a high-speed interface, and / or a video input block that can be used for camera and related pixel input functionality. The SoC 604 can also include an input / output controller that can be controlled by software and can be used to receive I / O signals that are not assigned a specific role. For example, the CSI 623 can include a MIPI CSI-2 connector, such as a 16-lane MIPI CSI-2 connector, D-PHY 2.1 (up to 40 Gbps) and C-PHY 2.0 (up to 164 Gbps) for supporting 16 virtual channels and 6 or more cameras, an 8-lane MIPI CSI-2 connector, D-PHY 2.1 (up to 20 Gbps for supporting 8 virtual channels and 4 or more cameras, and / or a 2x MIPI CSI-2, 22-pin camera connector, depending on the embodiment and implementation.
[0190] The accelerator 614 (e.g., hardware acceleration cluster) can include an on-chip computer vision network (CVNOC) 663 and SRAM for providing high-bandwidth, low-latency SRAM for the accelerator 614. In some examples, the on-chip memory can include at least 4 MB of SRAM, such as but not limited to being comprised of eight field-programmable memory blocks, accessible by the PVA 607, OFA 611, DLA 609, and / or other accelerator 614. Each pair of memory blocks can include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory 615 can be used. The PVA 607, OFA 611, DLA 609, and / or other accelerator 614 can access memory via a backbone that provides high-speed memory access for the accelerator 614. The backbone can include an on-chip computer vision network that interconnects the accelerator 614 to the memory (e.g., using an APB).
[0191] The CVNOC 663 can include an interface that determines whether the accelerator 614 provides a ready and valid signal before transmitting any control signals / addresses / data. Such an interface can provide separate stages and separate channels to transmit control signals / addresses / data, as well as burst-style communication for continuous data transmission. This type of interface can comply with ISO 26262 or IEC 61508 standards, although other standards and protocols can be used.
[0192] SoC 604 can include a data store 616 and / or a memory 615. The data store 616 can be an on-chip memory 615 of the SoC 604 that can store neural networks and / or other algorithms to be executed on the CPU 606, GPU 608, and / or one or more accelerators 614. In some examples, the data store 616 can be large enough in capacity to store multiple instances of a neural network for redundancy and safety. The data store 616 can include, for example, L2 and / or L3 caches 612. The memory 615 can include SRAM, LPDDR5, and / or other memory types. For example, the memory 615 can include 4 MB SRAM, 32 GB and / or 64 GB 256-bit LPDDR5 (204.8 GB / s), 8 GB and / or 16 GB 128-bit LPDDR5 (102.4 GB / s), and / or other memory types and sizes. References to the data store 616 can include references to memory associated with the PVA 607, OFA 611, DLA 609, and / or other accelerators 614, as described herein.
[0193] The data store 616 can include various storage types, such as eMMC, NVMe, and / or the like. For example, the SoC 604 can include storage in the form of an embedded Multi- Media Card (eMMC) (e.g., 64 GB eMMC 5.1) and / or an SD card slot, with external NVM Express (NVMe) capabilities, such as through an M.2 Key M. The data store 616 and / or other storage can be accessed via, for example, NVMe, using PCI Express (PCIe), RDMA, TCP, and / or other protocols.
[0194] SoC 604 can include one or more processors 610 (e.g., embedded processors). The processors 610 can include a boot and power management processor (BPMP) 653, which can be a specialized processor and subsystem for handling boot power and management functions and related security enforcement. The BPMP 653 can be part of a boot sequence of the SoC 604 and can provide runtime power management services. The BPMP 653 can provide clock and voltage programming, assistance with system low power state transitions, management of SoC 604 thermal and temperature sensors, and / or management of SoC 604 power states. Each temperature sensor can be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 604 can use the ring oscillator to detect the temperature of the CPU 606, GPU 608, accelerator 614, and / or other components. If it is determined that the temperature exceeds a threshold, the BPMP 653 can enter a temperature fault routine and place the SoC 604 into a lower power state and / or place the machine 600 into a driver safe stop mode (e.g., safely stop the machine 600).
[0195] The processors 610 can also include a set of embedded processors that can be used as an audio processing engine (APE) 655. The APE 655 can be an audio subsystem capable of full hardware support for multi-channel audio through multiple interfaces, as well as a wide and flexible set of audio I / O interfaces. In some examples, the APE 655 is a specialized processor core with a digital signal processor with dedicated RAM.
[0196] The processors 610 can also include an always-on processor engine (AOPE) 657, which can provide the necessary hardware features to support low-power sensor management and wake-on use cases. The AOPE 657 can include a processor core, tightly coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0197] The processors 610 can also include a security processor 613 (or “security island 613”), which can include a security cluster engine that includes a specialized processor or processor subsystem for handling security management for automotive, robotics, and / or other applications. The security processor 613 and / or security cluster engine can include two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In a security mode, the two or more cores can run in lockstep mode and act as a single core with comparison logic to detect any differences between their operations. In some embodiments, the security processor 613 can include a discrete processor such that a failure of other system components can not impact the performance and availability of the security processor 613.
[0198] The processor 610 can also include a real-time or near real-time sensor engine (SE) 659, which can include a dedicated processor subsystem for managing real-time or near real-time camera, LiDAR, RADAR, and / or other sensor modalities.
[0199] The processor 610 can also include one or more image signal processors (ISPs) 627, which can include high dynamic range signal processors and / or hardware engines as part of one or more sensor processing pipelines.
[0200] The processor 610 can include a video image compositor (VIC) 661, which can be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to produce the final image for the player window. The VIC 661 can perform lens distortion correction for wide-angle camera 668B, surround camera 668D, in-cabin monitoring camera sensors, and / or other camera sensors with distorted fields of view.
[0201] The VIC 661 can include enhanced temporal noise reduction for spatial and temporal noise reduction. For example, when motion occurs in a video, the noise reduction appropriately weights spatial information, reducing the weight of information provided by adjacent frames. When an image or portion of an image does not contain motion, the temporal noise reduction performed by the video image compositor can use information from previous images to reduce noise in the current image.
[0202] The VIC 661 can also be configured to perform stereo correction on input stereo lens frames. The video image compositor can also be used for user interface composition when the operating system desktop is in use and the GPU 608 does not need to continuously render new surfaces. The video image compositor can also be used to offload the GPU 608, even if the GPU 608 is powered on and actively performing 3D rendering, to improve performance and responsiveness.
[0203] The SoC 604 can also include various peripheral interfaces for input / output (I / O) 625, e.g., to enable communication with peripherals, audio codecs, power management, and / or other devices. The SoC 604 can be used to process data from a camera (e.g., over a Gigabit Multimedia Serial Link and / or an Ethernet connection), data from sensors (e.g., LiDAR sensors 664, RADAR sensors 660, etc. that can be connected over an Ethernet connection), data from the bus 602 (e.g., speed of the machine 600, steering wheel position, etc.), data from a GNSS sensor 658 (e.g., over an Ethernet or CAN bus connection). The SoC 604 can also include a dedicated high-performance mass storage controller, which can include its own DMA engine, and can be used to free the CPU 606 from regular data management tasks. In some embodiments, the SoC 604 I / O 625 can include connectors (e.g., 40-pin connectors or 40-pin expansion connectors) that support Universal Asynchronous Receiver / Transmitter (UART), Serial Peripheral Interface (SPI), Inter-Integrated Sound (I2S), Inter-Integrated Circuit (I2C), Controller Area Network (CAN), Pulse-Width Modulation (PWM), Digital Microphone Interface (DMIC), Digital Speaker Station (DSPK), General Purpose I / O (GPIO), etc., automation connectors (e.g., 12-pin automation connectors), audio panel connectors (e.g., 10-pin audio panel connectors), Joint Test Action Group (JTAG) connectors (e.g., 10-pin JTAG connectors), fan connectors (e.g., 4-pin fan connectors), RTC battery backup connectors (e.g., 2-pin battery backup connectors), microSD slots, DC power jacks, power, force, restore, and reset buttons, one or more display connectors (e.g., DisplayPort (DP) such as DP 1.4A (+MST), eDP 1.4 1, HDMI 2.1, and / or 4K30 Multi-Mode DP 1.2 (+MST) connectors), and / or other I / O 625 elements, components, or features.
[0204] The SoC 604 can include machine-intra-network capabilities using, e.g., Ethernet (e.g., automotive Ethernet), SERDES, Controller Area Network (CAN), FlexRay, Local Interconnect Network (LIN), Low-Voltage Differential Signaling (LVDS), Media Oriented Systems Transport (MOST), another network type, and / or combinations thereof. For example, the SoC 604 can include an RJ45 connector with up to 10 GbE, a 1 GbE connector, and / or other network connector types.
[0205] SoC 604 can include one or more digital signal processors (DSPs) 643. For example, the DSPs 643 can include special-purpose or specialized microprocessor chips optimized for digital signal processing (e.g., audio signal processing, telecommunications, digital image processing, RADAR, SONAR, LiDAR, and / or other sensor processing, speech recognition, and / or other applications).
[0206] SoC 604 can include one or more video encoders 619 and / or one or more video decoders 621. For example, the video encoders 619 can include hardware-based (e.g., as part of the GPU 608) video encoders (e.g., supporting H.264, H.265, etc., and conforming to the HEVC standard, such as NVIDIA’s NVENC) that can process image inputs (e.g., as YUV, RGB, etc.) to generate video bitstreams. The video decoders 621 can include video decoder engines that can provide fully accelerated hardware video decoding functionality (e.g., supporting decoding bitstreams in various formats, such as AV1, H.264, H.265, VP8, VP9, MPEG-1, MPEG-2, MPEG-4, VC-1, etc., and conforming to the HEVC standard, such as NVIDIA’s NVDEC). In some examples, the video decoders 621 can be hardware-based (e.g., as part of the GPU 608).
[0207] SoC 604 can include one or more general compute acceleration clusters (GCACs) 629. For example, the GCACs 629 can include various processor types that can be used to accelerate computations, such as one or more vector microcode processors (VMPs) 633, one or more multi-threaded processing clusters (MPCs) 631, one or more programmable macro arrays (PMAs) 635, and / or one or more other processor types. For example, the GCACs 629 can include a PMA 635, two VMPs 633, and 2 MPCs 631.
[0208] SoC 604 can include one or more vector microcode processors (VMPs) 633. In embodiments, the VMPs 633 can include wide vector (very long instruction word (VLIW) and single instruction multiple data (SIMD)) machines that perform various operations, such as short integral type operations common in computer vision and deep learning algorithms.
[0209] SoC 604 can include one or more multi-threaded processing clusters (MPCs) 631. MPCs 631 can include processing clusters that are more general purpose than GPUs and more efficient than CPUs in embodiments. For example, MPCs 631 can include multi-threaded processors that allow multiple threads to share resources and execute instructions concurrently.
[0210] SoC 604 can include one or more programmable macro arrays (PMAs) 635. PMAs 635 can include a coarse-grained reconfigurable architecture (CGRA) dataflow machine with a unique architecture that can provide strong performance on dense computer vision and deep learning algorithms that can not be achievable in classic digital signal processing (DSP) architectures.
[0211] SoC 604 can include one or more display processing units (DPUs) 645 for performing hardware-accelerated image processing. For example, DPUs 645 can retrieve pixel data from memory 615 and send it to display peripherals through a standard interface. Thus, DPUs 645 can handle display processing and rendering for in-machine and / or on-machine displays.
[0212] SoC 604 can include one or more application processing units (APUs) 639. For example, APUs 639 can include quad-core or dual-core processors with 48KB / 32KB LI caches (with parity and ECC), and 1 MB L2 cache with ECC. APUs 639 can support NEON instructions as well as single- and double-precision floating point operations.
[0213] SoC 604 can include one or more real-time processing units (RTPUs) 669. RTPUs 669 can include dual-core processors with 32KB / 32KB LI caches, and 256KB TCM with ECC. RTPUs 669 can support single- and double-precision floating point operations.
[0214] SoC 604 can include one or more built-in self-test (BIST) components 637. For example, BIST components 637 can include memory BIST (MBIST) for testing memory of the system and / or logic BIST (LBIST) for testing logic of the system. BIST components 637 can include embedded logic for directly testing logic and / or memory of the system.
[0215] The SoC 604 can include one or more dynamic reconfigurable processors (DRPs) 671. For example, the DRPs 671 can be used to accelerate various computing operations. For example, in embodiments, the DRPs 671 can be combined with MAC units to function as an AI accelerator. In embodiments, the DRPs 671 can execute an application while dynamically switching the circuit connection configuration of arithmetic units (e.g., ALUs) on the chip at each operation clock according to the content to be processed. Since only necessary arithmetic circuits are used, the DRPs 671 can consume less power than a CPU and can achieve higher speed. Also, compared to a CPU, since accessing an external memory frequently due to cache misses and other reasons degrades performance, the DRPs 671 can construct the necessary data path in hardware in advance, thereby reducing performance degradation and operation speed variation (jitter) due to memory access. The DRPs 671 can include a dynamic loading function that switches the circuit connection information at each algorithm change, thereby enabling processing with limited hardware resources even in robot / car applications that require processing of multiple algorithms.
[0216] In some embodiments, the accelerator 614 can include an OpenCV accelerator to accelerate processing of OpenCV, which is an open-source industry standard library for computer vision processing. In some embodiments, the combination of one or more DRPs 671 deployed as AI accelerators and the OpenCV accelerator can enhance AI computing and image processing algorithms, thereby enabling complex and computationally intensive operations such as visual simultaneous localization and mapping (SLAM).
[0217] Compared to conventional systems, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, the techniques described herein allow multiple neural networks to be executed simultaneously (e.g., at least partially in parallel) and / or sequentially, and the results to be combined together to implement level 2-5 autonomous driving functionality and / or autonomous robot motion, control, planning, and / or navigation operations. Moreover, since the SoC 604 can include various computing engines (e.g., the processors 610, the CPUs 606, the GPUs 608, the accelerators 614, etc.), tasks can be distributed among the computing engines, which in some cases, due to the discrete footprint of the computing engines, are less susceptible to common cause failures. Furthermore, since the SoC 604 can include a dedicated safety processor 613 (or safety island 613), critical safety or redundant operations can be performed without common cause failures of the main processing components or computing engines of the SoC 604. Due to these features, the SoC 604 and / or the underlying system of the machine 600 can be able to meet higher safety levels - e.g., automotive safety integrity level (ASIL) D of the ISO 26262 standard.
[0218] Figure 6E According to some embodiments of this disclosure, cloud-based servers (e.g., servers such as those described herein in a data center) and Figure 6A The following is a system diagram illustrating communication between an example autonomous or semi-autonomous vehicle or machine 600. System 676 may include server 678, network 690, and machine 600. Server 678 may include multiple GPUs 684(A)-684(H) (collectively referred to herein as GPU 684), switches 682(A)-682(H) (e.g., PCIe 4.0 / 5.0 switches, M.2 slots, Thunderbolt, USB4, NVIDIA's NVLink, NVIDIA's NVSwitch, GPUDirect RDMA, GPUDirect Storage, etc.), CPUs 680(A)-680(B) (collectively referred to herein as CPU 680), accelerators, and / or other processor types. GPU 684, CPU 680, and PCIe switch 682 may interconnect with high-speed interconnects, such as, but not limited to, NVIDIA-developed NVLink interface 688 and / or PCIe connection 686. In some examples, the GPU 684 is connected via NVLink and / or NVSwitchSoC, and the GPU 684 and PCIe switch 682 are connected via PCIe interconnect. While the figure shows eight GPUs 684, two CPUs 680, and two PCIe switches, this is not limiting. According to embodiments, each server 678 may include any number of GPUs 684, CPUs 680, and / or PCIe switches. For example, each server 678 may include eight, sixteen, thirty-two, and / or more GPUs 684.
[0219] The server 678 can receive sensor data from the machine 600 over the network 690 indicating information about new locations or previously unexplored locations, and / or sensor data indicating changes to previously seen / stored locations (e.g., unexpected or changed road conditions, such as road construction that recently started). The server 678 can transmit a neural network 692, an updated neural network 692, map information 694, etc., including information about traffic and road conditions, to the machine 600 over the network 690. Updates to the map information 694 can include updates to the HD map 622, SD maps, navigation maps, etc., such as information about construction sites, potholes, detours, flooding, and / or other obstacles. In some examples, the neural network 692, the updated neural network 692, the map information 694, and / or other information can be from new training and / or experience represented in data received from any number of machines 600 in the environment, and / or based on training performed at a data center (e.g., using the server 678 and / or other servers).
[0220] The server 678 can be used to train machine learning models (e.g., neural networks) based on training data. The training data can be generated by the machine 600, and / or can be generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., the neural network benefits from supervised learning) and / or otherwise pre-processed, while in other examples, the training data is not labeled and / or pre-processed (e.g., the neural network does not require supervised learning). The training can be performed according to any one or more machine learning techniques, including but not limited to the following categories, such as: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including spare dictionary learning), rule-based machine learning, anomaly detection, and any variants or combinations thereof. Once the machine learning model is trained, the machine learning model can be used by the machine 600 (e.g., transmitted to the machine 600 over the network 690, and / or the machine learning model can be used by the server 678 to remotely monitor and / or control the machine 600.
[0221] In some examples, the server 678 can receive data from the machine 600 and apply the data to a latest real-time neural network for real-time intelligent inference. The server 678 can include a deep learning supercomputer and / or a specialized AI computer driven by GPUs 684, such as the DGX and DGX Station machines developed by NVIDIA. However, in some examples, the server 678 can include a deep learning infrastructure of a data center that is driven by CPUs only.
[0222] The deep learning infrastructure of server 678 may be capable of rapid, real-time inference and can use this capability to assess and verify the health of the processor, software, and / or related hardware in machine 600. For example, the deep learning infrastructure may receive periodic updates from machine 600, such as image sequences and / or objects located by machine 600 in the image sequence (e.g., through computer vision and / or other machine learning object classification techniques). The deep learning infrastructure may run its own neural network to identify objects and compare them with the objects identified by machine 600. If the results do not match and the infrastructure concludes that the AI in machine 600 has malfunctioned, server 678 may send a signal to machine 600 instructing the fail-safe computer of machine 600 to take over control, notify the occupants, and perform safety maneuvers or operations, such as slowing down, returning control to the driver, stopping, and / or pulling over / closing the vehicle.
[0223] For inference, server 678 can include GPU 684 and one or more programmable inference accelerators (such as NVIDIA's TensorRT). The combination of GPU-driven servers and inference acceleration enables real-time response. In other examples, such as in scenarios where performance is less critical, servers powered by CPUs, FPGAs, and other processors can be used for inference. The computing ecosystem for generating, training, and deploying AI.
[0224] Figure 7 This is a system diagram illustrating a three-computer ecosystem 700 according to at least some embodiments of the present disclosure, including a first computing system 702 for generating or creating artificial intelligence (AI) (e.g., AI training and validation data), a second computing system 704 for training the AI, and a third computing system 706 for deploying AI at the edge (which may include or correspond to...). Figures 6A-6E (SoC 604). For example, to develop and deploy materialized or physical AI, a three-computer ecosystem 700 can be used, comprising three accelerated computer systems to handle physical AI training, simulation, and runtime (e.g., edge deployment). These systems can use scalable, physical-based simulations of the machine 600 and its world to generate training data and train multimodal base models (and / or other model types). By doing so, simulations of the machine 600 can be performed at scale, allowing skills (e.g., robotic skills) to be improved, tested, and optimized in virtual worlds that simulate the laws of physics (e.g., using NVIDIA's OMNIVERSE), thereby helping to reduce the cost of real-world data acquisition and ensuring that the machine 600 can operate safely in a controlled environment.
[0225] The computing system 704 (e.g., NVIDIA’s DGX platform) can be used to train and fine-tune powerful foundation and generative AI models. Models such as general-purpose foundation models (e.g., NVIDIA’s Project GROOVT) can be used to enable robots and other machines 600 to understand natural language and mimic actions by observing human actions. The computing system 704 can include a platform that brings software, infrastructure, and expertise into a modern, unified AI development and training solution. The computing system 704 can include individual computing devices 710 (e.g., NVIDIA’s DGX B200, H200, etc.) and / or any number of computing devices 710 in a data center infrastructure 712 (e.g., NVIDIA’s DGX SuperPOD).
[0226] For example, a single computing device 710 can include a GPU (e.g., 8 GPUs, 1,440 GB of total GPU memory) and a CPU (e.g., 2 CPUs, 112 cores total, 2.1 GHz or 4 GHz (with boost)) that provide up to 72 petaFLOPS of training capability and 144 petaFLOPS of inferencing capability. The computing device 710 can include memory (e.g., 4 TB of memory) and storage (e.g., 2 x 1.9 TB NVMe M.2 of OS storage and 8 x 3.84 TB NVMe U.2 of internal storage). The computing device 710 can include various networking and network management components, such as OSFP ports (e.g., 4 OSFP ports) for serving single-port intelligent host channel adapters (e.g., 8 single-port ConnextX-7 virtual protocol interconnects (VPis)) that provide up to 400 GB / s of Infiniband / Ethernet. The computing device 710 can also include, for example, dual-port quad small form-factor pluggable (QSFP) data processing units (DPUs) (e.g., 2 dual-port QSFP112 DPUs - such as NVIDIA’s BlueField-3 DPUs) that provide up to 400 Gb / s of InfiniBand / Ethernet. The computing device 710 can include onboard network interface cards (NICs) (e.g., 10 Gb / s onboard NICs with RJ45), dual-port Ethernet NICs (e.g., 100 GB / s dual-port Ethernet NICs), and / or host board management controllers (MBCs) (e.g., with RJ45). In some embodiments, the NICs for the computing device 710 can include SuperNICs (e.g., NVIDIA’s ConnectX-8 SuperNIC) in order to provide up to 800 Gb / s of data throughput for in-network computing acceleration engines, providing the performance and powerful feature set needed to support exa-parameter scale AI factory and scientific computing workloads. In other embodiments, the computing device 710 can include intelligent host channel adapters (HCAs) (e.g., NVIDIA’s ConnectX-7) in order to provide ultra-low latency, 400 Gb / s of throughput for in-network computing acceleration engines.
[0227] The data center infrastructure 712 can include any number of computing devices 710 and an operating system (OS) (e.g., DGX OS extension of a Linux distribution) for maximizing system uptime, security, and reliability, network / storage acceleration libraries and management for accelerating end-to-end infrastructure performance, cluster management for scaling and managing one node (e.g., one computing device 710) to thousands of nodes, job scheduling and orchestration for ensuring smooth execution of jobs for each developer, AI workflow management and machine learning operations (MLOps) for moving more models from prototyping to production, and enterprise software to speed developer success.
[0228] The computing system 702 (e.g., NVIDIA’s OVX server) can provide a development and simulation platform for testing and optimizing physical AI using APIs and frameworks for simulation (e.g., NVIDIA’s DriveSIM, ISAAC Sim, ISAAC Gym, ISAAC Lab, etc.). The computing system 702 allows developers to simulate and validate robot models using simulation frameworks, and / or generate large amounts of physics-based synthetic data to guide model training. The computing system 702 can support learning frameworks powered for robot reinforcement learning and imitation learning to accelerate robot policy training and refinement. For example, the computing system 702 can be used to generate any number of simulations 708, such as in NVIDIA’s OMNIVERSE. The computing system 702 can be used to optimize acceleration across the entire software stack, from training, fine-tuning, and deploying generative AI to support for industrial digitization within content collaboration platforms (including APIs, software development kits (SDKs), and services) that allow integration of OpenUSD, ray-tracing rendering technology (e.g., NVIDIA’s RTX), and generative physics AI into existing software tools and simulation workflows, such as for industrial and robotics use cases (e.g., NVIDIA’s OMNIVERSE). Thus, the computing system 702 can host or support a native OpenUSD software platform that enables enterprises to connect 3D pipelines and develop advanced real-time 3D applications for industrial digitization. With powerful ray-tracing acceleration AI and graphics capabilities, the computing system 702 can provide strong performance for workloads such as extended reality (XR), multi-user design collaboration, and digital twins. This allows for the creation of physically accurate models with high-fidelity ray-tracing and path-tracing material rendering, large-scale, AI-enabled simulation operations, and realistic 3D synthetic data for training. The computing system 702 can include a single computing device 714 (e.g., NVIDIA’s OVX L40S server) and / or any number of computing devices 714 (e.g., NVIDIA’s OVX system) in a data center infrastructure 716.
[0229] The computing devices 714 (which can include servers) can include CPUs (e.g., 2 CPUs each with 32 cores) and GPUs (e.g., 4 or 8 GPUs each including 48 GB GDDR6 with ECC memory, 864 GB / s memory bandwidth, PCIe Gen4 x 16: 64 GB / s bidirectional interconnect interface, 18,176 CUDA cores, 142 ray-tracing (RT) cores, and 568 Tensor cores). The computing devices 714 can include various networking and network management components, such as intelligent host channel adapters (HCAs) (e.g., 2 or 4 single-port ConnextX-7, each port 200 Gb / s, providing up to 800 Gb / s Infiniband / Ethernet), one or more DPUs (e.g., dual-port QSFP112 DPUs, such as NVIDIA BlueField-3 DPUs), providing up to 400 Gb / s InfiniBand / Ethernet. In some embodiments, the NICs for the computing devices 714 can include SuperNICs (e.g., NVIDIA’s ConnectX-8 SuperNIC) in order to provide up to 800 Gb / s of data throughput for in-network computing acceleration engines, providing the performance and powerful feature set needed to support exa-parameter scale AI factories and scientific computing workloads. In other embodiments, the computing devices 714 can include intelligent host channel adapters (HCAs) (e.g., NVIDIA’s ConnectX-7) in order to provide ultra-low latency, 400 Gb / s throughput for in-network computing acceleration engines. The computing devices 714 can include host memory (e.g., 384 Gb DDR5 ECC for 4 GPUs, or 768 Gb DDR5 ECC for 8 GPUs), and can include dual in-line memory module (DIMM) slots, host boot drives (e.g., 1 TB NVMe), and / or host storage (e.g., 24 TB NVMe).
[0230] Similar to the data center infrastructure 712, the data center infrastructure 716 can allow for any number of computing devices 714 to be combined into a cluster configuration in accordance with the reference architecture.
[0231] The computing systems 706 can be used to deploy trained AI models on runtime computers, such as the SoC 604 described herein. For example, these computing systems 706 can be designed for compact on-board computing needs, including model collections of control policies, vision and language models, and the like deployed on power-efficient on-board edge computing systems 706. Reference can be made to Figures 6A-6E Details of the components, features, and capabilities of the computing systems 706 are described in greater detail.
[0232] Example generative models
[0233] In at least some embodiments, language models such as large language models (LLMs), visual language models (VLMs), multi-modal language models (MMLMs), visual language action (VLA) models, and / or other types of generative artificial intelligence (AI) can be implemented. These models can be capable of understanding, summarizing, translating, and / or otherwise generating text (e.g., natural language text, code, etc.), images, videos, computer-aided design (CAD) assets, OMNIVERSE and / or METAVERSE file information (e.g., USD format, such as OpenUSD), and / or the like based on context provided in an input prompt or query. In embodiments, these language models can be considered “large” in that these models are trained on massive datasets and have architectures with a large number of learnable network parameters (weights and biases) - e.g., millions or billions of parameters. LLMs / VLMs / MMLMs / etc. can be implemented for summarizing textual data, analyzing data (e.g., text, images, videos, etc.), and extracting insights from data (e.g., text, images, videos, etc.), as well as generating new text / images / videos / etc. in a user-specified style, tone, and / or format. In embodiments, the LLMs / VLMs / MMLMs / etc. of the present disclosure can be specialized for text processing, while in other embodiments, multi-modal LLMs can be implemented to accept, understand, and / or generate text and / or other types of content like images, audio (sounds, synthesized speech, etc.), 2D and / or 3D data (e.g., USD format), and / or videos. For example, a visual language model (VLM) or more specifically a multi-modal language model (MMLM) can be implemented to accept image, video, sensor, audio, text, 3D design (e.g., CAD), and / or other input data types and / or generate or output image, video, audio, text, 3D design, and / or other output data types.
[0234] Various types of LLM / VLM / MMLM / etc. architectures can be implemented in various embodiments. For example, different architectures can be implemented that use different techniques to understand and generate output (e.g., text, audio, video, images, 2D and / or 3D design or asset data, etc.). In some embodiments, LLM / VLM / MMLM / etc. architectures (e.g., recurrent neural networks (RNNs) or long short-term memory networks (LSTMs)) can be used, while in other embodiments, transformer architectures (e.g., architectures that rely on self-attention and / or cross-attention (e.g., between contextual data and textual data) mechanisms) can be used to understand and recognize relationships between words or tokens and / or contextual data (e.g., other text, video, images, design data, USD, etc.). One or more generative processing pipelines that include LLM / VLM / MMLM / etc. can also include one or more diffusion blocks (e.g., denoisers). LLM / VLM / MMLM / etc. of the present disclosure can include encoder and / or decoder blocks. For example, discriminative or encoder-only models (e.g., BERT (Bidirectional Encoder Representations from Transformers)) can be implemented for tasks involving language understanding (e.g., classification, sentiment analysis, question answering, and named entity recognition). As another example, generative or decoder-only models (e.g., GPT (Generative Pretrained Transformer)) can be implemented for tasks involving language and content generation (e.g., text completion, story generation, and dialogue generation). LLM / VLM / MMLM / etc. that include both encoder and decoder components (e.g., T5 (Text-to-Text Transformer)) can be implemented to understand and generate content, such as for translation and summarization. These examples are not intended to be limiting, and any architecture type (including, but not limited to, the architecture types described herein) can be implemented depending on the particular embodiment and the tasks being performed using the LLM / VLM / MMLM / etc.
[0235] In various embodiments, LLMs / VLMs / MMLMs / etc. can be trained using unsupervised learning, where the LLMs / VLMs / MMLMs / etc. learn patterns from large amounts of unlabeled text / audio / video / image / design / USD / etc. data. As a result of extensive training, in embodiments, the models can not require task-specific or domain-specific training. LLMs / VLMs / MMLMs / etc. that are extensively pre-trained on large amounts of unlabeled data can be referred to as base models, and can be good at a variety of tasks, such as question answering, summarization, filling in missing information, translation, image / video / design / USD / data generation. Some LLMs / VLMs / MMLMs / etc. can be customized for specific use cases using techniques such as prompt tuning, fine-tuning, retrieval-augmented generation (RAG), adding adapters (e.g., custom neural networks and / or neural network layers for tuning or adjusting prompts or tokens to bias the language model toward a particular task or domain), and / or using optimization models for specific tasks and / or for other fine-tuning or customization techniques within a particular domain.
[0236] In some embodiments, LLMs / VLMs / MMLMs / etc. of the present disclosure can be implemented using various model alignment techniques. For example, in some embodiments, guardrails can be implemented to identify incorrect or unwanted inputs (e.g., prompts) and / or outputs of the models. In this process, the system can use guardrails and / or other model alignment techniques to prevent particular unwanted inputs from being processed using the LLMs / VLMs / MMLMs / etc., and / or to prevent outputs or presentations (e.g., displays, audio outputs, etc.) of information generated using the LLMs / VLMs / MMLMs / etc. In some embodiments, one or more additional models (or layers thereof) can be implemented to identify problems with inputs and / or outputs of the models. For example, these “guard” models can be trained to identify “safe” or otherwise ok or wanted inputs and / or outputs and / or “unsafe” or otherwise unwanted inputs and / or outputs for a particular application / implementation. Thus, LLMs / VLMs / MMLMs / etc. of the present disclosure can be less likely to output language / text / audio / video / design data / USD data / etc. that can be offensive, vulgar, inappropriate, unsafe, out of scope, and / or otherwise unwanted for a particular application / implementation.
[0237] In some embodiments, the LLM / VLM / etc. can be configured to or have access to or use one or more plugins, application programming interfaces (APIs), databases, data stores, repositories, etc. For example, for certain tasks or operations for which the model is not ideally suited, the model can have instructions for accessing one or more plugins (e.g., third-party plugins) to assist in processing the current input (e.g., as a result of training, and / or based on instructions in a given prompt). In such examples, when at least a portion of the prompt is related to restaurants or weather, the model can access one or more restaurant or weather plugins (e.g., via one or more APIs) to retrieve relevant information. Another example is if at least a portion of the response requires mathematical calculations, the model can access one or more mathematical plugins or APIs to assist in solving the problem, which can then be used in the output of the model from the response of the plugin and / or API. This process can repeat (e.g., recursively) for any number of iterations and using any number of plugins and / or APIs until a response to the input prompt can be generated that addresses each inquiry / question / request / process / operation / etc. Thus, the model can rely not only on its own knowledge gained from training on large datasets, but also on the expertise or optimized properties of one or more external resources (e.g., APIs, plugins, etc.).
[0238] In some embodiments, multiple language models (e.g., LLMs / VLMs / MMLMs / etc., multiple instances of the same language model, and / or multiple prompts provided to the same language model or instance of the same language model) can be implemented, executed, or accessed (e.g., using one or more plugins, user interfaces, APIs, databases, data stores, repositories, etc.) to provide outputs responsive to the same query or responsive to separate portions of the query. In at least one embodiment, the same input query and prompts (e.g., a set of constraints, condition generators, etc.) can be provided to multiple language models (e.g., language models with different architectures, language models trained on different (e.g., updated) data corpora). In one or more embodiments, the language models can be different versions of the same base model. In one or more embodiments, at least one language model can be instantiated as multiple agents, e.g., more than one prompt can be provided to constrain, guide, or otherwise influence the style, content, or character of the output provided, etc. In one or more example non-limiting embodiments, the same language model can be required to provide outputs corresponding to different roles, perspectives, characters, or having different knowledge bases, etc. as defined by the prompts provided.
[0239] In any of such embodiments, the outputs of two or more (e.g., each) language models, two or more versions of at least one language model, two or more instantiations of at least one language model, and / or two or more prompts provided to at least one language model can be further processed, e.g., aggregated, compared, or filtered, or used to determine (and provide) a consensus response. In one or more embodiments, the output from one language model (or version, instantiation, or agent) can be provided as input to another language model for further processing and / or validation. In one or more embodiments, a language model can be required to generate or otherwise obtain an output regarding an input source material, where the output is associated with the input source material. Such association can include, for example, generating an embedding (e.g., as metadata) a caption or text portion within an input source text or image. In one or more embodiments, the output of a language model can be used to determine the validity of an input source material for further processing or inclusion in a dataset. For example, a language model can be used to assess the presence (or absence) of a target word in a text portion or the presence (or absence) of an object in an image, where the text or image is annotated to indicate such presence (or lack thereof). Alternatively, a determination from a language model can be used to determine whether a source material should be included in a curated dataset, for example, but not limited to.
[0240] Figure 8 is a block diagram of an example generative language model system 800 suitable for implementing at least some embodiments of the present disclosure. In Figure 8 In the example shown, the generative language model system 800 includes a retrieval-augmented generation (RAG) component 892, an input processor 805, a tokenizer 810, an embedding component 820, a plug-in / API 895, and a generative language model (LM) 830 (which can include a LLM, a VLM, a MMLM, a VLA model, etc.).
[0241] At a high level, input processor 805 can receive input 801 that includes text and / or other types of input data (e.g., audio data, video data, image data, sensor data (e.g., LiDAR, RADAR, ultrasound, etc.), 3D design data, CAD data, Universal Scene Description (USD) data (e.g., OpenUSD, etc.), depending on the architecture of generative LM 830 (e.g., LLM / VLM / MMLM / etc.). In some embodiments, input 801 includes plain text in the form of one or more sentences, paragraphs, and / or documents. Additionally or alternatively, input 801 can include sequences of numbers, pre-computed embeddings (e.g., word or sentence embeddings), and / or structured data (e.g., table format, JSON, or XML). In some implementations where generative LM 830 is capable of processing multi-modal input, input 801 can combine text (or can omit text) with image data, audio data, video data, design data, USD data, and / or other types of input data such as, but not limited to, the data described herein. Taking the example of raw input text, input processor 805 can prepare the raw input text in various ways. For example, input processor 805 can perform various types of text filtering to remove noise (e.g., special characters, punctuation, HTML tags, stop words, portions of images, portions of audio, etc.) from the relevant textual content. In examples involving stop words (commonly used words that tend to have little semantic meaning), input processor 805 can remove stop words to reduce noise and cause generative LM 830 to focus on more meaningful content. Input processor 805 can apply text normalization, e.g., by converting all characters to lowercase, removing diacritics, and / or handling special cases (such as abbreviations or contractions) to ensure consistency (e.g., converting 1 / 4 to one-quarter). Similarly, input processor 805 and / or post-processor can perform inverse text normalization (ITN) in order to convert plain language back to canonical or other forms (e.g., converting one-quarter to 1 / 4). These are just a few examples, and other types of input and / or output processing can be applied.
[0242] In some embodiments, the RAG component 892 (which can include one or more RAG models, and / or which can use the generative LM 830 itself to perform) can be used to retrieve additional information to be used as part of the input 801 or prompt. The RAG can be used to augment the input to the LLM / VLM / MMLM / etc. with external knowledge in order to make the answer to a particular question or query or request more relevant, e.g., in cases where specific knowledge is needed. The RAG component 892 can obtain this additional information (e.g., base information, such as base text / images / video / audio / USD / CAD / etc.) from one or more external sources, and can then feed it to the LLM / VLM / MMLM / etc. along with the prompt in order to improve the accuracy of the model’s response or output.
[0243] For example, in some embodiments, the input 801 can be generated using the query or model input (e.g., question, request, etc.) in addition to data retrieved using the RAG component 892. In some embodiments, the input processor 805 can analyze the input 801 and communicate with the RAG component 892 (or in embodiments, the RAG component 892 can be part of the input processor 805) in order to identify relevant text and / or other data to provide to the generative LM 830 as additional context or source of information from which to identify a response, answer, or output 890. For example, when the input indicates that the user is interested in the required tire pressure for a particular make and model of vehicle, the RAG component 892 can use a RAG model to perform a vector search, e.g., in an embedding space, to retrieve tire pressure information or text corresponding thereto from a digital (embedded) version of the user manual for that particular vehicle make and model. Similarly, when the user accesses a chatbot related to the sale or service of a particular product again, the RAG component 892 can retrieve the previously stored dialog history (or at least an abridged version thereof) and provide the previous dialog history along with the current inquiry / request as part of the input 801 to the generative LM 830.
[0244] The RAG component 892 can use various RAG techniques. For example, a naive RAG (RAG) can be used in which documents are indexed, chunked, and applied to an embedding model to generate embeddings corresponding to the chunks. A user query can also be applied to this embedding model and / or another embedding model of the RAG component 892, and the embeddings of the chunks can be compared to the embedding of the query to identify the most similar / most relevant embeddings to the query, which can be provided to the generative LM 830 to generate an output.
[0245] In some embodiments, more advanced RAG techniques can be used. For example, chunks can undergo pre-retrieval processes (e.g., routing, rewriting, metadata analysis, expansion, etc.) before being passed to the embedding model. Furthermore, post-retrieval processes (e.g., re-ranking, hint compression, etc.) can be performed on the output of the embedding model before generating the final embedding, which is then used for comparison with the input query.
[0246] As a further example, modular RAG techniques can be used, such as those similar to Naive RAG and / or Advanced RAG, but also including features such as hybrid search, recursive retrieval and query engines, StepBack methods, subqueries and hypothetical document embeddings.
[0247] As another example, Graph RAG can use a knowledge graph as a source of context or factual information. Graph RAG can be implemented using a graph database as a source of contextual information sent to LLM / VLM / MMLM / etc. Instead of providing the model with data chunks extracted from larger documents (which may result in a lack of context, factual accuracy, linguistic accuracy, etc.) (or anything other than providing the model with data chunks extracted from larger documents), Graph RAG can also provide structured entity information to LLM / VLM / MMLM / etc. by combining structured entity text descriptions with their many attributes and relationships, thus giving the model deeper insights. In implementing Graph RAG, the systems and methods described herein use graphs as content stores and extract relevant document chunks, requiring LLM / VLM / MMLM / etc. to use them to answer questions. In such embodiments, the knowledge graph may contain relevant textual content and metadata about the knowledge graph, or it may be integrated with a vector database. In some embodiments, Graph RAG can use the graph as a subject matter expert, where descriptions of concepts and entities relevant to the query / hint can be extracted and passed to the model as semantic context. These descriptions may include relationships between concepts. In other examples, the graph can be used as a database where a portion of a query / hint can be mapped to a graph query, the graph query can be executed, and LLM / VLM / MMLM / etc. can aggregate the results. In such examples, the graph can store relevant factual information and can be used for queries (natural language queries) and entity links to graph query tools (NL to graph query tools). In some embodiments, the graph RAG (e.g., using a graph database) can be combined with standard (e.g., vector database) RAGs and / or other RAG types to benefit from a variety of approaches.
[0248] In any embodiment, the RAG component 892 can implement plugins, APIs, user interfaces, and / or other functionality to perform RAG. For example, LLMs / VLMs / MMLMs / etc. can use a graph RAG plugin to run queries on a knowledge graph to extract relevant information to feed into a model, and can use a standard or vector RAG plugin to run queries on a vector database. For example, a graph database can interact with a REST interface of the plugin, such that the graph database can be decoupled from the vector database and / or embedding model.
[0249] The tokenizer 810 can segment (e.g., processed) textual data into smaller units (tokens) for subsequent analysis and processing. Depending on the implementation, tokens can represent individual words, subwords, characters, portions of audio / video / image / etc. Word-based tokenization divides text into individual words, treating each word as a separate token. Subword tokenization breaks down words into smaller meaningful units (e.g., prefixes, suffixes, stems), enabling the generative LM 830 to understand morphological variations and more effectively process out-of-vocabulary words. Character-based tokenization represents each character as a separate token, enabling the generative LM 830 to process text at a fine-grained level. The choice of tokenization strategy can depend on factors such as the language being processed, the task at hand, and / or characteristics of the training dataset. Accordingly, the tokenizer 810 can transform (e.g., processed) text into a structured format according to a tokenization scheme implemented in a particular embodiment.
[0250] The embedding component 820 can transform discrete tokens into a (e.g., dense, continuous vector) representation of semantic meaning using any known embedding technique. For example, the embedding component 820 can use pre-trained word embeddings (e.g., Word2Vec, GloVe, or FastText), one-hot encoding, Term Frequency-Inverse Document Frequency (TF-IDF) encoding, one or more embedding layers of a neural network, and / or others.
[0251] In some implementations in which input 801 includes image data / video data / etc., input processor 805 can resize the data to a standard size compatible with the format of the respective input channel and / or can normalize the pixel values to a common range (e.g., 0 to 1) to ensure consistent representation, and embedding component 820 can encode the image data using any known technique (e.g., using one or more convolutional neural networks (CNNs) to extract visual features). In some implementations in which input 801 includes audio data, input processor 805 can resample the audio files to a consistent sampling rate for uniform processing, and embedding component 820 can extract and encode audio features using any known technique, e.g., in the form of a spectrogram (e.g., a mel spectrogram). In some implementations in which input 801 includes video data, input processor 805 can extract frames or apply resizing to extracted frames, and embedding component 820 can extract features such as optical flow embeddings or video embeddings and / or can encode temporal information or sequences of frames. In some implementations in which input 801 includes multi-modal data, embedding component 820 can fuse representations of different types of data (e.g., text, image, audio, USD, video, design, etc.) using techniques such as early fusion (concatenation), late fusion (sequential processing), attention-based fusion (e.g., self-attention, cross-attention), etc.
[0252] The generative LM 830 and / or other components of the generative LM system 800 can use different types of neural network architectures depending on the implementation. For example, a transformer-based architecture (such as used in GPT, etc. models) can be implemented, and it can include a self-attention mechanism that weights the importance of different words or tokens in the input sequence and / or a feed-forward network that processes the output of the self-attention layers, which applies a non-linear transformation to the input representation and extracts higher-level features. Some non-limiting example architectures include transformers (e.g., encoder-decoder, decoder-only, multi-modal), RNNs, LSTMs, fusion models, diffusion models, cross-modal embedding models that learn a joint embedding space, graph neural networks (GNNs), hybrid architectures that combine different types of architectures adversarial networks such as generative adversarial networks or GANs or adversarial autoencoders (AAEs) for joint distribution learning, linear-time sequence modeling that employs a selective state space modeling (SSM) architecture (e.g., Mamba LLM architecture), etc. Thus, depending on the implementation and architecture, the embedding component 820 can apply the encoded representation of the input 801 to the generative LM 830, and the generative LM 830 can process the encoded representation of the input 801 to generate an output 890, which can include response text and / or other types of data.
[0253] As described herein, in some embodiments, generative LM 830 can be configured to access or use (or be able to access or use) plugins / APIs 895 (which can include one or more plugins, application programming interfaces (APIs), databases, data stores, repositories, etc.). For example, for certain tasks or operations that are not ideally suited for generative LM 830, the model can have instructions (e.g., as a result of training, and / or based on instructions in a given prompt, such as instructions retrieved using RAG component 892) for accessing one or more plugins / APIs 895 (e.g., third-party plugins) to help process the current input. In such examples, when at least a portion of the prompt is related to restaurants or weather, the model can access one or more restaurant or weather plugins (e.g., via one or more APIs), send at least a portion of the prompt related to the particular plugin / API 895 to the plugin / API 895, the plugin / API 895 can process that information and return an answer to generative LM 830, which can use that response to generate output 890. This process can repeat (e.g., recursively) for any number of iterations and with any number of plugins / APIs 895 until output 890 can be generated that addresses each query / question / request / process / operation / etc. from input 801. Thus, the model can rely not only on its own knowledge obtained from training on large datasets and / or from data retrieved using RAG component 892, but also on the specialized knowledge or optimized properties of one or more external resources (e.g., plugins / APIs 895).
[0254] In some embodiments, one or more converter engines (TEs) can be implemented. Converter engines can use micro-tensor scaling to optimize performance and accuracy, such as enabling 16-bit floating-point (FP16), 8-bit floating-point (FP8), and / or 4-bit floating-point (FP4) AI processing. For example, a converter engine can use 16-bit or 8-bit floating-point precision and 8-bit or 4-bit floating-point data formats, combined with software algorithms, to improve AI performance and capabilities. By reducing mathematical operations to 8 or 4 bits, TEs can train larger networks faster without compromising accuracy. For example, TEs can include libraries for accelerating converter models on processing devices such as GPUs to provide better performance in training and inference with lower memory utilization. When TEs are combined with other technologies, such as high-speed interconnects between nodes (e.g., using NVLink switches) and tensor cores (which enable mixed-precision computation, such as micro-scaling precision support), server clusters can be more capable of training massive networks (e.g., billions of parameters) at high speeds. Therefore, it can support tensor core precision for FP64, TF32, BF16, FP16, FP8, INT8, FP6 and FP4, as well as CUDA core precision for FP64, FP32, FP16 and BF16.
[0255] The LLM / VLM / MMLM / VLA and other architectures described herein are intended to be examples only, and other suitable architectures may be implemented within the scope of this disclosure.
[0256] Example computing device
[0257] Figure 9is a block diagram of an example computing device 900 suitable for implementing some embodiments of the present disclosure. The computing device 900 can include an interconnection system 902 coupling the following components: a memory 904, one or more central processing units (CPU) 906, one or more graphics processing units (GPU) 908, a communication interface 910, an input / output (I / O) port 912, an input / output component 914, a power supply 916, one or more presentation components 918 (e.g., one or more displays, one or more speakers, etc.), and one or more logic units 920. In at least one embodiment, one or more computing devices 900 can include one or more virtual machines (VMs), and / or any component thereof can include virtual components (e.g., virtual hardware components). For a non-limiting example, one or more of GPUs 908 can include one or more vGPUs, one or more of CPUs 906 can include one or more vCPUs, and / or one or more of logic units 920 can include one or more virtual logic units. As such, one or more computing devices 900 can include discrete components (e.g., a full GPU dedicated to computing device 900), virtual components (e.g., a portion of a GPU dedicated to computing device 900), or a combination thereof.
[0258] Although Figure 9 various blocks of are shown as connected through an interconnection system 902 using a bus, this is not intended to be limiting and is for clarity only. For example, in some embodiments, a presentation component 918, such as a display device, can be considered an I / O component 914 (e.g., if the display is a touchscreen). As another example, a CPU 906 and / or GPU 908 can include memory (e.g., memory 904 can represent a storage device in addition to memory of GPU 908, CPU 906, and / or other components). Thus, Figure 9 computing devices of are merely illustrative. There is no distinction, in terms of Figure 9 scope, between the computing devices of as “workstations,” “servers,” “laptops,” “desktops,” “tablet computers,” “client devices,” “mobile devices,” “handheld devices,” “gaming consoles,” “electronic control units (ECUs),” “virtual reality systems,” and / or other device or system types, as all are contemplated within the scope of
[0259] The interconnection system 902 can represent one or more busses or links, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnection system 902 can include one or more bus or link types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. As an example, the CPU 906 can be directly connected to the memory 904. Further, the CPU 906 can be directly connected to the GPU 908. Where there are direct or point-to-point connections between components, the interconnection system 902 can include a PCIe link to perform the connection. In these examples, a PCI bus need not be included in the computing device 900.
[0260] The memory 904 can include any of a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by the computing device 900. Computer-readable media can include both volatile and nonvolatile media, removable and non-removable media. By way of example, and not limitation, computer-readable media can comprise computer storage media and communication media.
[0261] Computer storage media can include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules, and / or other data types. For example, the memory 904 can store computer readable instructions (e.g., representing one or more programs and / or one or more program elements, such as an operating system). Computer storage media can include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing device 900. As used herein, computer storage media does not include signals per se.
[0262] Computer storage media can embody computer readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave or other transport mechanism, and includes any information delivery media. The term “modulated data signal” can refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, computer storage media can include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Combinations of the any of the above should also be included within the scope of computer readable media.
[0263] The CPUs 906 can be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 900 to perform one or more of the methods and / or processes described herein. The CPUs 906 can each include one or more cores capable of handling numerous software threads concurrently (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.). The CPUs 906 can include any type of processors, and can include different types of processors depending on the type of computing device 900 being implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of computing device 900, the processors can be Advanced RISC Machines (ARM) processors implemented using Reduced Instruction Set Computing (RISC) or x86 processors implemented using Complex Instruction Set Computing (CISC). The computing device 900 can include one or more CPUs 906 in addition to, or as an alternative to, one or more microprocessors or co-processors such as math co-processors.
[0264] One or more GPUs 908 can be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 900 to perform one or more of the methods and / or processes described herein, in addition to or instead of one or more CPUs 906. One or more of the GPUs 908 can be integrated GPUs (e.g., with one or more of the CPUs 906) and / or one or more of the GPUs 908 can be discrete GPUs. In embodiments, one or more of the GPUs 908 can be a co-processor of one or more of the CPUs 906. The GPUs 908 can be used by the computing device 900 to render graphics (e.g., 3D graphics) or to perform general-purpose computing. For example, the GPUs 908 can be used for general-purpose computing on GPUs (GPGPU). The GPUs 908 can include hundreds or thousands of cores capable of handling hundreds or thousands of software threads concurrently. The GPUs 908 can generate pixel data for output images in response to rendering commands (e.g., received from the CPUs 906 via a host interface). The GPUs 908 can include graphics memory (e.g., display memory) for storing pixel data or any other suitable data (e.g., GPGPU data). The display memory can be included as part of the memory 904. The GPUs 908 can include two or more GPUs operating in parallel (e.g., via a link). The link can connect the GPUs directly (e.g., using NVLINK) or can connect the GPUs through a switch (e.g., using an NVSwitch). When combined together, each GPU 908 can generate pixel data or GPGPU data for a different portion of an output or for a different output (e.g., a first GPU for a first image and a second GPU for a simulated image). Each GPU can include its own memory or can share memory with other GPUs.
[0265] In addition to or in place of CPU(s) 906 and / or GPU(s) 908, logic unit(s) 920 can be configured to execute at least some of the computer-readable instructions to control one or more components of computing device 900 to perform one or more of the methods and / or processes described herein. In embodiments, one or more CPU(s) 906, one or more GPU(s) 908, and / or one or more logic unit(s) 920 can perform any combination of the methods, processes, and / or portions thereof, discretely or jointly. One or more of logic unit(s) 920 can be part of one or more of CPU(s) 906 and / or integrated in one or more of CPU(s) 906 and / or one or more of logic unit(s) 920 can be discrete components or otherwise external to CPU(s) 906 and / or GPU(s) 908. In embodiments, one or more of logic unit(s) 920 can be a co-processor of one or more of CPU(s) 906 and / or one or more of GPU(s) 908.
[0266] Examples of logic unit(s) 920 include one or more processing cores and / or components thereof, such as a data processing unit (DPU), a tensor core (TC), a tensor processing unit (TPU), a pixel vision core (PVC), a vision processing unit (VPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multi-processor (SM), a tree traversal unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), a deep learning accelerator cluster (XNN), a neural processing unit (NPU), a neural network accelerator (NNA), a programmable vision accelerator (PVA) (which can include one or more direct memory access (DMA) systems), one or more vision or vector processing units (VPUs), one or more pixel processing engines (PPEs) (e.g., including a 2D array of processing elements each in north, south, east, west communication with one or more other processing elements in the array), one or more decoupled accelerators or units (e.g., a decoupled lookup table (DLUT) accelerator or unit), etc., a vision processing unit (VPU), an optical flow accelerator (OFA), a field programmable gate array (FPGA), a neuromorphic chip, a quantum processing unit (QPU), an associative processing unit (APU), an arithmetic logic unit (ALU), an application-specific integrated circuit (ASIC), a floating-point unit (FPU), an input / output (I / O) element, a peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) element, etc.
[0267] The communication interface 910 can include one or more receivers, transmitters, and / or transceivers that enable the computing device 900 to communicate with other computing devices via electronic communication networks, including wired and / or wireless communications. The communication interface 910 can include components and functionality to enable communication over any of a number of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., through Ethernet or InfiniBand), low power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, the logic unit 920 and / or the communication interface 910 can include one or more data processing units (DPUs) to transfer data received over the network and / or through the interconnect system 902 directly to the one or more GPUs 908 (e.g., memory of the one or more GPUs 908).
[0268] The I / O ports 912 can enable the computing device 900 to be logically coupled to other devices including I / O components 914, one or more presentation components 918, and / or other components, some of which can be built into (e.g., integrated with) the computing device 900. Illustrative I / O components 914 include a microphone, mouse, keyboard, joystick, game pad, game controller, dish satellite antenna, scanner, printer, wireless device, etc. The I / O components 914 can provide a natural user interface (NUI) that processes air gestures, speech, or other physiological inputs generated by a user. In some cases, inputs are transmitted to an appropriate network element for further processing. A NUI can implement any combination of speech recognition, pen / mouse recognition, facial recognition, biometric recognition, gesture recognition both on screen and adjacent to the screen, air gestures, head and eye tracking, and touch recognition (as described in more detail below) associated with a display of the computing device 900. The computing device 900 can include a depth camera, such as a stereoscopic camera system, infrared camera system, RGB camera system, touchscreen technology, and combinations of these, for gesture detection and recognition.
[0269] The power supply 916 can include a hard-wired power supply, a battery power supply, or a combination thereof. The power supply 916 can provide power to the computing device 900 to enable the components of the computing device 900 to operate.
[0270] The presentation component 918 may include a display (e.g., a monitor, touchscreen, television screen, head-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. The presentation component 918 may receive data from other components (e.g., GPU 908, CPU 906, etc.) and output the data (e.g., as images, videos, sounds, etc.).
[0271] Example network environment
[0272] A network environment suitable for implementing embodiments of this disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may be... Figure 9 It is implemented on one or more instances of one or more computing devices 900—for example, each device may include similar components, features, and / or functions of one or more computing devices 900. Furthermore, in the case of implementing back-end devices (e.g., servers, NAS, etc.), the back-end devices may be included as part of a data center, such as, but not limited to, those described herein.
[0273] Components of a network environment can communicate with each other via a network, which can be wired, wireless, or both. A network can include multiple networks or one of multiple networks. For example, a network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or the Public Switched Telephone Network (PSTN)), and / or one or more private networks. Where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connectivity.
[0274] A compatible network environment may include one or more peer-to-peer network environments (in which case the server may not be included in the network environment) and one or more client-server network environments (in which case one or more servers may be included in the network environment). In a peer-to-peer network environment, the functionality described herein for the server can be implemented on any number of client devices.
[0275] In at least one embodiment, a network environment can include one or more cloud-based network environments, distributed computing environments, combinations thereof, and the like. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which can include one or more core network servers and / or edge servers. The framework layer can include a framework that supports one or more software layers and / or one or more applications of an application layer. The software or applications can respectively contain network-based service software or applications. In embodiments, one or more client devices can use the network-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer can be, without limitation, an open-source software web application framework as can be used for large-scale data processing (e.g., “big data”) using a distributed file system.
[0276] A cloud-based network environment can provide cloud computing and / or cloud storage that performs any combination of the computing and / or data storage functions described herein (or one or more portions thereof). Any of these different functions can be distributed across multiple locations from central or core servers (e.g., one or more data centers that can be distributed across a state, a region, a country, globally, and the like). If a connection with a user (e.g., a client device) is relatively close to an edge server, a core server can designate at least a portion of a function to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), can be public (e.g., available to many organizations), and / or combinations thereof (e.g., a hybrid cloud environment).
[0277] One or more client devices can include at least some of the components, features, and functionality of one or more example computing devices 900 described herein with respect to Figure 9 As examples and not by way of limitation, a client device can be implemented as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, video camera, surveillance device or system, vehicle, boat, spaceship, virtual machine, drone, robot, handheld communication device, hospital device, gaming device or system, entertainment system, vehicle computer system, embedded system controller, kiosk, remote control, appliance, consumer electronic device, workstation, edge device, any combination of these depicted devices, or any other suitable device.
[0278] Example Paragraph
[0279] 1. In some embodiments, a method comprising: superimposing a voxel grid over a three-dimensional (3D) geometry of a first component; determining a set of contact surfaces on the first component based on contact between a subset of grid cells in the voxel grid and the 3D geometry; and generating a second component configured to be coupled to the first component within an assembly based on the set of contact surfaces.
[0280] 2. The method of paragraph 1, further comprising: updating at least one of the set of contact surfaces on the first component or a corresponding set of contact surfaces on the second component based on a gap distance between the first component and the second component.
[0281] 3. The method of any of paragraphs 1-2, wherein updating at least one of the set of contact surfaces or the corresponding set of contact surfaces comprises: generating an occupancy grid corresponding to at least one of the set of contact surfaces or the corresponding set of contact surfaces; and removing one or more grid cells falling within the gap distance from the occupancy grid.
[0282] 4. The method of any of paragraphs 1-3, further comprising: inputting a representation of the first component into a machine learning model; determining one or more attributes associated with the first component via execution of the machine learning model; and determining a top portion of the 3D geometry of the first component based on the one or more attributes.
[0283] 5. The method of any of paragraphs 1-4, further comprising: prior to superimposing the voxel grid over the top portion of the 3D geometry, rotating the 3D geometry of the first component based on the one or more attributes.
[0284] 6. The method of any of paragraphs 1-5, wherein the one or more attributes comprise at least one of: a description of the first component, a type of the first component, an assembly axis, or an assembly direction.
[0285] 7. The method of any of paragraphs 1-6, wherein the machine learning model comprises a visual language model.
[0286] 8. The method of any of paragraphs 1-7, further comprising: updating one or more parameters of a machine learning model based on the first component, the second component, and one or more training objectives associated with assembling the first component and the second component to produce a trained machine learning model.
[0287] 9. The method of any of paragraphs 1-8, further comprising: assembling the assembly using the first component and the second component via execution of a robot.
[0288] 10. The method of any of paragraphs 1-9, wherein generating the second component includes adjusting a denoising process associated with a diffusion model on the set of contact surfaces.
[0289] 11. In some embodiments, at least one processor comprising: processor circuitry to perform operations comprising: projecting a voxel grid over a three-dimensional (3D) geometry of a first component; determining a set of contact surfaces on the first component based on contact between a subset of grid cells in the voxel grid and the 3D geometry; and generating a second component within an assembly coupled to the first component based on the set of contact surfaces.
[0290] 12. The at least one processor of paragraph 11, wherein the operations further comprise: generating an occupancy grid corresponding to at least one of the set of contact surfaces or a corresponding set of contact surfaces on the second component; and removing one or more grid cells from the occupancy grid that fall within a gap distance between the first component and the second component to generate at least one of an updated first component corresponding to the first component or an updated second component corresponding to the second component.
[0291] 13. The at least one processor of any of paragraphs 11-12, wherein the operations further comprise: providing a rendering of the 3D geometry of the first component and one or more instructions to describe the first component as input to a machine learning model; determining one or more attributes associated with the first component via execution of the machine learning model; and determining the set of contact surfaces based on the one or more attributes.
[0292] 14. The at least one processor of any of paragraphs 11-13, wherein the operations further comprise: rotating the 3D geometry of the first component based on the one or more attributes prior to superimposing the voxel grid over the 3D geometry.
[0293] 15. The at least one processor of any of paragraphs 11-14, wherein the one or more attributes comprise at least one of a description of the first component, a type of the first component, an assembly axis, or an assembly direction.
[0294] 16. The at least one processor of any of paragraphs 11-15, wherein generating the second component includes adjusting a denoising process associated with a diffusion model on the set of contact surfaces.
[0295] 17. The at least one processor of any of paragraphs 11-16, wherein the operations further comprise updating one or more parameters of a machine learning model based on the first component, the second component, and one or more training objectives associated with the assembly of the first component and the second component to produce a trained machine learning model.
[0296] 18. The at least one processor of any of paragraphs 11-17, wherein the at least one processor is included in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing optical transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using edge devices; a system implemented using robots; a system for performing one or more generative AI operations; a system for performing operations using one or more large language models (LLMs); a system for performing operations using one or more visual language models (VLMs); a system for performing operations using one or more multi-modal language models (MMLMs); a system for performing operations using one or more visual language action (VLA) models; a system for using or deploying one or more inference microservices; a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system including one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
[0297] 19. In some embodiments, a system comprising: one or more processors to perform operations comprising: projecting a voxel grid over a three-dimensional (3D) geometry of a first component; determining a set of contact surfaces on the first component based on contact between a subset of grid cells in the voxel grid and the 3D geometry; and generating a second component within an assembly coupled to the first component based on the set of contact surfaces.
[0298] 20. The system of paragraph 19, wherein the one or more processors are contained in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing optical transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using edge devices; a system implemented using robots; a system for performing one or more generative AI operations; a system for performing operations using one or more large language models (LLMs); a system for performing operations using one or more visual language models (VLMs); a system for performing operations using one or more multi-modal language models (MMLMs); a system for performing operations using one or more visual language action (VLA) models; a system for using or deploying one or more inference microservices; a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system containing one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
[0299] The disclosure can be described in the general context of computer code or machine-useable instructions, including computer-executable instructions such as program modules, being executed by a computer or other machine, such as a personal data assistant or other handheld device. Generally, program modules including routines, programs, objects, components, data structures, and the like, refer to code that performs particular tasks or implements particular abstract data types. The disclosure can be practiced in a variety of system configurations, including hand-held devices, consumer electronics, general- purpose computers, more specialty computing devices, and the like. The disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network.
[0300] As used herein, the term “and / or,” with respect to a listing of two or more elements, means that one or more of the listed elements can be included. For example, “element A, element B, and / or element C” can include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or element A, B, and C. Furthermore, “at least one of element A or element B” can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, “at least one of element A and element B” can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0301] The subject matter of the present disclosure is described with specificity herein to meet statutory requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors have contemplated that the claimed subject matter might also be embodied in other ways, to include different steps or combinations of steps similar to the ones described in this document, in conjunction with other present or future technologies. Moreover, although the terms "step" and / or "block" might be used herein to connote different elements of methods employed, the terms should not be interpreted as implying any particular order among or between various steps herein disclosed unless and except when the order of individual steps is explicitly described.
Claims
1. A method comprising: Overlay the voxel mesh on top of the three-dimensional 3D geometry of the first component; The set of contact surfaces on the first component is determined based on the contact between a subset of mesh cells in the voxel mesh and the 3D geometry; and A second component, configured to be coupled to the first component, is generated based on the set of contact surfaces.
2. The method of claim 1, further comprising: updating at least one of the set of contact surfaces on the first component or the corresponding set of contact surfaces on the second component based on the gap distance between the first component and the second component.
3. The method of claim 2, wherein updating at least one of the set of contact surfaces or the corresponding set of contact surfaces comprises: Generate an occupation mesh corresponding to at least one of the set of contact surfaces or the corresponding set of contact surfaces; and Remove one or more grids from the occupied grid that fall within the gap distance.
4. The method of claim 1, further comprising: The representation of the first component is input into the machine learning model; The execution of the machine learning model determines one or more attributes associated with the first component; and The top of the 3D geometry of the first component is determined based on one or more of the attributes.
5. The method of claim 4, further comprising: rotating the 3D geometry of the first component based on one or more properties before superimposing the voxel mesh on the top of the 3D geometry.
6. The method of claim 4, wherein one or more attributes include at least one of the following: a description of the first component, a type of the first component, an assembly axis, or an assembly direction.
7. The method of claim 4, wherein the machine learning model comprises a visual language model.
8. The method of claim 1, further comprising: updating one or more parameters of a machine learning model based on the first component, the second component, and one or more training objectives associated with assembling the first component and the second component to generate a trained machine learning model.
9. The method of claim 1, further comprising: assembling the component using the first component and the second component via execution by a robot.
10. The method of claim 1, wherein generating the second component comprises: The denoising process is adjusted to be associated with the diffusion model on the set of contact surfaces.
11. At least one processor, comprising: Processor circuitry for performing operations, said operations including: Project the voxel mesh over the three-dimensional 3D geometry of the first component; The set of contact surfaces on the first component is determined based on the contact between a subset of mesh cells in the voxel mesh and the 3D geometry; and A second component, coupled to the first component, is generated within the component based on the set of contact surfaces.
12. The at least one processor of claim 11, wherein the operation further comprises: Generate an occupation grid corresponding to at least one of the set of contact surfaces or the set of corresponding contact surfaces on the second component; and Remove one or more grids from the occupied grid that fall within the gap distance between the first component and the second component to generate at least one of the following: an updated first component corresponding to the first component or an updated second component corresponding to the second component.
13. The at least one processor of claim 11, wherein the operation further comprises: Provides rendering of the 3D geometry of the first component and one or more instructions for describing the first component as input to a machine learning model; The execution of the machine learning model determines one or more attributes associated with the first component; and The set of contact surfaces is determined based on one or more of the aforementioned attributes.
14. The at least one processor of claim 13, wherein the operation further comprises: rotating the 3D geometry of the first component based on one or more properties before superimposing the voxel mesh on the 3D geometry.
15. The at least one processor of claim 13, wherein the one or more attributes include at least one of the following: a description of the first component, a type of the first component, an assembly axis, or an assembly direction.
16. The at least one processor of claim 11, wherein generating the second component comprises: adjusting a denoising process associated with a diffusion model on the set of contact surfaces.
17. The at least one processor of claim 11, wherein the operation further comprises: updating one or more parameters of a machine learning model based on the first component, the second component, and one or more training objectives associated with the components of the first component and the second component to generate a trained machine learning model.
18. The at least one processor according to claim 11, wherein the at least one processor is included in at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system for performing one or more simulation operations; A system for performing one or more digital twin operations; A system for performing optical transmission simulation; A system for performing collaborative content creation for 3D assets; A system for performing one or more deep learning operations; Systems implemented using edge devices; Systems implemented using robots; A system for performing one or more generative AI operations; A system for performing operations using one or more large language model LLMs; A system for performing operations using one or more visual language models (VLMs); A system for performing operations using one or more multimodal language models (MMLMs); A system for performing operations using one or more Visual Language Action (VLA) models; A system for using or deploying one or more inference microservices; A system for performing one or more conversational AI operations; A system for generating synthetic data; A system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.
19. A system comprising: One or more processors for performing operations, said operations including: Project the voxel mesh over the three-dimensional 3D geometry of the first component; The set of contact surfaces on the first component is determined based on the contact between a subset of mesh cells in the voxel mesh and the 3D geometry; and A second component, coupled to the first component, is generated within the component based on the set of contact surfaces.
20. The system of claim 19, wherein the one or more processors are at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system for performing one or more simulation operations; A system for performing one or more digital twin operations; A system for performing optical transmission simulation; A system for performing collaborative content creation for 3D assets; A system for performing one or more deep learning operations; Systems implemented using edge devices; Systems implemented using robots; A system for performing one or more generative AI operations; A system for performing operations using one or more large language model LLMs; A system for performing operations using one or more visual language models (VLMs); A system for performing operations using one or more multimodal language models (MMLMs); A system for performing operations using one or more Visual Language Action (VLA) models; A system for using or deploying one or more inference microservices; A system for performing one or more conversational AI operations; A system for generating synthetic data; A system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.